The Citation Ledger Is Fine. The Citation Currency Is Dying.
When I need to get oriented in a literature I don't know, I now start by interrogating a language model. Twenty minutes of questions does roughly what two days of skimming used to do. The knowledge of fifty papers reaches me: their framings, their standard objections, their load-bearing results, the names of the three papers everyone argues about. I will end up citing maybe three of them, the ones I actually open, because something specific needs verifying: a number I'm about to repeat, a method I'm about to borrow, a claim that sounded too convenient.
A decade ago, most of those fifty would have gotten at least a skim from me, and a handful would have earned a citation for the skim. The knowledge transfer still happened this time. The toll was not paid.
I want to argue that this small, private change in how researchers read is quietly splitting apart two functions that citation has bundled together for three centuries, and that the split has consequences for how scientific reputation will be denominated. I also want to be upfront that the empirical evidence is currently a mess that points in two directions at once, and that the strongest objection to my argument came from examining my own reading workflow. Then I will make three falsifiable predictions with dates attached, because arguments in this genre are cheap and getting cheaper, and the only honest way to write one is to specify how it loses.
Update, July 2026: the measurements are in, and the mechanism I bet on lost. The predictions below are left exactly as written; what came back is reported after them.
Citation is two things wearing one coat
This part is textbook, and I want to be precise about what is and is not my claim, because the distinction itself belongs to other people.
In the standard model of scholarly communication, formulated by Roosendaal and Geurts in 1997, the system performs distinct functions: registration (fixing who found what, when, in an auditable public record), certification, awareness, and archiving, with later extensions adding reward as its own function. Merton's sociology of science supplies the reward half its texture: scientists give away their findings and are paid in acknowledgment; citations operate as a symbolic currency within the reward system of science. A 2026 review of citation theory by Shu and Jia restates this consensus without needing to revise it: citations are acknowledgment bestowed by peers, and that acknowledgment is what careers are built from.
So one physical act, one string in a reference list, does two unrelated jobs:
- A ledger entry. Registration, priority, provenance. The paper exists, it is timestamped, this specific claim traces to it. If a dispute arises about who showed what first, the ledger settles it.
- A unit of currency. Citation counts aggregate into reputations, and reputations convert into jobs, grants, tenure, invitations. The h-index is a bank statement. Entire national evaluation systems disburse money against it.
Everyone who thinks about scholarly infrastructure knows these are different functions. What is easy to miss, because it has been true for so long that it reads as a law of nature, is why they have always moved together.
The coupling was an accident
The two functions were never designed to share a vehicle. They share one because, for the entire history of modern science, there was exactly one channel through which knowledge moved from a paper into a new researcher's head: that researcher, or someone in their immediate epistemic neighborhood, read it. Reading was the transport layer. And the citation act sat directly on top of reading: you read, you used, you cited. The ledger entry and the currency transfer happened in the same keystroke because the behavior that justified both was the same behavior.
This means the correlation we all rely on, "highly cited" as a proxy for "widely used and influential," is not a property of citation. It is a property of reading being the only way in. Every reputational structure built on citation counts (the h-index, impact factors, evaluation exercises, promotion dossiers) inherits an unexamined assumption: that knowledge flow leaves citation tracks because it physically must.
For three hundred years that assumption was safe. It is now false.
Language models are a new transport layer
LLMs move knowledge from papers to researchers without the researcher touching the paper. The ideas arrive synthesized, anonymized, and blended, through training data or through retrieval. Every stage of the old chain still exists: the paper, the DOI, the priority claim, the reference list format. But the stage that minted currency, a human reading and then citing, fires less often per unit of knowledge transferred. My fifty-papers-to-three-citations morning is one researcher's worth of the drain. Multiply it by every researcher who has moved their literature triage into a chat window.
Two clarifications, because there are nearby arguments this is not.
First, this is a reader-side problem, and it is distinct from the writer-side problem that already has a literature. Earp and colleagues argued in Nature Machine Intelligence that LLM-assisted writing breaks the chain of scholarly credit: your text may reproduce distinctive ideas from sources you never saw, so citations fail to get attached during composition, with no intent to deceive anywhere in the loop. That is a real and well-described failure. Mine is one step earlier in the pipeline and, I think, larger: the consumption of knowledge no longer generates the impulse to cite, because consumption no longer involves the source. The writer-side problem corrupts the citations that get written. The reader-side problem prevents them from being conceived.
Second, this is not the "AI makes us understand less" argument. Messeri and Crockett's warning about illusions of understanding is about what happens inside the researcher's head. My concern is about what happens in the accounting layer regardless of whether the understanding is real. Even if LLM-mediated reading produced perfect comprehension, the ledger would still stop recording it.
There is a theoretical ancestor worth naming. Merton described "obliteration by incorporation": a finding becomes so absorbed into common knowledge that people stop citing its origin (nobody cites Darwin for evolution). Blechinger recently extended this to generative AI, arguing that LLMs perform obliteration at industrial scale, dissolving authorship into parameters. I agree, with one amendment: obliteration by incorporation was historically the end state of the most successful ideas, a retirement honor reached after decades. LLMs apply it to everything, immediately, including the merely useful middle of the distribution that used to live on routine citations.
The web already ran this experiment
If you want to know what happens when an intermediary starts answering questions instead of routing traffic, you do not have to speculate. Search engines used to be traffic-routing intermediaries; LLM-augmented search turned them into providers of synthesized answers. Gholami and colleagues, studying this transition, report that roughly 60 percent of Google searches now end without any click to the underlying web. Publishers call the phenomenon zero-click search, and industry measurements of the referral collapse, whatever their individual reliability, all point the same direction: double-digit declines in traffic to the pages the answers were synthesized from.
The scholarly version of this observation is already in print. Roohi Ghosh, writing on The Scholarly Kitchen in May 2026, transposed the zero-click frame to research discovery and asked the question this essay is trying to answer properly: if papers are not actually being read, what exactly are citations measuring?
My answer: they are still measuring registration, perfectly well. They are ceasing to measure use. Those used to be the same measurement. They are not anymore, and nearly everything downstream of that fact is unexamined.
What the data actually says, which is: something confusing
I spent a while doing prior-art archaeology before writing this (forty-some searches, seventeen full texts), and the empirical picture is genuinely unsettled in a way worth laying out rather than smoothing over.
The largest relevant study is Kusumegi and colleagues in Science (December 2025). Working with 2.1 million preprints and, critically, 246 million online views and downloads of scientific documents linked to citation records, they found that researchers who adopt LLMs produce 24 to 89 percent more output, and cite more diversely: younger work (by about 0.4 years), less-cited work (about 2.3 percent lower citation impact), more books. If you expected LLM adoption to concentrate citation on famous papers, this is evidence against you.
Meanwhile Algaba and colleagues, analyzing 274,951 references generated by GPT-4o, found the opposite tendency in the model itself: when a model proposes citations, it over-selects already-highly-cited papers, a Matthew-effect amplifier that survives controls for publication year, venue, title length, and author count.
These results are not actually contradictory; they sit in different cells of a two-by-two. Any LLM-mediated citation decomposes into two mechanical questions: where does the candidate pool come from (live retrieval, or the model's parametric memory), and who commits the final selection (the human, or the model). Kusumegi measured retrieval plus human commitment. Search costs fall hardest on the long tail, because famous papers were already free to find, so the human gets carried toward more diverse sources. Algaba measured parametric memory plus model commitment. Parametric memory is grown from corpus frequency, so sampling from it reproduces the prestige distribution and then amplifies it. Neither study touched the off-diagonal cells where most real usage now lives: an agent that retrieves and commits inherits the bias of whatever ranking function the search index uses, and citation-count-ordered rankers quietly rebuild the Matthew effect inside nominally "grounded" systems; a human who commits but draws candidates from model memory is anchored to the head of the distribution before choosing has even begun. The net effect on the citation ledger is therefore a horse race between mediation modes, and anyone confidently telling you "LLMs will concentrate citations" or "LLMs will democratize citations" is reporting their prior, not the evidence.
It gets worse for clean narratives. The one citation trend everyone wants to attribute to ChatGPT turns out to predate it by a decade. Smyth and Cunningham, working across 16.7 million papers in 23 fields, documented that the citation premium of review articles has been decaying since roughly 2010-2015: systematic reviews' normalized citation index fell from about 1.04 to 0.86, meta-analyses from 1.40 to 0.81, while ordinary research papers gained ground. Their explanation is supply-side: review-production tooling made reviews abundant, and abundance diluted the premium. Generative AI accelerates a collapse that automated search started. If I want to claim a reader-side effect (people stop citing reviews because a chatbot now does what a review did), I have to find it as an additional break on top of a fifteen-year trend that has nothing to do with my mechanism. That is a real statistical burden and I will carry it explicitly.
One more data point, small but suggestive. Shen and colleagues tried to build an impact metric from LLM parametric memory: how strongly do 17 models remember 549 computer science papers? The correlation between model memory and citation counts was 0.15. Nearly zero. Ideas can be thoroughly absorbed by the models everyone now reads through, while the citation record barely notices. That is the decoupling, photographed from an odd angle.
Anyone telling you a clean story about LLMs and citations is currently ahead of the evidence. Including me. Hence the predictions section.
The claim
Here it is in one sentence: citation will not die; it will retire into pure registration, the way a patent number works, while the reward function migrates to other signals.
A patent number is a perfectly healthy institution. It establishes priority, it is legally load-bearing, it is consulted when disputes arise, and it will outlive all of us. Nobody's reputation is denominated in patent-number lookups. Gold is still an anchor of value and a settlement layer; once day-to-day transactions stopped being conducted in coins, the social meaning of holding gold changed anyway. The ledger survives. The currency stops circulating through it.
Notice what this claim is not. It is not "citation counts will go to zero." Registration alone generates citations: every paper still must anchor its claims somewhere, priority disputes still need adjudication, and (as I will argue below) verification of machine-generated syntheses may generate a durable floor of citations all by itself. The claim is that the meaning of the count changes underneath the number: from "this many acts of human reading and endorsement" to "this many acts of record-keeping." Evaluation systems that keep paying out against the number will increasingly be paying for bookkeeping.
Where the reward function goes: three candidates, all weak
If reward detaches from citation, it has to land somewhere; reputation abhors a vacuum. The candidates I can see, with their problems stated as plainly as their promise:
Artifact reuse. Your dataset, benchmark, or library as a live dependency in other people's pipelines. This is the most credible candidate: reuse is costly to fake, automatically logged (download counts, dependency graphs, forks), and closer to "actual use" than citation ever was. There is even a working implementation of dependency-as-credit: Sochat's CiteLang derives credit shares from package dependency trees, no DOI required. Two problems. Reuse is brutally power-law concentrated, worse than citations: on Hugging Face, 82 datasets account for roughly 80 percent of all dataset downloads, and Koch and colleagues showed ML research clustering onto a shrinking set of legitimizing benchmarks. A currency more concentrated than the one it replaces is not obviously an improvement. And note the direction of the existing movement here: the software-citation and data-citation initiatives (FORCE11 and descendants) have spent a decade trying to pull artifacts into the citation system, to get code and data cited like papers. The migration I am describing runs the other way.
Training-data attribution. Shapley-style allocation of an LLM answer's value back to the source documents it drew on. The mechanism exists: Ye and Yoganarasimhan built exactly this for LLM search platforms, motivated by exactly the zero-click drain, to route revenue shares to news and content sources. Nobody has run it for scholarly credit. And here is the detail I find most telling: a recent stakeholder framework for data attribution by Wührl and colleagues explicitly carves academics out of the compensation economy, reasoning that academics may be willing to have their texts used for free since "their main 'currency' is citations and social credit." The successor mechanism is being built, right now, on the assumption that the old currency still pays. That assumption is the thing dissolving. There is also a technical problem: attribution methods are numerically shaky at LLM scale (influence-function approximations degrade badly), so a currency built on them inherits their fragility, plus every gaming incentive that citation ever had, aimed at a more opaque target.
Attention. Social media reach, newsletter mentions, talk invitations, podcast circuits. The worst candidate on every dimension except availability. Notably, the altmetrics literature itself has never positioned altmetrics as citation's successor; fifteen years of that research frames them as a complement, or as an early predictor of future citations, which makes altmetrics parasitic on citation's authority rather than independent of it. If attention becomes the interim currency anyway, it will be by default, not by argument.
The uncomfortable synthesis: the old currency can drain faster than any successor matures. Fields have lived through reputation-system transitions before, and the interim period rewards whoever is best at the crudest available signal. In this list, the crudest available signal is attention. That prospect should worry people who are good at science and bad at posting.
The obvious counter-move, and why I think it fails
The incumbents are not asleep, and their response deserves engagement rather than a strawman.
In June 2026, Cambridge University Press, COUNTER, and NISO convened about forty experts from libraries, publishers, funders, and AI companies to design exactly the fix you would expect: make AI usage of scholarly content traceable and measurable. Persistent identifiers embedded in machine-readable chunks, content hashes, digital signatures, provenance metadata, COUNTER usage reporting extended to AI agents. Their diagnosis is verbatim mine: AI systems retrieve, summarize, and synthesize across papers without a researcher ever clicking through, and that is unmeasured value extraction.
Same diagnosis, opposite prognosis. They believe the drain is a measurement gap; I believe it is a regime change. Here is the crux.
Telemetry restores measurability. It does not restore currency. What made a citation worth accumulating was never the count itself; it was what the count certified: a scarce, costly act of human endorsement. A researcher spent limited reading time, exercised judgment, and attached their name, in public, in a document they signed, to the statement that this specific source mattered to this specific work. Scarcity, cost, and staked reputation are what let citation counts function as money. A COUNTER report saying an AI agent retrieved your paragraph 40,000 times last month has none of the three. It is not scarce (machine retrieval is free and infinite), not costly (nobody spent anything), and nobody staked anything (no human judgment attaches to any individual retrieval). You cannot mint currency out of it; you can only thicken the ledger.
Thickening the ledger is genuinely useful! Licensing negotiations, copyright enforcement, infrastructure funding cases: all better with telemetry. But if the plan is to feed machine-usage counts into reputation systems as citation's successor, the plan reinvents the web-traffic economy inside science, and we have already seen what optimizing for machine-legible engagement does to an information ecosystem.
The strongest objection I know
The best counterargument to all of this came from watching my own workflow more honestly, and it deserves its own section rather than a footnote.
Recall the fifty-to-three morning. Why three, rather than zero? Because model output is probabilistic and I know it. Any number I am about to repeat, any method I am about to build on, any claim that will bear load in my own argument, I go to the source and check, and having checked, I cite. Verification demand puts a floor under citation flow. As long as LLMs hallucinate at any nonzero rate, and as long as frontier and contested claims matter most, the reading-then-citing loop cannot fully close. You could sharpen this into a rebuttal: the currency is not draining, it is concentrating into the citations that were always the real ones, and the forty-seven papers I no longer cite were mostly ritual anyway (background padding, related-work etiquette, reviewer appeasement). Perhaps LLMs are not killing the currency but burning off its inflation.
I take this seriously, and here is why I think it rescues less than it seems to. Look at which function the surviving citations perform. Verification-driven citation is registration: I checked this specific claim against this specific source, here is the pointer. That is ledger behavior, auditable and essential, and it is exactly what I predict survives. What it does not do is what the currency required: map breadth of intellectual influence onto breadth of acknowledgment. The forty-seven papers that shaped my understanding without individually bearing load get nothing, where the old regime paid at least some of them for the skim. A reward system fed only by verification pays authors for being checkable, not for being formative. Those are different virtues, and the second one is what the currency was supposed to track.
But the objection generates something better than a rebuttal: a measurable signature that distinguishes the two stories. If the verification-floor story is right, citation composition should shift before citation volume does: load-bearing citations (methods used, data reused, specific claims tested) should hold steady while narrative citations (background, framing, related-work courtesy) decay, and the shift should be steepest where LLM adoption is highest. Citation-context classification is a mature enough tool to test this. I have added it to the measurement plan below, and I note for the record that the objection came from the workflow of the person writing this essay, which is either good epistemics or a very small sample size.
Three predictions
A position like this is worthless unless it can lose. Here is how mine loses, with the measurement approach for each, because I am running these rather than proposing them.
P1, the horse race. Citation concentration will move in the direction of whichever mediation mode dominates. Where selection stays with humans choosing over retrieved candidates, concentration falls (the Kusumegi direction); where commitment is delegated to models, or to citation-ranked retrieval, concentration rises (the Algaba direction). Operationally: concentration time-series (Gini, top-percentile shares) per field-cohort, conditioned on field-level LLM penetration. I have a pilot of this running on OpenAlex data now (its per-work, per-year citation counts make the series cheap to build); full cohorts 2015-2024 this autumn. If concentration is flat and independent of penetration through 2028, the mediation mechanism is wrong.
P2, the mechanism split. The review-article citation decline has a documented pre-LLM, supply-side trend. If reader-side drain is real, there should be an additional slope break after 2023, ordered by field-level penetration, on top of the 2010s trajectory Smyth and Cunningham measured. Difference-in-slopes on their design, extended past the ChatGPT boundary. If the decay curve continues smoothly on its old trajectory, the reader-side story loses to the boring supply-side one, and I will report that.
P3, the core bet. The correlation between artifact reuse and citation counts will fall, earliest and fastest in high-LLM-penetration fields. Today the two are tightly and positively coupled: papers with popular repositories are cited more, not less, and the coupling has historically strengthened with repository popularity. Prior work runs entirely against me here, which is what makes the bet informative. One confound I have to handle honestly: İlter documented that about 17 percent of citations in AI-assisted survey papers fail to resolve to any real document, and that phantom rate concentrates in exactly the high-penetration fields where I predict the effect, so any measured decoupling has to survive controls for citation-record corruption. If reuse-citation coupling holds firm through 2028, this whole essay is wrong, and you may cite it against me, which would at least be one more citation.
And the fourth measurement, promoted from the objection above: citation composition shifts before volume. Load-bearing citation types persist, narrative types decay, ordered by penetration. If composition stays flat too, then nothing is draining and I have written a long essay about my own reading habits.
I am starting the full measurement runs this autumn, OpenAlex first, artifact-reuse data from Papers with Code, Hugging Face, and GitHub dependency graphs. If you run infrastructure at a publisher, an index, or a preprint server and hold linked readership-citation data (the kind of thing the Kusumegi team had), I would genuinely like to talk.
What the measurements returned
Added July 2026. Everything above is unchanged.
P3 was the core bet, and it is the one I ran. I linked 20,529 arXiv papers from 2015 to 2025 to the GitHub repositories their authors designated as official, measured reuse by annual fork counts and citation by annual citation counts, and estimated the rank association between the two flows in every cohort-by-year cell.
The phenomenon is there. The association fell from roughly 0.45–0.50 in the late 2010s to roughly 0.25 by 2024. Prior work ran against me on this and prior work was wrong about the direction of travel. That is the part of the bet I won.
The mechanism is dead, and it was my mechanism. Three separate results kill it:
The decline is not dated to language models. The post-2023 shift is −0.002, with a 95 percent interval of [−0.061, +0.057], against a decade-long decline of about 0.23. That interval excludes any break larger than about a quarter of the total movement. There is no discontinuity where I said there would be one.
It does not sort by field-level exposure. The interaction between period and field-level LLM penetration comes back at p = 0.82. The gradient I predicted, earliest and fastest where adoption is highest, is absent.
Worst of all for my story, the variation is on the calendar-period axis, not the publication-cohort axis. Holding the set of papers fixed, the 2015 cohort declines from 0.49 to 0.04 over its own lifetime, as steeply as anything recent. With period in the model the cohort coefficient is −0.0017 (p = 0.74). Whatever is happening is happening to papers published in 2015 in the same year and to the same degree as to papers published in 2023. A reader-side behavioral shift that arrived with ChatGPT does not have that shape. Something that moves all vintages simultaneously, every year, for a decade, is not a technology adoption curve.
The fourth measurement, the one the objection generated, also lost. I predicted composition would shift before volume: load-bearing citations holding while narrative ones decay. The share of substantively influential citations is flat across the decade, to within a percentage point, and both types weaken at the same rate. That one stings a little more, because it was the prediction I derived from the objection I was proudest of having taken seriously.
And a caution I earned the hard way. A naive version of that composition test showed substantive citations weakening significantly faster, exactly as predicted. It was an artifact of cohort-varying zero inflation. Raising the threshold collapsed the difference monotonically and then reversed its sign. I had a result that confirmed my own hypothesis and it was a pipeline artifact, which is the failure mode this whole essay should have made me most afraid of.
One result I did not predict, and it cuts against the essay's practical claim harder than the mechanism refutation does. The decoupling does not reach the top of the distribution. Rather than sampling, I enumerated the entire frame, all 160,150 papers from 2015 to 2024 with an author-designated repository, and looked at works that are both substantially reused and substantially cited. Among those, across all ten cohorts, the association has been flat at about 0.28 for a decade: a trend of −0.0031 per year, 95 percent interval [−0.0156, +0.0094], n = 3,885. That interval excludes the full-sample estimate of −0.0241. The test has 80 percent power against a slope of 0.0151, and the population-level effect is 1.6 times that, so this is a null with power behind it rather than thin data.
Think about what that does to the argument above. I claimed the reputation currency is dying. But hiring, promotion and funding are argued over exactly that tail, and in the tail the two signals still track each other as well as they did in 2015. The currency is not being devalued where it is actually spent. What has come apart is the relationship in the bulk of the distribution, which matters for field-level aggregates and dashboards and any statistic dominated by the many rather than the few. That is a real finding and a narrower one than the essay I wrote.
P2 and P1, honestly. P2 as literally specified, extending Smyth and Cunningham's review-article design past the ChatGPT boundary, I did not run. What I did run is the same mechanism claim on my own measure, and it failed there. P1, the concentration horse race, I did not settle. There is an exploratory probe in the replication package and its own memo says it is not a conclusion; I am not going to dress that up as a tested prediction.
So where does this leave the essay. The split I described is real and measurable, and I would still write the first half. The causal story I attached to it is not supported, and if I had only run the cohort-axis version of the test I would have reported a confirmation, because on the wrong axis the cohort gradient looks like exactly what I predicted. Two accounts survive and neither is mine: that the population reusing artifacts is separating from the population that publishes and cites, and that a GitHub fork no longer measures what it did in 2016. Distinguishing those is the next study.
I said a position like this is worthless unless it can lose. It lost. I would rather report that than the version where I quietly stop mentioning the predictions.
Two disclosures
First: parts of this argument were developed in dialogue with language models, which is either fitting or damning depending on your priors. One bias correction worth applying to me and to every essay in this genre: the training corpus oversamples complaint. The ninety percent of the scholarly system that works fine does not blog about it. A model's picture of academia, and therefore mine after long exposure, skews structurally pessimistic. Discount accordingly; the predictions are there so the discount has something to grind against.
Second: the registration layer is already creaking under the flood this essay describes, and I have the receipts personally. arXiv announced in October 2025 that position papers and surveys in CS will only be accepted with proof of prior peer review, because LLM-generated ones were arriving faster than moderators could process. Three of my own empirical papers recently sat in the moderation queue for weeks, collateral damage of that flood. And this essay, being a position piece, cannot go on arXiv at all until some journal accepts it. Think about what that policy is: the ledger, overwhelmed, is outsourcing verification to journals and defending registration by restricting entry. The registration function is being defended at the expense of throughput, exactly the priority ordering my argument predicts institutions will choose when the two functions come apart. I would prefer corroboration that did not cost me three moderation queues, but you take your data where you find it.
Links: Roosendaal & Geurts 1997 · Shu & Jia 2026 · Earp et al., Nature MI 2025 · Messeri & Crockett, Nature 2024 · Blechinger, KULA 2026 · Ghosh, Scholarly Kitchen 2026 · Gholami et al. 2026 · Kusumegi et al., Science 2025 · Algaba et al. 2025 · Smyth & Cunningham 2025 · Shen et al. 2026 · Sochat, JOSS 2022 (CiteLang) · Koch et al., NeurIPS 2021 · HF dataset concentration (arXiv 2401.13822) · Ye & Yoganarasimhan 2026 · Wührl et al. 2025 · İlter 2026 · Cambridge/COUNTER/NISO workshop report 2026 · arXiv policy 2025-10-31
This essay was the blog stage of Weakening in Real Time, which reports the measurements above. The essay is left as written; the paper is where the result lives.
Citation
Citing the research rather than this note? Cite the paper: 10.5281/zenodo.21452779. That DOI always resolves to the current version.
If you refer to this note, please cite it as:
@misc{an2026citation,
title = {The Citation Ledger Is Fine. The Citation Currency Is Dying.},
author = {An, Tao},
year = {2026},
howpublished = {\url{https://tao-hpu.github.io/articles/citation-decoupling}},
note = {Personal research notes}
}