Undiscovered Public Knowledge

Some of science's most important findings go unrecognized for decades. Artificial intelligence is changing who wakes them, and how many more there will be.

BY GEOFFREY W. SMITH
SEPTEMBER 29, 2026

The jellyfish paper

In the summer of 1961, a Japanese chemist named Osamu Shimomura began catching jellyfish off the docks at Friday Harbor, Washington. He wanted to understand the chemistry behind the faint green glow of Aequorea victoria. The first report of the green fluorescent protein (GFP) appeared in 1962 in the Journal of Cellular and Comparative Physiology, and over the summers of the following 19 years, ~85,000 of the jellyfish were caught off Friday Harbor to isolate enough of the protein to characterize it more fully, as Ulkar Aghayeva recounts in Works in Progress. The work was meticulous, and for a long time it was nearly useless. Asked about it many years later, Shimomura admitted, “I didn't know any use of . . . that fluorescent protein, at that time.”
‍

Then, three decades on, the world changed around the paper. The fluorescing protein's gene was cloned in 1992 and passed to the biologist Martin Chalfie, who had the idea of expressing GFP in E. coli and in C. elegans worms, and thus showed that it could serve as a fluorescent marker in living organisms. When Chalfie and colleagues published in Science in 1994, they cited Shimomura's 1962 and 1979 papers, which suddenly drew a flood of new citations. A curiosity about jellyfish became one of the most-used tools in cell biology.


Bibliometricians have a name for papers like Shimomura's; they call them “sleeping beauties.” A sleeping beauty is published, ignored, and then, sometimes a generation later, kissed awake. The fairy-tale framing is a little twee, but it points at something unsettling about how knowledge grows. We tend to picture the scientific literature as a frontier that moves forward. The sleeping beauties suggest it is closer to an attic, where things of great value lie in boxes that no one has thought to open. And we are now building machines that can open all the boxes at once.
‍

The term comes originally from Anthony van Raan, a professor of quantitative studies of science at Leiden University. In a 2004 paper in Scientometrics, he defined a “Sleeping Beauty in Science” as a publication that goes unnoticed for a long time and then, fairly abruptly, draws a great deal of attention, and he reported what he believed was the first extensive measurement of how often this happens. He scored papers on three criteria: how long they slept, how deeply (the average number of citations during the dormant period), and how intensely they woke (the citations received in the four years after). Using somewhat arbitrary thresholds, he found sleeping beauties at a rate of about 0.01 percent of papers published in a given year (Aghayeva, Works in Progress).
‍

That number made the phenomenon look like a rarity, a handful of mispriced assets in an otherwise efficient market. It did not hold up. In 2015, a team at Indiana University led by Qing Ke and Alessandro Flammini returned to the question with far more data. They introduced a single measure capturing both the intensity of a paper's recognition and the length of its sleep, and concluded that sleeping beauties are far from exceptional. They noted that earlier estimates of scarcity depended heavily on arbitrary thresholds applied to small or single-discipline datasets. Drawing on 384,649 papers from American Physical Society journals and more than 22 million from Web of Science, they found a wide, continuous range of delayed recognition in every scientific field, and raised the estimated share of sleeping beauties at least 100-fold over van Raan's (Aghayeva, Works in Progress). Delayed recognition, in other words, is not a curious exception. It is a normal part of how science works.

‍

Why good papers sleep

The canonical example of a sleeping beauty is also the most famous. Ke and his colleagues describe the 1935 Einstein, Podolsky, and Rosen “paradox” paper as the model sleeping beauty. The so-called EPR paper was hardly obscure. It set off intense debate and even made a New York Times headline, but its citations stayed well below what one would expect, because its central question needed an experimental test that was not feasible for a long time. John Bell supplied necessary theoretical additions in 1964; but it was only in the early 1980s, thanks to progress in laser physics, that conclusive experiments became possible. And even then the EPR paper's citation spike did not come until the late 1980s (Aghayeva, Works in Progress). The idea had been ready for half a century. The instruments had not.
‍

So why do good papers sleep? Aghayeva sorts the reasons into roughly three kinds. Sometimes contemporaries lack the tools to test an idea (the EPR case); sometimes the community does not understand what has been found, perhaps for want of theory; and sometimes the reason is more mundane: the paper appeared somewhere obscure and never reached the right readers.


The examples in each category are striking. Herbert Freundlich's 1907 paper on adsorption in solutions began to be cited regularly only in the early 2000s, because it bore on new water-purification technologies. William Hummers and Richard Offeman's 1958 “Preparation of Graphitic Oxide” also woke in the 2000s, because it proved relevant to graphene. In 1911, the pathologist Francis Peyton Rous reported that a filtered tumor extract from a cancerous chicken could cause sarcoma in a healthy one. The significance of that work went unrecognized until after 1951, when a murine leukemia virus was isolated and tumor virology began in earnest; Rous received the Nobel Prize in Medicine in 1966, 55 years after publishing (Aghayeva, Works in Progress).
‍

Aghayeva is careful about one point, and it deserves emphasis: citations are a rough proxy for attention. Karl Pearson's 1901 paper describing what became principal components analysis appears to have slept for a century before a citation surge around 2002. Yet the technique itself was widely used throughout that period. The same holds for Fisher's exact test and for the 1949 Monte Carlo paper by Nicholas Metropolis and Stanislaw Ulam, whose methods were already having an enormous effect even while the papers themselves looked dormant. Some beauties were never asleep. They were simply living under another name.

‍

Dormancy beyond the lab

The pattern extends beyond scientific publishing. The evolutionary biologist Andreas Wagner, of the University of Zurich, has argued that biology itself is full of sleeping beauties. In his 2023 book Sleeping Beauties: The Mystery of Dormant Innovations in Nature and Culture, he explains that novel traits sometimes have to wait for the environment to change before they become useful. He describes antibiotic-resistance genes in ancient bacteria as “solutions in search of a problem,” proteins that were potent but useless until the right enemy arrived. He applies the same lens to lineages such as mammals and grasses which diversified explosively only millions of years after they first appeared.
‍

Technology has its own sleepers. In a 2024 essay for the newsletter Scope of Work, Natasha Balwit-Cheung traces the history of soda ash, the alkali behind glass and soap. Ernest Solvay's industrial process made him enormously wealthy, but his first patent application was denied, partly because the chemical reaction at its core had been known since 1811, discovered decades earlier by a young Augustin-Jean Fresnel. She notes other long dormancies: the ninth-century cryptologist Al-Kindi used letter frequencies to break ciphers centuries before Pascal and Fermat laid the foundations of probability, and Mendelian inheritance was proposed and then forgotten for decades.
‍

Seen this way, a sleeping beauty is less a property of a paper than of the relationship between a paper and its world. What wakes it is a change in context: a new instrument, a new theory, a new problem, or a reader who happens to hold the missing half of the puzzle. That last possibility turns out to be the one machines are best placed to supply.

‍

Undiscovered public knowledge

In the mid-1980s, a University of Chicago information scientist named Don Swanson noticed that two bodies of medical literature seemed to be talking past each other. One showed that patients with Raynaud's disease, a circulatory disorder, had elevated blood viscosity and platelet aggregation. The other showed that dietary fish oil lowered both. His 1986 essay set out to identify two literatures that were logically connected but did not interact, each barely acknowledging the other, and to suggest that important relationships might be escaping their authors' notice. Swanson proposed that fish oil might help Raynaud's patients. The hypothesis was later confirmed by DiGiacomo and colleagues in 1989. He called what he had found undiscovered public knowledge: knowledge that is published somewhere yet remains largely unknown.


Swanson's method, which he later partly automated, founded a small field called literature-based discovery. For decades it stayed a niche, limited by crude text processing and a flood of spurious connections. Then the tools improved.


In 2019, a team at Lawrence Berkeley National Laboratory led by Vahe Tshitoyan and Anubhav Jain trained a simple language model on materials-science abstracts. They collected 3.3 million abstracts from more than 1,000 journals published between 1922 and 2018. Without being told any chemistry, the model learned relationships between compounds and properties. Then the researchers ran a test that amounted to a controlled experiment in waking sleeping beauties. They trained the model only on older literature and asked it which materials might be good thermoelectrics. They reported that the model would have predicted some of the best thermoelectric materials found over the previous decade, such as CuGaTe2, years before they were actually discovered. The knowledge was already in the corpus, spread across papers no single person had been able to read together.


The sociologist James Evans and his collaborator Jamshid Sourati took the idea a step further. Their insight was that the obstacle to discovery lies partly in the texts and partly in the people reading them, who tend to cluster around familiar problems and familiar colleagues. In a 2023 paper in Nature Human Behaviour, they showed that modeling the distribution of human expertise improved AI predictions of future discoveries by up to 400 percent, especially where the relevant literature was sparse. Then they inverted the model. By tuning the system to “avoid the crowd,” they generated promising hypotheses that humans were unlikely to imagine or pursue until the distant future. These “alien” inferences had three notable features: humans rarely find them; when humans do, it is many years later; and on average they were better than the human ones, likely because researchers exhaust existing approaches before trying new ones.


That is a formal description of a sleeping beauty's prince: a reader positioned outside the crowd, holding a combination the crowd has not tried.


A vivid recent case came from Imperial College London. José Penadés and his colleagues had spent years working out how a family of genetic parasites called cf-PICIs manage to spread across bacterial species. Before publishing, they gave the question to Google's AI “co-scientist,” a system built on large language models. They posed a question that had taken years to resolve experimentally but was still unpublished: how could these elements spread across bacterial species? The system's top-ranked hypothesis matched the experimentally confirmed mechanism: cf-PICIs hijack diverse phage tails to widen their range of hosts. Penadés was startled enough that, as he later told BBC Radio 4, he emailed Google to ask, “Do you have access to my computer?” It did not.


The episode should be read carefully as the caveats are instructive. According to the peer-reviewed paper in Cell, the researchers provided only a single input document made up of previously published, openly available information, and their own group had published the cf-PICI capsid mechanism in 2023. As one critical review of the case notes, the test was retrospective and not blind; the evaluation was designed and scored by scientists who already knew the answer, working with the Google team that built the system; and the group's earlier published work gave the model the setup and the anomaly, though not the conclusion. But that caveat is the point. The pieces were already in the public record. The model put them together. This was Swanson's fish oil but at machine speed.

‍

More beauties, deeper sleep

So AI can act as a prince. The harder question is whether it will do so systematically, and here the evidence cuts in two directions.


The first complication is that language models left to themselves reproduce the attention patterns of the literature they learned from. A 2024 study by Andres Algaba and colleagues at the Vrije Universiteit Brussel asked GPT-4 to suggest references for papers published after its training cutoff, and found that the model's citation patterns closely resembled human ones, but with a stronger bias toward already highly cited work, even after controlling for year, venue, and other factors. The authors warn that such models may amplify existing biases, including the so-called Matthew effect, the tendency for attention to flow to those who already have it. A sleeping beauty is by definition a paper the Matthew effect has passed over. A research assistant that echoes the crowd will put it deeper to sleep. The difference between a model that finds sleepers and one that buries them deeper lies less in the model than in how it is asked. Sourati and Evans had to build the crowd into their system in order to steer away from it.


The second complication concerns production rather than discovery. The literature was already outgrowing its readers before generative AI arrived. Mark Hanson and colleagues found that articles indexed in Scopus and Web of Science in 2022 were about 47 percent more numerous than in 2016, outpacing the limited growth in the number of working scientists during the same period. Language models are now adding to that supply. A 2025 analysis by Weixin Liang and colleagues of more than 1.1 million papers and preprints estimated that up to 22 percent of computer-science papers showed signs of LLM modification, with lower levels in mathematics and Nature portfolio journals. That modification was more common among first authors who post preprints frequently, in more crowded research areas, and in shorter papers. Assisted writing is not bad science. But a larger haystack means more needles lost in it. If Ke's finding holds, that delayed recognition is a continuum across all fields, then an expanding literature is also an expanding population of future sleepers.


There is a subtler form of production too. The hypotheses that AI systems now generate, the “alien” inferences that Sourati and Evans describe, are sleeping beauties made to order: ideas proposed ahead of the community that could test or appreciate them. The co-scientist gave Penadés's team more than one answer. It proposed four further hypotheses, all of which the researchers judged plausible, and one they had never considered and have since begun to pursue. Scale that up across every laboratory with a subscription, and the bottleneck in science moves. The scarce input is no longer the idea. It is the experiment, the instrument, the grant, and the hours of human judgment needed to decide which ideas deserve them. EPR waited fifty years for lasers. The machine-generated hypotheses of the 2020s may wait for bench time.

‍

A receptive environment

If there is a lesson here, it concerns institutions more than algorithms. Knowledge, once published, is a classic public good: one person's reading of a paper does not reduce anyone else's. Attention is not. The sleeping beauty is what we get when a non-rivalrous good depends on a scarce complement. Wagner's version of the point, about biological innovation, applies equally to other scientific findings. As Works in Progress summarizes his argument, no innovation prospers unless it finds a receptive environment.


AI can make the attic searchable. It can read the Freundlich paper next to the water-purification literature, the Raynaud's papers next to the fish-oil trials, the 1958 synthesis next to the 2004 graphene breakthrough. But whether those connections become discoveries depends on things AI does not supply: whether the underlying papers are open to be read, whether funders will back hypotheses that come from outside a field's crowd, and whether anyone is paid to do the unglamorous work of testing ideas that no one yet cares about. Aghayeva makes the point about access plainly: freely available knowledge could speed progress for the mundane reason that more scientists can read it and connect it to what they already know.


Balwit-Cheung ends her soda-ash essay underground in a Wyoming trona mine. She describes salty icicles on the mine ceilings so fragile that they melt as soon as they are brought to the surface, and reflects that not everything survives being brought to the surface. Some sleepers will turn out to be wrong, and some will turn out to be Pearson's PCA, never really asleep. Shimomura's glowing protein waited thirty years for a biologist who knew what to do with it. The machines now reading the literature are likely to shorten waits like that one. They may also create many more of them.♦


‍

‍