The Commons Is Not A Pre-training Subsidy
AI companies have treated the mathematical record as training material, benchmark, and marketing asset. The Leiden Declaration asks a harder question: what happens when systems extract from a public order of knowledge without accepting its standards of attribution, verification, and responsibility?
In June 2026, sixteen mathematicians from fifteen institutions across three continents published the Leiden Declaration on Artificial Intelligence and Mathematics. and thousands of mathematicians and researchers have since signed it, including Fields Medalist Peter Scholze, a Fields Medal recipient. The International Mathematical Union endorsed it. It will be formally presented at the International Congress of Mathematicians in Philadelphia this July.
The declaration contains no lawsuit. No demand for licensing fees. No request for compensation.
That is what makes it worth reading carefully.
Writers took AI companies to court over scraped books. Hollywood went on strike, extracting AI clauses from studios after months on the picket line. Musicians pursued licensing deals, some settling for money, others still fighting. Each time, the complaint was the same: something personal had been taken, and someone wanted payment. The AI industry has grown comfortable with this kind of opposition. It knows how to negotiate it, delay it, and absorb it.
Mathematics is a different kind of opponent. A theorem does not belong to its author. A proof does not expire. The mathematical record, comprising papers, textbooks, proof libraries, databases, seminar notes, and errata, belongs to no one in particular. It is the accumulated output of a civilization of reason, maintained without payment, without copyright, without corporate sponsorship, for centuries. Mathematicians want no licensing fee. They want to know what happens to the floor.
That floor runs under everything, mostly invisibly. The encryption protecting a wire transfer rests on number theory developed by people who had no idea what a wire transfer was. The algorithm reading a CT scan for tumours rests on linear algebra worked out in the nineteenth century for reasons that had nothing to do with medicine. When engineers design a bridge, they are applying mathematics that was worked out in seminar rooms by people who published their results openly, with no expectation that a car would one day cross on the strength of their reasoning. The GPS signal in our pocket depends on corrections derived from general relativity. Relativity required the mathematics that preceded it by half a century, geometry developed for geometry’s sake, equations whose authors had no destination in mind. The distance between pure mathematics and its applications is not a pipeline. It is a delay, sometimes of decades, sometimes of centuries. Mathematics works by proof, not by patch release. Errors do not get caught in beta. They surface when the bridge moves wrong, or not at all.
Now consider what the AI industry has done to it.
It trained on the mathematical commons: papers, textbooks, formal proof libraries, open databases, at a scale and speed no previous technology approached. It built products. It filed for market valuations. Then it issued press releases.
The press releases repay careful study. They tell you everything except what happened.
A model wins gold at the International Mathematical Olympiad. A system generates novel proofs across dozens of open problems. Another disproves an eighty-year-old conjecture. The announcement arrives with benchmarks, demos, carefully lit product launches. The pattern is legible: conquering mathematics is the proof of intelligence, the headline that moves markets, the signal that justifies the next funding round. The benchmark is the advertisement. The literature is the extraction site.
In May 2026, OpenAI announced that an internal model had disproved the unit distance conjecture, a problem posed eighty years ago by Paul Erdős. Erdős offered cash prizes for solutions to problems he cared about, a practice so characteristic of him that the amounts became a kind of index of how hard he thought a problem was. The unit distance conjecture was one he had been thinking about since the 1940s. The announcement arrived as a corporate product launch, but it did not arrive without mathematics. OpenAI released a proof, an abridged account of the model’s reasoning, and a companion note written by external mathematicians who had checked and digested the argument. What remained opaque was not whether the final argument could be inspected. It was the system that produced it: the training corpus, the selection process, the failed attempts, the human framing, and the debts embedded in the model’s path to the result. Mathematician Melanie Matchett Wood noted that the result failed to cite a history of closely related ideas already in the literature. The process was opaque. The contributions were uncredited.
The conclusion entered the public record. It is available to be built upon, cited, inherited by the next result, which will rest on this one. Each generation of discovery now layers on a foundation whose integrity no one can verify.
Knowledge gets poisoned slowly. Not with a single false theorem, but with a chain of results whose provenance is unknown, whose debts are unacknowledged, and whose errors, if any exist, are buried too deep to find. The declaration warns that current automated techniques can produce plausible but unreliable arguments that are difficult to distinguish from correct proofs. Models trained on published work frequently return outputs that do not properly attribute the human contributions they drew on. The problems that look tractable to AI become the problems that get funded, hired around, and recognised, displacing, in the declaration’s words, “expert judgment of their deeper significance.”
Think about what mathematical error actually looks like at scale. It does not arrive as a retraction or a red flag. It arrives as a citation, then as a foundational assumption in the next paper, then as a building block in the next proof, each step removing it further from scrutiny. A wrong lemma can travel for decades before the structure above it shows a crack. Mathematicians already know this: the literature contains errors that propagated for years before anyone caught them, in fields with full peer review, full attribution, full methodological transparency. Now imagine that process with the provenance stripped out and the volume multiplied.
Mathematics is the language through which we reach for the universe. The equations that will guide us to other planets, the models that will let us read signals from distant stars, the proofs that will unlock the next century of physics: all of it stands on the accumulated, verified, attributed work of mathematicians across generations. On a corrupted foundation, we do not reach the universe. We reach a convincing simulation of reaching it.
Mathematics is also how humanity checks its own reasoning. A civilisation that can no longer tell a proof from the appearance of a proof has lost something more than a tool.
Elinor Ostrom spent her career studying what happens to commons. In Governing the Commons, she showed that a commons survives when the people who draw from it also maintain it. The covenant is unwritten but binding: you graze the pasture, you help fix the fence. What she found, studying fishing grounds and irrigation systems and alpine meadows across the world, was that communities had developed intricate informal rules for preventing collapse, rules that outsiders often failed to see until they were gone. The commons did not fail because people were selfish. It failed when the scale of extraction outran the community’s ability to enforce its own norms.
The AI enterprise entered the mathematical commons as one participant among many. Then it brought every herd in the world to graze overnight, packaged the yield, and sold it back. The fence is gone. The grass is gone. The shepherds who tended it are now invited to subscribe. And something has been left in the soil. Not every result is wrong. But some are. Nobody knows which, or where, or how far the roots have spread. Ostrom studied extraction. This is extraction plus contamination: the commons looted and poisoned at the same time.
The model can make mistakes. Please verify important outputs. Not, one notices, when the press release announces that AI has disproved an eighty-year-old conjecture. Then the caveats are omitted and the superlatives are not. Afterwards: the model hallucinated, the system has limitations, the architecture is opaque. These are postmortem labels. They do not sign the report. Someone approved the training pipeline. Someone deployed the system. Someone called it a breakthrough. The model does not have a legal team. The institution does.
A black box makes a poor moral agent. A nondisclosure agreement proves nothing. Technical opacity launders nothing.
The ethics are straightforward, once the extraction is named.
Disclosure means provenance, not a data dump. A company that hands over five hundred terabytes of training data and says “it’s all in there” has disclosed nothing useful. What the mathematical community requires is attribution at the level of argument: when a model produces a proof step, a lemma, a line of reasoning, it must be possible to trace what it drew on and how. This is not a fantasy. Attention-based attribution is an existing technical direction. It can be required.
Verifiability and contribution are separate matters. A proof can be formally correct and published while contributing nothing to the actual frontier of mathematics. The question of whether a result matters, where it sits in the landscape of open problems, what it opens, what it forecloses, belongs to the mathematical community. Benchmarks measure what companies choose to measure. Volume and clean formatting earn nothing by themselves.
Responsibility and credit cannot be separated. The industry’s preferred move is the disclaimer: this system is for research assistance only; all outputs require independent expert verification; the company accepts no liability. Fine. Then the company does not get the press release. If AI has independent reasoning capacity, treat it as an agent, and agents are held to the same standards as any other participant in the commons. If AI is a tool, then the mathematician who used it owns the result, and the company that built it holds the receipt. Nobody has ever seen a ruler issue an earnings call.
There is a harder problem underneath all of this, and it has not been solved.
A model trained on the mathematical commons does not merely retrieve. It selects. After training, it develops preferences for certain kinds of arguments, certain proof strategies, certain classes of problems that look tractable given what it has absorbed. These preferences carry weight. They reflect the data it was trained on, the objectives it was optimised for, and the commercial priorities of the company that built it. When a mathematician uses such a system, the system answers questions and quietly shapes which questions seem worth asking.
Imagine a field where a generation of graduate students runs their half-formed conjectures through the same system before deciding which ones to pursue. The system scores some directions as promising, others as intractable. Students follow the promising ones. Funding follows the students. The field moves. Nobody decided this. Nobody announced a research agenda. The agenda was set by a training objective and a market timeline, distributed into a thousand individual choices that each looked like independent judgment.
A calculator has no research preferences. A database does not find certain conjectures more natural than others. These systems do, at a scale that could reshape an entire field without anyone making a single explicit decision.
The same question reaches beyond mathematics. Wherever AI systems suggest the next step, in legal reasoning, in medical diagnosis, in scientific hypothesis generation, the same reckoning waits: whose priorities are encoded in the suggestion? Who decided what looks promising? Who profits when that direction gets pursued?
These questions cannot be answered by mathematicians alone, or by AI companies alone, or by regulators alone. They require all three, in the open, with the mathematical community at the table as a principal, not as a data source, not as a benchmark, not as a marketing prop. The Leiden Declaration is an invitation to that conversation.
Sixteen mathematicians, from fifteen institutions, with nothing to sell and nothing to protect except the integrity of human knowledge, said plainly what is being taken and what the consequences will be.
The theorem was public. The tollbooth was not. One cannot claim the theorem and abandon the proof. One cannot take the priority and disclaim the error. One cannot extract from the record and then refuse the record’s standards. Generation launders nothing. A chain of contribution remains a chain of contribution, whatever the press release calls it.
The mathematical commons belongs to everyone. It is part of humanity’s public order of knowledge, one of the ways human beings store, test, correct, and transmit reason across generations. If this commons can be enclosed, turned into product, and sold back without durable commitments, every other commons has received its warning. The mathematical record, the literary archive, the scientific literature, the public domain: the same logic reaches all of them.
The AI industry has achieved remarkable fluency in the price of compute, data centres, and market share. Its acknowledgment of the public order that made its systems worth selling has been, to put it charitably, less developed. That asymmetry was chosen.
It is the business model.
The alarm has been sounded. The question now is who picks up the fight.