This is Part 4 of AI Grew Up on Git, a series on the Git infrastructure that raised AI and what AI is now doing to it.
In February 2018, Science put a name on something machine learning researchers had been muttering about for years. The headline read “Artificial intelligence faces reproducibility crisis,” and the subtitle carried the diagnosis: unpublished code and sensitivity to training conditions were making many of the field’s claims hard to verify. The hook was a conference the week before in New Orleans, where reproducibility had been on the agenda, with some teams diagnosing the problem and one presenting tools to fight it.
The evidence had come from inside the house. A group at McGill, Peter Henderson and colleagues, among them a professor named Joelle Pineau, had published a paper that year with a title that reads like a warning label: “Deep Reinforcement Learning that Matters.” What they showed was quietly devastating. Take a published reinforcement-learning algorithm. Run it on the same task, with the same settings, changing nothing but the random seeds, the arbitrary numbers that initialize a training run. The results moved. Sometimes they moved enough to flip the conclusion. Two runs of the same idea could support different papers, and the difference between them was luck.
A field that could not tell its findings from its luck had a problem, and the old publication culture was structurally unable to catch it. A paper described its method in prose. The code, if you wanted it, was “available upon request,” a phrase of great diplomatic beauty that in practice frequently meant no. You could not rerun what you could not obtain, so most claims were never rerun, and a field that measured itself by leaderboard numbers kept rewarding results that were new over results that were checkable.
What rescued machine learning, to the extent it was rescued, did not come from journals. It came from the culture this series has been tracing: the one that keeps its work in repositories.
Paper, meet repository
Machine learning had already left the journals behind in one important way. The field lived on arXiv, the preprint server, where a paper appears in days rather than surviving months of review. Speed was the point, and the same appetite for speed had made the field an early and total adopter of Git and GitHub. Through the 2010s these two cultures fused, and out of the fusion came a norm that nobody decreed and nobody owns: a serious paper ships with a repository.
The link changed what a paper was. The repository stopped being supplementary material, the dusty appendix nobody opened, and became half the publication. And inside the repository, a humble file quietly took over one of the paper’s oldest jobs. The README, the plain-text front page of a repository, became the real methods section: here is the environment, here are the dependencies, here are the commands, here is what you should see when it works. A prose methods section gestures at what was done. A README executes. One is a description of an experiment; the other is the experiment, one command away.
That difference rewired verification. For most of scientific history, checking a claim meant rebuilding the work from its description, a labor so heavy it was rarely performed. Now checking a machine-learning claim, in the best case, meant something a graduate student could do before lunch: clone the repository, check out the exact version the authors ran, and run it. Reproducibility collapsed, in the favorable cases, into a checkout. The Part 1 machinery made this trustworthy in a specific way: a commit hash is an address for a precise moment in a codebase’s history, tamper-evident and unambiguous, so “the code we ran” stopped being a claim and became a coordinate.
The norm needed one more thing: a map. With thousands of papers arriving monthly, knowing which ones had code, and which code matched which claim, was itself a research problem. Two researchers made it their side project.
Two researchers and a side project
In July 2018, Robert Stojnic and Ross Taylor, working under a small company called Atlas ML, launched a website called Papers with Code. The idea fit in its name. The site matched machine-learning papers on arXiv to their implementations on GitHub, and extracted the reported results into leaderboards, so that “does this paper have code” became a searchable fact and “what is the state of the art on this task” became a page instead of a rumor.
The community adopted it the way dry ground adopts rain. By late 2019 the index covered more than eighteen thousand papers with code and over fifteen hundred leaderboards, and in December of that year Facebook AI acquired it, with a public commitment that the resource would remain, in its own words, neutral, open and free. The bigger institutional blessing came in October 2020, when arXiv itself, in partnership with the site, added a Code tab to its article pages. The paper-to-repository link was no longer a community habit. It was part of the paper’s official record.
What happened to the two founders next is the part of the story that turns it from a pleasant tale about tooling into this series’ argument. Inside Meta, Stojnic and Taylor moved from indexing research to building models. In November 2022, two weeks before ChatGPT arrived, they released Galactica, a large language model trained on scientific literature, papers, reference works, knowledge bases, with Taylor as first author and Stojnic as last. The public demo survived about three days. Users found that the model, asked for scholarship, would generate confident citations to papers that did not exist, and Meta pulled the demo in the storm that followed. A model raised on papers had hallucinated the papers. The failure was loud, but the work did not die: by the founders’ own accounts, its lessons fed directly into how Meta built and released what came next. Taylor went on to lead the reasoning team working on Llama 2 and Llama 3; Stojnic’s own bio reads simply, Llama 2 and Llama 3 technical leadership.
Hold the arc up to the light. Two researchers built the index of open research. Then they built a model trained on scientific text. Then they helped lead the models trained on everything. The people who mapped the paper-and-code commons became the people who fed it to machines, and there is no cleaner illustration of what that commons had quietly become.
The checklist
While the index grew, the institutions moved, and the person at the center of the institutional reform was the same professor whose lab had helped expose the crisis.
In 2019, NeurIPS, the field’s premier conference, created a new role on its program committee: Reproducibility Chair. Joelle Pineau held it, with Koustuv Sinha, and the program they rolled out had three parts. A code submission policy, encouraging authors to submit the code behind their claims. A community-wide reproducibility challenge, in which volunteers picked papers and tried to make them true again. And, on every single submission, a mandatory Machine Learning Reproducibility Checklist, piloted the year before: did you describe your setup, report your seeds and variance, specify your data, link your code. The checklist spread outward from NeurIPS the way good paperwork does, adopted in various forms by ICML, ICLR, AAAI and beyond, until a discipline invented at one conference had become the field’s default furniture.
The program’s own report, published in the Journal of Machine Learning Research in 2021, is candid in a way program reports rarely are: checklists and challenges improved the community’s norms, and improving norms is not the same as guaranteeing truth. Results still vary with seeds. Code still rots as dependencies drift. The crisis was blunted, not solved. But the reform rested on an assumption so deep it was almost invisible, and the assumption is this series’ subject. The checklist could demand code because a result now had an address. A commit hash, a tag, a release: the tamper-evident chain that Part 1 described, built to protect a kernel from a licensing dispute, had become the referencing system of a science. Pinning history was the thing Git had always done. Science just started using it.
The end of the index, and what research became
The index itself did not survive. On July 24, 2025, Meta shut Papers with Code down, five and a half years after promising to keep it neutral, open and free, a promise that, to be fair, it had kept until it stopped. The community did what this culture does with anything it values: it forked the data. JSON dumps landed on GitHub, an archive went up on Hugging Face. And the day after the shutdown, Hugging Face’s chief technology officer, Julien Chaumond, a name readers of this series have met twice already, announced a partnership with Meta to build a successor. The index died the way infrastructure dies now, with its contents rescued and its function reassigned within twenty-four hours.
The norm the index served never wavered. Papers still ship with repositories. The Code tab is still on arXiv. Reproducibility-as-checkout, partial and imperfect, remains the field’s working standard, and no one seriously proposes going back to prose and polite requests.
But the decade of paper-and-repository pairs did something bigger than improve verification, and the Galactica-to-Llama arc already showed it. A paper with a repository is not just checkable by a stranger. It is readable by a machine. Method and implementation, claim and check, the idea in prose and the idea executing, linked pair by pair across hundreds of thousands of works: a culture built so humans could rerun each other’s experiments had, as a side effect, arranged the whole of open research into training data. Nobody designed that. It was assembled the way everything in this series is assembled, one reasonable convenience at a time, by people solving nearer problems. Part 2 ended with code and conversation as an accidental corpus. Part 4 adds science itself to the pile. The next part follows the money: what it meant when one company decided that the platform all of this lived on was worth seven and a half billion dollars.
The Working Tree
This part’s commands are about pinning a moment in history, the mechanics under “the code we ran.” They assume Part 1’s snapshots and Part 2’s remotes.
Every commit has a hash, and a hash is an address. Checking out that address puts your working copy into the exact state the authors ran, down to the byte.
git checkout a1b2c3d
Hashes are precise and unmemorable, so Git lets you name one. A tag is a human-readable label pinned to a single commit, the conventional way to mark the version a paper reports.
git tag v1.0-paper
git checkout v1.0-paper
On GitHub, a release wraps a tag with notes and downloadable files, which is why “the version in the paper” is usually one click on a repository’s Releases page.
A research repository is also defined by what it leaves out. The .gitignore file lists what Git should never record: datasets too large to version, secrets, the litter of local runs. It keeps the record clean, which is part of keeping it honest.
echo "data/" >> .gitignore
And the README, as this part argued, is the methods section that executes. The convention is stable across the field: what this is, how to install it, the exact commands to reproduce the headline result, and what you should see. When those four things are present, a stranger can check the claim before lunch. That is the entire revolution, written in plain text at the top of a repository.
commit 4/8
What changed: Checking a scientific claim stopped meaning rebuilding the work from prose and started meaning cloning it and pinning history to one commit.
Why: A result with an address can be rerun by a stranger, and machine learning's papers grew up with addresses.
This is Part 4 of AI Grew Up on Git, a series on the infrastructure that raised artificial intelligence and what AI is now doing to it. Part 5 follows the money: what Microsoft understood it was buying when it paid $7.5 billion for GitHub, and why Copilot was the receipt.


