This is Part 5 of AI Grew Up on Git, a series on the Git infrastructure that raised AI and what AI is now doing to it.
On June 4, 2018, Microsoft announced it would pay $7.5 billion, in its own stock, for a company that had never turned a profit and had spent a decade being the thing open source trusted precisely because Microsoft was not involved in it.
The number alone was a surprise. GitHub had last been priced by private investors at $2 billion, back in 2015. The rumor mill in the weeks before the deal had settled on something closer to $5 billion. Microsoft paid roughly fifty percent above the rumor and nearly four times the last real valuation anyone had put on the company. Whatever Microsoft thought it was buying, it was willing to pay a premium most acquirers reserve for certainty, not for a hunch.
The people changed with the ownership. Nat Friedman, who had sold his own company, Xamarin, to Microsoft two years earlier, took over as GitHub’s CEO. Chris Wanstrath, the co-founder from Part 2 who had built the site’s Rails application on Saturdays over Vietnamese egg rolls, stepped back to become a Microsoft technical fellow, working on strategic projects instead of running the place he’d started. The founder generation handed the keys to the acquirer’s generation, the way it does.
And the open-source world, or a loud and credible part of it, was afraid. The fear had a name, borrowed from Microsoft’s own history: embrace, extend, extinguish, the old accusation that Microsoft absorbed open standards, layered proprietary extensions on top, and let the original wither. GitHub, the home of the culture Part 2 of this series described, the platform that had made forking and pull requests the default grammar of open collaboration, now belonged to the company that culture had spent twenty years keeping at arm’s length.
Microsoft, for its part, seemed almost unaware of the irony sitting in its own announcement. When the acquisition finally closed that October, Nat Friedman published the news under a headline built entirely out of the platform’s own vocabulary: “Pull request successfully merged. Starting build…” The company whose social invention this series has already traced, the pull request that made contributing to someone else’s work feel like a conversation instead of an intrusion, had just used that exact phrase to describe absorbing its own founders into a much larger machine.
What Microsoft said it was buying
Read what Microsoft said at the time, and the reasoning has nothing to do with artificial intelligence.
The press release, filed the same day with securities regulators, framed the deal around developers and cloud. The two companies would, in Microsoft’s own words, “empower developers to achieve more at every stage of the development lifecycle, accelerate enterprise use of GitHub, and bring Microsoft’s developer tools and services to new audiences.” Satya Nadella called Microsoft “a developer-first company” and framed the purchase as strengthening “developer freedom, openness and innovation.” On the balance sheet, GitHub would be folded into the Intelligent Cloud segment, the same bucket as Azure.
Read plainly, this was an on-ramp. Microsoft had tens of millions of developers who had spent years ignoring its cloud, and here was a way to put Azure in front of all of them at once, on a platform they already trusted and used every day.
Nowhere in that reasoning is a training corpus, and there is a simple explanation for the omission: in June 2018, nothing existed that could have used one at this scale. GPT-3, the model that would make large language models a mainstream concept, was two years away. Codex, the specific model that would turn GitHub’s contents into a coding assistant, was three years away. The single most valuable thing about this acquisition, from where the industry sits now, was not on the list of reasons Microsoft gave for making it, because the reason had not been invented yet.
This is the same shape this series keeps finding under different names. GitHub’s own founders built a corpus of code and conversation while trying to make collaboration pleasant. Hugging Face built the infrastructure for open models while trying to comfort teenagers. Papers with Code built a map of open research while trying to fix a reproducibility problem. Now the pattern shows up at the scale of a corporate acquisition: Microsoft bought a developer platform and a cloud on-ramp, and years later discovered it had also bought the thing that mattered more.
The corpus nobody named
In 2021, GitHub and OpenAI shipped the thing that made the accident visible: Copilot, an AI pair programmer, built on OpenAI’s Codex model, itself a version of GPT-3 retrained specifically to write code.
Codex’s own technical paper is specific about where its fuel came from. The training set was collected in May 2020 from fifty-four million public GitHub repositories, filtered down to a final corpus of one hundred fifty-nine gigabytes of Python. Strip the euphemism and name the thing plainly: the training data was GitHub. Not a licensed dataset assembled for the purpose, not a corpus purchased from a vendor who had cleared the rights, but the public repositories of the platform Microsoft had already owned for three years, for reasons that had nothing to do with this.
Copilot’s own disclosed numbers, once it reached general availability in 2022, gave a sense of what that training data was worth in practice. GitHub itself acknowledged that roughly one percent of Copilot’s suggestions matched the training data closely enough, at lengths over about a hundred and fifty characters, to be flagged as near-duplicates. One percent sounds small. At the scale of millions of suggestions delivered every day, it is not.
The lawsuit
Some of the people whose code trained Copilot did not think one percent, or any percent, was fine.
In November 2022, a group of anonymous programmers, represented by the lawyer Matthew Butterick and the Joseph Saveri Law Firm, filed a class action in the Northern District of California against GitHub, Microsoft, and OpenAI, later consolidated under the name Doe v. GitHub. The claim, in outline: that Copilot had been trained on open-source code without honoring the licenses attached to it, stripping the attribution and copyleft terms that licenses like the GPL require, and that doing so violated the part of the Digital Millennium Copyright Act protecting a work’s copyright management information. The plaintiffs described the scale of what they alleged as without real precedent.
The defendants’ position has held steady since the first filing: that training a model on code that was already public, and generating new output that is not identical to any single source, is not copyright infringement, and that the relevant DMCA provision requires a level of identity between the output and the original that Copilot’s suggestions do not reach.
The case has moved unevenly. In 2024, a federal judge dismissed most of the claims, including the core DMCA count, ruling that the code Copilot produced was not similar enough to the plaintiffs’ copyrighted work to trigger the statute. The contract and remaining claims survived and continued toward trial. The dismissal itself then went up on an interlocutory appeal to the Ninth Circuit, where judges heard oral argument in February 2026 on a narrow but decisive question: whether that DMCA provision requires an exact, identical copy to count as a violation at all. As this is written, the appeal has not been decided, and the underlying case has not gone to trial. It is a live, contested legal question, and this article is not positioned to settle it. What can be said plainly is what each side is arguing, and that neither side has yet won.
The receipt
Whatever the courts eventually decide about the code, the market had already rendered its own verdict on the value of what Copilot had learned from.
On Microsoft’s earnings call on July 30, 2024, Satya Nadella gave this whole story its punchline. Copilot, he told analysts, had accounted for over forty percent of GitHub’s revenue growth that year, and GitHub’s annual revenue had reached a two-billion-dollar run rate. Then he went further: Copilot alone was “already a larger business than all of GitHub was when we acquired it.” The acquisition had been priced, in 2018, on the value of a code-hosting business. Copilot alone, six years later, was worth more than that entire business had been at the moment Microsoft bought it.
The growth did not stop there. By May 2026, Nadella was describing GitHub as seeing “unprecedented growth driven by the proliferation of agentic coding,” with Copilot Enterprise adoption nearly tripling year over year. By early August 2026, GitHub Copilot had crossed fifty million users, with revenue still growing more than sixty percent quarter over quarter. Whatever else is true about this acquisition, it has not slowed down.
Trace the throughline this series has been building. Part 2 showed GitHub’s corpus of code and conversation assembling by accident, a byproduct of making collaboration pleasant. Part 3 showed the identical accident happen to model weights on Hugging Face. Part 4 showed it happen again to published research, papers quietly becoming training data because they shipped with runnable code. Part 5 shows the accident get priced, in public, on an earnings call, seven years after a deal made for entirely unrelated reasons. Nobody bought GitHub for this. Somebody, eventually, got it anyway.
What all four of these accidents have shared so far is that the material stayed still long enough to be gathered: code, weights, papers, sitting in repositories, waiting. The next part turns to a harder question. Git was built to show you exactly what changed between one version and the next. Model weights, the thing this whole corpus was eventually for, are the one artifact in this story that Git still cannot really compare at all.
The Working Tree
This part’s commands are about a record you can question, and about undoing a decision without pretending it never happened. They sit on top of everything Parts 1 through 4 built.
When you want to know who wrote a given line, and when, Git can tell you. Blame walks every line in a file back to the commit that introduced it.
git blame filename.py
Undoing a change in Git does not have to mean deleting the evidence that the change happened. Revert creates a brand-new commit that reverses an earlier one, so both the mistake and its correction stay in the permanent record.
git revert a1b2c3d
Reset is a sharper tool, and it comes in two forms that matter to keep straight, because the difference between them is really a difference in how much you trust yourself. A soft reset moves your branch back to an earlier commit but keeps all your changes staged, ready to be recommitted differently.
git reset --soft HEAD~1
A hard reset moves back the same way and then throws the changes away entirely, as if they had never been made.
git reset --hard HEAD~1
Soft reset is a decision you’re revisiting. Hard reset is a decision you’re pretending you never made. Git will let you do either one. It just won’t let you lie about which one you did, because the reflog remembers for thirty days regardless of what you told it to forget.
commit 5/8
What changed: A code-hosting business was bought for its hosting, and turned out to be worth more for what it had been quietly recording the entire time.
Why: Nobody in 2018 had built a model that could learn from a platform's history. Three years later, somebody had.
This is Part 5 of AI Grew Up on Git, a series on the infrastructure that raised artificial intelligence and what AI is now doing to it. Part 6 turns to the file that started this whole detour: why a model’s weights still refuse to do the one thing every other file in Git can do, show you exactly what changed.


