This is Part 6 of AI Grew Up on Git, a series on the Git infrastructure that raised AI and what AI is now doing to it.
Every part of this series so far has been a story about Git’s tamper-evident history getting stretched over something it was never built for. A social layer for strangers. A home for gigabyte-sized weights. A record of scientific claims. A $7.5 billion acquisition. The strain has a precise technical cause, and it deserves a full look on its own terms.
Git’s whole value, since Part 1, rests on one promise: any two points in its history can be compared, honestly, by anyone, and the comparison will tell you exactly what changed. For one specific kind of file, that promise quietly stops applying. Not because the file is inconvenient. Because the question “what changed” stops having a stable answer.
What Diff Does
Start with the tool itself, because the failure only makes sense once the mechanism is precise.
Git’s diff traces back to a 1986 paper by Eugene Myers in the journal Algorithmica. Myers took two old, separately known problems, finding the longest stretch two sequences have in common, and finding the shortest set of edits to turn one sequence into the other, and showed they were the same problem in disguise. Both could be modeled as finding the shortest path across a grid, an “edit graph,” and he gave a fast way to search that grid. That search is what runs, in some refined form, every time you type git diff.
The algorithm operates on sequences, and by convention, Git feeds it sequences of lines. Each line of a text file becomes one unit to compare, which is why a single changed line shows up in a diff as one deletion and one insertion rather than a “modification.” There is no concept of an edited line in this output, only lines that vanished and lines that appeared to replace them. This line-by-line habit is a choice about what to feed the algorithm, not a limit built into the algorithm itself, and Git can rerun the same logic at the level of words or characters when you ask it to.
On a binary file, none of this runs. Git checks for line structure, finds none, and rather than attempting a comparison it cannot make sense of, it prints two words and stops: “Binary files differ.” That message is not a failed diff. It is Git declining to try.
One precision matters here, because the rest of this argument depends on it: Git is not incapable of comparing bytes. A flag, git diff --binary, produces a genuine byte-level patch, encoded so it can be transmitted and reapplied. The capability exists. What matters is what happens once you have that comparison in hand for a specific kind of file, and discover it told you nothing true.
The file that breaks the premise twice
A model’s weights are close to the worst possible input for any of this, and the reasons stack rather than substitute for each other.
The first reason is the one Part 3 already covered. A weight file is one enormous binary blob, often many gigabytes, with no line structure for Myers’s algorithm to even attempt. This alone would earn the “Binary files differ” treatment and nothing more.
The second reason is deeper, and it is the part almost nobody outside machine learning knows. Even a perfect, complete, byte-for-byte comparison between two weight files can tell you something false. Neural networks have a property, identified as early as 1990 and studied continuously since, called permutation symmetry. Take any hidden layer in a network and swap two of its neurons, and correspondingly swap the rows and columns of weights feeding into and out of them. The network computes the exact same function afterward. Not an approximately similar function. The identical function, on every possible input, forever.
The consequence for comparison is severe. Two models can behave identically, agree on every input, make every decision the same way, and share zero bytes in common, because one is an internally relabeled version of the other. Two models can also look nearly identical byte for byte and compute something quite different, if a small, meaningful change happens to land somewhere the relabeling doesn’t touch. A byte-level diff between the first pair would report total disagreement about a distinction that, functionally, does not exist. Diff would be lying to you with perfect technical accuracy.
Researchers have built tools to fight exactly this. A 2023 paper, “Git Re-Basin: Merging Models modulo Permutation Symmetries,” by Samuel Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa, permutes one model’s neurons to align with a reference model before any comparison happens, hunting for the real difference hiding underneath the relabeling. The paper was accepted as an oral presentation at ICLR that year, and by the lead author’s own account it scored first out of 4,849 submissions to the conference. The researchers were not making a joke with that title. They were solving Git’s problem, for weights, by hand, because nothing automated could.
State the conclusion the way it deserves to be stated. It is not that model weights are hard to diff. It is that “what changed” is not always a well-defined question to ask of them.
Two workarounds, and what each one gives up
Two different responses have grown up around this gap, and they give up on different things.
Git LFS, from Part 3, replaces the giant file with a small pointer and lets the real bytes live untouched on separate storage. Diffing was not solved by this. It was made invisible, which is a different achievement and a smaller one.
DVC, released in 2017 by a company called Iterative, extended the same trick with real ambition, adding pipelines, experiment tracking, and a way to treat data and models as first-class citizens of a Git-based workflow. Underneath the tooling, it is still the pointer trick, generalized. DVC does not diff a model any more than LFS does. It gives you a better-organized way to work around the fact that it can’t.
DVC’s own recent history repeats a shape already visible in Part 4’s story of an open-research index. Starting in 2024, Iterative’s own attention shifted toward a new project on unstructured data, called DataChain, and DVC stopped being the center of its creators’ roadmap. In November 2025, the open-source project found a new steward, acquired by a company called lakeFS, with Iterative’s founders publicly welcoming the move. A tool built to version other people’s models outlived its own founders’ interest in maintaining it, the same way an open-research index got forked and rehomed within a day of shutting down.
The second response gives up on something larger, and is more honest for it. Model registries, tools like MLflow’s model registry or Weights & Biases, do not attempt to compare two weight files at all. They track lineage instead: what code produced this model, what data trained it, what hyperparameters were set, what it scored when it was evaluated. The question changes shape entirely, from “what is different inside this file” to “what recipe produced this file,” and the second question happens to be one that Git-style tooling can answer well.
The honest state of the art
Newer tools reach further than either workaround.
A small open-source project called diffai computes tensor statistics and parameter-level summaries across two model files, describing itself as a semantic diff, built explicitly against the plain failure of ordinary diff on the same files. Read what that delivers, precisely. A statistic can tell you a tensor’s average value shifted, or that its distribution widened. That is real information. It cannot tell you, the way a line-by-line diff tells you, exactly what changed and exactly where, because “where,” for a set of numbers with no fixed meaning outside the whole network around them, frequently is not a coherent question to ask.
The industry still has no answer. Not a slow answer, arriving eventually. Every existing tool takes one of two paths. Either it hides the file from Git entirely, the LFS and DVC path, or it swaps the original question for a different, useful, but not equivalent one, the registry path. Nobody has handed Git’s core promise back to the one artifact model weights represent.
What this means for everything after it
Return to where Part 1 started. Git’s value was never really about storage capacity or speed. It was that any two points in history could be compared, honestly, by anyone, without needing to trust whoever made the change. Diff is not a convenience layered on top of that system. It is close to the entire point of building it.
Weights sit inside that system without ever fully joining it. They are named by hash, committed, chained into history exactly as Part 1 described. But they are not, in any meaningful sense, reviewable the way a changed line of code is reviewable, because there is no stable notion of what a “changed line” would even mean for a gigabyte of numbers with permutation symmetry hiding inside it.
This gap has gone mostly unnoticed for a simple reason: until recently, a human always chose what to commit, and could describe, in a message, what the change was supposed to accomplish. The commit message did the work that diff couldn’t. The next part asks what happens to review, to blame, to the entire meaning of a commit, once the thing proposing a change is not a person who can be asked to explain it, and the thing being changed is a file that nobody, human or tool, can show you the difference on.
The Working Tree
This part’s Working Tree has no new commands to learn. It looks more closely at commands already covered, and treats DVC and model registries as concepts only, with no command syntax for either, because the point here is understanding a limitation, not adding a tool.
Ask Git to compare a text file, and you get the familiar line-by-line output.
git diff file.py
Ask it to compare a binary file, and you get the two words that mean it declined to try.
git diff model.bin
Force a byte-level comparison, and Git will actually produce one, a real patch of the raw difference, transmittable and reversible.
git diff --binary model.bin
That patch is a genuine technical achievement and a genuine dead end for weights: it will faithfully show you that every byte changed, even in the one case where nothing about the model’s behavior did.
DVC, conceptually, sits one layer above this. It keeps a small metadata file in Git, similar in spirit to an LFS pointer, and manages the real data or model file outside Git’s diff machinery entirely. It never claims to diff the file. It claims to track which version of the file went with which version of the code, which is a more modest and more honest promise.
Model registries sit at a different layer still. Rather than versioning a file at all, they log a record: this exact code, this exact dataset, these exact settings, produced a model that scored this well. Comparing two entries means comparing that record, not the weights themselves. It answers a different question, well, instead of answering the original question, badly.
commit 6/8
What changed: Git's core promise, that any two points in history can be honestly compared, quietly stopped applying to model weights.
Why: A model's weights have no lines to compare, and worse, two functionally identical models can share zero bytes in common.
This is Part 6 of AI Grew Up on Git, a series on the infrastructure that raised artificial intelligence and what AI is now doing to it. Part 7 turns to the author of the commit itself, and what happens to review, blame, and trust once that author is no longer a person.


