This is the overture to Intelligence Is Compression, a seven-part series exploring the core ideas in Principles and Practice of Deep Representation Learning, a new open-source textbook from Yi Ma‘s group at UC Berkeley. The book argues that a single principle, learning compressed representations, unifies the major architectures of modern AI.
“Just as the constant increase of entropy is the basic law of the universe, so it is the basic law of life to be ever more highly structured and to struggle against entropy.” Václav Havel wrote that about life in general. Yi Ma’s lab at Berkeley took it as a research program.
Compression here does not mean shrinking a file. It means finding the small set of numbers that explains a mountain of raw data: the handful of directions that vary across a face, the underlying structure that lets a sentence be predicted from its first few words. Ma’s claim is that intelligent systems, biological or artificial, are all in the business of finding those numbers, and that this business can be stated as mathematics rather than intuition.
In 2023, while most of the field was scaling language models by throwing more compute at the problem, Ma published a paper with an unusually provocative subtitle: “Compression Is All There Is?” The question mark was doing real work. He was not claiming to have solved intelligence. He was asking whether a single mathematical objective, compressing raw, high-dimensional data into structured, low-dimensional representations, could explain why the major architectures of modern deep learning look the way they do. The full argument, co-written with Sam Buchanan, Druv Pai, and Peng Wang, arrived in 2025 as an open-source textbook. Its subtitle is bolder still: A Mathematical Theory of Memory.
Two years after the question mark, the answer reads less like a hypothesis and more like an indictment. Transformers, diffusion models, CLIP, DINO, sparse coding: methods developed independently, by different teams, for different purposes, all turn out to be approximate solutions to the same underlying problem.
That convergence is an uncomfortable claim for a field that prides itself on empirical discovery over theory. It means the engineering did not outrun the mathematics after all. The mathematics was there the whole time, waiting for someone to write it down and show that the transformer Vaswani’s team found by trial and error in 2017 is close to the architecture you get if you start from a compression objective and derive it, layer by layer, from scratch.
This series does not walk through the textbook chapter by chapter. It follows the argument the way it earns its force: the provocation first, then the machinery underneath it, then what the machinery still cannot touch.
The Seven Parts
Seven parts, in the order the argument builds, all on the Intelligence Is Compression shelf:
Part 1: One Principle: why one Berkeley lab thinks the field has been solving the same equation without knowing it.
Part 2: What Does a Neural Network Remember?: why a photograph with millions of pixels can be described in a few dozen numbers, and what that has to do with your visual cortex.
Part 3: Learning by Denoising: the trick behind every image generator, and the uncomfortable question of whether it creates or copies.
Part 4: The Information Game: the single objective that explains why a method with no labels learned to see nearly as well as one trained on millions of them.
Part 5: Building the White Box: what happens when you derive the transformer from mathematics instead of finding it by trial and error.
Part 6: When Theory Meets the Real World: whether a theory built on clean examples survives contact with images, language, 3D scenes, and human motion.
Part 7: What We Still Don’t Understand: three tests for the kinds of intelligence this framework cannot yet explain, named after the men who might have designed them.
If This Is You
If you build with transformers, diffusion models, or CLIP and have never come across a convincing account of why any of them work, start with Part 1: One Principle and Part 5: Building the White Box. The second one derives the architecture you use every day from an equation, not from years of trial and error.
If you have been following the argument over whether image generators copy their training data or generalize from it, Part 3: Learning by Denoising makes the sharpest version of that case you will find outside the textbook itself, with a real answer for when a model does which.
Maybe your interest runs the other direction, from brains toward machines rather than the reverse. Part 2: What Does a Neural Network Remember? traces a straight line from a 1996 discovery about the visual cortex to the compression objective the rest of this series is built on.
Perhaps what you want is the big question: whether intelligence itself reduces to compression. Part 1 states the claim at full strength, and Part 7: What We Still Don’t Understand is the one place in the series that tells you, precisely, where the claim stops being true.
And if you are the kind of reader who wants to watch a theory survive contact with real data before you believe a word of it, Part 4: The Information Game and Part 6: When Theory Meets the Real World put the framework against images, language, 3D scenes, and human motion, and admit where it still falls short. Wherever you enter, you arrive back at the same equation.
The struggle continues. Start wherever the door opens for you.



Having the voices in the head discussing this all day, one way to consider the limits of compression would be as how it is expressed by the digestive system. If compression was the purpose, that distilled little pile of fertilizer would be the goal. Yet it is the entropy. The structure being broken down and the energy stored in it radiated back out to power the body.
Then consider how energy is expressed as consciousness, as all the emotions, desires, needs propelling us on. In fact the intellect is often trailing along behind, trying to formulate some degree of logical structure to the resulting effects.
If it was truly the intellect and knowledge that ruled, where are the philosopher kings? Why isn't academia the apex of control? Why is it the military and the financial system that seems to be running the show?
Maybe if we better understood the larger dynamic, we might start to make some sense of it and actually ride that wave, not constantly tumbled by it.
Focus and field.
The nervous system is the focus. The circulation system is the field.
Then there is the gut, processing the energy driving it all.
The Field needs to step back and not just obsess over the function of the nervous system.
By itself, it's a rabbit hole. Context matters.
Galaxies are structure coalescing in, as light radiates out, so yes it is all compression, if you only consider structure.
Efficiency is to do more with less, until we reach peak efficiency and can do everything with nothing.
We are linear, goal oriented creatures in a cyclical, circular, reciprocal, feedback generated reality. It's like we really haven't come to terms with the implications of the world being round, not flat.