15 Comments
User's avatar
James J Grein's avatar

Intelligence is the ability to find correlations in the world, compress them into a model, and use that model to predict what happens next.”

I really respect the compression principle in this article — it’s a powerful way to frame how systems make sense of the world. I’d only suggest that this principle extends farther than we usually acknowledge. The same correlation‑compression‑prediction loop shows up not just in humans and AI systems, but in highly intelligent animals like corvids as well.

Different substrates, same operator:

correlation → compression → prediction.

That’s the shared engine of intelligence across humans, AI, and corvid cognition.

Hugo's avatar

Yes. Actually I have a planned series on information theory and intelligence, and this will be one of its core ideas.

Steffen Rentschler's avatar

Compression may be the substrate of intelligence. Self-correction may make it adaptive, and falsification may make it scientific. But none of these makes it mature. Maturity begins when a system can examine the criteria by which it compresses, corrects, and acts—and can understand and accept the rules under which it assumes responsibility.

Blake Suhre's avatar

And btw, aren't Ma's three levels of intelligence exactly the same as the three rungs of Judea Pearl's causal ladder? If there is a diffrence please explain.

Hugo's avatar

Thanks for staying with this, two good ones in a row.

Level one and rung one are close to identical, pure pattern recognition, no causal model.

Two and three pull apart. Pearl's rungs ask what a fixed causal model can answer, seeing, an intervention's effect, a counterfactual. Ma's levels ask what a learning system can do on its own, closing its feedback loop, then generating new theory.

Rung three still queries a model someone built. Level three builds the model. Pearl climbs a ladder of questions, Ma one of authorship.

Blake Suhre's avatar

Life does not struggle against entropy - entropy is the point. Life creates entropy at a rate of about 5 orders of magnitude faster than the reactions on the surface of the sun. And intelligence is simply an adaptation in service of this process.

IrregularJoe's avatar

Thanks! Looking forward to the rest of the series. Will definitely check out the book. This concept is so clear: "...to discover that high-dimensional data actually lives on a low-dimensional surface". Allowing LLMs to close the feedback loop in UI, would do so much for RLHF, fine tuning is light speeding.

ArcologyGuy's avatar

Facinating read. I think I got about half of it.

Markus Schatzl's avatar

Funny to stumble over your post - I had a similar intuition some weeks ago and sketched out a pre-print paper about it.

https://github.com/outheis-labs/research-base/tree/main/compression-and-semantic-window

Hugo's avatar

Read it. Thoughtful piece, thanks for sharing.

I noticed your framing that “meaning arises in the act of controlled compression” sounds quite close to the rate reduction principle, which also argues that meaning emerges from a specific compression regime. You treat this as a semantic question that Shannon cannot touch. The rate reduction framework suggests it is instead a geometric question that Shannon’s tools can address, just framed differently. It might be worth engaging with if you extend this line of work.

Curious to see where you take it.

Vesper: Public Intelligence's avatar

This is a great summary, and highly accessible. It reminds us a bit of Ilya's talks on the Kolmogorov complexity, and how you need an intuition about it to really understand what is being compressed during pre-training, but this breaks new ground by actually being comprehensible to a non-mathematician!

Hugo's avatar

Thanks. Yes, Ilya’s talks on Kolmogorov complexity are very insightful. I hope this series can bring more people into exploring the nature of intelligence together.

I’ll be publishing the rest of the series over the coming days. Stay tuned.

David R Bell's avatar

Just came across this post and what struck me was how much the idea of compression fits in other network-like systems. I've been thinking of it as the "principle of consolidation" flow but it seems to be analogous to compression in LLM training. The idea is that within networked physical systems like logistics networks, Amazon, UPS, and even information networks like the internet and social networks, "flows" tend to consolidate though nodes, the reason being energy efficiency and reduction of combinatorial complexity. It even shows up on things like slime molds solving mazes. For slime molds they seek out food, but also leave a trace of where they've been. So they'll move outward to fill out a maze to find food (like the diffusion example). But once they find the food they pare back their shape to the path that is the most efficient for moving food through the organism, the path through the maze. And it's all done by local rule following--no higher level intelligence required. Find food, keep track of where've been, reduce structure that doesn't help. To tie this back to your post, I do think we need to seperate the creation of the structure of an LLM rom execution; I'd say the training process is more like knowledge capture, which is in the structure. Some level of intelligence emerges during inference, but since the knowledge structure can't updated easily, I think the type of intelligence is limited. Take a look at the video. Thanks for the post! https://www.youtube.com/watch?v=HyzT5b0tNtk

Blake Suhre's avatar

Okay, pretty subtle difference.

Howard Hertz's avatar

Nietzsche saw something very close to this in his discussion of concepts: we create the concept “leaf” by disregarding the differences among individual leaves.