Robonaissance

Robonaissance

The Thermodynamics of Intelligence, Part 7: The Stall Every Neural Network Hits First

Every neural network stalls the same way before it breaks through. In 1995 two physicists solved the equations for why. Those equations still explain how AI learns today.

Hugo's avatar
Hugo
Aug 11, 2026
∙ Paid

Every Hidden Unit Learns the Same Thing First

Every instalment of this series runs the same interrogation: what is the order parameter, what is the control parameter, where is the critical point. Usually the answers arrive after the fact, calculated once a phenomenon is already known to need explaining.

This one is different in a specific way. In 1995, David Saad and Sara Solla wrote down a set of differential equations for a training plateau, solved them exactly, and published the solution years before plateaus became something most of machine learning had occasion to worry about. The equations describe a transition with an order parameter that is exactly zero on one side and grows continuously from zero on the other, a control parameter with a precise meaning, and a critical point that turns out not to be a single number at all but a law: how long the plateau lasts as a function of how big the problem is. Three decades later, the same equations are the ones a 2026 paper reaches for to explain why some heads in a transformer’s attention mechanism specialize and others do not.

The escape from the plateau, it turns out, is symmetry breaking in the exact sense physics uses the term. Not resembling it. Being it.

User's avatar

Continue reading this post for free, courtesy of Hugo.

Or purchase a paid subscription.
© 2026 Robonaissance · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture