The Thermodynamics of Intelligence, Part 8: The Emergent Ability That Might Not Be There
Every critical point so far was calculated, seen, or proven absent. This one has an award-winning paper arguing it never existed. The dispute still has not settled.
A Test That Has Not Finished Running
Seven instalments of this series have run the same interrogation and reached a verdict each time. A number, calculated: 0.138, α_c ≈ 0.03, a temperature Onsager solved for exactly. A boundary, seen in a training log, with a formula for how long the approach to it takes. A logical form, examined and found to permit no critical point at all, which was itself a finding rather than a failure. Even the disagreements resolved: Part 5 opened with two papers reaching opposite conclusions about the same training curve and closed with both of them right, measuring different things.
This one is not clean, and it will not close that way. In 2022 a group of researchers described dozens of tasks on which large language models jump from near-random performance to strong performance as they scale up, unpredictably and all at once. In 2023 a different group won a top award arguing the jump is not there, that it is an accounting trick performed by the choice of scoring rule. Neither side has conceded. The argument has continued past both papers, new evidence keeps arriving on both sides of it, and the most recent contributions as of this writing do not resolve it either. This instalment does not get to end with a formula, because the field it is reporting on has not produced one.
Dozens of Cliffs
The paper is Emergent Abilities of Large Language Models, sixteen authors including Jason Wei, Percy Liang and Jeff Dean, published in 2022 in Transactions on Machine Learning Research.
Their definition is precise: an ability is emergent if it is absent in smaller models and present in larger ones, which means it cannot be predicted by extrapolating the performance of smaller models. You cannot see it coming from where you are standing.
They catalogue dozens of examples. Three-digit addition and subtraction. Unscrambling a shuffled word back into English. Answering questions posed in Persian. Following an instruction format the model was never explicitly trained on. The arithmetic case is a representative shape. Small models score near zero on the exact-match version of the task, not gradually worse than large models but essentially uniformly wrong, with no visible upward drift anywhere in the small-model range. Then, at a scale threshold that differs from task to task and that nothing about the smaller models foreshadowed, accuracy climbs sharply toward competence. The paper’s own graphs are the evidence: dozens of near-identical cliffs, each one arriving at a different, unforeseeable point on the scale axis, with nothing in the shape of the curve below the cliff that would let you extrapolate where it starts.
The stakes are not academic. A 2022 paper from Anthropic, Predictability and Surprise in Large Generative Models, lays out the safety version of the worry: these models combine a broadly predictable overall capability, tracked reliably by scaling laws, with specific capabilities that can appear all at once and without warning. If a dangerous capability behaves like the tasks in Wei’s catalogue, nobody gets advance notice. The paper’s own explanation for how this is possible is worth holding onto, because it will matter shortly: any single narrow task is a tiny slice of a model’s full output distribution, and a tiny slice can move abruptly even while the whole distribution moves smoothly.



