Anything You Can Score, a Machine Can Win
A move no human would play. A fifty-year problem solved in a weekend. A score built out of our own judgment. The frontier was never difficulty. It was measurement.
From Whatever You Ask For, a series on the one job machines cannot do for us.
One in Ten Thousand
On the tenth of March, 2016, in a hotel in Seoul, a machine played a move that no one in the room understood, and the person it was playing against was not there to see it.
Lee Sedol had stepped out for a cigarette. He was thirty-three, holder of eighteen world titles, and by most accounts the strongest Go player of his generation. He had lost the first game of the five-game series the day before, and he had begun the second in a more careful frame of mind than the confidence with which he had entered the match. While he was outside, AlphaGo selected its thirty-seventh move, and Aja Huang, the DeepMind researcher placing stones on the machine’s behalf, set it quietly on the board.
The stone went on the fifth line.
To anyone who does not play Go, this means nothing. To anyone who does, it was close to nonsense. In the opening and middle stages of a game, stones played that far from the edge are held to be inefficient, too high to secure territory, a gift of free points to the opponent. It is the kind of move a teacher corrects in a beginner. In the commentary booth, Michael Redmond, a professional of the highest rank calling the match live, thought at first that the board feed had made an error. He picked up a stone, and put it down again.
Lee came back in, sat down, and looked at the board without moving for a long time. Accounts of how long vary between roughly twelve and fifteen minutes. Fan Hui, the European champion whom AlphaGo had beaten five months earlier, was watching in the building. His verdict, once he had stared at it long enough, was that <q>it’s not a human move. I’ve never seen a human play this move.</q>
He was right in a way that turned out to be measurable. DeepMind’s system carried an internal estimate, learned from a large corpus of games between human players, of how likely a human was to play any given move in any given position. For move 37 that estimate was roughly one in ten thousand.
The machine played it anyway, and it was the move that won the game.
There is a comfortable way to read this story, in which a machine studied human masters very hard and eventually caught up with the best of them. That reading is wrong, and the one in ten thousand is exactly what makes it wrong. A system that learns to imitate human play has a ceiling, and the ceiling is human play. What happened in Seoul was that a system stopped being bounded by that ceiling, because what it had been given to pursue was not the approval of human masters. It was winning.
That distinction leads somewhere more useful than admiration. If you want to know which human activities have already been lost on the axis of raw capability, and which are next, the question to ask is not how hard the activity is, or how creative, or how much intuition it takes. The question is far more boring than that.
The question is whether success can be scored.
One Board After Another
Chess had fallen nineteen years earlier, and it fell in a completely different way.
Deep Blue beat Garry Kasparov in 1997 with an architecture that was, at bottom, an enormous act of human articulation. Its evaluation function, the component that looked at a position and produced a number saying how good it was, had been assembled and tuned with the help of grandmaster consultants. Human chess knowledge, painstakingly extracted from human chess players, was written into the machine. What the machine added was search: the ability to look further ahead, faster, without fatigue, without the lapses of attention that lose games.
This was a real victory and it deserves to be counted as one. But notice its shape. Deep Blue’s understanding of chess was human understanding. Its superiority was speed. If you had asked why it evaluated a position the way it did, the answer would eventually bottom out in a person who had told it so.
Twenty years later, DeepMind released a system called AlphaZero, and the shape changed.
AlphaZero was given the rules of chess. It was not given an opening book, the catalogue of studied first moves that every serious program relied on. It was not given endgame tables. It was not given a single human game. It played against itself, starting from random moves, and adjusted itself according to what won.
After about four hours of this, DeepMind estimated its rating had passed that of Stockfish 8, the strongest conventional program of the day. After about nine hours, it played Stockfish a hundred-game match under time control and won twenty-eight while losing none, drawing the remaining seventy-two. The paper reports superhuman play in chess, shogi, and Go within twenty-four hours.
Those numbers get quoted without the other half of the account, and the other half matters. The self-play games were generated on five thousand first-generation tensor processing units, with a further sixty-four second-generation units training the network, all running in parallel. Four hours of wall-clock time was four hours of an amount of computation that no individual and few institutions could assemble. The achievement is not that learning chess is easy. The achievement is that when you have that much computation, human knowledge stops being the thing you need.
Kasparov, who had more reason than anyone alive to take the result personally, wrote about it in Science with striking generosity. His observation was that AlphaZero did not play the dry, cautious, drawing-oriented chess that everyone had assumed a perfect machine would play. It preferred activity to material, giving up pieces for positions that to his eye looked risky. Conventional programs, he noted, carry the priorities and prejudices of the people who wrote them. AlphaZero wrote itself. He drew the conclusion that its style therefore reflects something truer than a programmer’s taste, and observed that it outplayed the world’s best conventional engine while examining far fewer positions per second.
That last detail is the one worth holding onto. The conventional engine was searching enormously more possibilities per second and losing. Whatever advantage AlphaZero had was not in looking harder. It was in knowing where to look, and that knowledge had been assembled by nothing but millions of games against itself and a rule for who won.
The shogi results were stranger still to the people qualified to read them. Shogi is a Japanese game in which captured pieces return to play in the hands of the capturer, which makes the position volatile in ways chess never is, and it has its own centuries of accumulated theory about how to keep a king safe. Strong players who went through the machine’s games reported that the openings violated known theory and that kings wandered into the centre of the board at moments when every book says to tuck them into a corner. The games were, to trained eyes, close to unreadable, and they were also winning games.
Put the two systems side by side and the pattern is clear enough to be uncomfortable. Deep Blue was human knowledge plus machine speed. AlphaZero was machine speed plus a definition of winning, with the human knowledge deliberately removed. The version with the human knowledge removed was better.
None of this means the machines of 2016 were flawless, and the same match in Seoul contains the proof.
In the fourth game, three losses down and playing for pride, Lee Sedol wedged a stone into the middle of the board between two of AlphaGo’s groups. It has been called the divine move since. AlphaGo’s own estimate of the probability that a human would play it was, by an odd symmetry, also about one in ten thousand. And the system did not handle it. Its assessment of its own winning chances collapsed over the following moves, its play deteriorated, and it lost the game.
So in March of 2016 there was still a hole in the machine, and a human being found it under maximum pressure with the world watching. That is worth stating plainly rather than quietly leaving out. It is also worth stating what happened afterward, which is that holes of this kind have been steadily closed, that the successors to that system no longer lose games of this sort, and that no human has taken a game from the top engines in a serious setting for a long time. The 2016 result is a snapshot of a transition, not a permanent balance of power.
What travels from these boards to everything else is not the winning. It is the mechanism. Chess and Go were always going to be the first to go, and the reason is embarrassingly simple. In a game, success is defined with complete precision by the rules. You win or you do not, and the scoring is free, instant, and beyond dispute. A system can play forty million games against itself over a weekend and receive forty million unambiguous verdicts.
Which suggests where to look next. Not for tasks that are easy. For tasks that come with a scoreboard.
A Fifty-Year Problem
Proteins are chains of amino acids that fold, within moments of being made, into intricate three-dimensional shapes. The shape determines what the protein does. The sequence determines the shape. Working out the second from the first had been an open problem in biology since roughly the early 1970s, and it was open in a way that resisted every kind of assault: physical simulation from first principles, statistical analysis of evolutionary relatives, decades of accumulated structural intuition.
In 1994, a group of researchers led by John Moult did something about the fact that everybody in the field was claiming progress and nobody could check.
They created an examination. Every two years since, the Critical Assessment of Structure Prediction has taken proteins whose structures have just been determined experimentally, or in some cases are still being determined, and released the sequences to the world. Any group may submit predictions. Nobody has access to the answers, because for some targets the answers do not yet exist. When the experimental structures come in, predictions are scored against them.
The scoring uses a measure called the Global Distance Test, which runs from zero to one hundred and can be thought of loosely as the percentage of the chain that ends up close enough to where it actually goes. Moult has said that a score of around ninety is informally regarded as competitive with determining the structure in a laboratory. For most of the history of the assessment, the best predictions hovered somewhere around sixty.
At CASP14, in 2020, an entry registered as group 427 scored a median of 92.4 across all targets. On the hardest category, the targets with no useful structural relatives to lean on, it scored a median of 87.0. Its average error was around 1.6 angstroms, which is roughly the width of a single atom.
Moult, who has chaired the assessment since he started it, walked his audience through the history of the competition before showing the graph. The graph is the whole story in one image. Two decades of lines crawling upward in the vicinity of sixty, and then one line standing somewhere the others had never been.
Group 427 was AlphaFold2. Moult announced that the problem, for single protein chains, had been solved.
The qualifier belongs there and belongs in every honest account. Single chains are not all of protein science. How proteins assemble into complexes, how they move, how they behave inside a living cell, all of that remained open. The grand challenge that closed was a specific, precisely stated one.
And that precision is the point of telling the story here. What made structure prediction fall was not that it turned out to be easy, since half a century of failure says otherwise. What made it fall is that in 1994 the field built itself a scoreboard.
Consider what CASP supplied, viewed as engineering rather than as science. It supplied an unambiguous measure of success, applicable to any prediction, computable in seconds. It supplied a ground truth that could not be gamed, since the answers were physically determined by somebody else in a laboratory. It supplied a stream of fresh problems on a fixed schedule. It supplied thirty years of scored historical attempts, which is to say a graded record of what better and worse look like in this domain.
A field that has done all of that has, without intending to, prepared its own problem for automated attack. The scoreboard is the precondition. Everything else is engineering and computation, and both of those have been getting cheaper every year for a long time.
This generalizes with uncomfortable ease. Wherever a discipline has agreed on a benchmark, it has published a target. Machine translation had scored benchmarks. Speech recognition had scored benchmarks. Image classification had a labelled set of a million-odd photographs and an annual competition, and everyone who watched that competition knows how the story ends. The pattern is not that the hard things fall first or the easy things fall first. The pattern is that the measured things fall first.
Which raises the obvious question about everything that has not been measured.
Scoring the Unscoreable
The last defence was supposed to be the things that cannot be scored.
Writing is the standard example. There is no procedure that takes a paragraph and returns a number for how good it is. Two competent editors disagree. The same editor disagrees with herself on a different day. Quality in language is entangled with context, audience, purpose, and taste, none of which reduce to a measurement. If the machines need a scoreboard, and language has no scoreboard, then language is safe.
That argument was sound. What happened to it is the reason nothing else in this account matters as much.
The field did not find an objective measure of good writing. There is still no such thing. What it did instead was manufacture a score out of the only material available, which was human judgment itself.
The method has a longer history than most people assume. In 2008, Knox and Stone described a system called TAMER in which a person watched an agent act and gave evaluations, and those evaluations were used to train a model that predicted what the person would say. The model, rather than the person, then supplied the training signal. In 2017, Christiano and colleagues published the work that most people treat as the direct ancestor of current practice, applying the idea to agents in Atari games. In 2019, Ziegler and colleagues carried it across to language models, and much of the vocabulary in use today was fixed in that paper. In 2022, Ouyang and colleagues demonstrated it at industrial scale on instruction-following, and shortly afterwards most people on earth met the results without being told what they were meeting.
The mechanism at the centre is worth seeing concretely, because it is much simpler than its reputation.
A person sits in front of two pieces of text, both produced by the model in response to the same prompt. The task is not to grade them. Grading is precisely what humans do badly, since one person’s seven is another’s five and neither is stable across an afternoon. The task is only to say which of the two is better. That is a judgment people make reliably, and the mathematics for turning a pile of such comparisons into a consistent scale is older than the field it is being used in, going back to work by Bradley and Terry in 1952 on paired comparisons.
Consider the working conditions in which that judgment gets made, because they end up mattering. The person is doing this many times an hour, for pay, against a written guideline explaining what the client considers a better answer. Some of the pairs are close to identical. Some involve subjects the person knows nothing about, where the more confident and better organized of the two answers will tend to be chosen whether or not it is the more accurate one. This is not a criticism of the people. It is a description of what any human being does under those conditions, and the resulting choices are the raw material.
Collect enough of these choices and you can train a second model whose job is to predict which of two texts a human would prefer. That second model outputs a number. And a number is all that was ever missing.
Look carefully at what has been built here, because it is easy to walk past. The score being optimized is not a fact about the world. It is a model of us. It is a compressed, learned imitation of the judgments of a particular set of people, hired at a particular time, working under particular instructions, tired in the afternoon like everybody else. Where CASP scored predictions against physical reality determined in a laboratory, this scores text against a statistical portrait of human approval.
The portrait is useful. It is also, necessarily, an approximation, and it differs from the thing it portrays in ways nobody has fully mapped. Whatever the model of our preferences fails to capture about our actual preferences is not in the objective, and therefore is not being optimized for, and therefore is subject to whatever happens to a quantity that a powerful optimizer has no reason to protect.
That is a fact about the construction, stated here without any claim about how bad the consequences are. The immediate point is narrower and harder to argue with. The category of unscoreable things is smaller than it looked. If a domain resists measurement, it is possible to build a measurement out of human preference and proceed. The defence held only as long as nobody thought of manufacturing the score.
What Falls Next
Which leaves a practical question, and a usable answer.
Take any activity you like: a job, a craft, a profession, a task you spent years learning to do well. Ask three questions about it.
First, can success and failure be told apart reliably? Not perfectly, and not by everyone, but consistently enough that competent judges agree most of the time. Games clear this trivially. Protein structure prediction clears it because a laboratory eventually produces the answer. Writing does not clear it in any absolute sense, but clears it in the relative sense that people can usually say which of two attempts is better.
Second, can that judgment be produced cheaply and at scale? This is the question that decides timing rather than possibility. A verdict that is free and instant, as in a game, means a system can generate its own training data by the million. A verdict that requires a laboratory is slower but still tractable, given a thirty-year archive of scored attempts. A verdict that requires paid human beings reading text is expensive, which is exactly why so much effort has gone into training a model to imitate those human beings and remove them from the loop.
Third, and this is the one that does not decide when a domain falls but decides what happens after: is the score actually the thing you want? A game’s score is the thing you want, because in a game winning is by definition the entire point. A protein structure measured against experiment is very close to the thing you want. A learned model of what annotators prefer is not the thing you want. It is a portrait of the thing you want, and there is a gap between a portrait and a face.
Run the three questions over your own work. Most people find the first answer arrives quickly and is uncomfortable. Most of what we call skill turns out to be gradeable, at least in the relative sense, by someone competent looking at two attempts side by side. The second question is where the honest uncertainty lives, since cost and scale are moving targets and they have been moving in one direction for a long time. The third question is the one almost nobody asks, and it is the one that determines what your field looks like on the other side.
It is worth doing this slowly on something ordinary. Take radiology. Can success and failure be told apart? Yes, and precisely, because a diagnosis is eventually confirmed or refuted by what happens to the patient. Can the judgment be produced cheaply and at scale? Largely yes, since hospitals have been accumulating scored examples in the form of images with confirmed outcomes for decades. Is the score the thing you want? Here the answer gets complicated, because what is scored is agreement with a recorded diagnosis, and what is wanted is a patient who does well, and those two things run together most of the time and come apart in exactly the cases that matter most.
Now take something that looks safer. Take management. The first question is already hard, since competent people disagree about whether a given manager is good and the disagreement does not resolve quickly. The second is harder, since the outcome of a management decision arrives years later, entangled with everything else that happened. The third is hardest of all, because every proxy anyone has proposed for good management, from retention to engagement scores to output per head, is obviously not the thing itself. On this diagnostic, management is not safe because it is deep. It is unmeasured, which is a different and less flattering kind of protection.
The machines have taken the axis of optimization. On any well-specified target, with enough computation, the search for good moves is no longer a contest, and the results have a way of being not merely stronger than ours but stranger, arriving at places our traditions taught us not to look. That is not a prediction. It is a description of chess, Go, shogi, protein structure, and a growing list of things that used to be on the list of what only people could do.
What remains is the target itself. Somebody decides what the score is. In every case described above, that decision took an afternoon, and every one of the extraordinary results is downstream of it, faithful to it, and indifferent to anything it left out.
So the last question is the one to sit with. In your work, is there a score yet? And if there is, who wrote it, and were they thinking about you when they did?
This is Part 2 of Whatever You Ask For, a series on the last thing machines will ever need from us, and how badly we do it.


