Machines Can Learn What We Do, Not What We Want
Twenty-five years of trying to automate the asking. A robot that wants your coffee. A proof that behaviour cannot reveal intent. Every attempt to outsource the goal hits one wall.
From Whatever You Ask For, a series on the one job machines cannot do for us.
Asking It Backwards
Everything so far has assumed that a person writes the objective. A human being decides what counts as success, hands it to a system, and lives with the consequences of having written a finite sentence around an infinite space. The obvious question, and one that a serious research community has spent a quarter of a century on, is whether that step can be automated too. Can a machine work out what we want, so that we do not have to spell it out?
The most elegant version of the idea reverses the usual direction of the problem. Ordinarily you are given an objective and you search for the behaviour that best satisfies it. Stuart Russell, in a short conference paper in 1998 and then in a 2000 paper with Andrew Ng that gave the field its name, asked what happens if you run the arrow the other way. Suppose you are given the behaviour, and you search for the objective that would make it optimal. This is inverse reinforcement learning, and its appeal is immediate. It seems to dissolve the entire problem of this book. You would not have to write down what you want. You would only have to act, and be watched.
The appeal is not naive. It is how humans transmit most of what cannot be said. An apprentice watches a master and absorbs a thousand judgments the master could never articulate. A child learns what a household values less from what it is told than from what it sees rewarded and ignored. We infer goals from behaviour constantly, effortlessly, and mostly correctly, in exactly the way the idea proposes a machine might. If people can do it, the reasoning goes, the process must be learnable.
Watch a joiner teach an apprentice and you can see why the idea is so seductive. The master rarely explains. He planes an edge, the apprentice planes an edge, the master glances and grunts, the apprentice tries again. Almost nothing is stated. The standard for a good edge passes from one to the other through a channel that carries no sentences, only examples and the faint signal of approval or its absence, which is precisely the channel inverse reinforcement learning proposes to build. If a boy can pick up a craft this way, surely a machine with unlimited patience and perfect memory can pick up ours.
For twenty-five years the field has pursued that intuition down several roads. The roads are worth walking, because they do not arrive where they were expected to, and where they arrive is the same place.



