The Attention Wager, Part 1: The Opening Bet
A bottleneck the whole field had agreed to live with. Eight researchers bet it could go entirely. The mechanism they reached for was already a century old.
This is The Attention Wager, a close reading of Attention Is All You Need, one section and one bet at a time.
In 1985, Jeffrey Moran and Robert Desimone lowered a microelectrode into the visual cortex of a macaque and waited.
The monkey had been trained to do something simple and strange. Two stimuli would appear inside the receptive field of a single neuron, one at a location the animal had been cued to attend to, one at a location it had been told to ignore. The cell could see both. The question was whether the cell cared.
It cared enormously. When the monkey attended to one of the two stimuli, the neuron in area V4 responded roughly as though the other stimulus were simply not there. The response to the unattended object was, in the authors’ word, dramatically reduced. Cells one stage earlier, in striate cortex, showed nothing of the kind. Attention was not a metaphor for what the animal was thinking about. It was a measurable event in a single cell, and its mechanism was subtraction. The brain was not turning up what mattered. It was turning down everything else.
Ninety-five years earlier, without a microelectrode or a monkey, William James had said the same thing in the same order. In the eleventh chapter of The Principles of Psychology, he wrote that every one knows what attention is, and then defined it anyway: the taking possession by the mind of one out of several simultaneously possible objects. Then came the clause that matters most. Attention, James wrote, implies withdrawal from some things in order to deal effectively with others.
Withdrawal. Not amplification. The oldest definition in the field and the first single-cell recording agree that attention is what you subtract.
Hold that thought. It comes back in about two thousand words, wearing a different notation.
The cost everyone had agreed to pay
By 2016, machine translation had settled into a shape. You took a sentence, ran it through a recurrent neural network one word at a time, accumulated a hidden state that carried everything seen so far, and used that state to produce a translation one word at a time on the other side. The dominant variants were LSTMs and gated recurrent units. They worked. Google had put one into production.
The architecture had a property that the field discussed constantly and had nonetheless stopped treating as a problem. A recurrent model computes the hidden state at position t as a function of the hidden state at position t minus one. That is not an implementation detail. It is the definition. Position five cannot be computed until position four exists, which cannot be computed until position three exists.
The Transformer paper’s own introduction states the consequence with unusual bluntness for a section whose job is throat-clearing. This inherently sequential nature, the authors write, precludes parallelization within training examples, and the problem becomes critical at longer sequence lengths, because memory limits how many examples you can batch together to compensate.
Consider what that means when the hardware in the room is a GPU. A GPU is a machine for doing thousands of things at once. A recurrent network is a machine for doing one thing after another. Running the second on the first is like hiring an orchestra and handing them a piece written for solo violin. Everyone shows up. Most of them wait.
The field knew. There were fixes. The paper cites factorization tricks and conditional computation, both of which improved efficiency, one of which improved quality as well. And then it delivers the sentence that sets up everything that follows: the fundamental constraint of sequential computation, however, remains.
This is what a field looks like when it has priced a cost into its worldview. Nobody thought recurrence was free. Everybody thought it was necessary.
Attention arrives, as a helper
Attention entered machine translation in 2014, and it entered as a repair.
The problem it repaired was specific, and it is worth being precise about it, because the shape of this first fix determined how the field thought about attention for the next three years. In the encoder-decoder models of the time, the encoder read the entire source sentence and compressed it into a single fixed-length vector. The decoder then produced the translation from that vector alone. Every word of a forty-word sentence had to survive inside one array of numbers whose size did not change no matter how long the sentence got.
Dzmitry Bahdanau, then at Jacobs University Bremen, working with Kyunghyun Cho and Yoshua Bengio at Montreal, named this the bottleneck. Their proposal was to stop forcing the compression. Let the encoder keep a state for every input position, and let the decoder, at each output step, compute a set of weights over all of those states and take a weighted sum. The model would learn where to look. Their title states the resulting capability plainly: neural machine translation by jointly learning to align and translate. Alignment, which statistical translation systems had needed hand-built machinery to approximate, fell out of the network as a learned byproduct.
The mechanism worked, and it spread. A year later Minh-Thang Luong, Hieu Pham, and Christopher Manning published a set of refinements that made it cheaper and more general. By 2016 attention was standard equipment in essentially every competitive translation system.
And in every one of those systems, it sat on top of a recurrent network.
This is the detail on which the entire story turns, and the Transformer paper flags it in a single sentence in its introduction. Attention mechanisms, the authors note, had become integral to sequence modeling, allowing dependencies to be modeled without regard to distance. In all but a few cases, however, such mechanisms were used in conjunction with a recurrent network.
Read that again with an eye for what it is not saying. It is not saying attention was underrated. It is saying that for three years the field had possessed a mechanism that relates any two positions directly, in one step, regardless of how far apart they are, and had bolted it onto the very architecture whose defining weakness is that it cannot do that. Attention was the passenger. Recurrence was the car.
Nobody asked whether the car was load-bearing.
The man who asked
Almost nobody, anyway.
Jakob Uszkoreit did not set out to work on language. He backed into it. His father, Hans Uszkoreit, was a computational linguist who as a teenager had spent fifteen months in an East German prison for protesting the Soviet invasion of Czechoslovakia, escaped to the West after his release, studied computers and linguistics in Berlin, and eventually found his way to an artificial intelligence lab in Menlo Park, where his son was born. The family returned to Germany. Jakob went to university there, took an internship at Google’s Mountain View office, and landed in the translation group. He has described this as ending up in the family business. He never finished the PhD.
In 2012 he joined a Google team building a system that could answer questions directly on the search page. Apple had just shipped Siri, and Google’s leadership had decided this was an emergency. Uszkoreit’s view, in hindsight, is that the panic was unfounded. But it put resources behind a group working on machines that could hold something resembling a conversation, and that is where he ran into the wall.
Recurrent networks, even with long short-term memory bolted on, could not hold a long passage together. The canonical demonstration is a sentence in which a fact stated early determines the meaning of a phrase stated late, and the model, marching left to right, has to still be carrying the early fact when it arrives at the late phrase. LSTMs made this possible for longer spans than plain recurrence had allowed. They did not make it work at the scale Uszkoreit wanted. His own assessment of the field’s toolkit at that moment was that the methods he was applying were, in his words, “basically Band-Aids.”
Around 2014 he began developing an alternative he called self-attention: let a model translate a word by referring directly to any other part of the passage, weighting each one according to how much it clarifies the word at hand. He suspected two things about it. The first was that it might simply work better. The second was less obvious and turned out to matter more. Self-attention looks at many inputs simultaneously rather than in sequence, which meant it was shaped exactly like the parallel processing chips the machine learning boom was producing in volume. The mechanism and the hardware were built for each other. Nobody had designed it that way.
The reception was cold. Dropping recurrence meant discarding the architecture that the entire field had spent years perfecting, and Uszkoreit’s account is that people raised their eyebrows at the suggestion. The skeptics included his father, with whom he was, by his own description, not entirely seeing eye to eye across the dinner table.
He persuaded a few colleagues to run experiments anyway. The results were promising enough to publish in 2016, on a small scale, using only short spans of text. That paper is Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit, “A Decomposable Attention Model,” presented at EMNLP 2016. It performed natural language inference with attention and no recurrence at all.
Then his collaborators moved on. The technique was good enough to deploy, so they went and deployed it, across Google search and eventually ads. It was, by the standard measure, a success. Uszkoreit was the only one who thought the small experiment implied a large one.
Which is why, when "Attention Is All You Need" appeared a year later and its introduction conceded that attention was almost always used in conjunction with a recurrent network, the citation attached to the phrase in all but a few cases was that 2016 paper. The exception to the rule, sitting in the bibliography at reference twenty-seven, was the work of the man about to break it.
A café, a corridor, an espresso machine
The Transformer was assembled out of overheard conversations. This is not a flourish; it is what the participants describe.
At lunch in a Google café in 2016, Uszkoreit heard Illia Polosukhin complain about the team building direct answers for the search page, which had to return results in milliseconds and was not succeeding. Uszkoreit’s suggestion was self-attention. Polosukhin sometimes worked with Ashish Vaswani, who had come from a doctorate in machine translation at USC to Google Brain and was looking for a large problem. Vaswani’s building sat next door to Polosukhin’s. He heard about the idea and joined.
The three of them wrote a design document. The name in its title was chosen at the outset, on the theory that the mechanism transforms the information passing through it, and also because Uszkoreit had owned two Transformer toys as a small child. The document closed with a drawing of six of them shooting lasers at each other in the mountains. Its opening sentence informed the reader that the authors were awesome.
Niki Parmar joined from Google search, where she had been building model variants with Uszkoreit. Llion Jones heard about self-attention from a colleague named Mat Kelcey and came aboard; Kelcey, briefed on the project later, told Jones he doubted it would work, and now describes this as the most incorrect prediction of his life. Łukasz Kaiser arrived from a separate Google Brain effort on language models, bringing his intern, Aidan Gomez, an undergraduate who had talked his way into a position reserved for doctoral students and would not find this out for months. Kaiser and Gomez decided to merge their project into the other one.
Polosukhin left Google in early 2017 to start his own company, which is why the paper’s byline carries no affiliation for him.
Then the group hit a wall. They built a self-attention translation model, measured it against the standard benchmark, and found it landed roughly level with the best LSTM systems of the day. Level. Not ahead. A radical architecture that discards the field’s accumulated priors and returns parity is not a result. It is an expensive way to stay where you were.
The wall came down because a man walked past a doorway.
Noam Shazeer had been at Google since 2000 and had a reputation dating back to the company’s early advertising system. He had spent five years on deep learning and had become interested in language models, which in his estimation were nowhere near capable of the fluency he thought achievable. Walking down a corridor in Building 1965, he passed Kaiser’s workspace and overheard Vaswani and Parmar talking with some heat about self-attention. He found recurrent networks irritating. The proposal to replace them struck him as both correct and fun.
Uszkoreit’s account of what Shazeer contributed is worth stating carefully, because it describes something the field talks about rarely. A mechanism that is sound in theory, he has said, often requires very careful implementation by a small number of experienced people before it shows any signs of working at all. Shazeer did not debug the team’s code. He read the idea and wrote his own implementation from scratch, checked in with Kaiser occasionally, disappeared for a while, and came back with a version that worked. His colleagues describe what he did with words like magic and alchemy. Jones simply calls him a wizard.
The specific things Shazeer added are the subject of the next three articles in this series. The paper’s credit footnote assigns him scaled dot-product attention, multi-head attention, and the parameter-free position representation, which are the anchor texts of Parts 2, 3, and 4 respectively. One person, in a sprint, produced three of the four components this series takes apart.
What the bet actually was
The Abstract states the position in one sentence. The authors propose a new simple network architecture based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
The word doing the work there is entirely.
Removing recurrence was not, by itself, the radical part. Others were trying. Section 2 of the paper names them: the Extended Neural GPU, ByteNet, and ConvS2S, all of which replaced recurrence with convolution to compute representations for all positions in parallel. The goal of reducing sequential computation, the authors acknowledge, was already the foundation of that work.
But convolution bought parallelism at a price, and the paper is precise about the price. In these models the number of operations required to relate two arbitrary positions grows with the distance between them: linearly for ConvS2S, logarithmically for ByteNet. Still not free. Two words at opposite ends of a long sentence remain expensive to connect, and the paper cites the standard result that longer paths make long-range dependencies harder to learn.
Self-attention collapses that to a constant. Any position to any other position, one step, whatever the distance. Table 1 lays the three layer types side by side and the column that matters is maximum path length: order n for recurrent, order log n for convolutional, order one for self-attention.
That is the bet. Not that attention is useful, which the field already believed. The bet was that constant path length was worth so much that you could discard every other structural prior the field had accumulated and still come out ahead.
The paper does not hide what it gave up. Two sentences after the constant-operations claim comes an admission that rarely survives summary: the constant cost arrives at the price of reduced effective resolution, because attention averages over weighted positions. Multi-head attention exists to counteract that. The paper is telling you, on page two, that its headline mechanism blurs things and that one of its most celebrated components is a patch. Part 3 will return to this.
And there is a second price, stated in Table 1 and defended in Section 4, which the paper treats as a reasonable trade and which the following decade has treated as the central problem in AI infrastructure. Complexity per layer for self-attention is order n squared times d. The paper’s defense is that n is usually smaller than d for sentence-length inputs in translation, which was true in 2017 and stopped being true the moment anyone wanted to feed a model a book. Part 7 collects on that.
Twelve hours
The deadline was May 19, for the December conference. The last two weeks were spent in Building 1965, partly because some of the team had desks there and partly because the espresso machine was better than the one in 1945. Gomez, the intern, has described a period of continuous debugging in which nobody slept much, systematically removing components to see whether the model still worked without them. Much of what is now called the Transformer is the residue of that process: the parts that could not be removed.
Then the numbers came in.
The base model trained for one hundred thousand steps on a single machine with eight NVIDIA P100 GPUs. Twelve hours of wall-clock time. That model surpassed every previously published model and ensemble on English to German. The big model, three and a half days on the same eight-GPU machine, reached 28.4 BLEU, beating the previous best, ensembles included, by more than two full points. Uszkoreit celebrated by opening a bottle of champagne he had been keeping in his mountain expedition truck.
Now look at Table 2, the quietest table in the paper and the most damaging. It lists training cost in floating point operations. The competitive ensembles of the day sit between 1.1 times ten to the twenty-first and 7.7 times ten to the nineteenth. The Transformer big model sits at 2.3 times ten to the nineteenth, and the base model an order of magnitude below that. The new state of the art was set by the cheapest system in the table.
A field that had spent years buying quality with scale was shown a model that bought it with structure instead. And because that structure mapped cleanly onto parallel hardware, the same architecture would turn out to be the ideal vehicle for spending compute later, once there was compute worth spending. The 2017 result reads as an efficiency story. In retrospect it was a capacity story wearing an efficiency story’s clothes.
The paper was submitted with roughly two minutes to spare. Parmar has said the English to French numbers arrived about five minutes before submission, while she sat in the micro-kitchen waiting for the last one. This may explain a small inconsistency that survives in the published paper to this day: the Abstract and Table 2 report an English to French score of 41.8, and Section 6.1 reports 41.0. The most influential paper in modern machine learning contains a number that disagrees with itself, because the experiment finished after the prose did.
The title arrived a couple of nights before the deadline. Jones, who is Welsh, pointed out that the team had rejected the field’s accepted practice in favor of a single technique, and that the Beatles had already written the sentence for them. He has said the thought took about five seconds and that he did not expect anyone to use it.
They used it.
Back to the monkey
What the Transformer computes for each position is a weighted sum over all other positions, with the weights produced by a softmax. A softmax does not merely rank. It suppresses. Raise one score and every other score in the distribution falls, because the total is pinned to one. The mechanism at the center of the architecture is competitive, and its output is defined as much by what it drives toward zero as by what it lifts.
Which is what Moran and Desimone found in a V4 neuron with two objects in its receptive field. Which is what James described as withdrawal from some things in order to deal effectively with others.
It would be too much to claim the 2017 authors were implementing James. They were not. Queries, keys, and values come from information retrieval, not psychology, and Part 2 will trace that lineage properly. But the word they chose was not an accident either, and the field had been sharpening the concept behind it for a century. When Anne Treisman and Garry Gelade published feature integration theory in 1980, they argued that features are registered early and in parallel, and that attention is the operation binding them into objects. Parallel registration, then selective binding. That is a description of a Transformer layer written thirty-seven years early by people studying human vision.
Whether that convergence is a real fact about intelligence or an accident of vocabulary is the question this series ends on. Part 8 will try to answer it.
There is a better ending available, though, and it happened in December 2017 at the conference poster session. The room stayed full for four hours. Security eventually had to clear it out. At some point in that evening a man came up to the poster to tell the authors he was impressed by the work, and the man was Sepp Hochreiter.
Hochreiter co-invented long short-term memory. He had come to congratulate the people who had just made it obsolete.
The Bet
The stake. Recurrence is not necessary for sequence modeling. An architecture built entirely on attention, with no recurrent or convolutional layers anywhere, can match and exceed the state of the art, and the constant path length between any two positions is worth more than every structural prior being discarded.
The return. Won, and faster than the authors expected. Within roughly two years the architecture had displaced recurrent models across natural language processing. The frontier language models built since are its descendants. The paper’s own stated ambition, in its conclusion, was to extend the approach to other modalities and to make generation less sequential. It got an industry.
Status. Closed, with one asterisk. The bet was won on a trade the paper made explicitly and priced correctly for 2017: constant path length in exchange for quadratic cost in sequence length. Every long-context problem in AI today is the interest payment on that trade. Part 7 collects it.
Part 2 reads Section 3.2.1, where three matrices and one softmax turn a lookup table into something that can be trained.


