The 2017 Transformer Architecture That Changed How Machines Learn
Published October 1, 2026
Eight P100 GPUs and 3.5 days were enough to beat the BLEU record. The paper did it by deleting the recurrence the field once considered mandatory.

In June 2017, eight authors from Google Brain and the University of Toronto uploaded an eight-page paper to arXiv. Its title was plain: "Attention Is All You Need". There was no company name to announce, no preview program, no keynote. The result was an architecture that deleted the component the field considered mandatory, and that now sits underneath almost every large AI model in use.
The paper did something the field had never managed in one step: better translation quality with substantially less computation. Their large model trained for 3.5 days on eight P100 GPUs and reached a BLEU score of 41.8 on WMT 2014 English-to-French. The previous best models cost many times more compute, and several of them only reached that score by ensembling ten models at once.
What matters is not the numbers but the architecture. Before 2017, the dominant way to process a sentence was a recurrent neural network: the model read one word at a time, carrying everything it knew about previous words in a single state it kept updating. That approach had one fatal flaw that more waiting could not fix. To process word 1,000, the RNN had to finish word 999 first. No number of GPUs changes that, because the dependency is sequential.
An architecture that deletes the sequence
The Transformer solves this the way it sounds: let every word see every other word directly, in the same layer.
This is self-attention. Each word becomes a vector representing its content. Every pair of words then computes how related they are, and attention weight is allocated according to that relevance. The word "cat" directs more attention to "satin" beside it than to "the" two paragraphs away. Same layer, all words simultaneously.
The decisive word is simultaneously. With no dependency between time steps, the entire computation runs in parallel across thousands of GPU cores. An RNN takes 1,000 sequential steps. A Transformer processes all 1,000 words in one step. On hardware built for parallelism, that is the difference between impossible and trivial.
Google was not the first to use attention. Attention had been in use since 2014 to connect encoders to decoders, but always alongside an RNN. The 2017 contribution was removing the RNN entirely and adding a multi-head attention sublayer that handles several relationships in the same pass.
Each number above is a different model, ordered from the laggard: a small Transformer at 28.4, the large Transformer at 41.0, and the 41.8 reported in the paper's abstract for the variant with adjusted dropout. All on the same benchmark.
The numbers that made reviewers reconsider
The most-cited part of this paper is not the architecture but its cost table. The authors counted floating point operations for each training run and put them side by side.
| Model | BLEU English-German | Training cost (FLOPs) |
|---|---|---|
| ByteNet | 23.75 | not listed in the main table |
| GNMT + RL | 24.6 | 2.3 x 10^19 |
| ConvS2S | 25.16 | 9.6 x 10^18 |
| Transformer (base) | 27.3 | not listed in the main table |
| Transformer (big) | 28.4 | 3.3 x 10^19 |
Look at both columns together: the model that wins on BLEU is not necessarily the cheapest. ConvS2S is the most frugal and scores lowest among the modern entries. What the Transformer delivers is the top-right corner of that table: the highest score and lower cost than its closest predecessor.
A sentence in the abstract sums it up. The authors wrote that their model reached a new record "at a small fraction of the training costs of the best models from the literature". That is not a technical claim, it is an economic one. If one training run is cheaper, researchers run more experiments, and more experiments is what moves a field.
For the large model the arithmetic is concrete: 300,000 steps at roughly one second each, on a single machine with eight P100 GPUs. For the small model, 100,000 steps and 12 hours. That 12-hour figure is what made the paper reproducible by a small lab, which may be its most important long-term effect.
From translation to language models
The Transformer changed role from translator architecture to general architecture. Current large language models use essentially no recurrence or convolution. GPT, BERT, and the whole family descending from them are built from the same blocks: multi-head attention, residual connections, layer normalization, and a position-wise feed-forward network.
Two structural changes made that transformation possible. First, positional encoding: because attention has no notion of order by itself, position must be injected as an extra vector. Without it the model would see "cat sat mouse" and "mouse sat cat" as identical. Second, residual connections and layer normalization made it possible to stack far deeper networks without instability, which was impossible to do repeatedly with RNNs.
Note where the cost went. For short-sequence classification, sequential processing was sufficient. For long documents, RNNs hit a ceiling. The Transformer did not merely suit long inputs; it removed the structural reason RNNs were still used. Once that switch happened, the one advantage RNNs kept was streaming long input, which is now slower and no longer the point.
Since 2017 the context window has become the axis of competition, and every increase in it is a modification of the same decision. Models with a 4,096-token window in 2019, 100,000-token windows by 2023, and a million-token window by 2024. Each of those pushes traces back to an architectural decision eight authors made in 2017.
What the Transformer genuinely gave up
There is one thing the Transformer does not offer, and it is worth naming so nobody reads the paper as a claim about something it is not.
The Transformer in its original form has no memory outside its input window. Every document is processed from zero. In an RNN, information from the first word is still available in the internal state when the last word is processed, for as long as it fits. In a Transformer, once the window fills, whatever fell outside is gone.
This is why a model with a 100,000-token window still has real problems on a 500,000-token document, and why teams have had to choose between a larger window and better use of the window they have. Both are expensive, and both have a different bill.
One other structural difference: the Transformer computes attention for every pair of words, so its cost is quadratic in document length. At 1,000 words that is around a million pairwise computations. At 100,000 words it is ten billion. That number does not disappear; it moves somewhere else, into the problem that FlashAttention would address in 2022.
Why 3.5 days changed everything
Back to the training figure. Eight P100 GPUs, 3.5 days, one paper. No cluster. No infrastructure program, no institutional lab, no dedicated budget.
Before 2017, infrastructure cost often decided who could solve what. Labs with thousands of GPUs ran experiments that were simply out of reach for anyone working alone. The Transformer shifted that equation: the hardware required per experiment dropped sharply, and the field became something a single person could compete in.
Another detail often forgotten: the large model used a dropout rate of just 0.1 for English-to-French, not the 0.3 used for the small model. On paper it is a trivial detail, but it shows something larger. The authors did not simply follow the existing recipe. They changed it because a larger model trained on more data overfits more slowly. It took years for the field to absorb that lesson.
Nine years on, no mainstream AI architecture has left this family. Image generation, protein structure prediction, self-driving cars, and the chatbot you are reading this through all use attention. What survived from the RNN is not the architecture but the idea: processing a sequence with attention that can be parallelised.
The question now is not whether the Transformer survives, but how long it can be patched without a successor. Every limitation found since 2017, quadratic attention, positional bias, wasted computation, has been fixed by improving the paradigm rather than replacing it. Whether that continues is the open question.
Deeper dive: why this architecture survived almost nine years
One question rarely gets asked: why has a nearly nine-year-old architecture not been replaced?
Strip one layer out of the Transformer and what remains was already enough to overturn the paper's reviewers' assumptions. Here is the comparison they made themselves:
| Approach | Core idea | Result in the paper |
|---|---|---|
| ByteNet | Convolution for sequence modelling, no recurrence | BLEU 23.75, cost not listed |
| ConvS2S | Convolution only, no attention at all | BLEU 25.16, cost 9.6 x 10^18 FLOPs |
| GNMT + RL | RNN with attention bolted on | BLEU 24.6, cost 2.3 x 10^19 FLOPs |
| Transformer base | Attention only, no recurrence and no convolution | BLEU 27.3, cost far below all of the above |
| Transformer big | Multi-head attention with adjusted dropout | BLEU 28.4, cost 3.3 x 10^19 FLOPs |
The last row is the argument. The best model is not the cheapest and not the simplest; the best model is the one with the unnecessary parts removed. ConvS2S is simpler than an RNN. GNMT + RL adds reinforcement learning on top and ends up costing twice as much as ConvS2S for a lower score. Every approach before 2017 added machinery, while the Transformer achieved more with less of it. The answer is not that nobody tried. People tried repeatedly, and most of them failed in the same way.
The failed attempts reduced complexity instead of adding it. Linear attention and Performer swap the softmax attention for a form that can be computed linearly. In theory that sounds sound and looks obvious. In practice the FLOP count drops sharply while wall-clock time barely moves, because FLOPs are not the real cost. What is expensive is not matrix multiplication but moving data from memory into the compute units. Reducing multiplications does not touch that.
The successful attempts added structure instead of removing it. FlashAttention is the clearest case. Its principle does not simplify attention at all: the attention matrix stays N by N and is still computed exactly. What changed is the order in which memory is accessed. The result is 2 to 4 times faster with no quality traded away. This pattern repeats across the field: tractable optimisation wins, and exact reduction does not.
There is one more thing rarely mentioned, and it is about money. Large Transformers sit under unavoidable cost pressure: training cost rises close to linearly with model size while capability does not rise linearly with it. That pressure is precisely what the architectural decisions in this paper were trading against, and it is why the 3.5-day training figure matters as much as the BLEU score.
Which raises the question the paper could not have asked in 2017: is self-attention still the right choice for contexts measured in millions of tokens, or will this architecture eventually look as primitive as RNNs look to us now?
Two possible answers. The first is that this architecture gets polished indefinitely and there is no successor at all. The second is that a successor arrives from an unexpected direction, not from a larger model but from a smaller and more specialised one. That pattern already happened once, in 2017, which is why an eight-page paper still sets the terms of the field.
Neither answer can be known yet. What can be known is this: an eight-page paper trained on eight GPUs changed the direction of the field, and that direction is still visible in every model that exists today.
Related tools
Free browser tools that apply to this topic.
Share this article
Share to
Related articles

October 1, 2026
Why AI Misreads Long Documents, and What Actually Fixes It
A pattern repeats across models: accuracy peaks when the relevant information sits at the start or end of a document, and collapses in the middle.

September 30, 2026
When Did AI First Appear? Not 1956, but 13 Years Earlier
The name was coined at Dartmouth in 1956, but the idea was written down in 1943. AI's history shows a pattern that keeps repeating.

September 30, 2026
20 Countries Cap Social Media for Kids, Indonesia Moves First
Australia from December 2025, Indonesia from March 2026. At least 20 countries have similar rules, but only some are actually in force.



