The masked self-attention inside every GPT-1 block, worked through by hand and in PyTorch.
The takeaway first
TL:DR;
In 2018, most NLP systems were built task by task. GPT-1 showed that you could train one Transformer on raw books to predict the next word, then reuse the whole network for almost any language task with only small changes.
The paper calls its approach semi-supervised: an unsupervised pretraining stage on unlabeled text, followed by a supervised fine-tuning stage on each labeled task. Today we would just say "pretrain, then fine-tune". GPT-1 is where that recipe was first shown to work this well with a Transformer.
Generative pretraining
A 12-layer Transformer decoder learns to predict the next token over 7,000+ unpublished books. No labels at all.
Discriminative fine-tuning
The same network, plus one small linear layer, is trained briefly on each labeled task: entailment, QA, similarity, classification.
Learn an entire reusable language model.
The results backed this up. GPT-1 beat the previous state of the art on 9 of the 12 datasets it was tested on, often beating models designed specifically for that task. In the ablation study, removing pretraining caused by far the largest drop in performance.
Motivation
Labeled data is scarce, raw text is not
Supervised deep learning needs lots of labeled examples, and good labels are expensive. Someone has to read each premise and hypothesis pair and decide whether one implies the other, or read a passage and write the correct answer. Most NLP datasets of that time had thousands or tens of thousands of examples.
Unlabeled text, on the other hand, is nearly free and practically unlimited. The question the paper asks is simple:
The authors name two open problems that had held this back:
- Which objective to pretrain on? Language modeling, machine translation and discourse coherence had all been tried, and it was unclear which one transferred best.
- How to transfer what was learned? Earlier methods often needed task-specific architecture changes, complex learning schemes, or extra auxiliary objectives for each new task.
GPT-1's answer to both is short: use language modeling as the objective, use a Transformer as the model, and transfer the whole network by turning every task into a sequence of tokens.
Why a Transformer, not an LSTM?
Recurrent networks squeeze everything they have read into one hidden state that is passed along word by word, so information from far back fades. Self-attention lets every token look directly at every earlier token. The authors argue that this gives a more structured memory for long-range dependencies, and that this is what makes the representations transfer well. The ablations in section 10 test that claim directly.
Related work
What transfer learning looked like before
Reusing knowledge from unlabeled text was not new. What was new was how much got reused. It helps to see what earlier approaches transferred.
Word-level embeddings
Word2Vec and GloVe learn one fixed vector per word:
That vector is the same whether the sentence is about a river bank or a savings bank. Only the first layer of the downstream model gets pretrained knowledge. Everything above it still learns from scratch on the small labeled set.
Phrase- and sentence-level representations
Other work learned to compress a whole sequence into a single vector, usually with averaging, CNNs, RNNs or LSTMs:
This captures more than single words, but the downstream task gets one summary vector and nothing else.
THE GPT Framework · Stage 1
Generative pretraining
Given an unlabeled corpus of tokens \(\mathcal{U} = \{u_1, \ldots, u_n\}\), the model maximizes the standard left-to-right language-modeling likelihood:
- \(k\)
- the context window, 512 tokens in GPT-1
- \(\Theta\)
- all of the network's parameters
- \(P\)
- the probability the model assigns to the true next token
Simply said, given the previous tokens, make the true next token as probable as possible. Averaged over billions of predictions, this one objective forces the model to pick up grammar, facts, sentiment, coreference and some reasoning, because all of these help it guess what comes next.
Forward Pass
The model is a decoder-only Transformer: masked multi-head self-attention followed by a position-wise feed-forward network, stacked 12 times. The paper writes the whole forward pass in three lines:
Read bottom to top. Token ids become vectors, pass through 12 masked self-attention blocks, and the final hidden state is projected back to the vocabulary with the transpose of the same embedding matrix.
- \(U\)
- the context of token ids, one-hot encoded
- \(W_e\)
- the learned token-embedding matrix
- \(W_p\)
- the learned position-embedding matrix (not the sine waves of the original Transformer)
- \(n\)
- the number of layers, 12
The GPT Framework · Stage 2
Supervised fine-tuning
Fine-tuning reuses the whole pretrained network and adds just one linear layer, \(W_y\). Feed in a labeled example \(x^1, \ldots, x^m\), take the top block's hidden state at the last token, \(h_l^m\), and predict the label from it:
Every task becomes a sequence of tokens
Classification fits this directly, but entailment has two sentences and multiple-choice QA has a passage, a question and several answers. Instead of building a custom architecture for each, GPT-1 changes the input: it joins the pieces into one sequence with a start token, a delim token between segments, and an extract token at the end, whose hidden state goes to \(W_y\).
For similarity, both orderings run separately and their hidden states are added. For multiple choice, each answer gets its own sequence and a softmax picks one.
The training objective
The task loss \(L_2\) is the log-likelihood of the correct labels. The paper also keeps the language-modeling loss \(L_1\) running on the task's own text, at half weight:
This auxiliary term acts as a regularizer and speeds up convergence: the model specializes while it keeps practising language. Training is short and gentle: learning rate \(6.25 \times 10^{-5}\), batch size 32, usually just 3 epochs. The only new parameters are \(W_y\) and the embeddings for the special tokens.
Experiments · setup
The recipe: data, cleaning and hyperparameters
BookCorpus
The pretraining corpus is BookCorpus, over 7,000 unique unpublished books in genres such as adventure, fantasy and romance. (The full corpus as originally described has around 11,000 books, 74 million sentences and roughly a billion words.)
Cleaning with ftfy
Raw text from the web and from scanned books is full of mojibake, text that was decoded with the wrong
character encoding. ftfy ("fixes text for you") repairs it before anything else happens:
François → François
The full preprocessing order: raw text → ftfy cleanup → punctuation and whitespace standardization → spaCy tokenizer → byte-pair encoding with 40,000 merges → language-model training.
Model and training
| Component | GPT-1 |
|---|---|
| Architecture | 12-layer decoder-only Transformer |
| Hidden size · heads | 768 · 12 heads (64 dims each) |
| Feed-forward size | 3,072 |
| Context length | 512 tokens |
| Vocabulary | BPE, 40,000 merges |
| Activation | GELU |
| Position encoding | Learned (not sinusoidal) |
| Optimizer · peak LR | Adam · 2.5 × 10⁻⁴ |
| LR schedule | Linear warm-up for 2,000 steps, then cosine decay to 0 |
| Batch · epochs | 64 sequences × 512 tokens · 100 epochs |
| Regularization | Dropout 0.1, modified L2 (w = 0.01) |
| Initialization | N(0, 0.02) |
| Fine-tuning | LR 6.25 × 10⁻⁵, batch 32, 3 epochs, λ = 0.5 |
The pretrained model reached a very low token-level perplexity of 18.4 on BookCorpus. These numbers seem small today, but they became the template: GPT-2 and GPT-3 kept this shape and mostly scaled it up.
Experiments · downstream tasks
Twelve benchmarks, and what they measure
GPT-1 was fine-tuned on four kinds of task: natural language inference, question answering and commonsense reasoning, semantic similarity, and text classification. A few of these needed explaining before the results made sense to me.
Textual entailment (natural language inference)
Given a premise, does a hypothesis follow from it, contradict it, or neither?
| Premise | Hypothesis | Label |
|---|---|---|
| A man is playing a guitar. | A person is playing an instrument. | Entailment |
| A woman is running in a park. | Nobody is in the park. | Contradiction |
| A woman is running in a park. | She runs every morning. | Neutral |
It is hard because it needs lexical knowledge (a guitar is an instrument), coreference, and some reasoning. Two datasets come up again and again:
- SNLI (Stanford Natural Language Inference): sentence pairs written about image captions.
- MultiNLI (Multi-Genre NLI): the same task across many genres, such as fiction, government reports, speech transcripts and letters. "Matched" tests on genres seen in training, "mismatched" on unseen ones.
The benchmarks, decoded
Each benchmark is a public dataset with a fixed test set, so different models can be compared on the same questions. Here is what every abbreviation stands for and what it asks the model to do:
| Name | Stands for | What the model must do |
|---|---|---|
| Natural language inference | ||
| MNLI | Multi-Genre Natural Language Inference | Decide if a hypothesis follows from, contradicts, or is unrelated to a premise, across many genres. "-m" is matched genres, "-mm" mismatched. |
| SNLI | Stanford Natural Language Inference | The same task on sentences describing photos. |
| SciTail | Science entailment dataset | Decide if a statement follows from a sentence, using school science exam questions. |
| QNLI | Question Natural Language Inference | Decide if a Wikipedia sentence contains the answer to a question. |
| RTE | Recognizing Textual Entailment | Decide if one passage implies another. Very small: about 2,500 training examples. |
| Question answering & commonsense reasoning | ||
| Story Cloze | Story Cloze Test ("cloze" means fill in the blank) | Read a four-sentence story and pick the right ending out of two. |
| RACE | ReAding Comprehension from Examinations | Answer multiple-choice questions on passages from English exams for Chinese students. "-m" is middle school, "-h" high school. |
| Classification & semantic similarity | ||
| CoLA | Corpus of Linguistic Acceptability | Decide if a sentence is grammatical. |
| SST-2 | Stanford Sentiment Treebank, 2 classes | Decide if a movie-review sentence is positive or negative. |
| MRPC | Microsoft Research Paraphrase Corpus | Decide if two news sentences mean the same thing. |
| STS-B | Semantic Textual Similarity Benchmark | Score how similar in meaning two sentences are, from 0 to 5. |
| QQP | Quora Question Pairs | Decide if two questions asked on Quora are duplicates. |
| Overall | ||
| GLUE | General Language Understanding Evaluation | Not a single task: a bundle of nine (including CoLA, SST-2, MRPC, STS-B, QQP, MNLI, QNLI and RTE), averaged into one overall score. |
Results against the previous best
Two datasets from each category, with every method the paper compares against on them (Tables 2–4). GPT-1 is in red, the strongest earlier result in black, and other earlier methods in grey. The badge is GPT-1's gap to the best earlier result.
Scores as reported in the paper. Bars start at zero, and every panel uses a 0–100 scale. "(5x)" and "(9x)" mark ensembles of 5 or 9 models. CoLA is Matthews correlation × 100; GLUE is the benchmark's overall score.
GPT-1 is ahead on 9 of the 12 datasets, often against ensembles, while being a single model. The biggest gains are where long-range context matters most: +8.9 on Story Cloze (choosing the right ending to a story) and +5.7 on RACE (reading comprehension from school exams). CoLA jumps from 35.0 to 45.4, and the GLUE overall score from 68.9 to 72.8.
The three misses are RTE, a small entailment set where a multi-task biLSTM wins (61.7 vs 56.0), and SST-2 and MRPC, where specialized single-task systems remain ahead.
CoLA, the Corpus of Linguistic Acceptability, asks whether a sentence is grammatical. It is imbalanced, so accuracy is misleading. It is scored with MCC, which uses all four cells of the confusion matrix:
| MCC | Meaning |
|---|---|
| +1 | Perfect prediction |
| 0 | No better than random |
| −1 | Exactly inverted prediction |
GPT-1's CoLA score of 45.4 means an MCC of about 0.454 on the \([-1, 1]\) scale.
Analysis · layers transferred
Where does the knowledge live?
If pretraining only produced good word vectors, transferring the embedding layer would capture almost all of the benefit. To test this, the authors transfer only the first \(n\) pretrained layers into the downstream model and train the rest from scratch:
The setup: the first cell is the embedding layer, the other 12 are the Transformer blocks.
On MultiNLI and RACE, transferring the embeddings helps, as expected, and each extra Transformer layer helps further, up to about 9% for full transfer on MultiNLI. There is no plateau where the upper layers stop mattering:
Redrawn from Figure 2 (left) of the paper. Each curve keeps rising with every layer transferred. The full-transfer scores and the ~9-point MultiNLI gain come from the paper; the in-between points are approximate, so read the shape rather than exact values.
WHAT TRAINS zero-shot behaviours
Abilities that appear in pretraining
Why does language-model pretraining help at all? The authors suggest that, to model language well, the model has to learn to do many of these tasks implicitly. To test this, they take the pretrained model with no fine-tuning, solve four tasks using simple hand-written heuristics, and track how performance changes as pretraining goes on.
On all four, zero-shot performance rises steadily during pretraining. The authors also ran the same test on an LSTM, which did worse and varied more. This suggests that the Transformer's inductive bias is part of why its representations transfer well.
Analysis · ablation studies
What actually matters? Take it apart
The ablation table removes one ingredient at a time and reports the average score across tasks. It is the most useful table in the paper, because it ranks the contributions:
Average score across tasks. Bars start at zero and the scale runs to 80.
Removing pretraining: 74.7 → 59.9
A drop of 14.8 points on average, by far the largest. Train the same Transformer directly on each task, with no pretraining, and it does much worse. The architecture alone does not explain the results; the pretraining does.
Swapping the Transformer for an LSTM: 74.7 → 69.1
A single-layer 2048-unit LSTM, trained in the same pretrain-then-fine-tune way, loses 5.6 points on average. The LSTM only wins on one dataset (MRPC). Self-attention clearly produces better transferable representations.
Removing the auxiliary LM objective: 74.7 → 75.0
The average actually goes up slightly. Looking at individual tasks, the auxiliary objective helps on the larger datasets (the NLI tasks and QQP) and hurts on the smaller ones. The authors' reading: it is useful when there is enough data, but it is not fundamental.
Conclusions
My Learnings
- Predicting the next token forces a model to learn syntax, meaning and some reasoning, because all of them help it predict.
- Pretraining helps the model learn capabilities even before task specific tuning
- Transformers transfer better than LSTMs. Attention's structured memory beats a single recurrent state, both after fine-tuning and zero-shot.
- Zero-shot abilities start appearing during pretraining. Abilities grow with pretraining alone. GPT-2 and GPT-3 later developed this into zero-shot and few-shot prompting.