GPT-1 Paper Notes Transformers Pretraining Fine-tuning

The takeaway first

TL:DR;

In 2018, most NLP systems were built task by task. GPT-1 showed that you could train one Transformer on raw books to predict the next word, then reuse the whole network for almost any language task with only small changes.

The paper calls its approach semi-supervised: an unsupervised pretraining stage on unlabeled text, followed by a supervised fine-tuning stage on each labeled task. Today we would just say "pretrain, then fine-tune". GPT-1 is where that recipe was first shown to work this well with a Transformer.

Stage 1 · unsupervised

Generative pretraining

A 12-layer Transformer decoder learns to predict the next token over 7,000+ unpublished books. No labels at all.

Stage 2 · supervised

Discriminative fine-tuning

The same network, plus one small linear layer, is trained briefly on each labeled task: entailment, QA, similarity, classification.

Don't learn only word embeddings.
Learn an entire reusable language model.

The results backed this up. GPT-1 beat the previous state of the art on 9 of the 12 datasets it was tested on, often beating models designed specifically for that task. In the ablation study, removing pretraining caused by far the largest drop in performance.

Not a new idea, a scaled-up one Three years earlier, Dai & Le (2015) pretrained an LSTM on unlabeled text as a language model, transferred the whole model (embeddings and recurrent weights), and fine-tuned it on a supervised task. Pretrained LSTMs beat randomly initialized ones, and extra unlabeled data could stand in for labeled data. GPT-1 is a direct descendant: the same pretrain → transfer → fine-tune recipe, with a Transformer and far more data.

Motivation

Labeled data is scarce, raw text is not

Supervised deep learning needs lots of labeled examples, and good labels are expensive. Someone has to read each premise and hypothesis pair and decide whether one implies the other, or read a passage and write the correct answer. Most NLP datasets of that time had thousands or tens of thousands of examples.

Unlabeled text, on the other hand, is nearly free and practically unlimited. The question the paper asks is simple:

The question Can a model learn something general about language from raw text alone, and can that knowledge then be moved into a model that solves labeled tasks with far less data?

The authors name two open problems that had held this back:

  • Which objective to pretrain on? Language modeling, machine translation and discourse coherence had all been tried, and it was unclear which one transferred best.
  • How to transfer what was learned? Earlier methods often needed task-specific architecture changes, complex learning schemes, or extra auxiliary objectives for each new task.

GPT-1's answer to both is short: use language modeling as the objective, use a Transformer as the model, and transfer the whole network by turning every task into a sequence of tokens.

Why a Transformer, not an LSTM?

Recurrent networks squeeze everything they have read into one hidden state that is passed along word by word, so information from far back fades. Self-attention lets every token look directly at every earlier token. The authors argue that this gives a more structured memory for long-range dependencies, and that this is what makes the representations transfer well. The ablations in section 10 test that claim directly.

Related work

What transfer learning looked like before

Reusing knowledge from unlabeled text was not new. What was new was how much got reused. It helps to see what earlier approaches transferred.

Word-level embeddings

Word2Vec and GloVe learn one fixed vector per word:

$$e_{\text{bank}} \in \mathbb{R}^{d}$$

That vector is the same whether the sentence is about a river bank or a savings bank. Only the first layer of the downstream model gets pretrained knowledge. Everything above it still learns from scratch on the small labeled set.

Phrase- and sentence-level representations

Other work learned to compress a whole sequence into a single vector, usually with averaging, CNNs, RNNs or LSTMs:

$$s = f(x_1, x_2, \ldots, x_n)$$

This captures more than single words, but the downstream task gets one summary vector and nothing else.

GPT-1's difference It does not produce a fixed word or sentence embedding at all. Every token gets a contextual hidden state at every layer, and higher layers mix in more syntax, meaning and long-range context. The whole stack is transferred, not just its bottom layer.

THE GPT Framework · Stage 1

Generative pretraining

Given an unlabeled corpus of tokens \(\mathcal{U} = \{u_1, \ldots, u_n\}\), the model maximizes the standard left-to-right language-modeling likelihood:

$$L_1(\mathcal{U}) = \sum_i \log P(u_i \mid u_{i-k}, \ldots, u_{i-1}; \Theta)$$
\(k\)
the context window, 512 tokens in GPT-1
\(\Theta\)
all of the network's parameters
\(P\)
the probability the model assigns to the true next token

Simply said, given the previous tokens, make the true next token as probable as possible. Averaged over billions of predictions, this one objective forces the model to pick up grammar, facts, sentiment, coreference and some reasoning, because all of these help it guess what comes next.

Forward Pass

The model is a decoder-only Transformer: masked multi-head self-attention followed by a position-wise feed-forward network, stacked 12 times. The paper writes the whole forward pass in three lines:

softmax → P(next token) h₁₂ · We T Transformer block 12 ⋮ Transformer block 2 Transformer block 1 h₀ = U · We + W p The cat sat on U: token ids (up to 512) same W_e (tied)

Read bottom to top. Token ids become vectors, pass through 12 masked self-attention blocks, and the final hidden state is projected back to the vocabulary with the transpose of the same embedding matrix.

\(U\)
the context of token ids, one-hot encoded
\(W_e\)
the learned token-embedding matrix
\(W_p\)
the learned position-embedding matrix (not the sine waves of the original Transformer)
\(n\)
the number of layers, 12
Weight tying The matrix that maps tokens into vectors is reused, transposed, to map vectors back into tokens. This saves parameters, and it means "the next token should be cat" is literally "the hidden state should point in the same direction as cat's embedding".

The GPT Framework · Stage 2

Supervised fine-tuning

Fine-tuning reuses the whole pretrained network and adds just one linear layer, \(W_y\). Feed in a labeled example \(x^1, \ldots, x^m\), take the top block's hidden state at the last token, \(h_l^m\), and predict the label from it:

$$P(y \mid x^1, \ldots, x^m) = \operatorname{softmax}(h_l^m W_y)$$

Every task becomes a sequence of tokens

Classification fits this directly, but entailment has two sentences and multiple-choice QA has a passage, a question and several answers. Instead of building a custom architecture for each, GPT-1 changes the input: it joins the pieces into one sequence with a start token, a delim token between segments, and an extract token at the end, whose hidden state goes to \(W_y\).

Classification sentiment, acceptability
starttextextract→ Transformer → Linear
Entailment premise → hypothesis
startpremisedelimhypothesisextract→ Transformer → Linear
Similarity no natural order, so use both
starttext 1delimtext 2extract→ Transformer ↘
starttext 2delimtext 1extract→ Transformer ↗ add → Linear
Multiple choice QA, commonsense reasoning
startcontextquestiondelimanswer 1extract→ Transformer → Linear ↘
startcontextquestiondelimanswer Nextract→ Transformer → Linear ↗ softmax

For similarity, both orderings run separately and their hidden states are added. For multiple choice, each answer gets its own sequence and a softmax picks one.

Change the task to fit the model, not the model to fit the task

The training objective

The task loss \(L_2\) is the log-likelihood of the correct labels. The paper also keeps the language-modeling loss \(L_1\) running on the task's own text, at half weight:

$$L_3(\mathcal{C}) = \underbrace{L_2(\mathcal{C})}_{\text{task}} + \lambda \cdot \underbrace{L_1(\mathcal{C})}_{\text{next token}}, \qquad \lambda = 0.5$$

This auxiliary term acts as a regularizer and speeds up convergence: the model specializes while it keeps practising language. Training is short and gentle: learning rate \(6.25 \times 10^{-5}\), batch size 32, usually just 3 epochs. The only new parameters are \(W_y\) and the embeddings for the special tokens.

Experiments · setup

The recipe: data, cleaning and hyperparameters

BookCorpus

The pretraining corpus is BookCorpus, over 7,000 unique unpublished books in genres such as adventure, fantasy and romance. (The full corpus as originally described has around 11,000 books, 74 million sentences and roughly a billion words.)

Why books? Books contain long, coherent stretches of text: dialogue, recurring characters, and plot that depends on something said chapters earlier. The alternative, the 1B Word Benchmark used by ELMo, is similar in size but shuffled at the sentence level, which destroys exactly the long-range structure a 512-token Transformer can learn from.

Cleaning with ftfy

Raw text from the web and from scanned books is full of mojibake, text that was decoded with the wrong character encoding. ftfy ("fixes text for you") repairs it before anything else happens:

ftfy.fix_text
François   →   François

The full preprocessing order: raw text → ftfy cleanup → punctuation and whitespace standardization → spaCy tokenizer → byte-pair encoding with 40,000 merges → language-model training.

Model and training

Component GPT-1
Architecture 12-layer decoder-only Transformer
Hidden size · heads 768 · 12 heads (64 dims each)
Feed-forward size 3,072
Context length 512 tokens
Vocabulary BPE, 40,000 merges
Activation GELU
Position encoding Learned (not sinusoidal)
Optimizer · peak LR Adam · 2.5 × 10⁻⁴
LR schedule Linear warm-up for 2,000 steps, then cosine decay to 0
Batch · epochs 64 sequences × 512 tokens · 100 epochs
Regularization Dropout 0.1, modified L2 (w = 0.01)
Initialization N(0, 0.02)
Fine-tuning LR 6.25 × 10⁻⁵, batch 32, 3 epochs, λ = 0.5

The pretrained model reached a very low token-level perplexity of 18.4 on BookCorpus. These numbers seem small today, but they became the template: GPT-2 and GPT-3 kept this shape and mostly scaled it up.

Experiments · downstream tasks

Twelve benchmarks, and what they measure

GPT-1 was fine-tuned on four kinds of task: natural language inference, question answering and commonsense reasoning, semantic similarity, and text classification. A few of these needed explaining before the results made sense to me.

Textual entailment (natural language inference)

Given a premise, does a hypothesis follow from it, contradict it, or neither?

Premise Hypothesis Label
A man is playing a guitar. A person is playing an instrument. Entailment
A woman is running in a park. Nobody is in the park. Contradiction
A woman is running in a park. She runs every morning. Neutral

It is hard because it needs lexical knowledge (a guitar is an instrument), coreference, and some reasoning. Two datasets come up again and again:

  • SNLI (Stanford Natural Language Inference): sentence pairs written about image captions.
  • MultiNLI (Multi-Genre NLI): the same task across many genres, such as fiction, government reports, speech transcripts and letters. "Matched" tests on genres seen in training, "mismatched" on unseen ones.

The benchmarks, decoded

Each benchmark is a public dataset with a fixed test set, so different models can be compared on the same questions. Here is what every abbreviation stands for and what it asks the model to do:

Name Stands for What the model must do
Natural language inference
MNLI Multi-Genre Natural Language Inference Decide if a hypothesis follows from, contradicts, or is unrelated to a premise, across many genres. "-m" is matched genres, "-mm" mismatched.
SNLI Stanford Natural Language Inference The same task on sentences describing photos.
SciTail Science entailment dataset Decide if a statement follows from a sentence, using school science exam questions.
QNLI Question Natural Language Inference Decide if a Wikipedia sentence contains the answer to a question.
RTE Recognizing Textual Entailment Decide if one passage implies another. Very small: about 2,500 training examples.
Question answering & commonsense reasoning
Story Cloze Story Cloze Test ("cloze" means fill in the blank) Read a four-sentence story and pick the right ending out of two.
RACE ReAding Comprehension from Examinations Answer multiple-choice questions on passages from English exams for Chinese students. "-m" is middle school, "-h" high school.
Classification & semantic similarity
CoLA Corpus of Linguistic Acceptability Decide if a sentence is grammatical.
SST-2 Stanford Sentiment Treebank, 2 classes Decide if a movie-review sentence is positive or negative.
MRPC Microsoft Research Paraphrase Corpus Decide if two news sentences mean the same thing.
STS-B Semantic Textual Similarity Benchmark Score how similar in meaning two sentences are, from 0 to 5.
QQP Quora Question Pairs Decide if two questions asked on Quora are duplicates.
Overall
GLUE General Language Understanding Evaluation Not a single task: a bundle of nine (including CoLA, SST-2, MRPC, STS-B, QQP, MNLI, QNLI and RTE), averaged into one overall score.

Results against the previous best

Two datasets from each category, with every method the paper compares against on them (Tables 2–4). GPT-1 is in red, the strongest earlier result in black, and other earlier methods in grey. The badge is GPT-1's gap to the best earlier result.

GPT-1 best earlier result other earlier methods

Scores as reported in the paper. Bars start at zero, and every panel uses a 0–100 scale. "(5x)" and "(9x)" mark ensembles of 5 or 9 models. CoLA is Matthews correlation × 100; GLUE is the benchmark's overall score.

GPT-1 is ahead on 9 of the 12 datasets, often against ensembles, while being a single model. The biggest gains are where long-range context matters most: +8.9 on Story Cloze (choosing the right ending to a story) and +5.7 on RACE (reading comprehension from school exams). CoLA jumps from 35.0 to 45.4, and the GLUE overall score from 68.9 to 72.8.

The three misses are RTE, a small entailment set where a multi-task biLSTM wins (61.7 vs 56.0), and SST-2 and MRPC, where specialized single-task systems remain ahead.

Side note · Matthews correlation coefficient (MCC)

CoLA, the Corpus of Linguistic Acceptability, asks whether a sentence is grammatical. It is imbalanced, so accuracy is misleading. It is scored with MCC, which uses all four cells of the confusion matrix:

$$\text{MCC} = \frac{TP \cdot TN - FP \cdot FN}{\sqrt{(TP+FP)(TP+FN)(TN+FP)(TN+FN)}}$$
MCC Meaning
+1 Perfect prediction
0 No better than random
−1 Exactly inverted prediction

GPT-1's CoLA score of 45.4 means an MCC of about 0.454 on the \([-1, 1]\) scale.

Analysis · layers transferred

Where does the knowledge live?

If pretraining only produced good word vectors, transferring the embedding layer would capture almost all of the benefit. To test this, the authors transfer only the first \(n\) pretrained layers into the downstream model and train the rest from scratch:

The setup: the first cell is the embedding layer, the other 12 are the Transformer blocks.

On MultiNLI and RACE, transferring the embeddings helps, as expected, and each extra Transformer layer helps further, up to about 9% for full transfer on MultiNLI. There is no plateau where the upper layers stop mattering:

405060708090 024681012 pretrained layers transferred dev accuracy (%) ≈ +9 MultiNLI RACE

Redrawn from Figure 2 (left) of the paper. Each curve keeps rising with every layer transferred. The full-transfer scores and the ~9-point MultiNLI gain come from the paper; the in-between points are approximate, so read the shape rather than exact values.

Interpretation The transferable knowledge is not confined to the input embeddings. Every layer of the pretrained stack holds something useful for the target task, which is exactly why transferring the whole network beat transferring word vectors.

WHAT TRAINS zero-shot behaviours

Abilities that appear in pretraining

Why does language-model pretraining help at all? The authors suggest that, to model language well, the model has to learn to do many of these tasks implicitly. To test this, they take the pretrained model with no fine-tuning, solve four tasks using simple hand-written heuristics, and track how performance changes as pretraining goes on.

On all four, zero-shot performance rises steadily during pretraining. The authors also ran the same test on an LSTM, which did worse and varied more. This suggests that the Transformer's inductive bias is part of why its representations transfer well.

$$\text{More LM pretraining} \;\Rightarrow\; \text{Better zero-shot task behaviour}$$

Analysis · ablation studies

What actually matters? Take it apart

The ablation table removes one ingredient at a time and reports the average score across tasks. It is the most useful table in the paper, because it ranks the contributions:

Average score across tasks. Bars start at zero and the scale runs to 80.

Removing pretraining: 74.7 → 59.9

A drop of 14.8 points on average, by far the largest. Train the same Transformer directly on each task, with no pretraining, and it does much worse. The architecture alone does not explain the results; the pretraining does.

Swapping the Transformer for an LSTM: 74.7 → 69.1

A single-layer 2048-unit LSTM, trained in the same pretrain-then-fine-tune way, loses 5.6 points on average. The LSTM only wins on one dataset (MRPC). Self-attention clearly produces better transferable representations.

Removing the auxiliary LM objective: 74.7 → 75.0

The average actually goes up slightly. Looking at individual tasks, the auxiliary objective helps on the larger datasets (the NLI tasks and QQP) and hurts on the smaller ones. The authors' reading: it is useful when there is enough data, but it is not fundamental.

Pretraining ≫ architecture ≫ auxiliary objective

Conclusions

My Learnings

  1. Predicting the next token forces a model to learn syntax, meaning and some reasoning, because all of them help it predict.
  2. Pretraining helps the model learn capabilities even before task specific tuning
  3. Transformers transfer better than LSTMs. Attention's structured memory beats a single recurrent state, both after fine-tuning and zero-shot.
  4. Zero-shot abilities start appearing during pretraining. Abilities grow with pretraining alone. GPT-2 and GPT-3 later developed this into zero-shot and few-shot prompting.
The key lesson Learning to model language itself produces a general-purpose representation that transfers to many different NLP problems. GPT-2, GPT-3 and the models after them are, at heart, this paper with far more data, parameters and compute.

Related reading

Attention, from scratch

The masked self-attention inside every GPT-1 block, worked through by hand and in PyTorch.