Two commands, and a list of what to change. Nothing here is required — you have already built the thing. This is the road onward.
Five numbers in the settings, and the model is three times the size. Everything else in gpt.py stays exactly as it is.
cd ~/llm-from-scratch && source .venv/bin/activate && python -c "
import torch, torch.nn as nn
def size(n_layer, n_head, n_embd, block, vocab=65):
per_block = 4*n_embd*n_embd + 4*n_embd + 8*n_embd*n_embd + 5*n_embd
return vocab*n_embd + block*n_embd + n_layer*per_block + 2*n_embd + n_embd*vocab + vocab
print('what you built : layers 4, heads 4, width 256, context 128')
print(' parameters :', f'{size(4,4,256,128):,}')
print()
print('the scaled version: layers 6, heads 6, width 384, context 256')
print(' parameters :', f'{size(6,6,384,256):,}')
print()
print(' that is', round(size(6,6,384,256)/size(4,4,256,128),1), 'times bigger')
"
Prints the size of what you built next to the size of the scaled version.
Ours reads one character at a time. GPT-2 reads chunks of words. Same sentence, far fewer steps.
cd ~/llm-from-scratch && source .venv/bin/activate && uv pip install -q tiktoken && python -c "
import tiktoken
enc = tiktoken.get_encoding('gpt2')
msg = 'To be, or not to be, that is the question'
ids = enc.encode(msg)
print('sentence :', msg)
print()
print('our way :', len(msg), 'characters, one number each')
print('GPT-2 way:', len(ids), 'tokens ->', ids[:12], '...')
print()
print('the pieces:', [enc.decode([i]) for i in ids[:10]])
"
Installs tiktoken and encodes one sentence both ways.
Open gpt.py and change the settings at the top. Expect a validation loss near 1.5, and tens of minutes rather than three.
n_layer = 6
Six blocks instead of four.
n_head = 6
Six attention heads instead of four.
n_embd = 384
A wider vector for each character.
block_size = 256
Twice as much context to look back over.
max_iters = 5000
More steps, because there is more to learn.