Three commands. You build the model, look inside it, and push one batch of real text through it before it has learned anything.
Three and a quarter million numbers, and almost all of them are in the four blocks.
cd ~/llm-from-scratch && source .venv/bin/activate && python -c "
import torch, gpt
m = gpt.GPT()
total = sum(p.numel() for p in m.parameters())
print('total parameters:', f'{total:,}')
print()
for name, mod in [('token embedding', m.tok_emb), ('position embedding', m.pos_emb),
('one block', m.blocks[0]), ('output head', m.head)]:
n = sum(p.numel() for p in mod.parameters())
print(f'{name:20} {n:>10,}')
"
Builds the model and prints how many parameters each part holds.
The same block, four times. That repetition is the whole architecture.
cd ~/llm-from-scratch && source .venv/bin/activate && python -c "
import torch, gpt
m = gpt.GPT()
print(m.blocks)
"
Prints the four stacked blocks and everything in them.
An untrained model should be no better than guessing. Check that it is not.
cd ~/llm-from-scratch && source .venv/bin/activate && python -c "
import torch, gpt
m = gpt.GPT().to(gpt.device)
x, y = gpt.get_batch('train')
logits, loss = m(x, y)
print('input :', tuple(x.shape))
print('output:', tuple(logits.shape), ' <- a score for each of 65 characters')
print('loss :', round(loss.item(), 3))
print()
import math
print('random guessing would be about', round(math.log(65), 3))
"
Pushes one batch through and prints the loss, next to what pure guessing would score.