Four commands. You download the text the model will learn from, then look at it properly before training on it.
About one megabyte of Shakespeare's plays. It arrives as a plain text file called input.txt.
cd ~/llm-from-scratch && curl -O https://raw.githubusercontent.com/karpathy/char-rnn/master/data/tinyshakespeare/input.txt
Fetches the file into your project folder.
Never train on data you have not read. Always open it first.
cd ~/llm-from-scratch && head -20 input.txt
Prints the first twenty lines.
Two numbers worth knowing before you go further.
cd ~/llm-from-scratch && wc -l input.txt && wc -c input.txt
Counts the lines, then the characters.
This is the number that decides the size of everything the model learns.
cd ~/llm-from-scratch && source .venv/bin/activate && python -c "
t = open('input.txt', encoding='utf-8').read()
chars = sorted(set(t))
print('characters in the file:', len(t))
print('different characters :', len(chars))
print('they are:', repr(''.join(chars)))
"
Prints how many characters there are, how many are different, and exactly which ones.