Three commands. You give every character a number, turn a sentence into numbers and back, then set some text aside to test with.
Sixty-five characters, sixty-five numbers. That pairing is the whole tokenizer.
cd ~/llm-from-scratch && source .venv/bin/activate && python -c "
t = open('input.txt', encoding='utf-8').read()
chars = sorted(set(t))
stoi = {c: i for i, c in enumerate(chars)}
itos = {i: c for c, i in stoi.items()}
for c in ['\n', ' ', 'A', 'a', 'z', '?']:
print(repr(c), '->', stoi[c])
print()
for i in [0, 1, 13, 39, 64]:
print(i, '->', repr(itos[i]))
\"
Prints the number for a few characters, and the character for a few numbers.
If it comes back the same, nothing was lost.
cd ~/llm-from-scratch && source .venv/bin/activate && python -c "
t = open('input.txt', encoding='utf-8').read()
chars = sorted(set(t)); stoi = {c:i for i,c in enumerate(chars)}; itos = {i:c for c,i in stoi.items()}
encode = lambda s: [stoi[c] for c in s]
decode = lambda ids: ''.join(itos[i] for i in ids)
msg = 'To be, or not to be'
print('text :', msg)
print('encoded:', encode(msg))
print('decoded:', decode(encode(msg)))
\"
Turns a line of Shakespeare into numbers and back again.
Ninety per cent to learn from, ten per cent kept hidden to test against.
cd ~/llm-from-scratch && source .venv/bin/activate && python -c "
import torch
t = open('input.txt', encoding='utf-8').read()
chars = sorted(set(t)); stoi = {c:i for i,c in enumerate(chars)}
data = torch.tensor([stoi[c] for c in t], dtype=torch.long)
split = int(0.9 * len(data))
train, val = data[:split], data[split:]
print('all data :', tuple(data.shape), data.dtype)
print('training :', tuple(train.shape))
print('held back:', tuple(val.shape))
print()
print('first 20 numbers:', train[:20].tolist())
\"
Splits the data and prints the sizes.