WebSpinner — a spinning top above the WebSpinner wordmark
Academy

Lesson 3 — Characters Into Numbers

Three commands. You give every character a number, turn a sentence into numbers and back, then set some text aside to test with.

First, open Terminal

  1. Hold ⌘ and press Space. A search box appears in the middle of the screen.
  2. Type Terminal and press Return. A window with plain text opens. That is Terminal.
  3. Click Copy beside a command below, click into the Terminal window, then press ⌘V to paste.
  4. Press Return to run it. Wait until the text stops moving before you do the next one.

1Give every character a number

Sixty-five characters, sixty-five numbers. That pairing is the whole tokenizer.

cd ~/llm-from-scratch && source .venv/bin/activate && python -c "
t = open('input.txt', encoding='utf-8').read()
chars = sorted(set(t))
stoi = {c: i for i, c in enumerate(chars)}
itos = {i: c for c, i in stoi.items()}
for c in ['\n', ' ', 'A', 'a', 'z', '?']:
    print(repr(c), '->', stoi[c])
print()
for i in [0, 1, 13, 39, 64]:
    print(i, '->', repr(itos[i]))
\"

Prints the number for a few characters, and the character for a few numbers.

2Encode a sentence, then decode it

If it comes back the same, nothing was lost.

cd ~/llm-from-scratch && source .venv/bin/activate && python -c "
t = open('input.txt', encoding='utf-8').read()
chars = sorted(set(t)); stoi = {c:i for i,c in enumerate(chars)}; itos = {i:c for c,i in stoi.items()}
encode = lambda s: [stoi[c] for c in s]
decode = lambda ids: ''.join(itos[i] for i in ids)
msg = 'To be, or not to be'
print('text   :', msg)
print('encoded:', encode(msg))
print('decoded:', decode(encode(msg)))
\"

Turns a line of Shakespeare into numbers and back again.

3Hold some text back

Ninety per cent to learn from, ten per cent kept hidden to test against.

cd ~/llm-from-scratch && source .venv/bin/activate && python -c "
import torch
t = open('input.txt', encoding='utf-8').read()
chars = sorted(set(t)); stoi = {c:i for i,c in enumerate(chars)}
data = torch.tensor([stoi[c] for c in t], dtype=torch.long)
split = int(0.9 * len(data))
train, val = data[:split], data[split:]
print('all data :', tuple(data.shape), data.dtype)
print('training :', tuple(train.shape))
print('held back:', tuple(val.shape))
print()
print('first 20 numbers:', train[:20].tolist())
\"

Splits the data and prints the sizes.