WebSpinner — a spinning top above the WebSpinner wordmark
Academy

Lesson 9 — Talking To It

Four boxes. One installs two libraries, one writes a file for you, and two of them start a conversation — first with the model you built, then with a much larger one.

First, open Terminal

  1. Hold ⌘ and press Space. A search box appears in the middle of the screen.
  2. Type Terminal and press Return. A window with plain text opens. That is Terminal.
  3. Click Copy beside a command below, click into the Terminal window, then press ⌘V to paste.
  4. Press Return to run it. Wait until the text stops moving before you do the next one.

1Two more libraries

These are what load a ready-made model. Everything else you already have.

cd ~/llm-from-scratch && source .venv/bin/activate && uv pip install transformers accelerate

Installs the two libraries that fetch and run a published model.

2Write the chat program

This box is long because the whole program is inside it. You do not have to read it — copy the box, paste it into Terminal and press Return, and it writes the file chat.py for you. If you do want to read it, it is right there.

cd ~/llm-from-scratch && cat > chat.py <<'PYEOF'
"""
Chat — the same loop, with a conversation wrapped around it.

Two models, one program:

  python chat.py --ours       the model you built in this course. It will not
                              chat with you, and the lesson is WHY.
  python chat.py --template   what a chat model is actually handed. Read this
                              one carefully; it is the whole trick.
  python chat.py              a real instruction-tuned model, running on your
                              own Mac, offline, for nothing.

Both models do the identical thing: look at everything so far, produce a
probability for every token that could come next, pick one, append it, repeat.
That is lesson eight, unchanged. Everything that makes the second one answer a
question sits AROUND that loop, never inside it.

Run:  python chat.py
"""

import argparse
import sys

import torch

# The instruction-tuned model. Small enough to download in a minute and to run
# on a laptop's own chip; large enough to hold a conversation. Ungated -- no
# account, no key, no licence to accept.
CHAT_MODEL = 'Qwen/Qwen2.5-1.5B-Instruct'

SYSTEM = 'You are a helpful assistant. Answer briefly and plainly.'


# ---------- 1. The model you built ----------
def chat_with_ours():
    """Ask our own Shakespeare model a question and watch it not answer."""
    import gpt as G                       # the file from lesson five

    model = G.GPT().to(G.device)
    model.load_state_dict(torch.load('model.pt', map_location=G.device))
    model.eval()
    n = sum(p.numel() for p in model.parameters())
    print(f'Your model: {n/1e6:.2f}M parameters, {G.vocab_size} characters in its whole vocabulary.')
    print('Ask it anything. Ctrl-C to stop.\n')

    while True:
        try:
            q = input('you  > ').strip()
        except (EOFError, KeyboardInterrupt):
            print(); return
        if not q:
            continue
        # Every character of the question must exist in Shakespeare, or the
        # model has no number for it. That alone is a hint at what is wrong.
        ids = [G.stoi[c] for c in q if c in G.stoi]
        if not ids:
            print('model> (it has never seen any of those characters)\n'); continue
        idx = torch.tensor([ids], dtype=torch.long, device=G.device)
        out = model.generate(idx, 180, temperature=0.8)[0].tolist()
        print('model> ' + G.decode(out[len(ids):]).replace('\n', '\n       ') + '\n')


# ---------- 2. What a chat model is really handed ----------
def load(cls, **kw):
    """Load from the local cache if it is there, and only reach out if it is not.

    Two reasons. A patron deserves to be told, once, that something large is
    about to be downloaded. And after that first run this program never touches
    the network at all -- which is the point worth making about running a model
    on your own machine.
    """
    try:
        return cls.from_pretrained(CHAT_MODEL, local_files_only=True, **kw)
    except Exception:
        print(f'First run: downloading {CHAT_MODEL} (about 3 GB). This happens once.')
        return cls.from_pretrained(CHAT_MODEL, **kw)


def show_template():
    """Print the raw text a chat model receives. This is the entire trick."""
    from transformers import AutoTokenizer

    tok = load(AutoTokenizer)
    turns = [
        {'role': 'system', 'content': SYSTEM},
        {'role': 'user', 'content': 'What is a language model?'},
    ]
    prompt = tok.apply_chat_template(turns, tokenize=False, add_generation_prompt=True)
    print('This is the text the model actually sees:\n')
    print('-' * 62)
    print(prompt, end='')
    print('-' * 62)
    print(f'\n{len(tok(prompt).input_ids)} tokens. There is no conversation in there.')
    print('There is a document, with special markers, that STOPS mid-sentence.')
    print('The model does what it has always done: it finishes the document.')


# ---------- 3. A real chat model, on your own machine ----------
def chat():
    from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer

    device = ('mps' if torch.backends.mps.is_available()
              else 'cuda' if torch.cuda.is_available() else 'cpu')
    print(f'Loading {CHAT_MODEL} on {device} …')
    tok = load(AutoTokenizer)
    model = load(AutoModelForCausalLM, dtype=torch.float16).to(device).eval()
    n = sum(p.numel() for p in model.parameters())
    print(f'{n/1e9:.2f} billion parameters — {n/3_221_569:.0f} times the size of the one you built.')
    print('Ask it anything. Ctrl-C to stop.\n')

    # The conversation is just a list that keeps growing. Nothing is remembered
    # inside the model between turns -- the whole history is re-sent every time.
    turns = [{'role': 'system', 'content': SYSTEM}]
    while True:
        try:
            q = input('you  > ').strip()
        except (EOFError, KeyboardInterrupt):
            print(); return
        if not q:
            continue
        turns.append({'role': 'user', 'content': q})
        batch = tok.apply_chat_template(turns, add_generation_prompt=True,
                                        return_dict=True, return_tensors='pt').to(device)
        print('model> ', end='', flush=True)
        streamer = TextStreamer(tok, skip_prompt=True, skip_special_tokens=True)
        with torch.no_grad():
            out = model.generate(**batch, max_new_tokens=160, do_sample=True,
                                 temperature=0.7, top_p=0.9,
                                 pad_token_id=tok.eos_token_id, streamer=streamer)
        answer = tok.decode(out[0][batch['input_ids'].shape[1]:], skip_special_tokens=True)
        turns.append({'role': 'assistant', 'content': answer})
        print()


if __name__ == '__main__':
    ap = argparse.ArgumentParser(description=__doc__)
    ap.add_argument('--ours', action='store_true', help='chat with the model you built')
    ap.add_argument('--template', action='store_true', help='print what a chat model is handed')
    a = ap.parse_args()
    if a.ours:
        chat_with_ours()
    elif a.template:
        show_template()
    else:
        chat()
PYEOF

Writes chat.py into your project folder.

3Ask the model you built

Type a question and press Return. It will not answer it. That is the lesson. Press Control and C together to stop.

cd ~/llm-from-scratch && source .venv/bin/activate && python chat.py --ours

Chats with your own 3.23 million parameter model.

4See what a chat model is really handed

The one box on this page worth reading the output of, slowly.

cd ~/llm-from-scratch && source .venv/bin/activate && python chat.py --template

Prints the exact text a chat model receives.

5Talk to a real one

The first run downloads about three gigabytes and takes a few minutes. Every run after that is offline and instant. Control and C stops it.

cd ~/llm-from-scratch && source .venv/bin/activate && python chat.py

Runs a 1.5 billion parameter instruction-tuned model on your own Mac.