Four boxes. One installs two libraries, one writes a file for you, and two of them start a conversation — first with the model you built, then with a much larger one.
These are what load a ready-made model. Everything else you already have.
cd ~/llm-from-scratch && source .venv/bin/activate && uv pip install transformers accelerate
Installs the two libraries that fetch and run a published model.
This box is long because the whole program is inside it. You do not have to read it — copy the box, paste it into Terminal and press Return, and it writes the file chat.py for you. If you do want to read it, it is right there.
cd ~/llm-from-scratch && cat > chat.py <<'PYEOF'
"""
Chat — the same loop, with a conversation wrapped around it.
Two models, one program:
python chat.py --ours the model you built in this course. It will not
chat with you, and the lesson is WHY.
python chat.py --template what a chat model is actually handed. Read this
one carefully; it is the whole trick.
python chat.py a real instruction-tuned model, running on your
own Mac, offline, for nothing.
Both models do the identical thing: look at everything so far, produce a
probability for every token that could come next, pick one, append it, repeat.
That is lesson eight, unchanged. Everything that makes the second one answer a
question sits AROUND that loop, never inside it.
Run: python chat.py
"""
import argparse
import sys
import torch
# The instruction-tuned model. Small enough to download in a minute and to run
# on a laptop's own chip; large enough to hold a conversation. Ungated -- no
# account, no key, no licence to accept.
CHAT_MODEL = 'Qwen/Qwen2.5-1.5B-Instruct'
SYSTEM = 'You are a helpful assistant. Answer briefly and plainly.'
# ---------- 1. The model you built ----------
def chat_with_ours():
"""Ask our own Shakespeare model a question and watch it not answer."""
import gpt as G # the file from lesson five
model = G.GPT().to(G.device)
model.load_state_dict(torch.load('model.pt', map_location=G.device))
model.eval()
n = sum(p.numel() for p in model.parameters())
print(f'Your model: {n/1e6:.2f}M parameters, {G.vocab_size} characters in its whole vocabulary.')
print('Ask it anything. Ctrl-C to stop.\n')
while True:
try:
q = input('you > ').strip()
except (EOFError, KeyboardInterrupt):
print(); return
if not q:
continue
# Every character of the question must exist in Shakespeare, or the
# model has no number for it. That alone is a hint at what is wrong.
ids = [G.stoi[c] for c in q if c in G.stoi]
if not ids:
print('model> (it has never seen any of those characters)\n'); continue
idx = torch.tensor([ids], dtype=torch.long, device=G.device)
out = model.generate(idx, 180, temperature=0.8)[0].tolist()
print('model> ' + G.decode(out[len(ids):]).replace('\n', '\n ') + '\n')
# ---------- 2. What a chat model is really handed ----------
def load(cls, **kw):
"""Load from the local cache if it is there, and only reach out if it is not.
Two reasons. A patron deserves to be told, once, that something large is
about to be downloaded. And after that first run this program never touches
the network at all -- which is the point worth making about running a model
on your own machine.
"""
try:
return cls.from_pretrained(CHAT_MODEL, local_files_only=True, **kw)
except Exception:
print(f'First run: downloading {CHAT_MODEL} (about 3 GB). This happens once.')
return cls.from_pretrained(CHAT_MODEL, **kw)
def show_template():
"""Print the raw text a chat model receives. This is the entire trick."""
from transformers import AutoTokenizer
tok = load(AutoTokenizer)
turns = [
{'role': 'system', 'content': SYSTEM},
{'role': 'user', 'content': 'What is a language model?'},
]
prompt = tok.apply_chat_template(turns, tokenize=False, add_generation_prompt=True)
print('This is the text the model actually sees:\n')
print('-' * 62)
print(prompt, end='')
print('-' * 62)
print(f'\n{len(tok(prompt).input_ids)} tokens. There is no conversation in there.')
print('There is a document, with special markers, that STOPS mid-sentence.')
print('The model does what it has always done: it finishes the document.')
# ---------- 3. A real chat model, on your own machine ----------
def chat():
from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer
device = ('mps' if torch.backends.mps.is_available()
else 'cuda' if torch.cuda.is_available() else 'cpu')
print(f'Loading {CHAT_MODEL} on {device} …')
tok = load(AutoTokenizer)
model = load(AutoModelForCausalLM, dtype=torch.float16).to(device).eval()
n = sum(p.numel() for p in model.parameters())
print(f'{n/1e9:.2f} billion parameters — {n/3_221_569:.0f} times the size of the one you built.')
print('Ask it anything. Ctrl-C to stop.\n')
# The conversation is just a list that keeps growing. Nothing is remembered
# inside the model between turns -- the whole history is re-sent every time.
turns = [{'role': 'system', 'content': SYSTEM}]
while True:
try:
q = input('you > ').strip()
except (EOFError, KeyboardInterrupt):
print(); return
if not q:
continue
turns.append({'role': 'user', 'content': q})
batch = tok.apply_chat_template(turns, add_generation_prompt=True,
return_dict=True, return_tensors='pt').to(device)
print('model> ', end='', flush=True)
streamer = TextStreamer(tok, skip_prompt=True, skip_special_tokens=True)
with torch.no_grad():
out = model.generate(**batch, max_new_tokens=160, do_sample=True,
temperature=0.7, top_p=0.9,
pad_token_id=tok.eos_token_id, streamer=streamer)
answer = tok.decode(out[0][batch['input_ids'].shape[1]:], skip_special_tokens=True)
turns.append({'role': 'assistant', 'content': answer})
print()
if __name__ == '__main__':
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument('--ours', action='store_true', help='chat with the model you built')
ap.add_argument('--template', action='store_true', help='print what a chat model is handed')
a = ap.parse_args()
if a.ours:
chat_with_ours()
elif a.template:
show_template()
else:
chat()
PYEOF
Writes chat.py into your project folder.
Type a question and press Return. It will not answer it. That is the lesson. Press Control and C together to stop.
cd ~/llm-from-scratch && source .venv/bin/activate && python chat.py --ours
Chats with your own 3.23 million parameter model.
The one box on this page worth reading the output of, slowly.
cd ~/llm-from-scratch && source .venv/bin/activate && python chat.py --template
Prints the exact text a chat model receives.
The first run downloads about three gigabytes and takes a few minutes. Every run after that is offline and instant. Control and C stops it.
cd ~/llm-from-scratch && source .venv/bin/activate && python chat.py
Runs a 1.5 billion parameter instruction-tuned model on your own Mac.