Session 11. State and memory — Mon 28 Sep

Follow along: today's class

Keep this page open during the session. Every step says what to open and what to run, in the order we run it in class.

Part What Time
0 Before we start: pull, sync, preflight green 5 min
A Ask the course: connect the course MCP, three questions, one empty answer 5 min
1 Where we are going: week 3, and a skill is a memory 5 min
2 Four kinds of memory, and four questions 10 min
3 The state wrapper, live, then a second user 25 min
D Demo 11: memory at scale, 20,000 memories, one index, one owner filter 10 min
4 Your exercise: a preference, a policy, and a store with owners 35 min
5 Break it on purpose: memory that was true once 10 min
6 Check, and read two refusal lists 10 min
7 Hand it in, and the exit ticket 5 min

0. Before we start

From your course folder:

git stash push -m "my work before today's update"
git pull
uv sync --extra projects --extra agents
uv run jupyter lab

Why the stash first. git pull stops when you have typed into a notebook it wants to update, and it says "your local changes would be overwritten". git stash puts your changes on a shelf and deletes nothing. To take one notebook back as you left it, see when git pull stops.

Name both extras, or you lose ChromaDB. uv sync makes the environment match exactly what you asked for. Ask for agents on its own and it uninstalls the projects extra, ChromaDB goes with it, and project 02's index stops building. Today's notebook needs neither.

Today needs no model and no key. Every cell runs on FakeLLM, in memory, with no database and no network.

A check that says it never ran means you did not save. bootcamp check reads the file on disk, not the kernel in your browser.

Open one file:

Run the first cell, the preflight, and wait until it is green. It never raises: if something is missing, it prints the command that fixes it.

No setup on your laptop? Open it in Colab. The first cell fetches the course for you.

Ask the course

A tool for the whole lab, and for every session after it. One URL gives your assistant the published course, so it answers from the pages instead of from memory. No key, no login, nothing to install.

Connect it. In Claude Code, run this once, in any folder:

claude mcp add --transport http dev3pack-course https://mcp.geckovision.tech/course/mcp

Then start claude and type /mcp. dev3pack-course is in the list. On Claude web or desktop: Settings, Connectors, Add custom connector, and paste the same URL. Any other app, and what to do with no MCP at all: search the course from your assistant.

Ask it three questions about today. Type each one to your assistant and ask it to use the course tools:

  1. "Why is the owner part of the key in a memory store?"
  2. "What does ch11-e2 check in my storage policy?"
  3. "When does a memory stop being true?"

Each answer should come from today's pages, for example units/en/unit3/session-11-state-and-memory/concepts-3 for the first. A hit looks like this, shortened:

{
  "page_id": "units/en/unit3/session-11-state-and-memory/concepts-3",
  "title": "Isolation, and memory that went stale [[isolation-and-staleness]]",
  "heading": "...",
  "text": "... are two memories that cannot collide, and there is no lookup anywhere in the class that can reach a row without naming its owner. ..."
}

What to do with page_id. It is the page's path in your clone, without .mdx. Open units/en/unit3/session-11-state-and-memory/concepts-3.mdx and check the answer against it. If your assistant did not name a page, ask it which one. An answer with no page is an answer you cannot check.

Now ask it something the course does not teach, in your own words: a recipe, a football score, anything off-topic. There is no example printed here on purpose. The moment a page quotes a question, the course contains that question, and searching for it finds the page that quoted it.

The search comes back with no hits and this note:

no passage in these pages matches; the course may not cover this yet

An empty result is an answer. It means the course does not cover it. Your assistant should say so, not fill the gap with a paragraph from its own memory. If it guesses anyway, tell it: "if the course search is empty, say the course does not cover it." In your course folder, the ask-the-course skill already says that.

What it serves. Only what is published to students: the lesson pages, projects and guides. No quizzes, no solutions, no instructor material.

1. Where we are going

Week 3 starts today. Last week your assistant learned to answer. This week it learns to hold things, serve things and ship. Today is the first of those: what it holds, and for how long.

You have You can say You cannot yet say
a dict that holds a preference it remembers whether the model ever saw it
a preference in the prompt, a cap of 5 it remembers, and forgets on schedule whose memory it is
a store keyed on the owner ana gets ana's, bruno gets bruno's nothing; that is the claim

Start with something you already built. Open the skill you wrote on Friday. It is a stored pattern your assistant applies without being told again. That is memory: procedural memory. It can go stale the same way a preference can. If the API it describes changes tomorrow, the skill keeps running, and it is now wrong.

2. Four kinds and four questions

Kind Holds Example Lives
Conversation state this exchange "answer briefly" one session
Semantic memory facts "this user reads Portuguese docs" until corrected
Episodic memory past interactions the last five questions until it rotates out
Procedural memory how-to patterns your skill from session 10 until the procedure changes

Every item answers four questions before it is stored:

  1. Why is it useful? Name the answer it improves.
  2. Who owns it? A memory with no owner is everybody's memory.
  3. How is it corrected? Overwrite, reset or delete, and the user can reach it.
  4. When does it stop being true? A date, a count, or the end of the session.

Fill the four answers for one thing your capstone wants to remember. Write them on paper or in a scratch cell. An item with no answer to question 4 does not go in.

3. The state wrapper, live, then a second user

Open the notebook and run sections 1 to 3 from the top. Three beats.

Section What happens
2. The wrapper SessionState: one preferences dict, one episodes list, one reset(). Written for you. What the application does with it is the exercise.
3. The preference The cell sets answer_style = "short", asks one question through a FakeLLM, and prints the end of the prompt the model received.
5. The second user MemoryStore takes a user_id and never uses it. Ana and bruno both store locale. Read what ana gets back.

As shipped, section 3 prints:

prompt the model saw: Question: How does chunking work in RAG?
episodes: ['Q: How does chunking work in RAG?']

The preference is set. The prompt does not carry it. That is a stored value the model never saw, and nothing raised to tell you.

As shipped, section 5 prints:

ana asked for her locale -> en-GB
bruno asked for his      -> en-GB

Ana asked for her locale and got bruno's. A test written for one user cannot fail on this, because one user cannot read anybody else.

The question for the room: how many stores have you written that take a user_id argument and never read it?

Demo 11: memory at scale

Ana just got bruno's locale from a store with two people in it. Now the same problem with 20,000 memories, in the kind of index most assistants use.

Open demos/11_memory_at_scale.ipynb and run it from the top. It takes about ten seconds: no key, no account, no model download. It uses ChromaDB, which uv sync --extra projects --extra agents installed in part 0. Your exercise stays in memory with no database; this demo shows what that store turns into when it grows.

Section What you see
2. The exact answer Comparing a question with every memory is always right, and five times the memories costs about five times the time.
3. HNSW The index answers several times faster by walking a graph instead of scanning a list. With few links and few candidates it finds only 3 or 4 of the true 10.
4. The setting that lies Raise ef_search after the index is built: the collection reports 200, and the recall does not move.
5. Whose memory is it? Without an owner filter, ana gets bruno's memories. With one, she asks for 10 and gets her 5. A filter is still the weaker form: the key is where the owner belongs.

The question for the room: section 4 and part 3 are the same bug. What is it? (A value was stored and the thing that should use it never did. Both are caught only by checking the result, never by reading what you set.)

For your capstone: every read names the owner, fewer results is a result, and the owner, the correction and the expiry go in docs/RETENTION.md.

4. Your exercise

In the session notebook. 300 marks, three exercises, in notebook order:

Challenge What you write Marks
ch11-e1 in answer_with_state: the preference reaches the prompt, and the episodes cap at 5 100
ch11-e2 your storage policy: five lines, each answered, the refusal list last 100
ch11-e3 MemoryStore: four fixes, so two users never share a memory 100
uv run bootcamp check ch11

As shipped it prints:

running ch11 (State and memory)…
ch11: 0/3 passed
❌ ch11-e1: (a) the 'short' preference did not reach the prompt the model saw
❌ ch11-e2: the policy needs every line filled, not just the labels
❌ ch11-e3: two users, one key: ana stored 'pt-BR' and bruno stored 'en-GB' under the same key; recall('ana', 'locale') returned 'en-GB' and must return 'pt-BR'; hint: the owner is part of the key — store on (user_id, key), never on key alone

That is where you start. check runs the whole notebook, so it takes a few seconds longer than the check cell in Jupyter.

The two that catch people:

5. Break it on purpose

Notebook section 6. Two memories, read on 1 November 2026:

locale  pt-BR      expires 2026-10-28 -> read it again from the source
plan    free tier  expires None       -> read it again from the source

Same verdict, two reasons. The locale expired, and you can see that it did. The plan never said when it stops being true.

After class, change today to date(2026, 10, 1) and run the cell again:

locale  pt-BR      expires 2026-10-28 -> quote it
plan    free tier  expires None       -> read it again from the source

The locale became quotable. The plan did not, and no date you pick will change that. A memory with no expiry is not fresh. It is unaudited, forever.

Then break your policy. Replace your EXPIRES: answer with n/a, run the check cell, and read the refusal. Put your answer back before you move on.

The question for the room: which item in your capstone is the plan row?

6. Check, and read two refusal lists

uv run bootcamp check ch11

Or review("ch11"), the last cell in the notebook. Aim for ch11: 3/3 passed.

Then two volunteers read their WE REFUSE TO REMEMBER: line out loud. Everybody else takes one item they had not thought of and adds it to their own.

For your capstone: the five lines go into docs/RETENTION.md in your repository, in the same words the check reads. See the capstone guide, row 11.

7. Hand it in

When the check is green:

uv run bootcamp submit ch11 --github <your-github-name> --push

Save the notebook first. submit reads the file on disk.

It goes to dev3pack-submissions, never to the course repository. No gh on your machine? The command prints the browser route, and how to submit has the rest.

Exit ticket. One thing that works now, one thing that is still unclear, your next action.

Homework. Add a "what we refuse to remember" section to your project README, and delete one field you are storing without a reason. Rerun notebook section 6 with date(2026, 10, 1) (part 5), and run demo 11 again with one thing changed: 32 links, 50,000 memories, or no where= on the last query.

Tomorrow

Session 12, MCP architecture and primitives. Everything you decided to store today becomes something a tool could return.