Field notes · the expertise layer for agents

Essays on the expertise layer for AI agents.

A series on what real memory for agents demands — provenance you can audit, time as a first-class dimension, identity as an access-control boundary, and structure that knows not just what is true, but what to do and when it stopped being true.

The series · one argument in 26 parts

The problem
01

Memory Is the New Attack Surface in Agentic AI

Agent memory is durable, self-modifying state nobody reviews, and it is the next real attack surface.

02

Why Context Windows Are Not Memory

A bigger context window is a roomier scratchpad, not memory. Durability and trust are a different problem.

What real memory needs
03

The Five Types of AI Memory

Someone saying a thing once is not the same as it being true, and a flat vector store throws that away.

06

Why Most Agent Failures Are System Failures

When an agent fails, suspect the system before the model. Corruption enters on the write path, upstream.

07

Retrieval Is a Trust Boundary

Retrieval's real question is not 'is this relevant?' but 'should I believe it?' Make that decision explicit.

08

Time Is the Missing Dimension of AI Memory

A stale fact is not a wrong fact. Most memory systems know what but forget when, and when is everything.

Proving it & the shape it takes
10

Identity Is Harder Than Memory

Entity resolution is an access-control mechanism. A bad merge is a cross-tenant leak, not a data-cleaning bug.

12

Why Everything Becomes a Graph Eventually

The value in accumulated memory was never in the rows. It lives in the edges, and edges want a graph.

The thesis
18

Cognitive Infrastructure Will Matter More Than the Next Model

The model is no longer the scarce resource. The expertise layer underneath it is where the value migrates.

More in the series
04

Claude Code vs Hermes Agent: The Comparison Is Mostly a Trap

Claude Code and Hermes Agent keep getting compared as rivals, but they are good at opposite jobs. One you choose to supervise diff by diff, even though it can run unattended. The other is a persistent agent built to keep running while you are gone.

05

Trust the Harness, Not the Agent

I wrote that the smart move is to compose Claude Code and Hermes, then realized I had skipped the hard part. Composition is not a vibe, it is a mechanism. There are two ways to trust work that happens without you: invest in the agent, or invest in the process. Hermes does the first. The two tools I built, Stratum and Compose, do the second.

09

Do Memory Types Earn Their Keep?

I built a typed memory system, so when a leading minimal baseline in this year's ICLR submissions asks whether all that structure is worth it, I have to take the question personally. This is me taking it personally, and then answering it with the rest of the literature.

11

The Tools We Had to Build

Before Stratum and Compose had names, they were a diary of things going wrong. The origin story of the harness, told from the notes we kept while it happened: the product that forced it, the day the plan flipped, and the agent that drifted off-plan while we were designing the anti-drift tool.

13

Associative Memory Is Back

The most interesting retrieval papers in this year's ICLR submissions are not vector search papers. They are association papers: graphs, multi-signal ranking, evidence paths you can audit. The work keeps circling an idea librarians and one famously systematic German sociologist had long before embeddings existed.

14

We Gave Our Agent a Spec Language

Before our agent writes code, it writes a spec: steps, dependencies, and a postcondition on every step that a server checks. The agent narrates in plain English while the contract runs underneath. Why the plan in the agent's head is the most dangerous artifact in agent-driven development, and what changed when we made it write the plan down in a language that can fail.

15

Reasoning Over Episodic Memory

The frontier in agent memory is not better storage, it is retrieval that thinks. Three ICLR 2026 submissions point the same direction: time-aware episodic memory, retrieval that iterates instead of looking up, and experience distilled into strategy. Together they sketch what memory looks like when it stops being a filing cabinet.

16

Every Step Gets Checked

An agent fixed the same endpoint seven times, and declared victory seven times, before the real bug surfaced. Another shipped an empty knowledge graph for months while every test stayed green. What those two failures taught us about the word done, and what it looks like when a server, not the agent, decides whether a step actually happened.

17

Memory Has a Lifecycle

An append-only store is not a memory, it is a landfill with an index. Four ICLR 2026 submissions make a related case: memory needs filtering, consolidation, decay, and offline maintenance, not just storage. One of them makes the sleep metaphor unusually concrete.

19

From Goal to Product

Postconditions catch broken steps. They cannot catch a project marching in the wrong direction, one perfectly verified step at a time. The lifecycle layer of our harness: phases as levels of concreteness, gates that can actually say no, competing architects whose first ideas are treated as samples rather than decisions, and a cockpit so project state never lives in terminal scrollback.

20

Agents Write Their Own Memory

What you tell an agent is only half its memory. The reasoning, opinions, plans, and decisions it produces need types too, and one of them outranks the rest.

21

Poisoning and Defending Agent Memory

An injection that lives only in a context window ends when the context is discarded. A memory injection can still be there next month, wearing the agent's own past. Security research on agent memory is arriving, and it sharpens the uncomfortable math: persistence can turn a one-shot attack into a durable one.

22

Decisions Are Not Facts

A decision stored as prose keeps the verdict and throws away the trial. Why decisions need to be a first-class memory type.

23

Subagents With a Model Budget

Once the lifecycle fans work out to many agents, a new question appears: which model gets which job? The obvious answer, cheap models for grunt work, is right about the jobs and wrong about the economics. Why per-token pricing inverts on unbounded tasks, how we route by role instead, and why the budget is what keeps a fleet honest.

24

The Decision Lifecycle

Facts rot quietly. Decisions age by challenge. How a memory system can know whether a commitment should still stand.

25

Should Memory Live in the Weights?

The strongest objection to everything I am building is that models will simply absorb memory into their parameters and the external store will go the way of the search index. A wave of ICLR 2026 submissions makes that case seriously. I want to steelman it properly, and then explain why I still think history belongs outside the model.

26

Memory That Can Show Its Work

Agents are starting to get asked the question humans get asked: how do you know that? We just shipped the machinery that lets a memory answer with lineage instead of vibes, and the design argument underneath it is that you seal the logbook, not the mind.

Check back weekly for new essays.