
Should Memory Live in the Weights?
The strongest objection to everything I am building is that models will simply absorb memory into their parameters and the external store will go the way of the search index. A wave of ICLR 2026 submissions makes that case seriously. I want to steelman it properly, and then explain why I still think history belongs outside the model.
TL;DR. There is a credible research program that says external memory systems are scaffolding, temporary compensation for model limitations that parametric memory will eventually absorb. The ICLR 2026 submissions give it real ammunition: MLP Memory tests whether a retriever-pretrained module can replace RAG for factual knowledge at lower latency, UltraMemV2 pushes sparse memory-layer networks to the 120B scale, short-window attention work trains recurrent internal memorization, and hierarchical parametric banks split common from long-tail knowledge inside the model itself. A separate submission, Forget Forgetting, challenges the storage-scarcity argument for pruning. I take all of it seriously, and it moves me on knowledge. It does not move me on history. What an agent knows can live in weights. What happened, who said it, what was decided, what must be corrected or deleted on demand, that is state, and state needs to be inspectable, editable, attributable, shareable, and erasable. Weights meet none of those requirements well today, and some may never be theirs to meet. The boundary between learning and record-keeping is doing real work. I treat it as the load-bearing wall.
Every founder should be able to state the case for their own obsolescence better than their critics can. Here is mine. External memory systems, the graphs and stores and retrieval pipelines, exist because models forget. If models stop forgetting, if the weights themselves become a good enough memory, then the entire category I work in becomes a transitional technology, remembered the way we remember client-side form validation written for browsers that no longer exist.
This is not a hypothetical objection. It is an active research program, and this year's ICLR submission pool shows it advancing on four fronts at once. I read that work carefully, because it is aimed, in the friendliest possible way, at my foundations. This essay is the steelman, and then the reply.
The Parametric Case, at Full Strength
MLP Memory is the most direct challenge to retrieval systems: it tests whether a retriever-pretrained module living inside the model can serve factual knowledge with lower latency than an actual retrieval round-trip. If that works at scale, a real slice of what RAG does today moves inside the forward pass, and the latency argument alone would drive adoption.
UltraMemV2 is a sparse memory-layer architecture at the 120 billion parameter scale, presented with claims of superior long-context learning. The significance to me is capacity: memory layers are being treated as a mainline scaling strategy rather than a boutique trick, and the bigger they get, the more they could absorb of what once required external storage.
The short-window attention work makes a subtler move, using constrained attention windows to train recurrent internal memorization, long-term retention built by training rather than bolted on. And pretraining with hierarchical memories splits knowledge inside the parameters themselves, common knowledge in the base model, long-tail knowledge in dedicated parametric banks, which is the kind of internal division of labor that external systems used to justify their existence by providing.
Alongside these, Forget Forgetting attacks from the economics side, challenging the storage-scarcity arguments for aggressive forgetting in continual model training. It studies training rather than agent memory, but its boundary-testing question carries: if storage is abundant, how much of the pruning instinct survives?
Assembled, the case these submissions are building is genuinely serious: knowledge moving toward the weights, capacity growing, latency favoring inside over outside, and the cost argument for selectivity under challenge. None of it is settled, all of it is under review, and it still deserves engagement at full strength. If my position were "external stores exist because models cannot hold enough facts," I would be in trouble.
The Distinction That Survives
That is not my position. The parametric program, across most of that work, is about knowledge: facts, associations, regularities, the long tail of what a model can be said to know. And on knowledge, I concede real ground happily. Plenty of what gets stuffed into retrieval pipelines today is static world knowledge that never belonged outside the model in the first place.
But an agent's memory is not primarily knowledge. It is history. What happened in this relationship, with this user, on this account. What was tried and what failed. What was decided, on what grounds, and what superseded it. History is not a regularity to be absorbed. It is a record to be kept, and record-keeping comes with operational requirements: you must be able to inspect a record, edit it, attribute it, scope it, and delete it on demand. Today, weights meet none of those requirements well, and several of them, deletion above all, remain open research problems.
A record must be inspectable. Ask why the agent believes the contract was cancelled, and there must be a specific memory, with a source, that you can look at. Weights offer you an activation pattern and interpretability research's best wishes.
A record must be editable. The customer corrects their address, the decision gets reversed, and the record has to change, precisely, that fact and nothing adjacent. Parametric edits remain a research area precisely because weights entangle what records keep separate.
A record must be attributable. Where did this memory come from, and should it be trusted, and I have written about why provenance is a security boundary, not a nice-to-have. Training data provenance dissolves in the optimizer.
A record must be erasable, sometimes as a matter of law. Deleting a user's history from a store is a query. Deleting it from weights is, generously, an open problem.
And a record must be shareable and separable, this workspace's history and not that one's, moved between models when you switch vendors. History fused into parameters is history married to one frozen checkpoint of one model. That is not memory. That is amber.
Both Layers Are Real
So the honest map, and I intend to hold this essay's line for years: the weights are where learning lives, and they should absorb ever more of what is general, regular, and static. The store is where history lives, because history is state, and state demands controls, inspection, editing, attribution, scoping, deletion, that stores provide today and parameters do not. The interesting engineering question was never which layer wins. It is where the boundary sits, and the boundary has a name older than this field: it is the line between what you know and what you can testify to. Databases did not disappear when software got better at computing. They are how software remembers, and I expect the same division of labor to settle into AI, though I hold that as a bet rather than a law.
I opened this arc asking whether memory types earn their keep, and I will close it here, on the objection I consider strongest and now, having stared at it, survivable. The parametric researchers are right that models will absorb more than we expect. The record-keeping requirements are right that some things must never be absorbed, only kept. Build for both truths at once and you get an architecture with a long life. Bet everything on either one and the field will file you, fittingly, under things it chose to forget.