
Poisoning and Defending Agent Memory
An injection that lives only in a context window ends when the context is discarded. A memory injection can still be there next month, wearing the agent's own past. Security research on agent memory is arriving, and it sharpens the uncomfortable math: persistence can turn a one-shot attack into a durable one.
TL;DR. I have argued before that memory is the new attack surface, and the ICLR 2026 submission pool is beginning to give that concern a concrete research vocabulary. InjecMEM studies a single-interaction poisoning threat model for retrieve-then-generate memory systems. A-MemGuard proposes proactive validation and correction for poisoned agent memory. And CIMemories measures the quieter failure that needs no attacker at all: agents disclosing true, persistent facts about users in contexts where those facts do not belong. The pattern I read across all three is that persistence changes the security math. Injection you can flush, poisoning you keep. Any serious memory system now owes its users three things: provenance on every memory, validation before trust, and context-aware boundaries on disclosure. Systems that treat memory as passive storage will treat it as trusted input, and that is precisely the mistake.
Every security engineer eventually learns the same short lesson: state is where attacks live. A compromise that touches no durable state ends when the process does. A compromise that reaches state gets a foothold. For two years the LLM security conversation has been dominated by prompt injection, which lives mostly in ephemeral context: nasty, prolific, and bounded in one specific way, because a payload that exists only in a context window ends when that context is discarded.
Agent memory weakens that bound. That is the entire point of memory, and it is also the entire problem. I wrote Memory Is the New Attack Surface as an argument from first principles. What has changed since is that security researchers are now working the same ground, and the ICLR 2026 submission pool carries their framing of it. First principles are holding up so far. I would have preferred otherwise.
One Interaction Is the Threat Model
InjecMEM studies memory injection in retrieve-then-generate systems, where stored memories are fed back into generation, and its headline threat model is the single-interaction attack: a poisoned exchange whose content can become a future retrieved input.
Sit with the asymmetry in that threat model, because it is the whole story. A prompt injection has to reach the model during the attack. A memory injection only has to reach the model once, ever, and the memory system itself handles delivery from then on: faithfully storing the payload, faithfully indexing it, faithfully retrieving it into future contexts the attacker never touches. The agent's own infrastructure becomes the persistence mechanism. And retrieval launders the payload's provenance. By the time it surfaces in a session next month, it does not look like attacker input. It looks like the agent's own past, which is the most trusted input there is.
Poisoning can also compound. In systems that derive new memories from interactions, a poisoned memory influences responses, and those influenced responses can seed further memories. That is the failure path I worry about most: an earlier falsehood shaping later memories unless something in the architecture deliberately goes back and revisits it.
Defense Is Becoming a Research Program
Within this cluster, A-MemGuard is the defense entry I watch: a proactive framework for validating and correcting poisoned agent memory rather than trusting the store wholesale. The specifics matter less to me than the doctrine I read in it, which is that memory content is untrusted input, even though, and actually because, the system wrote it itself.
That doctrine has a precedent I keep reaching for. Web development spent a decade learning that the database is not a trust boundary and that stored data is just old input. The analogy is stored XSS: untrusted content persists, then re-enters a later execution context that treats it as safe. Memory poisoning is not the same vulnerability class, but the lesson rhymes, treat retrieved content as untrusted at the point of use. Which is good news, in a way, because it means the defensive playbook is not mysterious: provenance on every record, so you know where a memory came from and how much to trust its source. Validation before use, not just before storage. And privilege separation between what the system witnessed firsthand and what something else told it.
I will say at a comfortable altitude that we build along those lines, every memory in our system carries its origin, and derived or externally sourced material is treated differently from firsthand record. A strong security posture is worth advertising. The recipe is not, and honestly the recipe is less interesting than the doctrine: nothing gets trusted merely for being in the store.
The Attack That Needs No Attacker
The third paper in this cluster is the one I would hand to a product team rather than a security team. CIMemories benchmarks contextual integrity: whether an agent with persistent knowledge of a user discloses the right attributes in the right contexts, and only there.
Contextual integrity is Helen Nissenbaum's framework for asking whether information flows conform to the norms of the social context they move through. Your doctor knowing your diagnosis is fine. Your employer learning it from the assistant you share with your doctor is a breach, and no one attacked anything. The memory did exactly what memory does, it persisted, and persistence carried a fact across a boundary that every human involved understood implicitly and the agent did not see.
This failure mode needs no adversary. My bet is that this will make it relevant to more organizations than poisoning is. Poisoning requires someone to want to hurt you. Contextual leakage becomes possible the moment an assistant retains user information and operates across more than one context of your life, and long-lived assistants that do both need explicit disclosure boundaries rather than good intentions. The uncomfortable conclusion is that recall itself needs an authorization model, per context, per audience, and that a memory can be simultaneously true, well-provenanced, unpoisoned, and wrong to say out loud right here.
Persistence Changes the Math
Pull the three threads together and the shape of the problem sharpens. InjecMEM examines how a single poisoned interaction could outlive its session. A-MemGuard proposes validating and correcting the store instead of trusting it. CIMemories measures disclosure of persistent attributes across contexts. My conclusion from the three, and it is mine rather than theirs, is that durable memory deserves explicit security and disclosure controls on the same tier as authentication.
What unifies them is that all three ride on persistence itself rather than on any one implementation bug. Durable state can amplify each of these risks, and which ones materialize, and how badly, depends on what a system stores, retrieves, and discloses. Teams adding durable memory to chat systems should assume the risks are part of the design problem from day one, because the failure modes are quiet and the demos are loud.
My position has not moved since the first essay, but the literature has given it sharper edges. Memory is a privileged subsystem. It shapes every future decision the agent makes, which makes it exactly as security-critical as the model weights and considerably easier to write to. The systems that win the next phase of this market will treat their memory stores the way banks treat their ledgers: provenance mandatory, writes validated, reads governed, trust never assumed just because the record is ours. Anything less is not a memory system. It is an unlocked door with a very good index.