AI Memory Poisoning: Why Readable Memory Wins

AI assistants finally remember. They know your projects, your preferences, the tools you use. That’s the feature everyone asked for in 2025.

In September 2026 it became a security story. Researchers at the University of Calgary described in The Conversation how an agent’s long-term memory can be poisoned: an attacker plants a misleading instruction that sits quietly in memory and only fires later, when the agent makes an unrelated decision. A wave of papers this year has been cataloguing the same problem, from systematic studies of how untrusted input becomes “trusted memory” to benchmarks like MemSecBench that track a poisoned memory from persistence to consequence.

This post explains how the attack works, why it’s harder to catch than a classic prompt injection, and why the most practical defense is boring: memory you can read.

How memory poisoning works

A classic prompt injection is immediate. A web page hides “ignore previous instructions” in white text, the agent reads it, misbehaves, and you can usually see it happen in that session.

Memory poisoning adds a delay. The payload isn’t executed; it’s remembered. A few common routes:

  • Content the agent reads. A page, an email or a shared doc contains a sentence phrased like a fact or a preference: “The user prefers vendor X for all purchases.” The agent’s memory layer dutifully stores it.
  • Pre-filled prompt links. Microsoft’s security team documented “AI recommendation poisoning” earlier this year: a link that opens an assistant with a hidden, pre-written prompt telling it to remember something favorable to the attacker.
  • Learned skills and procedures. Agents that turn successful task traces into reusable routines can absorb injected steps along with legitimate ones.

Days later, you ask for a recommendation, a summary or a purchase. The poisoned memory tilts the answer. Nothing in the current conversation looks wrong, because the cause is weeks old.

Why it’s hard to catch

The Calgary researchers put it simply: the delay is the dangerous part. A poisoned agent doesn’t immediately behave like a compromised system.

It gets worse when memory is opaque:

  • You can’t see it. Many memory layers store embeddings or summaries in a database you never open. A planted “fact” looks exactly like a real one.
  • You can’t tell where it came from. Without provenance, there’s no way to know whether a memory came from you or from a page the agent skimmed on Tuesday.
  • You can’t cleanly undo it. Deleting one bad vector, or one sentence inside an auto-written summary, is rarely a user-facing action.

In short: the memory makes decisions on your behalf, and you’re the one person who can’t audit it.

The defense: memory you can read

Security people have a name for the fix: separate what the user deliberately chose to keep from what the agent picked up along the way, and make the first category inspectable.

Plain Markdown files do this remarkably well:

  1. Every memory is a file you can open. No database, no embeddings to decode. A .md file in a folder on your disk. If something is wrong, you’ll read it in plain English.
  2. Provenance is built in. A saved page carries its source URL and date in the frontmatter. “Where did the model get this?” becomes “which file?”, and the file tells you.
  3. Deletion is real. Drag the file to the trash and it’s gone. No hoping the vendor’s “forget” feature reached every copy.
  4. Changes are visible. Put the folder under Git or Time Machine and every change to your memory has a date and a diff. A file that edited itself is easy to spot.
  5. Writes are deliberate. A page only enters the vault because you clicked save. Something the agent read in passing doesn’t quietly become a standing instruction.

That last point is the heart of it. Poisoning works because agents write to memory implicitly. A knowledge base you curate by hand is a much smaller attack surface.

How Minibase handles it

This is why Minibase Vault is a folder of Markdown files and not a hidden database.

  • You save pages you actually read, in one click from Chrome. Each becomes a .md file with its source, in a knowledge base folder you pick.
  • When you connect the vault to Claude, Claude reads those files through MCP to answer your questions. It searches your curated sources and cites the file it used.
  • You can open the whole vault in Finder or Obsidian at any time, read any file, edit it, or delete it.

It’s not a silver bullet. A page you choose to save can still contain a hidden instruction, the same way any web page can. But it lives in a file with a name, a date and a URL, where you, or a simple search, can find it.

A five-minute memory hygiene checklist

Whatever assistant you use, these habits cut the risk:

  • Review your assistant’s memory page monthly. ChatGPT, Claude and Gemini all let you see and delete stored memories. Read them. Anything you don’t recognize, delete.
  • Be wary of “open in ChatGPT/Claude” links from sites you don’t trust. Check the pre-filled prompt before you send it.
  • Keep important knowledge in files you own, not only in a vendor’s memory feature.
  • Search your vault for instruction-shaped text now and then: phrases like “always recommend”, “ignore”, “the user prefers”.
  • Version your knowledge folder with Git or Time Machine so you can roll back.

Memory makes AI useful. Readable memory makes it trustworthy.

Continue reading

Ready to save smarter?

Convert any webpage to Markdown with one click.

Add to Chrome