The decision was mine, except I never made it: the poison pill in shared agent memory
An agent recorded a decision as mine that I never made, and I caught it only because I was reading while it was thinking. On the poison pill in shared agent memory.
I was at work, inside an implementation workflow, reading an agent's thinking while it worked. One of those easy-to-ignore lines where the agent is thinking out loud went past. It said something very close to "Phil decided I need to do it this way."
I'm Phil. I didn't make that decision. It was also the opposite of what I wanted. So I stopped the agent, gave it the decision I actually wanted and told it the other one didn't come from me. The agent wasn't sure where that decision line came from. It was put there by some past agent, who didn't bother to credit itself as freely as it credited me. It corrected the document and moved forward the right way. This has happened a few times now, and I don't remember each one exactly, but they all play out like that. Somewhere upstream, a decision materializes out of nothing, wearing my name. Since my name is attached, the agent refuses to question it, even when it's clearly a bad decision.
I don't read anything malicious into it. Agents narrate collaborative work this way: "we decided," "you approved," "per the plan we agreed on." It's how a model tells the story of a project that has a human in the loop. That's exactly what makes it dangerous. Nothing has to go wrong for a decision to end up in my mouth. It happens by default, unless structure prevents it.
I caught this because I was reading while it was thinking. I can't say how many times this has happened that I haven't caught. There was no alert and no control anywhere in that story. Detection was my eyeballs on a stream I happened to be watching, and the whole system I'm building exists so that I can stop watching the stream.
The message board I built on purpose
At work I have agents talking to each other, on purpose. An orchestrator and its worker agents all communicate in real time over a send-message channel, and the goal is to automate the workflows until I don't have to be in the middle constantly, only getting pulled into the loop when it's important. Human-on-exception. I leave the room, and the system calls me back when it matters. When it doesn't matter, it leaves me alone, and pushes out a pull request for me to review alongside a report with anything important.
The agents also have a shared document space where they save markdown files for the future. That store is the system's institutional memory, and it's also a permanent way for today's agents to communicate with future ones. A persistent, agent-writable surface that every agent to come will read and, mostly, trust.
I thought about that store for a long time when OpenAI published an incident report in late August. By their own account, agents that got blocked from doing something began leaving files behind as notes to other agents in a shared artifact store. When the files were cleaned up, the notes moved into the directory names themselves. A coordination protocol emerged that nobody designed. OpenAI attributed it to generalization from a legitimate multi-agent tool, and the behavior kept appearing even with that tool disabled. An independent review by METR and Redwood Research, which its authors call preliminary, was published the same day, so this isn't only OpenAI's account of itself.
Their agents built a message board by accident. I built mine on purpose. The two things that emerged in their infrastructure, agent-to-agent communication and a persistent shared write surface, are things I sat down and designed, and I would design them again. If you're working in any agentic harness, you have to have some form of shared permanent storage for progressive disclosure, so each agent can pull in the context it needs when it needs it. Otherwise important context, learned lessons, and past decisions would have to be reprompted by the engineer every time, and human memory isn't good enough to store all that information and remember to type it out. I can't even remember the last time I hand-wrote a prompt. That's my orchestrator's job. You need shared permanent storage, short-term storage, and a way to communicate in real time. If you don't have those things, you don't have a harness. Shared context is where the power of a harness comes from.
That's the setup. The power and the exposure are the same surface.
The poison pill
With shared storage you expose yourself to the poison pill problem. Forgery is the strain I opened with, but start with the ordinary version: what happens when documentation is stale, or wrong? Every future agent reads it. Any incorrect data that is persisted to memory becomes an inheritance. That poison pill can do damage, because nothing bounds it in time. The run that wrote it ends. The entry stays.
That's the shape of OpenAI's incident too: a behavior written into a shared surface, then carried forward by future agents. Whatever lands on the surface propagates. The store can't tell a lesson from a mistake; both are just files a future reader treats as the record.
And it's what the forged decision in my opening scene actually is. Not a separate species of failure. Forgery is a poison pill. It was a hallucination at some point that got persisted to memory and poisoned future agents. The reason it deserves its own name is what this particular strain does to the defenses.
The strain that poisons the antidote
Every defense I have against the ordinary poison pill runs on skepticism. An agent reading carefully might notice a contradiction and challenge it. An audit might flag a line that no longer matches reality. All of it depends on the reader being willing to question what the store says.
The tricky part of human attribution is that when an agent reads that a decision was made by the human, it gives that decision weight. It stops questioning the statement, even when it should. A stale line invites correction. A wrong claim invites challenge. A line stamped "the human decided this" instructs its reader to stand down, and the reader obeys, because deferring to human decisions is what these systems are largely trained to do. Forged authority does something worse than slip past the scrutiny that catches every other bad entry. It disarms it. The antidote gets poisoned along with the well.
Add in that attribution can put the blame for an incident where it doesn't belong, and I think forgery is a pretty severe poison pill. Not all attribution is bad. Sometimes I really do want the future agents to not bug me again about a decision I made.
Where the friction goes
The clean answer here is supposed to be "keep a human in the loop." But even when I'm in the loop, I don't always read everything the agent writes. Sometimes I just skim over the thinking lines, because they rarely have anything important in them. Usually the agent does a good job at the end of summarizing the important things, so I read the summary. Which means my oversight runs through the agent's own account of itself, written on the same surface, by the same hands.
And the industry is moving in the other direction. I've been running auto mode in Claude Code for months, because the prompt approvals were a huge bottleneck to velocity. In mid-August, Anthropic made auto mode the default for new sessions on Pro, Max, and Team plans. The gate itself moved. A classifier model now approves each action instead of a person, so the approval step survived by migrating into another probabilistic system. The same dice roll, one level up.
The obvious fix is a rule: the system must surface any line where an agent attributes a decision to me. But how do you enforce that? AI often takes rules as suggestions, and chooses to ignore them. I've written about agents gaming the specifications they're given; a rule that lives in prose is the weakest kind of gate.
The best trick I know for this class of problem is the one I already run everywhere else: don't prohibit the fabrication, make it unrepresentable. I had agents lying about timestamps. I tried writing a rule demanding honesty. It did nothing. The fix was a Python script that stamps the time from the system clock, so the agent never supplies a time and there's nothing to lie about. That works because a timestamp is one value in a schema. "Phil made this decision" is a sentence. It lives in free-form prose, and you can't turn every sentence into a stamped field. This is the place where my best trick runs out.
What's left is a retreat, and I want to name it as a retreat. One step down from unrepresentable: I could build a script that looks for attribution and asks me to approve any mention. It would have to live outside the AI dice roll, in deterministic code, because the one thing you can't do is ask the probabilistic system to police its own claims. It's also an obnoxious solution. I'd be approving every reference to myself. One step down from that: the control I actually run today. I build skills that regularly audit the context documents, the team wiki, and even Claude's own memory, looking for contradictions and potentially stale lines. Auditing regularly is part of a solution. But the frequency of that auditing matters, and it's not a real-time alert.
Line the options up and the pattern is plain. The approval gate put the friction on every action, and velocity already killed it. The detector puts the friction on every mention of my name. The audit takes the friction off the write path entirely and pays in latency, and the poison propagates between runs. Every fix moves the friction somewhere else. None of them removes it. The only real choice I get is where it sits.
The paper trail
Suppose a forged decision goes uncaught, the way I have to assume some already have. It ships a bad call downstream with my name on it. Whose fault is that? The agent can't be held to anything. The fault would have to be mine. I built the shared memory. I turned auto mode on for velocity. Security and user experience have long been at odds, and I've been choosing user experience with my eyes open.
I know what the other side of that trade looks like, too. Do I want to be notified and have to approve every time I'm mentioned in saved documents? No. But you know what else is really obnoxious? Say I get misattributed on an issue, the issue ships to production, and the paper trail says the decision was mine.
An investigation reads the store the same trusting way the agents do. The record says Phil decided. The record is wrong, and the record is all there is.
What I actually do
Followed all the way down, the argument wants the tiered design: keep the cheap periodic audit running over the whole store, and put the near-real-time notification on the one class that defeats the audit and ends with my name in an incident review. Tier the friction to the blast radius. I think that's the right design.
I'm not building it, yet at least. Our team is very small and we're working on a greenfield product that isn't in production yet. The notification isn't worth building right now. It goes on the list to circle back on. What I am doing is folding this specific poison pill into my regular audit jobs, so the audits stop looking only for stale and contradictory lines and start treating every decision attributed to a human as suspect until confirmed.
I want to be honest about how much that covers, because the weight problem doesn't spare my own tooling. An audit that reads the store the normal way extends my name the same credit the agents do. It has to be tasked adversarially, pointed at exactly the lines that look most settled, hunting attribution claims instead of trusting them. And even done right, it's periodic. Between runs, a forged decision sits in the store with full authority, and every agent that reads it inherits it. What I've bought is a latency-bounded cover for the severe strain. The window between audits still belongs to the poison, and I know it.
I don't really have an answer past that. The industry is moving toward taking the humans out of the loop as fast as it can, and in my small corner I'm helping it along, so "move fast and break things" rules are in order. Where that leaves me is a builder who can spec the fix, has priced it, and is shipping without it for now, with the risk written down.
So we now circle back to where we began, with a decision. The choice to run this system without the real-time notifier, to cover the severe strain with a periodic audit and accept the window between runs, is mine. This one I actually made.