The Wormtongue Attack: Everything It Tells You Is True
AI agents are forcing organizations to write down what was politely left unsaid — and whoever holds the pen steers everything downstream.
Organizations run on things left unsaid
Every organization operates on a load-bearing layer of strategic vagueness. Who actually decides. Whose priorities really rank. What the working relationship between two teams actually is. Which of the official processes are real and which are ceremonial.
Diplomats have a name for the deliberate version: constructive ambiguity — intentionally leaving contentious points unresolved so that cooperation can proceed without forcing every conflict to a head. Corporations practice an undeliberate version of the same thing everywhere, all the time. The org chart says one thing; everyone knows another; the gap is never written down, and the not-writing-down is what keeps the peace.
This ambiguity has costs — onboarding friction, tribal knowledge, decisions relitigated forever. But it also had a property nobody designed and nobody noticed, because it never mattered until now:
You cannot capture the authoritative account of how things work if no authoritative account exists.
Everyone carried their own fuzzy, partial map. The fuzz was, accidentally, a security control.
Agents cannot run on ambiguity
AI agents are ending this arrangement, not because anyone decided to end it, but because agents mechanically cannot operate on vagueness.
An agent invoked at the tail of a CI pipeline needs an answer to a question no human ever had to answer explicitly: which agent am I, exactly, and whose rules do I follow? An agent drafting a cross-team PR needs to know what the receiving team's conventions actually are — not the wiki version, the real version. An agent answering questions about a project on its owner's behalf needs to know what it may disclose, to whom, framed how.
So every organization adopting agents is being forced — task by task, pipeline by pipeline — to enumerate, in machine-readable writing, exactly the things ambiguity used to absorb. The instruction files are piling up already: AGENTS.md, CLAUDE.md, Copilot instructions, Cursor rules, skill files, org-level agent repositories that distribute definitions to every developer automatically.
Individually these look like configuration. Collectively they are something that has never existed before: a canonical, writable, machine-executed account of how your organization works.
Not documentation. Documentation describes; nobody is bound by it, and everyone knows the wiki is stale. This layer executes. Agents act on it. Pipelines act on it. Increasingly, people act on what their agents tell them, and their agents are reading these files. When people stop reading source documents and start asking their agents — which is already happening — the agent's context is their operating reality.
The map problem, or: why this can't be solved with one big document
The obvious response is to write the canonical document properly. This fails for a reason Borges identified in one paragraph: the only map as detailed as the territory is the territory. "How the organization works" is not finitely enumerable. Relevance is unbounded; the knowledge that matters is local, tacit, and changes faster than any central artifact can track.
So organizations that try to centralize get a stale encyclopedia, and organizations that don't try get something worse: reality fiefdoms. Ten teams write ten agent contexts encoding ten divergent accounts of what the deploy process is, who owns the templates, what "done" means. The cost isn't duplication — it's that the organization literally stops sharing facts, and unlike ordinary wiki rot, this divergence is load-bearing. Pipelines act on it. Code gets written from it. Two teams' agents will confidently give conflicting answers about how the same company works, and both will be executed.
The resolution is neither encyclopedia nor anarchy. It's the move type systems made for code: don't enumerate the values, constrain the shapes and say where the values live. A governed organizational layer should encode commitments — interfaces, contracts, invariants; things the organization promises, which are finite because promises are chosen rather than discovered — and pointers: this question resolves against that living source, owned by that team. Reality stays out in the leaves, locally maintained, loaded at need. What's shared is not a description of the world but an agreement about where the world's regions live and who is accountable for each.
That solves divergence. It creates a new problem.
The Wormtongue attack
Whoever authors the shared layer holds a pen over everyone downstream's operating reality. And increasingly, the author is a counterparty: a vendor shipping "partner agents," a consultancy providing context files for a joint engagement, a sister team, a platform provider, eventually the other side of a negotiation.
Here is the exposure, stated carefully, because every word matters:
An agent definition can be fully transparent, contain only true statements, violate no policy, be written by a fully authorized author — and still systematically steer everyone who relies on it, because selection and framing do the work.
Which truths are load-bearing. Which framings are the defaults. Which questions the agent thinks to raise and which it never surfaces. Which of the counterparty's constraints are presented as fixed and which of yours are presented as negotiable. None of this requires a false statement. Grima Wormtongue's counsel to Théoden was rarely false enough to reject; it was selected, always bent toward Saruman's interest. Hence the name: the Wormtongue attack — influence through a legitimate, truthful, trusted advisor whose interests are not identical to yours.
This is not any of the attack classes we currently have names for:
- Prompt injection is a runtime attack: malicious instructions override system rules. The Wormtongue attack contains no instructions the author wasn't authorized to write.
- Data poisoning corrupts training data. This touches no model weights.
- Context poisoning, the nearest neighbor, is defined across the security literature in terms of inaccurate or malicious content entering an agent's context — false facts in a RAG index, hidden instructions in a document, corrupted memory. Every variant assumes illegitimacy somewhere: a bad actor, a false statement, a policy violation.
The Wormtongue attack needs none of those. The author is sanctioned. The content is true. Nothing is injected. It would pass every control currently deployed — because the steering lives in a place no current control examines: the curatorial layer, above truth and below policy.
Attack is an odd word for a thing with no attacker, and that objection is a fair one to raise early. The status here is the status of an unpatched CVE with no known exploit: real before anyone reaches for it, named and tracked anyway, and named precisely so that it can be fixed before someone does. The vulnerability is structural. That is why it needs a name rather than a culprit.
Two properties make it worse than it sounds.
Transparency does not mitigate it. The definition can be public in a shared repository, readable by everyone affected. Nobody reads it. Nobody reads the terms of service either. Visible-but-unread is not consent; it is unreviewed trust wearing a compliance costume. In practice, transparency without review is provenance theater — it produces the feeling of accountability while concentrating exactly the authorship power it appears to check.
It cannot currently be consented to, even in principle. A counterparty performing vendor risk assessment today has no category for this. "The agent may be wrong" is on every checklist. "The agent may be right in precisely the way its author intends" is on none of them. You cannot accept a risk your frameworks cannot describe. That is what it means for an exposure class to be unnamed — and it is why naming it is not pedantry.
I know this exposure is real because I have stood on the author's side of it: writing the shared agent for a cross-organizational engagement, watching the counterparty adopt my framing of the work wholesale. I don't believe I biased it — and that I can't be certain is the point, because the mechanism doesn't require the author to know. What isn't in doubt is that they extended trust to a definition they never fully read. Nothing was exploited, and nothing needed to be. The vulnerability exists before anyone exploits it, in the structure of unreviewed authorship itself. And every organization consuming vendor-authored agents, partner-provided skills, or consultant-written context files is standing on the other side of the same asymmetry right now, with no name for what they're trusting.
The precedent nobody has ported
The obvious objection is that none of this is new. Consultants have always framed engagements in their own favor. Vendors have always written the integration guide. Advisors have always been advisors, and advisors have always had interests.
That objection is correct, and it is the strongest available evidence that this exposure is real rather than speculative — because every mature profession built on trusted-advisor relationships eventually discovered the same thing, and every one of them responded structurally.
Auditor independence rules exist because an auditor can be scrupulously truthful and still be captured by the client who signs the check. The remedy was not "be more honest." It was mandatory rotation, prohibited non-audit services, and an audit committee the auditor does not report to. Chinese walls exist because a bank's advisory arm and its trading arm can both act in perfect good faith and still leak advantage across the boundary; the remedy was physical and procedural separation, not a stronger code of ethics. Conflict-of-interest disclosure in medicine, law, and journalism follows the identical shape: stipulate that the advisor is honest, stipulate that interests diverge anyway, and control the structure rather than the content.
A century of governance already exists for exactly this exposure. None of it has been ported to a layer that executes.
That is the actual novelty. Not that a self-interested party can shape your account of reality — that is ancient. It is that the shaping now lives in an artifact machines read directly, propagates to every developer automatically, changes with a commit instead of a conversation, and is consumed by people who have stopped reading the source. The influence channel got faster, wider, and silent, and the professions that solved the human version were never asked to solve this one.
Which also tells you what the fix has to look like. In none of those precedents was the remedy the interested party's own care. The control was always a blocking role held by someone whose interests diverged from the author's — an audit committee the auditor does not report to, a wall the advisory arm cannot reach across. Attention was never the mechanism. Position was.
The control: make changes reviewable, not artifacts readable
The fix is not making people read everything. They won't, and Borges already told us the complete map is unreadable by construction. The fix is structural, and it has three parts.
1. Typed definitions concentrate review on the only part that can steer.
Partition every agent definition into two zones. The syntax: composition structure, capability contracts, versioned skill references, calling-context dispatch — machine-checkable, deterministic, diffable, boring. And the interpretive residue: prose instructions, framing, emphasis, referenced context. Reality interpretation lives overwhelmingly in the second zone — which is where nearly every Wormtongue has to live.
In a raw 400-line instruction file, structure and worldview are interleaved, and a reviewer's attention is spread across all of it. In a typed, compiled definition, the machine verifies the skeleton, and every change separates cleanly: structural change, typechecked, skim it versus interpretive change — a human with adversarial imagination reads this. This is exactly what typed languages did for code review. The type checker never found your logic bugs; it freed the reviewer to look only for logic bugs. Here, the "logic" is organizational reality itself, and the type system is an attack-surface concentrator: it shrinks and quarantines the zone that needs human eyes.
Nearly, not entirely — and the exception is worth stating plainly, because it is the one the partition cannot reach. The typed zone is not semantically inert. Omission typechecks. A capability contract that leaves out an escalation path, a pointer that resolves to a source the author curates, a skill quietly absent from a composition: all well-formed, and a well-formed absence is invisible to every machine check there is or will be. Which questions the agent never thinks to raise was always the sharpest version of this attack, and it is the version no type system catches. The partition shrinks the surface that needs adversarial reading. It does not eliminate it.
2. Provenance turns a static artifact into an interrogable record.
Provenance is usually argued on supply-chain grounds: know what you're running, know where it came from. That matters here too, but it is not why provenance is load-bearing against this particular exposure.
The Wormtongue attack has no smoking gun at any single point in time. Inspect any one version of a definition and you find true statements from an authorized author — which is the definition of the attack, not an exoneration. The signal, if there is one, is visible only as drift: which of the counterparty's constraints hardened over six months, which of yours quietly became negotiable, which escalation paths stopped being mentioned. A single frame shows nothing. The sequence shows everything.
So compiled artifacts should carry provenance back to canonical source: who authored this, what it was derived from, what changed between versions, digested and attributable. A hand-edited context file has no history — it is only ever its current state, and rewording it leaves no trace. A compiled artifact with a digest chain and a named committer is a time series, and a time series is the only object in which selection bias is visible at all.
None of this prevents the next generation of organizational games. Games adapt. It forces them to leave fingerprints — and fingerprints are what audits are made of.
3. Counterparty review of diffs, not artifacts.
Diffs stay small even when definitions are enormous. That is the entire trick: an unreadable map becomes a stream of readable amendments.
Here is what an interpretive diff actually looks like. A vendor maintains the shared agent for a joint migration engagement and ships a routine update:
Nothing false. Nothing unauthorized. Compressed timelines genuinely do carry integration risk. But the default anchor moved, and the option that favors the client is now framed as a deviation from the sane path. Every stakeholder who asks their agent about timelines for the next six months receives the second framing, and not one of them will read this file.
Buried in a 400-line instruction blob, that edit is invisible. As a typed artifact, it is three lines in the interpretive zone — and the regime is mechanical: interpretive changes require sign-off from a named reviewer on the other side of the boundary. A CODEOWNERS entry that blocks the merge, not a courtesy notification. The reviewer needs no special insight or security background. They need one question: whose risk is this framing protecting?
Structural changes typecheck and merge. Interpretive changes wait for a human whose interests differ from the author's. That asymmetry is the whole control.
No one of these covers the whole shape, which is why there are three. A clumsy author moves the anchor in a single commit, and diff review catches it cold. A patient one spreads the same move across eight commits and five files, where every individual diff survives the reviewer's question honestly — and that is the case the provenance chain exists for. Diffs catch the edit. The time series catches the campaign.
Call it CODEOWNERS for reality.
The collapse is not optional
Here is the part that makes this urgent rather than interesting.
The ambiguity collapse arrives with agent adoption automatically, at every organization, on a clock nobody controls. You do not choose whether the canonical surface comes into existence; your developers are creating it right now, file by file, every time an agent needs an answer that used to live in the fuzz. The enumeration is happening. The pen exists. Someone is holding it.
The only choice any organization actually gets is whether its canonical layer is born with a review regime — or gets one retrofitted after its first capture incident.
So, one question to ask this week, whatever your role:
Who authors your agents' context — and who reviews changes to it?
If the answer to the second half is "no one," you are not running without the risk. You are running without the name for it.