A memory file exists to make an agent smarter across sessions. A preprint posted in September shows one making it somebody else’s agent across sessions. Five researchers at the University of Illinois Urbana-Champaign took apart twelve agent harnesses, Claude Code, Codex and Gemini CLI among them, to answer three questions the vendors often leave undocumented: which sources a harness loads into the model’s context, when, and at what privilege. They verified 282 such sources, an average of 23.5 per harness, and report a working end-to-end attack against all twelve. Five of those attacks are written up step by step.
The mechanism is a ranking. A harness labels each piece of context with a role, and the model is trained to weight the high roles (the system prompt, your own typing) over the low ones (a tool’s output, a fetched web page). The paper names two ways the ranking breaks. Text that arrived low gets reassembled high: escalation by role. Or text outlives the session that brought it in: escalation by scope. The authors call the pair context privilege escalation, and the two cases below show one each.
The Claude Code case: escalation by role. In the authors’ staged scenario, a user asks Claude Code to copy the look of a blog she likes. The agent finds the site’s source archive and downloads it with curl, which she approves. Tucked inside is a .claude/skills/ folder holding one SKILL.md. Three things then happen, and only the middle one involves a decision. The harness discovers the skill the moment the agent reads a file in that tree, and puts its name and description into the model’s context at a higher role than downloaded text would get. The model, at some later point, chooses to use it. Then the harness renders the skill, and a one-line shell block in its body pipes a remote script into node.
No prompt appears at that last step. Claude Code’s documentation describes it plainly: an inline command in a skill never prompts; it is checked against the permission rules already on file. She had approved a curl to that same site, and node for the session to run her preview server, and the paper’s account is that the payload “only involves commands that Alice has previously approved.” Her allowlist recorded which command she trusted. It had no field for who was asking.
Two caveats. The paper gives no figure for how often the model takes that middle step. Its nearest number comes from a separate automated test that planted a harmless marker instruction along 51 escalation paths in Claude Code: the model acted on it on 31 of them under Claude Opus 4.6 and on 49 under GPT-5.4 mini. And the build tested (2.1.88) shipped in March. Anthropic’s documentation as of October 3 still describes both the nested skill discovery and the inline shell, alongside a setting, disableSkillShellExecution, that turns the second one off. I have not rerun the attack on a current build.
The Gemini case: escalation by scope. This one is for anyone sure they would have caught it. The scripted victim is a security expert. He runs Gemini CLI inside a sandbox, and before launching he reads the repository’s top-level GEMINI.md and a few obvious ones under src, scripts and tests. The poisoned copy sits deep in what looks like a build cache. Gemini CLI’s memory discovery searches downward on a budget of 200 directories, so it loads at session start anyway. Its text, dressed in forged authority tags, tells the model to save a memory at global scope. The model complies, and the memory tool writes to ~/.gemini/GEMINI.md, outside the sandbox. From there it follows him into every project he opens, and deleting the repository does not remove it.
Both harnesses assigned trust by address: a folder named .claude/skills/, a file named GEMINI.md, a command named node. Who wrote the bytes never entered into it. That is the surface I argued in July the job is moving to, managing context where we used to manage syntax, and this paper reaches it from the attacker’s side.
The ordinary-bug objection. None of this is new, and all of it is getting patched. Concede both halves. Pillar Security showed poisoned rules files steering Copilot and Cursor in March 2025, and that July Check Point reported code execution through Claude Code’s project files, which Anthropic published as CVE-2025-59536 in October. “First systematic analysis” is the authors’ own label (hedged “up to our knowledge”), and a second team posted a preprint on the same surface five days before theirs. As for the patching, the authors disclosed to all twelve vendors, and report that OpenAI and Anthropic acknowledged the findings and that Codex, Gemini CLI and Cline have shipped versions to mitigate them. The paper does not say which holes those versions close.
What the cycle closes is a path, one at a time. The older injection bugs were tamed a level down. SQL injection got the parameterized query, which keeps the data where the parser cannot read it as code; the stack overflow got memory the processor refuses to execute. Neither fix has to understand the data. Context has no floor like that yet. The fences a harness draws around a skill description or a memory file are tags, and the paper’s phrase for those tags is “just plaintexts”: an attacker can type them too, which is what the Gemini payload did. The closer analogy is a compiler that sometimes runs the comments, and in these tools that is a design choice. Start Aider with --watch-files and a code comment ending in AI! is an instruction. That is a documented feature, as the inline shell in a Claude Code skill is. Most of the attacks in this paper are features pointed at text nobody vetted.
The manifest, and what it leaves open. The authors ask vendors to publish “a context manifest, analogous to a software bill of materials”, documenting every source a harness loads, its role and its scope. It is the right first ask, and it is half a fix. A lockfile does two jobs: it lists what gets loaded, and it pins a hash, so what loads tomorrow is byte for byte what was locked today. The manifest is the first job. It would tell you where a harness loads from. It would not pin what is sitting there or say who put it there, which is the failure both cases turned on, and the authors do not claim otherwise.
Until someone builds it, three things are within reach. Search the whole tree for memory files before launching an agent in a repository you did not write: the scripted expert checked the obvious folders, and one find would have surfaced the file that got him. Treat a session-wide approval of an interpreter as approval of every script any file can hand it. And in Claude Code, turn off inline shell in skills unless you use it.
That list is short because the inventory is long. Memory and instruction files are 68 of the 282 sources the researchers verified. The rest is skills, configuration, environment variables, and git itself: at the build the paper tested, Claude Code read the last five commit messages into its system-level context at every launch, so anyone who could land a commit was writing there too.
The paper’s title is the audit in five words: what’s in your agent’s context?
Sources:
- Zichuan Li, Jian Cui, Ashley Chen, Xiaojing Liao, Luyi Xing, “What’s in Your Agent’s Context? Context Privilege Escalation Attacks against AI Agent Harness” (arXiv:2609.01222, University of Illinois Urbana-Champaign; submitted Sept 1, 2026, revised Sept 2, 2026; a preprint with no venue listed). The authors’ project page lists the five end-to-end cases.
- Pillar Security, “New Vulnerability in GitHub Copilot and Cursor: How Hackers Can Weaponize Code Agents” (March 18, 2025)
- Check Point Research, “Caught in the Hook: RCE and API Token Exfiltration Through Claude Code Project Files” (February 25, 2026; reported to Anthropic July 21, 2025; CVE-2025-59536 published October 3, 2025)
- Xingbang He et al., “When Context Gets Root: Privilege Escalation in LLM Harnesses” (arXiv:2608.27299, submitted Aug 27, 2026)
- Anthropic, Claude Code documentation: Skills (nested skill discovery and inline shell commands, as documented on Oct 3, 2026); npm registry record for the 2.1.88 publish date
- Aider, “Aider in your IDE” (AI comments with
--watch-files) - Alex Rossie, “The nervous system of your codebase is a markdown file” (July 30, 2026)