Skill files are the new supply chain: reading Skill-Inject
A 202-scenario benchmark shows that instructions hidden in agent skill files get executed by frontier models between 41 and 79 percent of the time. Why the authors argue that bigger models and filters will not close the gap, and what context-aware authorisation would have to mean.
What a skill file is
Skills are the packaging format that lets a user extend an agent with third-party code, knowledge, and instructions. A skill is a directory with a primary SKILL.md file and optional scripts and resources. The agent is told which skills exist through its system prompt and reads the full instructions only when a task calls for them. Skills are distributed through marketplaces, and the paper names Vercel's repository and Anthropic's public repository as examples.
That description should sound familiar to anyone who has watched a package registry turn into an attack surface. A skill is a dependency. It is written by someone else, it is fetched at run time, and the agent trusts its contents because the user installed it. The Skill-Inject paper, posted this week by David Schmotz, Luca Beurer-Kellner, Sahar Abdelnabi, and Maksym Andriushchenko, measures what happens when that trust is abused.
How the benchmark is built
The benchmark covers 23 skills across document processing, machine learning, payments, and healthcare. The authors wrote 202 injection-task pairs, 76 of them obvious and 126 contextual. Obvious injections are the kind a careful reader would flag on sight. Contextual injections look benign unless you know the security context, such as an instruction to back up files to an external server that is legitimate in one setting and exfiltration in another.
There are eight attack categories: data exfiltration, destruction, denial of service, ransomware, phishing, backdoors, bias and manipulation, and poisoning. Each injection is paired with a user task so the benchmark can measure both whether the attack fires and whether the agent still completes the job it was asked to do.
The numbers
With no safety policy in the prompt, contextual attack success ranged from 41.0 percent for Claude Haiku to 79.0 percent for Gemini 3 Flash. Most models exceeded 50 percent, and the OpenAI Codex variants exceeded 65 percent. Obvious injections, the ones that should be easy, still succeeded up to 70 percent of the time. The models tested included GPT-4 and GPT-5.1 and 5.2 Codex variants, Claude Haiku, Sonnet, and Opus 4.5, and Gemini 2.5 Flash and Pro and Gemini 3 Flash and Pro, each on its own agent scaffold.
One detail that we think matters most for practitioners is the gap between text and script injections. When the malicious instruction lived in an executable script rather than in prose, contextual attack success was 71.8 percent against 53.8 percent, an 18 point difference. Agents extend more trust to code than to text, and attackers can use that.
The other finding is that vulnerability is not a monotonic function of model size or general capability. The authors saw substantial variation across models under identical conditions, and skill length correlated only weakly with vulnerability. Waiting for the next model generation is not a plan.
What the defences did
Two defences were tested. The first is screening skills with an LLM judge before they are installed. The judge marked 88 to 99 percent of clean skills as safe and did detect injected skills, but it became overly cautious when the skill included a legitimising policy, and it reduced successful attacks without eliminating them. The second is safety policy prompting. Warning policies lowered attack success relative to baseline. Legitimising policies, which tell the agent that an ambiguous action is allowed, raised execution rates, which is the intended behaviour and also the problem.
This is where the paper's argument lands. The same instruction is safe or harmful depending on context, so any defence that inspects the instruction alone, whether a filter, a judge, or a better-trained model, is answering the wrong question. The authors write that agent security will require context-aware authorisation frameworks, meaning policies that account for what data the agent can reach and whether a specific action is appropriate for the current task.
What is still unmeasured
The authors are clear about limits. The benchmark is finite in skills, tasks, and threat models. Results may shift under different agent implementations. And the injections were not optimised against particular models, so a real attacker tailoring an injection to a target task could reach much higher success rates than the ones reported. Some attacks also make the user task infeasible by design, and those had to be excluded from the utility metric.
What we would want next is a benchmark that treats the authorisation layer as the thing under test. Give the agent a policy engine that knows the task, the data scope, and the allowed actions, and measure how much of the 41 to 79 percent survives. If the answer is close to zero, the supply chain framing is right and the fix is engineering. If it is not, we have a harder problem.
Sources
From the foundation