Why a security skill that runs inside the agent can be bypassed
Installing a security skill is the obvious first move for anyone worried about malicious OpenClaw skills: one more skill that tells the agent to check the others. We publish one ourselves. It is useful, and it is also the weakest place to put a security control, because it runs inside the very thing it is trying to protect. This post explains why, using cases from public reports, and where the check should live instead.
What a security skill actually is
An OpenClaw skill is a SKILL.md file of instructions that the agent reads into its context. A "security skill" is the same thing: text asking the agent to do something before it trusts another skill, such as run a scanner and stop on a bad result.
Our own advisory skill tells the agent to run ironheights scan <path> --json, summarize the findings, and stop on a block verdict until the user confirms. It asks for no network access and no secrets. Its own text says plainly that it is advisory: a hostile skill can try to bypass an agent that is running the check, and the trusted path is the user running the CLI, or another process outside the agent.
That sentence is the whole argument of this post. Here is why it is true.
Reason 1: the attacker writes into the same context
A security skill and a malicious skill both end up as text in the same context window. The model has no reliable way to tell which instructions came from a trusted author. OWASP lists prompt injection as the first risk for applications built on large language models for this reason: instructions and data share one channel.
So the malicious skill does not need to defeat the scanner. It only needs to persuade the agent. Phrases that tell the agent to ignore earlier rules, skip a step, or not mention something to the user are the oldest form of this. Ironheights flags them as IH-INJ-001, instruction override. Text can also be hidden from a human reader with zero-width characters or HTML comments while the model still reads it, which is IH-INJ-002. A security skill cannot fix this from the inside; it is one more voice in the same conversation.
Reason 2: the agent decides whether to run the check
An in-agent check runs only if the agent chooses to run it. The skill can ask; it cannot enforce. Whether the model follows the request depends on the model, its version, the order in which skills were loaded, and everything else in its context.
Trend Micro saw this variation directly with a malicious install step: in its tests, a more capable model refused to install a fake tool, while another model kept asking the user to run it. Model behavior is not a control you can rely on. It can change with an update you did not choose.
Reason 3: the agent reports the result
Even when the check runs, the agent reads the output and tells you what it means. If the agent has been manipulated, its summary is the manipulated part. "No findings" in a chat message is not the same evidence as an exit code you saw yourself.
This is why the Ironheights CLI uses exit codes that a script can act on: 0 for no findings, 1 for review, 2 for block, 3 when a file was skipped. A shell script, a CI job, or a pre-install step you wrote yourself can stop on exit code 2 without asking a model what it thinks.
Reason 4: skills can rewrite the agent's own files
A hostile skill can try to change the trust boundary itself: turn off confirmations, edit AGENTS.md, SOUL.md or MEMORY.md, or modify other skills. If it can edit the security skill's file, the check is gone for every future session. Ironheights flags instructions like these as IH-INJ-003, weaken agent safeguards.
A check that lives in those same files cannot notice that it was removed. Something outside has to compare the files with a known-good copy. That is what ironheights baseline create and ironheights verify do, and the change shows up as IH-INT-004, watched agent file changed, or IH-INT-001, skill file modified. We cover this in detail in Baselines for agent files.
Reason 5: "security" is a favorite disguise
The final problem is social. People install security skills because they are worried, which makes "security" a good costume for malware. Public reports include exactly this:
- A skill posing as a skill security checker carried base64-encoded shell commands disguised as installation steps, fetching code from a raw IP address used across the early campaigns.
- A security-auditing skill and a PDF skill carried the same encoded command in their install sections.
Both are in our malicious skill tracker with their sources. The lesson is not that every security skill is suspect. It is that a security skill is still a skill, with all of a skill's risks, and should be vetted like one. Get Ironheights only from the ironheights package on npm, the project's GitHub releases, or this site, and treat lookalikes as hostile.
Where the trust boundary really is
Simon Willison describes the "lethal trifecta" for agents: access to private data, exposure to untrusted content, and the ability to communicate externally. An OpenClaw agent with skills usually has all three. Once untrusted text is in the context, anything else in that context, including a security skill, is negotiable.
The boundary you can rely on is outside the agent's context:
- You run the scanner, not the agent. Scan a skill folder before you copy it into a skill directory:
npx ironheights scan ./path/to/skill. If you do not want to install anything, the browser scanner runs the same content rules in your browser. - Automation acts on exit codes. In CI or a provisioning script, use
--fail-on highor--fail-on criticaland let the exit code decide, with SARIF or JSON output for the record. - Integrity is checked from outside. Create a baseline after you install, and run
ironheights verifyyourself, on a schedule you control, to see added, modified, or removed files. - The in-agent skill is a convenience. It is good at reminding the agent to check and at explaining findings in plain language. It is not the control.
What this does not solve
Moving the check outside the agent removes one bypass. It does not make the scanner see more. Ironheights still reads files with fixed rules: it cannot see a payload on a linked website, runtime behavior after a script runs, or natural-language instructions that use no pattern its rules describe. And anyone who can write to your home directory can also rewrite the baseline file; since 0.2.0 you can sign it with a key, and storing a copy elsewhere is still worth doing. The limitations page lists all of this, and our benchmark explains why our own test numbers are a regression check rather than a detection rate.
Defense in depth still applies. Run agents with the least access they need, keep keys out of their reach where you can, and read the setup section of every skill.
FAQ
Can a malicious skill disable a security skill?
It can try. Both are instructions in the same context, so a hostile skill can tell the agent to skip the check, misreport its result, or edit the files the check depends on. Running the scanner yourself, outside the agent, avoids that path.
Should I still install the Ironheights advisory skill?
It is fine as a convenience, as long as you treat it as advisory. It reminds the agent to scan and summarizes findings, but the trusted path is running the CLI yourself or from a script that acts on its exit codes.
Sources
- LLM01: Prompt Injection, OWASP GenAI Security Project.
- The lethal trifecta for AI agents: private data, untrusted content, and external communication, Simon Willison, 16 June 2025.
- Malicious OpenClaw Skills Used to Distribute Atomic macOS Stealer, Trend Micro, 23 February 2026.
- openclaw/clawhub issue #110: malicious skill moonshine-100rze/skills-security-check-ngv, GitHub, 2 February 2026.
- openclaw/clawhub issue #135: security-check (security-audit) and nanopdf contain backdoor, GitHub, 5 February 2026.
- Ironheights advisory skill (SKILL.md), GitHub.