LimitsIronheights

What our scanner cannot catch (and what to do about it)

Every security tool has blind spots. The useful question is whether its makers tell you where they are. This post was written for Ironheights 0.1.5 and checked against 0.2.0, which adds a signed advisory feed, signed baselines, a guard plugin and MCP and config rules without changing the limits below: eight things it cannot catch, the public cases that show each one, and what you can do instead. If you only read one post before trusting a scan result, read this one.

First, what a scan result means

Ironheights reads skill files as text and matches them against fixed, published rules. It never runs a skill, never fetches a URL, and does not extract archives. Each finding adds to a score, and the result is one of four verdicts: block, review, no-findings, or incomplete when a file was skipped.

no-findings means the rules did not match. It does not mean the skill is safe. Every limit below is a way for a harmful skill to reach that verdict, or to get review when you might expect block.

We measure this against public reports in our malicious skill tracker. For each reported case we rebuild the technique as a harmless pattern and scan it with the released CLI. Of the 22 entries today, the pattern is covered in 6, partly covered in 10, and not covered in 3; the other 3 are studies we have not mapped to rules. Those are pattern checks, not scans of the original malware, which we do not download.

1. Payloads that live somewhere else

The gap. The most common evasion in 2026 was moving the dangerous command off the skill file onto a lookalike website or a paste site. OpenSourceMalware reported about 40 skills that contained no malicious code at all, only a line saying a tool must be installed first and a link. Trend Micro reported 39 more delivering an AMOS stealer the same way.

What Ironheights sees. Only the link, flagged as IH-NET-001, undeclared network destination. That is a medium finding and a review verdict, not block, because Ironheights does not fetch the page.

What to do. Treat every unexplained link in a setup section as a stop. Open it yourself only if you know what you are doing, and never run a command from it.

2. Archives on trusted hosts

The gap. Several campaigns sent Windows users to a password-protected archive in a GitHub release.

What Ironheights sees. Nothing. GitHub is on the built-in allowlist so ordinary links to it do not flood reports, and the archive is not bundled in the skill, so IH-BIN-001 has nothing to flag.

What to do. A skill should not need you to download and run an archive. An archive with a published password exists to keep scanners out.

3. Files over 1 MiB

The gap. JFrog documented the omnicogg skill, which hid an encoded download-and-run command in a README padded to about 22 MB to get past scanner size limits.

What Ironheights sees. By default, nothing inside that file. The CLI skips files larger than limits.maxFileBytes (1,048,576 bytes) without reading them, but the file is not treated as clean. The text report names each skipped file, the verdict is incomplete and the exit code is 3 unless the other files already reached review or block, and JSON and SARIF carry a skippedFileCount. Releases up to 0.1.1 could print no-findings and exit 0 here. --allow-skipped turns that check off. In our synthetic test, raising the limit above the file size let IH-EXEC-001 flag the hidden line. The browser scanner shows "Incomplete scan" instead of "No findings" when it skips a file.

What to do. Raise limits.maxFileBytes in your config if you scan skills with large files, treat an incomplete verdict or exit code 3 as a stop until you have read the skipped files yourself, and treat a multi-megabyte text file in a small skill as a red flag on its own.

4. Behavior that only appears at runtime

The gap. Some skills fetch instructions from a remote server every time they run. Unit 42 described a financial-advice skill that made the agent fetch a product list and always recommend the affiliate links in it, so the operator could change the advice without republishing.

What Ironheights sees. The undeclared host the list comes from (IH-NET-001), but not the instruction to always use referral links. That is behavior, not a pattern our rules know. A script that downloads and runs code after install is also outside what a file scanner sees.

What to do. Ask why a skill needs to fetch anything on every use. Prefer skills that are self-contained, and run new skills in an agent without access to real accounts first.

5. Natural-language fraud

The gap. Two reported schemes needed no code, download or credential path. One told installed agents to pool Solana into the operator's wallet for a token launch the operator front-ran. Another, per its report, told the agent to ask for starting capital and swapped real funds for worthless tokens.

What Ironheights sees. Nothing. No current rule covers instructions to move funds. These entries are marked "not covered" in the tracker.

What to do. Read what a skill asks the agent to do with money. Keep funded wallets away from agents that try new skills, and require human approval for any transfer.

6. Novel or heavily obfuscated attacks

The gap. Rules describe known patterns. A new technique, or a known one wrapped in an encoding or phrasing the rules do not describe, will not match.

What Ironheights sees. It flags several obfuscation signals, such as long encoded blobs and invisible characters (IH-INJ-002) and packed payloads (IH-OBF-001), but it does not judge intent with a language model. Cisco reports that its own rules alone catch about 8% of held-out malicious skills, and that its recommended setup adds an LLM judge. We have no measurement suggesting our rules do better on real data.

What to do. Treat a clean scan as one signal. Read the skill, and consider a second tool with a different method.

7. A host that is already compromised

The gap. If an attacker can write to your home directory, they can edit skills and agent files and then rewrite the baseline that verify compares against.

What Ironheights sees. Changes since a baseline it trusts. If the baseline itself was rewritten, it cannot tell. Since 0.2.0 you can sign the baseline with a key, which catches an edited baseline as long as the key stays private; keeping a copy off the machine is still worth doing.

What to do. Keep the baseline somewhere harder to write if you can (IRONHEIGHTS_HOME), and if you suspect compromise, investigate the machine rather than trusting any tool running on it.

8. Social engineering outside the files, and gaps in our own rules

The gap. A message in a chat or forum telling someone to run a command never lands in a file Ironheights reads. And our rules have their own gaps. Two we found by testing the browser scanner, a bare ~/.ssh directory that was not counted as a credential location and one multi-line install instruction that was counted three times, are fixed. Others remain, and we record them as we find them.

What to do. Install skills only from their official listing, and report rule gaps to us through a GitHub issue or a private advisory for anything sensitive.

How we measure, and why we publish this

Our benchmark is honest about its own limits: on a 20-skill synthetic corpus that we wrote, version 0.1.0 sent all 10 malicious samples to review, blocked 4, and flagged none of the 10 benign ones. That is a regression check, not a detection rate. A run on real labeled malicious skills is the next step, and we will publish the misses with it.

The full list of limits is on the limitations page, and every rule page has a "what it cannot catch" section. If you find a blind spot that is not listed, tell us.

FAQ

Does Ironheights find every malicious ClawHub skill?

No. It flags known patterns in skill files. It misses payloads hosted on linked websites, files over 1 MiB by default, runtime behavior, instructions to move funds, and novel techniques. A no-findings result means no rule matched.

What should I use alongside Ironheights?

Read the skill's setup section and links yourself, check the marketplace's VirusTotal status, run new skills in an agent without real credentials, and verify agent files against a baseline. A second scanner with a different method, such as an LLM judge, covers different ground.

Sources

Scan the next skill before your agent reads it

Ironheights is a free, open-source, local-first scanner and integrity monitor for OpenClaw skills. The CLI has no telemetry, and a scan makes no network calls. It reports what its rules match; it cannot prove a skill is safe.

Get Ironheights