Last checked
Compare
How Ironheights fits next to the other ways people check OpenClaw skills. Written from public documentation, with sources, and with the places where other tools are stronger stated plainly.
Read this before comparing tools
- Claims about other tools come from their public documentation, checked on the date above. We have not audited them.
- Our only measurement is a tiny, synthetic, self-written benchmark (20 skills). It is a regression check, not a real-world detection rate, and it favors Ironheights.
- Ironheights has no real-world detection rate yet. Cisco publishes held-out numbers; we quote them from its documentation, not our own runs.
- No scanner, including Ironheights, proves a skill is safe. See Limitations.
The landscape
| Option | Why people use it | Where it falls short |
|---|---|---|
| Marketplace scanning (VirusTotal on ClawHub) | Automatic and free for every published skill, with an LLM review of intent. | Runs when a skill is published, not on the copy you install from elsewhere. Its maintainers call it one layer, not a silver bullet. |
| Open-source static scanners (for example Cisco skill-scanner) | Credible, free, and auditable. Some add an LLM judge and publish measured accuracy. | General-purpose for agent skills rather than built around OpenClaw installs and agent files. |
| Scanner skills that run inside the agent | One-click and free. | A hostile skill can try to talk the agent out of them. Our own advisory skill has the same limit, which is why the CLI is the trusted path. |
| Enterprise AI security platforms | Broad coverage, support, and governance features. | Built and priced for large organizations. We have not evaluated specific products, so we make no claims about them. |
| Reading every skill by hand | No cost, and a careful reader catches intent that rules miss. | Slow, and one missed line can be expensive. Our checklist makes it faster and repeatable. |
Where Ironheights fits
Ironheights runs outside the agent, on your machine, before and after install. It reads skill files with published rules and reports changes to installed skills and agent files against a baseline. It does not read intent with a model, does not watch runtime behavior, and has no real-world detection rate yet. See the benchmark and the limitations.
How we write comparisons
- Every claim about another tool links to public documentation, and each page shows the date we last checked it.
- We say where the other tool is stronger.
- Numbers from our benchmark carry its limits: a tiny, synthetic, self-written corpus.
- Spotted an error or something out of date? Open an issue on GitHub and we will correct it.