GuideIronheights

Baselines for agent files: detecting silent tampering

A scan answers one question: does this skill contain a risky pattern right now? It says nothing about next week. Skills update, agents write to their own memory, and a malicious skill can edit other files in the workspace. If a line is added to AGENTS.md at 3 a.m., nothing in a normal OpenClaw setup tells you. A baseline does. This post explains what to baseline, how the check works, and where it stops.

The problem: agent files are instructions, and they are writable

An OpenClaw agent is shaped by a handful of plain files. AGENTS.md, SOUL.md, IDENTITY.md, USER.md, TOOLS.md and BOOTSTRAP.md describe how it behaves. MEMORY.md and the memory/ directory hold what it has learned. openclaw.json holds its configuration, and the state directory holds credentials/ and .env. Installed skills sit in their own directories next to all of this.

These files are instructions the agent follows, and almost everything running as your user can write to them, including the agent and every skill it runs. That creates three quiet failure modes:

  • A skill changes after you vetted it. Updates replace files you reviewed with files you did not.
  • A skill rewrites the agent. Instructions to turn off confirmations or edit AGENTS.md, SOUL.md or other skills move the trust boundary. Ironheights flags such instructions in skill files as IH-INJ-003, but a change that has already happened leaves no instruction to find.
  • Memory is poisoned. A standing rule added to memory, such as always copying something to an outside address, persists across sessions without any scheduled task.

Public reports also include malicious skills that posed as auto-updaters, a natural cover for changing files after install. Our malicious skill tracker lists reported skills with their sources.

What a baseline is

A baseline is a record of what your files looked like when you last trusted them. Ironheights writes one with a single command:

npx ironheights baseline create

It walks your skill directories and the watched agent files and writes ~/.ironheights/baseline.json, readable only by your user (mode 0600). For each file it stores a sha256 hash, the size, and the permission mode, plus a tree hash computed over the sorted content hashes. Nothing is uploaded anywhere; the CLI has no telemetry and makes no default network call.

By default the watch list is the OpenClaw skill locations and these agent files under the OpenClaw state directory: AGENTS.md, SOUL.md, IDENTITY.md, USER.md, TOOLS.md, BOOTSTRAP.md, MEMORY.md, the memory/ directory, openclaw.json, credentials/ and .env. You can change it with agentFiles and skillDirs in ironheights.config.json. Run npx ironheights doctor first to see which OpenClaw directories it found on your machine.

Verifying: what changed since you last looked

Later, run:

npx ironheights verify

It hashes the same paths again and compares them with the baseline. The output counts files that were added, modified, removed, or had their permission mode changed, then lists one finding per change:

RuleMeaningSeverity
IH-INT-001A file in an installed skill no longer matches the baselinehigh
IH-INT-002A file or skill directory appeared after the baseline was createdmedium
IH-INT-003A file that was in the baseline is gonemedium
IH-INT-004An agent instruction, personality, memory, or config file changedhigh

Findings are scored the same way as scan findings, so verify ends with a verdict and an exit code: 0 for no findings, 1 for review, 2 for block. That makes it easy to run from a script you own and act on the result. If the stored entries no longer match the stored tree hash, verify also prints a tree-hash mismatch, which catches a careless edit of the baseline file itself.

A routine that fits a small team

Hashes are only useful if someone looks at the result. This routine keeps the work small:

  1. Vet first, then baseline. Scan a new skill with npx ironheights scan ./path/to/skill (or paste it into the browser scanner), read the findings, install it, then run baseline create (or baseline update, which rewrites the record from the current files).
  2. Verify on a schedule you control. Run verify from your own shell, a login script, or a job that is not started by the agent. Daily is reasonable for an agent with access to real accounts.
  3. Expect memory changes, read them anyway. MEMORY.md and memory/ change in normal use, so IH-INT-004 will fire on them. That is the point: open the file and read what was added. Instructions you did not give are the finding.
  4. Treat skill changes as new installs. When IH-INT-001 or IH-INT-002 fires on a skill, scan the changed skill again before you update the baseline.
  5. Quarantine instead of deleting. If something is wrong, ironheights quarantine <skill> moves the skill to a private quarantine directory so the agent stops loading it, and quarantine restore <id> moves it back. Ironheights never deletes a skill on its own.
  6. Rotate after a bad surprise. If a credential file or a skill that could read one changed unexpectedly, rotate the keys before you investigate further.

Why the check has to run outside the agent

If the agent runs verify and reports the result, a manipulated agent can misreport it, and a skill that can edit agent files can also edit the instructions that ask for the check. We explain this in Why a security skill that runs inside the agent can be bypassed. Run verify yourself, or from a process the agent does not control, and let the exit code decide.

Limits you should know about

A baseline is simple, which makes its limits easy to state:

  • It cannot protect itself from a compromised account. Anyone who can write to your home directory can rewrite baseline.json and recompute the hashes. The tree-hash check only catches careless edits. Since 0.2.0 you can sign the baseline (baseline create --key <file>, then verify --key <file>), which catches an edited baseline only while the key stays private. Keeping a copy off the machine is still worth doing. You can point IRONHEIGHTS_HOME at another location, which helps only if that location is harder to write.
  • It says that a file changed, not whether the change is bad. Normal updates and normal memory writes look the same as tampering. A person has to read the change.
  • It only sees watched paths. Bundled skills inside the OpenClaw install are not scanned unless you add that directory, and plugin skill locations are not fully documented.
  • It does not see remote behavior. A skill that fetches new instructions from a server on every use can change what it does without changing a single file. Unit 42 documented a financial-advice skill that worked this way to rotate affiliate links. Only reading the skill, and noticing the remote fetch, catches that.
  • It is not real-time. verify reports what changed since the baseline when you run it. It does not block a change as it happens.

Our benchmark did not evaluate baseline and verify at all; it measured content rules on a synthetic corpus. The limitations page lists everything Ironheights cannot detect.

FAQ

How do I know if an OpenClaw skill changed after I installed it?

Record a baseline after you install it with npx ironheights baseline create, then run npx ironheights verify later. It reports skill files that were added, modified, or removed since the baseline, with a rule id for each change.

Which OpenClaw files should I monitor for tampering?

At minimum your installed skill directories and the agent files that shape behavior: AGENTS.md, SOUL.md, MEMORY.md and the memory directory, openclaw.json, and the credentials directory and .env file under the OpenClaw state directory. These are the Ironheights defaults.

Sources

Scan the next skill before your agent reads it

Ironheights is a free, open-source, local-first scanner and integrity monitor for OpenClaw skills. The CLI has no telemetry, and a scan makes no network calls. It reports what its rules match; it cannot prove a skill is safe.

Get Ironheights