Skip to main content
Contact PuglieseWeb on WhatsApp or Call
← All guides

Documentation

Runbooks as Agent Tools

Turn the scripts your team already runs into typed, allowlisted MCP tools with @systemdox/runbook-mcp, composed with SystemDox context in one configuration.

19 min read
Guide
Listen to this article

Context says what. Runbooks do it.

SystemDox puts your standards in the agent’s context window at write time: the project rule that says fetch before you research a checkout, the check that says never change an alarm threshold without its history. That is the context plane. The moment a rule says “run the sync first”, the agent needs a sanctioned way to actually run it — an execution plane.

Most teams answer that with a shell. A shell is an unbounded surface: any command, any argument, and no record of which ones the team sanctioned. The scripts your operators already keep — the sync, the clean-up, the re-index, the cache flush — are a much better answer. They encode how the team does things. They only need to be reachable as tools.

@systemdox/runbook-mcp does exactly that: you declare your scripts in a runbooks.json manifest and every one of them becomes an MCP tool with a typed, allowlisted input schema. No free text ever reaches your shell. It runs where your scripts live — on the developer’s machine or the ops box — next to the SystemDox server that supplies the context.

The difference

A declared surface, not an open one

The manifest is the contract. Agents cannot extend it; they can only call what it offers, with what it allows.

An agent with a shell

  • Runs anything it can type
  • Composes arguments as free text
  • No timeouts, unbounded output
  • Nothing tells the client what is safe
  • The operator learns what ran afterwards

An agent with runbooks

  • Runs only the scripts in the manifest
  • Fills typed parameters that must match a pattern or an enum
  • Per-runbook timeout; bounded output with the drop reported
  • readOnly / destructive annotations on every tool
  • Every result starts with the exact command line that ran

Step by step

Quick start

Three commands from nothing to a runbook an agent can call. Node.js 18+ is the only prerequisite.

1

Scaffold a manifest

Writes a runbooks.json with one example runbook and the tiny script it points at. Replace the example with your own scripts, one entry each.

npx @systemdox/runbook-mcp init
2

See what agents will get

validate checks the manifest against the rules below, confirms every script exists, and prints each tool with the exact command it would run. Fix what it reports before any agent sees it — the server refuses to start on an invalid manifest.

npx @systemdox/runbook-mcp validate
  sync_repos  [powershell]  ops/sync-repos.ps1
      params: reportOnly?: boolean
      $ powershell.exe -NoProfile -NonInteractive -ExecutionPolicy Bypass -Command "... & 'C:\ops\sync-repos.ps1'; exit $LASTEXITCODE"

  OK — 1 runbook(s) ready to serve.
3

Register it with your agent

Claude Code reads .mcp.json in the project root (Cursor: .cursor/mcp.json). Restart the tool, then ask it to list its runbooks.

{
  "mcpServers": {
    "runbooks": {
      "command": "npx",
      "args": ["-y", "@systemdox/runbook-mcp", "--manifest", "./runbooks.json"]
    }
  }
}

The package is @systemdox/runbook-mcp on the public npm registry — no credentials or registry configuration needed. It is a separate server from the SystemDox context server, on purpose: the context server is a thin client over the SystemDox API and never executes anything.

Anatomy of a runbook

A real manifest from a multi-repository workspace: one runbook that refreshes every checkout, one that deletes disposable git worktrees. Notice what the second one does not expose.

{
  "version": 1,
  "runbooks": [
    {
      "name": "sync_repos",
      "description": "Fetch every repo under the workspace and fast-forward clean default-branch checkouts. Dirty or unmerged trees are left alone and flagged. Run before researching a checkout.",
      "script": "ops/sync-repos.ps1",
      "interpreter": "powershell",
      "timeoutSeconds": 1200,
      "idempotent": true,
      "params": {
        "reportOnly": { "type": "boolean", "flag": "-ReportOnly", "description": "Fetch and report; change no working tree." }
      }
    },
    {
      "name": "clean_worktrees",
      "description": "Delete disposable _wt-* task worktrees. Worktrees with uncommitted changes are always held back. Defaults to a report; call with reportOnly=false to delete.",
      "script": "ops/clean-worktrees.ps1",
      "interpreter": "powershell",
      "timeoutSeconds": 1800,
      "destructive": true,
      "params": {
        "reportOnly": { "type": "boolean", "flag": "-ReportOnly", "default": true },
        "match":      { "type": "string",  "flag": "-Match", "pattern": "^[A-Za-z0-9_*?.-]{1,64}$", "description": "Wildcard after the _wt- prefix, e.g. sdx*." },
        "keep":       { "type": "array",   "flag": "-Keep",  "pattern": "^[A-Za-z0-9_*?.-]{1,64}$", "maxItems": 20, "description": "Folder names to leave alone." }
      }
    }
  ]
}

Patterns or enums, always

A string or array parameter must declare an anchored pattern or an enum. There is no “any string” type. Double quotes and control characters are refused whatever the pattern says.

Safe by default, explicit to mutate

reportOnly defaults to true on the destructive runbook, so a call with no arguments is a preview. The agent has to pass false to delete anything.

Expose only the switches an agent should hold

The clean-up script also has -Force and -SkipChecks. They are not in the manifest, so no agent can pass them. Allowlisting is not only about values — it is about which flags exist at all.

Annotations the client can act on

readOnly, destructive and idempotent are published as standard MCP tool annotations, so clients that confirm destructive calls with the user get the signal without guessing from the name.

Interpreters: powershell, pwsh, bash, sh, node, python and exe. POSIX-style interpreters receive an argument array directly — nothing goes through a shell. PowerShell scripts are invoked through -Command with every value single-quoted and embedded quotes doubled: inside single quotes PowerShell performs no expansion, so a value that passed its pattern cannot change the shape of the command, and array parameters arrive as real PowerShell arrays. The full field reference is in the package README.

In practice

What the agent sees, and gets back

The tool, as listed

clean_worktrees(reportOnly?: boolean, match?: string, keep?: string[])

“Delete disposable _wt-* task worktrees. Worktrees with uncommitted changes are always held back. Defaults to a report; call with reportOnly=false to delete.

Runs clean-worktrees.ps1 (powershell); timeout 1800s; destructive.”

annotations: readOnlyHint=false, destructiveHint=true, idempotentHint=false, openWorldHint=false

The result

The command line that ran comes first, then the tail of each stream. A non-zero exit, a timeout or a script that failed to start comes back as an MCP error result with the same layout — never a protocol error — so the agent can read the log and decide.

clean_worktrees: exit 0 (41.3s)
$ powershell.exe -NoProfile -NonInteractive -ExecutionPolicy Bypass -Command "... & 'C:\ops\clean-worktrees.ps1' -ReportOnly -Match 'sdx*'; exit $LASTEXITCODE"

--- stdout ---
Would delete 7 worktrees (3 held back: uncommitted changes)
  _wt-sdx-mcp-loopback   product-systemdox   clean
  ...

One configuration, the whole toolbelt

MCP clients run several servers side by side, and each one should do one thing. A support engineer’s agent, for example, carries four: SystemDox for context, the runbook server for execution, and two observability sidecars that expose Sentry and CloudWatch read-only. Credentials are referenced, never written down — Claude Code expands ${VAR} from the environment it was launched with.

{
  "mcpServers": {
    "systemdox": {
      "type": "http",
      "url": "https://<your-workspace>.mcp.systemdox.com/mcp",
      "headers": { "Authorization": "Bearer ${SYSTEMDOX_API_KEY}" }
    },
    "runbooks": {
      "command": "npx",
      "args": ["-y", "@systemdox/runbook-mcp", "--manifest", "./runbooks.json"]
    },
    "sentry": {
      "command": "npx",
      "args": ["-y", "@sentry/mcp-server@latest", "--host=de.sentry.io"],
      "env": { "SENTRY_ACCESS_TOKEN": "${SENTRY_ACCESS_TOKEN}" }
    },
    "cloudwatch": {
      "command": "uvx",
      "args": ["awslabs.cloudwatch-mcp-server@latest"],
      "env": { "AWS_REGION": "eu-west-2", "AWS_PROFILE": "${AWS_PROFILE:-default}" }
    }
  }
}
Put the file where the client reads it

Claude Code: .mcp.json in the project root, or user scope via claude mcp add. A file under .claude/ is not read — a silent failure worth checking with claude mcp list.

Secrets come from the environment

${VAR} and ${VAR:-default} are expanded from the process environment, not from a .env file. Export them in the shell that launches the tool.

Use the direct server, not a proxy

The deprecated awslabs.core-mcp-server exposes no CloudWatch tools unless a role variable is set; awslabs.cloudwatch-mcp-server is the one that answers alarm, log and metric questions.

Pin the absolute path when PATH is unreliable

Editor-spawned processes sometimes carry a stale PATH. If a server fails to start, give command the full path to node, npx or uvx.

Close the loop with SystemDox

The manifest says what an agent can run. The when belongs in your standards, where every agent in every repository reads it at write time through get_context. Keep the two apart and each stays small: the manifest is an operator’s file, the rules are the team’s.

A project rule that names the runbook

“Before researching or basing an implementation on a repository checkout, run the sync_repos runbook and work from the refreshed tree. Never trust origin/* refs that were not fetched in this session.”

Category process. Create it from the console or let the agent file it with create_project_rule the first time a stale checkout costs an afternoon.

A prompt template that sequences the sidecars

“Investigate a production error: (1) find the Sentry issue and its latest event; (2) query the matching CloudWatch log group for the five minutes around it and check active alarms; (3) correlate by timestamp and request id; (4) only then propose a fix, and if the fix is a known runbook, run it in report-only mode first and show the output.”

Stored once as a prompt template, served to every support session through get_prompt_templates. The tool names in it are the ones the sidecars and the manifest actually expose, so the sequence is executable, not aspirational.

A check that keeps the manifest honest

pattern: a string parameter in runbooks.json without “pattern” or “enum”  ·  fix: declare an anchored pattern  ·  verify: npx @systemdox/runbook-mcp validate

The server enforces this on its own; the check puts the same expectation in front of the agent that edits the manifest, and validate is the fitness test that proves it in CI.

Security model

What holds, and what is on you

Declared surface

Only manifest entries are callable. The agent cannot add scripts, flags or arguments; what is not declared does not exist.

Allowlisted values

Strings and arrays match an anchored pattern or an enum; numbers are bounded; booleans are flags. Double quotes and control characters are refused regardless of the pattern.

No shell

Argument arrays for POSIX-style interpreters; single-quoted tokens for PowerShell. The command shape is fixed by the manifest, not by the value.

Bounded

A hard timeout per runbook kills the whole process tree; per-stream output caps keep the tail and report the drop.

Fail closed

An invalid manifest or a missing script stops the server from starting. validate tells you first; RUNBOOK_MCP_DRY_RUN=1 rehearses every tool without running anything.

Your part

The scripts run with the permissions of whoever launched the client. Keep the manifest in a reviewed repository, expose only the switches an agent should hold, and put secrets in the client environment, not in the manifest.

Frequently Asked Questions

Why not just let the agent use a terminal?

For exploration on a developer’s own machine, a terminal is fine and most coding tools offer one. Runbooks are for the operations that have consequences and a right way to be done: they make the right way the only way an agent can do it, and they give the same capability to clients that have no terminal at all, such as desktop assistants and remote agents.

Can the agent change the manifest?

Not through the server — it has no tool for that, and it reads the manifest once at start-up. An agent that can edit files on the same machine could of course edit the file, which is why the manifest belongs in a repository under review, like any other operational configuration. The check above makes that review an explicit standard.

Does this replace CI/CD?

No. Deployments stay in pipelines with their own approvals and audit trail. Runbooks are the operator-side scripts around them: refreshing checkouts, cleaning workspaces, re-indexing, toggling a maintenance page, collecting diagnostics. If a script should only ever run from CI, leave it out of the manifest.

Is it hosted, like the SystemDox context server?

No, and it should not be: it runs scripts, so it runs where the scripts and their credentials live. The SystemDox context server is hosted because it only reads and writes your standards over an API. Keeping the two separate is what lets the context server stay a thin client that never executes anything.

Windows, macOS, Linux?

All three. On Windows, powershell means Windows PowerShell 5.1 and bash resolves to Git for Windows (override any interpreter path in the manifest’s interpreters block). Elsewhere powershell falls back to pwsh.

Context first, then hands

Connect one repository on the free tier and your agents read your standards at write time. Add a runbook manifest and they can act on them — within the lines you drew.