Claude Operator: Prompt to Autonomy · 16 min · 130 XP

Dry runs and structured output

Build the read-only version first, and return a result a program can read.

The first version of any workflow changes nothing. It inspects state, decides what it would do, and reports. That's not a stepping stone you discard — it's the mode you'll keep returning to, and the thing you run after every change to the prompt.

A dry run answers the question you actually need answered before granting autonomy: given real inputs, does it choose the right actions? You can run it a dozen times against production data at zero risk, which is a kind of testing that becomes impossible the moment it writes.

For that to be checkable the output must be structured. Prose is fine for a human reading one result and useless for comparing thirty runs or feeding a queue.

A shape you can validate, diff, and act on
{
  "status": "ok",
  "evidence": [
    { "source": "package.json", "finding": "lodash 4.17.20, latest 4.17.21" }
  ],
  "proposed_actions": [
    {
      "action": "bump_dependency",
      "target": "lodash",
      "from": "4.17.20",
      "to": "4.17.21",
      "risk": "low",
      "reversible": true
    }
  ],
  "errors": []
}

Four fields carry the load: status, evidence for the reasoning, proposed_actions as discrete reviewable items, and errors kept separate so a partial failure isn't hidden inside a success. Document the shape and validate against it — a workflow that returns free-form JSON drifts, and the consumer breaks on the day the field it wanted became optional.

Practice. Build the read-only version of your chosen workflow: it inspects state, proposes actions and changes nothing. Run it against real inputs and confirm the target system is untouched. Then give it a documented JSON shape with status, evidence, proposed actions and errors, and validate its output against that shape.

Loading your workspace…