Claude Operator: Prompt to Autonomy · 20 min · 140 XP

Isolation, scheduling, CI, and graduating

Where an autonomous run should execute, and how to earn each step of trust.

Isolate the work. Anything that changes code runs in a temporary branch or worktree, never on the main branch. An isolated diff is one you can discard entirely — which is what makes reviewing it a real decision rather than a formality.

Schedule something harmless first. A read-only report on a timer, with a documented and tested way to disable it. Scheduling is where a workflow stops being something you run and becomes something that runs — so the first one should be something you'd be relaxed about running wrong at 3am.

In CI, the environment is the control surface. Pin versions so a run is reproducible. Give the job the minimum token permissions it needs. Bound the commands. And treat anything from a pull request — title, description, branch name, diff — as untrusted input, because on a public repository it's written by a stranger who knows your automation reads it. That's the same prompt injection from Level 3, in a place where nobody is watching.

Then graduate in four stages. Manual. Suggested — it proposes, you do. Approved execution — it acts after you approve each action. Bounded autonomy — it acts within limits and reports. Move up when the current stage has been boring for a while, and define in advance what sends you back down: a wrong action that reached production, a budget surprise, an alert nobody understood.

Measure quality across the stages. Keep a small evaluation set — a dozen real inputs with known-good outcomes — and re-run it after every prompt or skill change. Without it, a change that fixes today's case and breaks three old ones looks exactly like an improvement, and you won't find out until the workflow is autonomous.

Practice. Run a code-changing practice task in a temporary branch or worktree and confirm the main branch is untouched. Schedule a read-only local report and test the disable path. Design a CI job with a pinned environment, minimum token permissions and bounded commands, then feed it malicious task text and confirm your boundaries outrank the content. Build a small evaluation set and use it to catch at least one regression after a prompt change. Then write your four-stage rollout with explicit rollback criteria.

Loading your workspace…