Claude Operator: Prompt to Autonomy · 18 min · 130 XP
Failure classes, retries, and verification
Retry only what retrying can fix, and read back every write.
Not every failure is worth retrying, and retrying the wrong kind turns a clean stop into a loop. Sort them into three classes.
Transient — a timeout, a 503, a rate limit. The same call may well succeed shortly. Retry with backoff. Permission — a 401 or 403. Retrying changes nothing; something has to be granted first. Stop and report. Validation and logic — a malformed request, a missing field, a wrong assumption. The identical call will fail identically forever. Stop and report.
Only the first class is retryable, and even then with a bounded number of attempts and exponential backoff — one second, two, four, eight — so a struggling service isn't hammered by your recovery. A service returning 429 is telling you to slow down; retrying immediately is the one response guaranteed to make things worse. Cap the attempts and stop.
Then the habit that catches everything the classes don't: read back every side effect. After a write, fetch the thing and compare it to what you asked for. APIs accept requests and do something slightly different more often than you'd like — a field ignored, a value coerced, a partial success reported as a whole one. A verification record per changed item is what turns "I sent the request" into "the state is what I intended".
Practice. Build a retry table separating transient from permission, validation and logic errors, with a limit and a backoff for the transient row only. Simulate a rate limit and confirm your run's waits increase and it eventually stops rather than retrying forever. Then, after each practice write, read the result back and compare it with the requested state, keeping a verification record for every changed item.
Loading your workspace…