// Capstone · assemble three patterns
Capstone: refund with judgment
Capstone work composes two or three patterns — not fifteen diagrams taped together. Build a stub support flow that prepares a refund reply, critiques it, and waits for a human before money moves.
// Pattern stack
What each pattern owns
- Orchestrator-Subagent
Dispatch research and draft specialists. The orchestrator only merges — it does not write the refund reply itself.
- Critic-Actor
A dedicated critic scores the merged draft against a refund rubric (no invented amounts, clear next step). Pass|revise with a retry budget of 2.
- HITL
Before any `execute_refund` stub runs, pause for human approve|reject|edit. No approval token → no side effect.
// Assembly sketch
One pipeline, three contracts
Ticket
│
▼
Orchestrator ──► research stub
│ draft stub
▼
Merge proposal
│
▼
Critic (rubric) ── revise? ──► Actor (≤2)
│ pass / needs_human
▼
HITL gate ── approve|reject|edit
│ approve
▼
execute_refund (stub) → doneCritic-as-runtime-scorer is not a CI eval harness. HITL is not a log line. Keep the contracts separate so you can swap stubs for real tools later without lying about what is automated.
// Mini-project
Stub runner checklist
Orchestrator + critic + HITL stubs · no real money
Ship a demo script that proves the gate order. Stubs are enough; live models are optional later.
- Accept a canned support ticket that implies a possible refund.
- Orchestrator delegates `research_policy` and `draft_reply` stubs; merge into one proposal JSON.
- Critic scores the proposal; on revise, feed notes as data and allow ≤2 actor retries.
- On critic pass (or retries exhausted with status needs_human), call `request_approval`.
- Stub human: edit amount once, then approve. Only then call `execute_refund`.
- Log the chain: delegate → critic rounds → approval id → execute. Assert execute never precedes approval.
Prefer underclaim: this capstone teaches assembly. It does not claim your entire agent fleet already runs production eval suites or auto-heals every failure.
// Failure mode
What usually breaks
Symptom: The orchestrator drafts, the critic is skipped, and a “preview” path calls execute without an approval token.
Fix: Fail closed on missing critic pass / needs_human handoff and missing approval id. Tests should assert execute count is zero until both gates clear.
// Reflect
One question
Which irreversible action in your own stack still has a specialist — but no critic and no human gate?