// Capstone · assemble three patterns

Capstone: refund with judgment

Capstone work composes two or three patterns — not fifteen diagrams taped together. Build a stub support flow that prepares a refund reply, critiques it, and waits for a human before money moves.

// Pattern stack

What each pattern owns

  • Orchestrator-Subagent

    Dispatch research and draft specialists. The orchestrator only merges — it does not write the refund reply itself.

  • Critic-Actor

    A dedicated critic scores the merged draft against a refund rubric (no invented amounts, clear next step). Pass|revise with a retry budget of 2.

  • HITL

    Before any `execute_refund` stub runs, pause for human approve|reject|edit. No approval token → no side effect.

// Assembly sketch

One pipeline, three contracts

Ticket
  │
  ▼
Orchestrator ──► research stub
      │         draft stub
      ▼
  Merge proposal
      │
      ▼
Critic (rubric) ── revise? ──► Actor (≤2)
      │ pass / needs_human
      ▼
HITL gate ── approve|reject|edit
      │ approve
      ▼
execute_refund (stub) → done

Critic-as-runtime-scorer is not a CI eval harness. HITL is not a log line. Keep the contracts separate so you can swap stubs for real tools later without lying about what is automated.

// Mini-project

Stub runner checklist

Orchestrator + critic + HITL stubs · no real money

Ship a demo script that proves the gate order. Stubs are enough; live models are optional later.

  1. Accept a canned support ticket that implies a possible refund.
  2. Orchestrator delegates `research_policy` and `draft_reply` stubs; merge into one proposal JSON.
  3. Critic scores the proposal; on revise, feed notes as data and allow ≤2 actor retries.
  4. On critic pass (or retries exhausted with status needs_human), call `request_approval`.
  5. Stub human: edit amount once, then approve. Only then call `execute_refund`.
  6. Log the chain: delegate → critic rounds → approval id → execute. Assert execute never precedes approval.

Prefer underclaim: this capstone teaches assembly. It does not claim your entire agent fleet already runs production eval suites or auto-heals every failure.

// Failure mode

What usually breaks

Symptom: The orchestrator drafts, the critic is skipped, and a “preview” path calls execute without an approval token.

Fix: Fail closed on missing critic pass / needs_human handoff and missing approval id. Tests should assert execute count is zero until both gates clear.

// Reflect

One question

Which irreversible action in your own stack still has a specialist — but no critic and no human gate?