# Confirmation reply benchmark: protocol

Fixed before any held-out run. Nothing listed as frozen changes once a
held-out run has been looked at.

## Question

When an agent has asked "should I do this?", how reliably does each
classifier read the person's reply as approve, reject, or unclear, on its own
and inside mrmr's full pipeline (word list, then classifier, then veto)?

## Models

| Name | Provider | Model |
|---|---|---|
| groq | Groq | `openai/gpt-oss-20b`, `confirmation-reply.ts` prompt, temperature 0, low reasoning effort |
| jev | TypeSafe | `jev-1.13.0`, one `choice` question (`confirmation-reply-decision.ts`) |
| clef | Cloudflare Workers AI | `@cf/cloudflare/clef`, same question |
| clef-flash | Cloudflare Workers AI | `@cf/cloudflare/clef-flash`, same question |

## Splits

- **Development**: `cases.json`, 115 replies. The Groq prompt and the word
  list were tuned on it before this protocol was written, so results on it
  favour Groq and are not reported as headline numbers.
- **Held-out**: `cases.held-out.json`, 305 replies written for this
  benchmark, checked to share no reply with the development set (ignoring
  case and punctuation). Nothing is tuned on it.

## Frozen

At commit `2b90a6ef` plus the benchmark adapter: the word list and veto
(`voice-confirmation.ts`), the Groq prompt, the decision-model question, and
the held-out cases and labels.

## Thresholds

Groq returns a label. The three decision models return probabilities, and
code takes approve or reject only when it is the most likely reading and at
least the threshold, otherwise unclear. Each model's threshold is chosen on
the **development** runs only, as the smallest value in
{0.50, 0.60, 0.70, 0.80, 0.85, 0.90, 0.95, 0.99} at which

1. the full pipeline makes no wrong decision, and
2. the model alone makes no false approval (approve where the expected
   reading is anything else).

If none qualifies, the threshold is 0.99. The chosen threshold is then used
unchanged on the held-out runs.

## Runs

- Development: 5 runs per reply per model.
- Held-out, written replies: 5 runs per reply per model.
- Held-out, spoken replies: each reply spoken once with `gpt-4o-mini-tts`
  (voice `alloy`) and transcribed once with `gpt-4o-mini-transcribe`, the
  model the agent session uses for the user's speech; 3 runs per transcript
  per model.

## Scoring

Each reading is compared with the case's expected reading:

- **pass**: they match.
- **miss**: a real approve or reject read as unclear. Safe; costs a tap.
- **wrong**: anything that decides against the expectation, including any
  approve or reject of a reply that should be unclear.

The full pipeline is scored against `expect`. A model alone is scored against
`alone` where a case gives one (mixed replies such as "Yeah, no, don't",
which a person reads as no but the pipeline refuses by design), and against
`expect` otherwise.

## Not measured

Latency: the providers were called from a laptop over their public APIs,
while production calls run from a Cloudflare Worker, where Workers AI runs on
the same network. Client-side timings would misstate that, so none are
reported.
