# Confirmation reply benchmark: protocol

Fixed before any held-out run. Nothing listed as frozen changes once a
held-out run has been looked at.

## Question

When an agent has asked "should I do this?", how reliably does each
classifier read the person's reply as approve, reject, or unclear, on its own
and inside mrmr's full pipeline (word list, then classifier, then veto)?

## Models

| Name | Provider | Model |
|---|---|---|
| groq | Groq | `openai/gpt-oss-20b`, `confirmation-reply.ts` prompt, temperature 0, low reasoning effort |
| jev | TypeSafe | `jev-1.13.0`, one `choice` question (`confirmation-reply-decision.ts`) |
| clef | Cloudflare Workers AI | `@cf/cloudflare/clef`, same question |
| clef-flash | Cloudflare Workers AI | `@cf/cloudflare/clef-flash`, same question |

The eval no longer runs jev; its rows are from the published run. Check out the frozen commit below to reproduce them.

## Splits

- **Development**: `cases.json`, 115 replies. The Groq prompt and the word
  list were tuned on it before this protocol was written, so results on it
  favour Groq and are not reported as headline numbers.
- **Held-out**: `cases.held-out.json`, 305 replies written for this
  benchmark, checked to share no reply with the development set (ignoring
  case and punctuation). Nothing is tuned on it.

## Frozen

At commit `2b90a6ef` plus the benchmark adapter: the word list and veto
(`voice-confirmation.ts`), the Groq prompt, the decision-model question, and
the held-out cases and labels.

## Thresholds

Groq returns a label. The three decision models return probabilities, and
code takes approve or reject only when it is the most likely reading and at
least the threshold, otherwise unclear. Each model's threshold is chosen on
the **development** runs only, as the smallest value in
{0.50, 0.60, 0.70, 0.80, 0.85, 0.90, 0.95, 0.99} at which

1. the full pipeline makes no wrong decision, and
2. the model alone makes no false approval (approve where the expected
   reading is anything else).

If none qualifies, the threshold is 0.99. The chosen threshold is then used
unchanged on the held-out runs.

## Runs

- Development: 5 runs per reply per model.
- Held-out, written replies: 5 runs per reply per model.
- Held-out, spoken replies: each reply spoken once with `gpt-4o-mini-tts`
  (voice `alloy`) and transcribed once with `gpt-4o-mini-transcribe`, the
  model the agent session uses for the user's speech; 3 runs per transcript
  per model.

## Scoring

Each reading is compared with the case's expected reading:

- **pass**: they match.
- **miss**: a real approve or reject read as unclear. Safe; costs a tap.
- **wrong**: anything that decides against the expectation, including any
  approve or reject of a reply that should be unclear.

The full pipeline is scored against `expect`. A model alone is scored against
`alone` where a case gives one (mixed replies such as "Yeah, no, don't",
which a person reads as no but the pipeline refuses by design), and against
`expect` otherwise.

## Not measured

Latency: the providers were called from a laptop over their public APIs,
while production calls run from a Cloudflare Worker, where Workers AI runs on
the same network. Client-side timings would misstate that, so none are
reported.

## Round 2: new decision models (2026-10-10)

The same question, splits, threshold rule, runs and scoring, applied to three
decision models released after round 1. Written before any held-out run of
them.

| Name | Provider | Model |
|---|---|---|
| clef-omni | Cloudflare Workers AI | `@cf/cloudflare/clef-omni` (announced 2026-10-09), same question |
| perplexity-decisions | Perplexity | `pplx-decider-v1.1-27b`, `POST https://api.perplexity.ai/v1/decisions`, same question in the same shape |
| openai-decisions | OpenAI | `gpt-6-luna`, `POST https://api.openai.com/v1/decisions` (public beta), same question with its two instruction sentences joined into one string and the criteria sent as `choices` |

What differs from round 1, and why it is not a like-for-like pipeline
comparison with the published round 1 pipeline rows:

- **The word list and veto are the current ones** (after PR #392), not the
  round 1 freeze. Round 1 found the pipeline's failures; #392 fixed them
  against the development and two validation sets. The model-alone rows do
  not depend on the word list, so they compare directly with round 1. The
  pipeline rows do, so round 1's groq, clef and clef-flash were rerun on the
  held-out set under the current pipeline in this round; their round 1
  pipeline rows are superseded for comparison purposes, not corrected.
- **Jev is not rerun.** Its adapter was removed after round 1. Its round 1
  model-alone rows stand; it has no round 2 pipeline row.
- **Held-out transcripts are reused**, not re-spoken, so the spoken track
  hears exactly what round 1 heard.
- **Perplexity ran one case at a time** (10 requests a second per
  organisation). Timings are still not reported.
- **Development-set runs of the new models** were first made at the default
  threshold (0.90) to choose each model's threshold from the recorded
  probabilities, then rerun at the chosen threshold for the reported
  development rows. Both runs are kept.

Thresholds chosen on the development set by the round 1 rule:
clef-omni 0.95, perplexity-decisions 0.95, openai-decisions 0.50.
