> ## Documentation Index
> Fetch the complete documentation index at: https://docs.stet.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Model comparison

> Compare a new model or reasoning setting on a matched frozen task slice

Use this flagship workflow to make a corpus-scoped routing choice, not a claim
about models everywhere. Keep the task slice and harness stable enough to
attribute a difference.

## Say this

```text wrap theme={null}
Use the Stet skill to compare a new model arm on this corpus. Reuse existing
evidence where valid, smoke before new spend, and run only the new arm on the
same task slice. Confirm the plan with me before model-token spend.
```

## What happens

Your agent should smoke the new configuration before investing in a full arm.
Then it runs only that new arm on the identical retained slice, freezes the
incumbent once as a named baseline, and compares the candidate against that
frozen baseline. Have it state expected time and spend before launch. The
frozen identity prevents hand-pairing directories or silently changing the
reference.

If compatible evidence is split across runs, have your agent combine it into
one durable result rather than hand-editing anything; combining and reanalyzing
existing artifacts spends nothing new. A reasoning-effort comparison is the
same shape: hold the model fixed and name the reasoning variant as the
treatment.

| Request                                          | Model-token spend    |
| ------------------------------------------------ | -------------------- |
| Smoke or run a new arm                           | Yes; approve first   |
| Freeze an existing incumbent                     | No new model attempt |
| Compare or combine existing compatible artifacts | No new model attempt |
| Reanalyze existing arms versus the field         | No new model attempt |
| Relaunch a changed arm or slice                  | Yes; approve first   |

Results stay scoped to the corpus, slice, harness, and time of the evidence.
Incomplete, stale, contaminated, or inspect-only evidence cannot support a
default or superiority claim.

## What a good answer sounds like

The report should name the frozen baseline, the candidate, how many tasks
counted, how strong each metric's evidence is, cost with its outliers, and any
confounds — every value from the actual Trial Result.

Ask: "Which metric is decision-grade?" "Were all cells retained?" "What does
this corpus-scoped result not establish?"

```text wrap theme={null}
Read the canonical Trial Result and compare evidence. Report the frozen baseline,
candidate, denominator, per-metric calibration, cost, validity, confounds, and
one routing action. State what this corpus-scoped evidence does not establish.
```
