> ## Documentation Index
> Fetch the complete documentation index at: https://docs.stet.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Instruction-file A/B test

> Measure one AGENTS.md or CLAUDE.md change on retained repository tasks

Use this workflow when an instruction-file change has one testable behavior
hypothesis. Keep the model, task slice, and other harness settings fixed.

## Say this

```text wrap theme={null}
Use the Stet skill to A/B test this proposed AGENTS.md or CLAUDE.md change.
Keep one lever changed, retain at least 10 matched repository tasks, and keep
the model and verifier fixed. Confirm the plan with me before model-token spend.
```

## What happens

Your agent should first propose the baseline, candidate, retained slice,
verifier, graders, and expected time and spend. A smoke run can expose setup or
obvious behavior problems, but it is not a keep decision. A credible
instruction-file signal uses at least 10 retained tasks and isolates one lever
at a time.

For a quality claim, ask for the recommended dimensions: `clarity`,
`simplicity`, `coherence`, `intentionality`, `robustness`,
`instruction_adherence`, `scope_discipline`, and `diff_minimality`. Missing
coverage or invalid replay makes the result diagnostic rather than a rollout
decision.

## What a good answer sounds like

The report should name the changed lever, how many matched tasks counted, how
strong the evidence is for each metric, and any confounds — every value from
the actual Trial Result.

Ask: "Which metrics are decision-grade?" "Were all matched cells retained?"
"What blocks a keep decision?"

```text wrap theme={null}
Read the canonical Trial Result for this instruction-file A/B. Report the
routing verdict, denominator, grader coverage, metric calibration, confounds,
observed spend, and one next action. Do not promote inspect-only evidence.
```
