> ## Documentation Index
> Fetch the complete documentation index at: https://docs.stet.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Skill evaluation

> Measure whether a repository-managed skill improves coding-agent behavior

Use this workflow for a new skill or a revision. The baseline must represent the
question honestly: a correct absent baseline for a new skill, or the committed
skill for a revision.

## Say this

```text wrap theme={null}
Use the Stet skill to evaluate this repository-managed skill. Choose an actually
absent baseline for a new skill, or the committed version for a revision. Keep
the repository tasks and verifier matched. Confirm the plan with me before spend.
```

## What happens

Your agent should inspect the baseline before launch rather than assuming that a
missing-looking file means absence, and state expected time and spend. It keeps
the task slice and behavior contract matched, then uses the `skill_workbench`
dimensions: `skill_routing`, `skill_actionability`, `skill_specificity`,
`skill_command_exactness`, `skill_tool_selection`, and `skill_regression_risk`.

A skill revision is not proven by tests alone. Coverage, replay validity, and the
correct baseline determine whether the result can support keeping the change or
only inspecting it.

## What a good answer sounds like

The report should name the baseline actually used, the candidate skill, how many
tasks counted, each `skill_workbench` dimension's result, and any regressions —
every value from the actual Trial Result.

Ask: "Was the baseline truly absent or committed?" "Which dimensions regressed?"
"Is coverage complete enough for this routing action?"

```text wrap theme={null}
Read the canonical Trial Result for this skill evaluation. Name the baseline
provenance, candidate, denominator, skill_workbench coverage, regressions,
calibration, observed spend, and one next action. Do not infer promotion from
missing or inspect-only evidence.
```
