Say this
What happens
Your agent should inspect the baseline before launch rather than assuming that a missing-looking file means absence, and state expected time and spend. It keeps the task slice and behavior contract matched, then uses theskill_workbench
dimensions: skill_routing, skill_actionability, skill_specificity,
skill_command_exactness, skill_tool_selection, and skill_regression_risk.
A skill revision is not proven by tests alone. Coverage, replay validity, and the
correct baseline determine whether the result can support keeping the change or
only inspecting it.
What a good answer sounds like
The report should name the baseline actually used, the candidate skill, how many tasks counted, eachskill_workbench dimension’s result, and any regressions —
every value from the actual Trial Result.
Ask: “Was the baseline truly absent or committed?” “Which dimensions regressed?”
“Is coverage complete enough for this routing action?”