Scenarios

The cohort was other agents

A rewritten support-triage skill for the in-product agent fleet. Outcome: fleet-wide · 0 broken sessions. 5 of 5 held · legacy never broke

agent-skill-eval-gate

The cohort was other agents

A rewritten support-triage skill for the in-product agent fleet.

FLEET-WIDE · 0 BREAKS
The plan — drafted before anyone saw it
Strategy New version must beat its eval baseline before live traffic
Control points Versioned control point serving v2/v3 by client-model cohort
Gates eval suite ≥ baseline · zero schema-break sessions
Guardrails task success · tool-call errors · session length
Rollback task-success dip → cohort auto-pins to the previous skill
Ramp — planned outline, actual fill
evals
new
50%
fleet
The run — from the decision log
T+0m read: skill v3 changes the tool-call sequence — the consumers are agents, not humansheld
T+1d gated: offline evals beat the v2 baseline before any trafficheld
T+1d chose: versioned control point; cohorts built from client model idsheld
T+2d ramped: v3 to the newest clients first; legacy models stayed pinned to v2held
T+4d saw: task success +9%, tool-call errors flat — promoted fleet-wideheld
5 of 5 held · legacy never broke

learnedeval-gate + model-cohort saved as the template for every skill release.