Scenarios
The cohort was other agents
A rewritten support-triage skill for the in-product agent fleet. Outcome: fleet-wide · 0 broken sessions. 5 of 5 held · legacy never broke
one control point a woven topology
“a one-line bugfix” ●shipped · 41 min“a CLI release to every laptop” ↶re-pinned in 51 s · then shipped“a mobile feature stuck behind store review” ●shipped · one binary“a rewrite that claims it’s 10× faster” ●proven · 2.1M comparisons“a risky change to a feature most orgs ignore” ●verdict · 5 h“a new skill for your agent fleet” ●fleet-wide · 0 broken sessions“a feature that spans API and UI” ●shipped · 3 days“one contract across three services” ●zero drift · one 40-min hold“a feature built on another feature” ●2 ships · 1 graph“the change that looked safe” ↶reversed · 43 s
agent-skill-eval-gate
FLEET-WIDE · 0 BREAKSThe cohort was other agents
A rewritten support-triage skill for the in-product agent fleet.
The plan — drafted before anyone saw it
Strategy New version must beat its eval baseline before live traffic
Control points Versioned control point serving v2/v3 by client-model cohort
Gates eval suite ≥ baseline · zero schema-break sessions
Guardrails task success · tool-call errors · session length
Rollback task-success dip → cohort auto-pins to the previous skill
Ramp — planned outline, actual fill
evals
new
50%
fleet
The run — from the decision log
T+0m read: skill v3 changes the tool-call sequence — the consumers are agents, not humansheld
T+1d gated: offline evals beat the v2 baseline before any trafficheld
T+1d chose: versioned control point; cohorts built from client model idsheld
T+2d ramped: v3 to the newest clients first; legacy models stayed pinned to v2held
T+4d saw: task success +9%, tool-call errors flat — promoted fleet-wideheld
5 of 5 held · legacy never broke
learnedeval-gate + model-cohort saved as the template for every skill release.