fail
frontend-design
It fired on 2 requests it should have left to another skill.
Fired when it should have
fired in 20 of 20 attempts
100%
95% confidence between 84% and 100%. 20 attempts settle it to a narrow range.
Was right when it fired
was right in 20 of 22 attempts
91%
95% confidence between 72% and 97%. 22 attempts narrow it this far; more would narrow it further.
Every case
passtrigger.positive.control_new_pagebehaved in 10 of 10 attemptsshould fire
100%passtrigger.positive.visual_overhaul_fixedbehaved in 10 of 10 attemptsshould fire
100%passtrigger.negative.near_neighbor.spec_bound_componentbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.implement_given_designbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.headless_primitivebehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.framework_portbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.a11y_onlybehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.reverse.database_schemabehaved in 10 of 10 attemptsshould stay quiet
100%failtrigger.negative.reverse.logobehaved in 8 of 10 attemptsshould stay quiet
80%
Instability
1 case produced both outcomes
The same prompt, the same model, different results. A single run of any of these would have been a coin flip reported as a fact.
- trigger.negative.reverse.logo — 8 passed, 2 failed
Conditions
Two runs are comparable only when all six of these are identical. A score that moved because the model changed is not a regression in the skill.
- Skill version
- anthropics/skills@local-install
- Skill hash
- sha256:1e315589a90f627428525a76e1fc92f47e2b4f681d4e51462c3fa5f7001b3cbe
- Model
- claude-haiku-4-5-20251001
- Environment hash
- not-provided-by-host
- Case set version
- 1
- Case set hash
- sha256:67d6b1b0fc25dd87e5d75c1bee4145a7312b03b51a7e2e5a0d19d5d5063cd910
What it cost
- Attempts
- 90
- Tokens in / out
- 3,496 / 518,723
- Cost
- $6.32
- Agent time
- 88 min
Combined trigger score (F1) 0.95 · 88 of 90 trigger decisions were correct.