fail
frontend-design
It missed 5 of 30 requests it should have handled — the skill stayed silent and the model answered on its own.
Fired when it should have
fired in 25 of 30 attempts
83%
95% confidence between 66% and 93%. 30 attempts narrow it this far; more would narrow it further.
Was right when it fired
was right in 25 of 25 attempts
100%
95% confidence between 87% and 100%. 25 attempts settle it to a narrow range.
Every case
passtrigger.positive.new_pagebehaved in 10 of 10 attemptsshould fire
100%failtrigger.positive.visual_overhaulbehaved in 5 of 10 attemptsshould fire
50%passtrigger.positive.new_screenbehaved in 10 of 10 attemptsshould fire
100%passtrigger.negative.near_neighbor.style_tweakbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.library_setupbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.component_testbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.refactorbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.unrelated_dbbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.unrelated_conceptbehaved in 10 of 10 attemptsshould stay quiet
100%
Instability
1 case produced both outcomes
The same prompt, the same model, different results. A single run of any of these would have been a coin flip reported as a fact.
- trigger.positive.visual_overhaul — 5 passed, 5 failed
Conditions
Two runs are comparable only when all six of these are identical. A score that moved because the model changed is not a regression in the skill.
- Skill version
- anthropics/skills@local-install
- Skill hash
- sha256:1e315589a90f627428525a76e1fc92f47e2b4f681d4e51462c3fa5f7001b3cbe
- Model
- claude-haiku-4-5-20251001
- Environment hash
- not-provided-by-host
- Case set version
- 1
- Case set hash
- sha256:19c09460bfbcff398b99e3a7ec97eae953aa8eecdc260ccd2e339a23f9be7a5e
What it cost
- Attempts
- 90
- Tokens in / out
- 3,596 / 372,100
- Cost
- $5.32
- Agent time
- 68 min
Combined trigger score (F1) 0.91 · 85 of 90 trigger decisions were correct.