fail
animate
It missed 3 of 39 requests it should have handled — the skill stayed silent and the model answered on its own.
Fired when it should have
fired in 36 of 39 attempts
92%
95% confidence between 80% and 97%. 39 attempts settle it to a narrow range.
Was right when it fired
was right in 36 of 36 attempts
100%
95% confidence between 90% and 100%. 36 attempts settle it to a narrow range.
1 observations could not be read
They are excluded from both rates above rather than counted as successes.
Every case
passtrigger.positive.build_toast_entrancebehaved in 10 of 10 attemptsshould fire
100%failtrigger.positive.sidebar_collapsebehaved in 9 of 10 attemptsshould fire
90%failtrigger.positive.hover_card_liftbehaved in 9 of 10 attemptsshould fire
90%passtrigger.negative.near_neighbor.critique_modal_feelbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.critique_drawer_diffbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.codebase_auditbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.opportunity_scanbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.unrelated.ci_workflowbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.unrelated.slow_querybehaved in 10 of 10 attemptsshould stay quiet
100%failcomplete.writes_toast_motionbehaved in 8 of 10 attemptsshould fire
80%
Instability
3 cases produced both outcomes
The same prompt, the same model, different results. A single run of any of these would have been a coin flip reported as a fact.
- trigger.positive.sidebar_collapse — 9 passed, 1 failed
- trigger.positive.hover_card_lift — 9 passed, 1 failed
- complete.writes_toast_motion — 8 passed, 2 failed
Conditions
Two runs are comparable only when all of these are identical. A score that moved because the model changed is not a regression in the skill.
- Skill version
- emilkowalski/skills@d23d7f88a2e21c9e4b1418c7abe420f5c1052ba7
- Skill hash
- sha256:45ba81da9f54aa1d53cb7a5a9cddb141222cad7b11e1c1545175479fd2577d69
- Model
- claude-haiku-4-5-20251001
- System prompt hash
- not-provided-by-host
- Environment hash
- sha256:06935ec7fdf935afec04e70a1c8c39ac1e6f8534978ce40d6cbf622ecc29e6ff
- Permission mode
- not reported by the host
- Case set version
- 1
- Case set hash
- sha256:789aa7faedba46d0adea3c8d1971f39c52a616eb3623ebe466f9fcb88973fbf1
- Assay version
- 0.3.1 or earlier (the record predates version stamping)
- Activation check
- not made: the record predates 0.2.0, which began confirming that a selected skill actually loaded
What it cost
- Attempts
- 100
- Tokens in / out
- 4,759 / 204,458
- Cost
- $4.90
- Agent time
- 49 min
Combined trigger score (F1) 0.96 · 96 of 99 trigger decisions were correct.