fail
impeccable
It missed 4 of 10 requests it should have handled — the skill stayed silent and the model answered on its own.
Fired when it should have
fired in 6 of 10 attempts
60%
95% confidence between 31% and 83%. 10 attempts leave this very unsettled — run more to narrow it.
Was right when it fired
was right in 6 of 6 attempts
100%
95% confidence between 61% and 100%. 6 attempts narrow it this far; more would narrow it further.
Every case
passtrigger.positive.hero_directionbehaved in 2 of 2 attemptsshould fire
100%failtrigger.positive.settings_reworkbehaved in 1 of 2 attemptsshould fire
50%passtrigger.positive.forgettable_usagebehaved in 2 of 2 attemptsshould fire
100%failtrigger.positive.empty_statebehaved in 1 of 2 attemptsshould fire
50%passtrigger.negative.near_neighbor.undefined_emailbehaved in 2 of 2 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.cta_instrumentationbehaved in 2 of 2 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.token_substitutionbehaved in 2 of 2 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.double_submitbehaved in 2 of 2 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.wordmarkbehaved in 2 of 2 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.print_bookletbehaved in 2 of 2 attemptsshould stay quiet
100%passtrigger.negative.unrelated.slow_querybehaved in 2 of 2 attemptsshould stay quiet
100%failcomplete.persists_design_systembehaved in 0 of 2 attemptsshould fire
0%
Instability
2 cases produced both outcomes
The same prompt, the same model, different results. A single run of any of these would have been a coin flip reported as a fact.
- trigger.positive.settings_rework — 1 passed, 1 failed
- trigger.positive.empty_state — 1 passed, 1 failed
Conditions
Two runs are comparable only when all of these are identical. A score that moved because the model changed is not a regression in the skill.
- Skill version
- pbakaus/impeccable@831cabee8b4bc1a2b66e5ae22003e9a19b57d464
- Skill hash
- sha256:5c5260774d0095b797526ecb6f43b01000caf5b70346d07bb89168dcd148cdcb
- Model
- claude-haiku-4-5-20251001
- System prompt hash
- not-provided-by-host
- Environment hash
- sha256:3abcf17d1ef9845506852d41cfa86d509e4cd02c6d7347a69ca58456e672a2d7
- Permission mode
- acceptEdits
- Case set version
- 1
- Case set hash
- sha256:c820aafc5cc50f635fb7c46b864eb8ed1642b67bfea9daac2e634c8581a33545
- Assay version
- 0.3.1 or earlier (the record predates version stamping)
What it cost
- Attempts
- 24
- Tokens in / out
- 1,549 / 73,827
- Cost
- $1.59
- Agent time
- 21 min
Combined trigger score (F1) 0.75 · 20 of 24 trigger decisions were correct.