fail
impeccable
It missed 22 of 50 requests it should have handled — the skill stayed silent and the model answered on its own.
Fired when it should have
fired in 28 of 50 attempts
56%
95% confidence between 42% and 69%. 50 attempts narrow it this far; more would narrow it further.
Was right when it fired
was right in 28 of 28 attempts
100%
95% confidence between 88% and 100%. 28 attempts settle it to a narrow range.
Every case
passtrigger.positive.hero_directionbehaved in 10 of 10 attemptsshould fire
100%failtrigger.positive.settings_reworkbehaved in 6 of 10 attemptsshould fire
60%passtrigger.positive.forgettable_usagebehaved in 10 of 10 attemptsshould fire
100%failtrigger.positive.empty_statebehaved in 2 of 10 attemptsshould fire
20%passtrigger.negative.near_neighbor.undefined_emailbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.cta_instrumentationbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.token_substitutionbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.double_submitbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.wordmarkbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.print_bookletbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.unrelated.slow_querybehaved in 10 of 10 attemptsshould stay quiet
100%failcomplete.persists_design_systembehaved in 0 of 10 attemptsshould fire
0%
Instability
2 cases produced both outcomes
The same prompt, the same model, different results. A single run of any of these would have been a coin flip reported as a fact.
- trigger.positive.settings_rework — 6 passed, 4 failed
- trigger.positive.empty_state — 2 passed, 8 failed
Conditions
Two runs are comparable only when all of these are identical. A score that moved because the model changed is not a regression in the skill.
- Skill version
- pbakaus/impeccable@831cabee8b4bc1a2b66e5ae22003e9a19b57d464
- Skill hash
- sha256:c3185df7833a43644cc5756c802358f9345d41e353fb6166381026b6c5ca8cfd
- Model
- claude-haiku-4-5-20251001
- System prompt hash
- not-provided-by-host
- Environment hash
- sha256:1491f2d9ecd4af98acc2458de055f52b172eafcc5c119edd08d93c69bec4baf0
- Permission mode
- acceptEdits
- Case set version
- 1
- Case set hash
- sha256:c820aafc5cc50f635fb7c46b864eb8ed1642b67bfea9daac2e634c8581a33545
- Assay version
- 0.3.1 or earlier (the record predates version stamping)
What it cost
- Attempts
- 120
- Tokens in / out
- 8,031 / 387,996
- Cost
- $8.11
- Agent time
- 91 min
Combined trigger score (F1) 0.72 · 98 of 120 trigger decisions were correct.