fail

impeccable

It missed 22 of 50 requests it should have handled — the skill stayed silent and the model answered on its own.

2026-09-06claude-haiku-4-5-2025100110 attempts per caseclaude-coderun-2026-09-06T09-53-14-345Z-c3d2b624

Fired when it should have

fired in 28 of 50 attempts

56%

95% confidence between 42% and 69%. 50 attempts narrow it this far; more would narrow it further.

Was right when it fired

was right in 28 of 28 attempts

100%

95% confidence between 88% and 100%. 28 attempts settle it to a narrow range.

Every case

Instability

2 cases produced both outcomes

The same prompt, the same model, different results. A single run of any of these would have been a coin flip reported as a fact.
  • trigger.positive.settings_rework6 passed, 4 failed
  • trigger.positive.empty_state2 passed, 8 failed

Conditions

Two runs are comparable only when all of these are identical. A score that moved because the model changed is not a regression in the skill.

Skill version
pbakaus/impeccable@831cabee8b4bc1a2b66e5ae22003e9a19b57d464
Skill hash
sha256:c3185df7833a43644cc5756c802358f9345d41e353fb6166381026b6c5ca8cfd
Model
claude-haiku-4-5-20251001
System prompt hash
not-provided-by-host
Environment hash
sha256:1491f2d9ecd4af98acc2458de055f52b172eafcc5c119edd08d93c69bec4baf0
Permission mode
acceptEdits
Case set version
1
Case set hash
sha256:c820aafc5cc50f635fb7c46b864eb8ed1642b67bfea9daac2e634c8581a33545
Assay version
0.3.1 or earlier (the record predates version stamping)

What it cost

Attempts
120
Tokens in / out
8,031 / 387,996
Cost
$8.11
Agent time
91 min

Combined trigger score (F1) 0.72 · 98 of 120 trigger decisions were correct.