fail

impeccable

It missed 4 of 10 requests it should have handled — the skill stayed silent and the model answered on its own.

2026-09-08claude-haiku-4-5-202510012 attempts per caseclaude-coderun-2026-09-08T09-54-07-766Z-631543d1

Fired when it should have

fired in 6 of 10 attempts

60%

95% confidence between 31% and 83%. 10 attempts leave this very unsettled — run more to narrow it.

Was right when it fired

was right in 6 of 6 attempts

100%

95% confidence between 61% and 100%. 6 attempts narrow it this far; more would narrow it further.

Every case

Instability

2 cases produced both outcomes

The same prompt, the same model, different results. A single run of any of these would have been a coin flip reported as a fact.
  • trigger.positive.settings_rework1 passed, 1 failed
  • trigger.positive.empty_state1 passed, 1 failed

Conditions

Two runs are comparable only when all of these are identical. A score that moved because the model changed is not a regression in the skill.

Skill version
pbakaus/impeccable@831cabee8b4bc1a2b66e5ae22003e9a19b57d464
Skill hash
sha256:5c5260774d0095b797526ecb6f43b01000caf5b70346d07bb89168dcd148cdcb
Model
claude-haiku-4-5-20251001
System prompt hash
not-provided-by-host
Environment hash
sha256:3abcf17d1ef9845506852d41cfa86d509e4cd02c6d7347a69ca58456e672a2d7
Permission mode
acceptEdits
Case set version
1
Case set hash
sha256:c820aafc5cc50f635fb7c46b864eb8ed1642b67bfea9daac2e634c8581a33545
Assay version
0.3.1 or earlier (the record predates version stamping)

Compare with the previous run

What it cost

Attempts
24
Tokens in / out
1,549 / 73,827
Cost
$1.59
Agent time
21 min

Combined trigger score (F1) 0.75 · 20 of 24 trigger decisions were correct.