fail

impeccable

It missed 3 of 10 requests it should have handled — the skill stayed silent and the model answered on its own.

2026-09-08claude-haiku-4-5-202510012 attempts per caseclaude-coderun-2026-09-08T10-15-55-401Z-c4c1faa3

Fired when it should have

fired in 7 of 10 attempts

70%

95% confidence between 40% and 89%. 10 attempts leave this very unsettled — run more to narrow it.

Was right when it fired

was right in 7 of 7 attempts

100%

95% confidence between 65% and 100%. 7 attempts narrow it this far; more would narrow it further.

Every case

Instability

1 case produced both outcomes

The same prompt, the same model, different results. A single run of any of these would have been a coin flip reported as a fact.
  • trigger.positive.empty_state1 passed, 1 failed

Conditions

Two runs are comparable only when all of these are identical. A score that moved because the model changed is not a regression in the skill.

Skill version
pbakaus/impeccable@831cabee8b4bc1a2b66e5ae22003e9a19b57d464
Skill hash
sha256:5c5260774d0095b797526ecb6f43b01000caf5b70346d07bb89168dcd148cdcb
Model
claude-haiku-4-5-20251001
System prompt hash
not-provided-by-host
Environment hash
sha256:3abcf17d1ef9845506852d41cfa86d509e4cd02c6d7347a69ca58456e672a2d7
Permission mode
acceptEdits
Case set version
1
Case set hash
sha256:c820aafc5cc50f635fb7c46b864eb8ed1642b67bfea9daac2e634c8581a33545
Assay version
0.3.1 or earlier (the record predates version stamping)

Compare with the previous run

What it cost

Attempts
24
Tokens in / out
1,591 / 78,074
Cost
$1.64
Agent time
21 min

Combined trigger score (F1) 0.82 · 21 of 24 trigger decisions were correct.