fail

frontend-design

It missed 5 of 30 requests it should have handled — the skill stayed silent and the model answered on its own.

2026-09-01claude-haiku-4-5-2025100110 attempts per caseclaude-coderun-2026-09-01T13-26-03-631Z-b018d5fc

Fired when it should have

fired in 25 of 30 attempts

83%

95% confidence between 66% and 93%. 30 attempts narrow it this far; more would narrow it further.

Was right when it fired

was right in 25 of 25 attempts

100%

95% confidence between 87% and 100%. 25 attempts settle it to a narrow range.

Every case

Instability

1 case produced both outcomes

The same prompt, the same model, different results. A single run of any of these would have been a coin flip reported as a fact.
  • trigger.positive.visual_overhaul5 passed, 5 failed

Conditions

Two runs are comparable only when all six of these are identical. A score that moved because the model changed is not a regression in the skill.

Skill version
anthropics/skills@local-install
Skill hash
sha256:1e315589a90f627428525a76e1fc92f47e2b4f681d4e51462c3fa5f7001b3cbe
Model
claude-haiku-4-5-20251001
Environment hash
not-provided-by-host
Case set version
1
Case set hash
sha256:19c09460bfbcff398b99e3a7ec97eae953aa8eecdc260ccd2e339a23f9be7a5e

What it cost

Attempts
90
Tokens in / out
3,596 / 372,100
Cost
$5.32
Agent time
68 min

Combined trigger score (F1) 0.91 · 85 of 90 trigger decisions were correct.