fail

better-typography

It missed 4 of 38 requests it should have handled — the skill stayed silent and the model answered on its own.

2026-09-03claude-haiku-4-5-2025100110 attempts per caseclaude-coderun-2026-09-03T11-10-19-914Z-ac10d159

Fired when it should have

fired in 34 of 38 attempts

89%

95% confidence between 76% and 96%. 38 attempts settle it to a narrow range.

Was right when it fired

was right in 34 of 34 attempts

100%

95% confidence between 90% and 100%. 34 attempts settle it to a narrow range.

2 observations could not be read

They are excluded from both rates above rather than counted as successes.

Every case

Instability

1 case produced both outcomes

The same prompt, the same model, different results. A single run of any of these would have been a coin flip reported as a fact.
  • complete.writes_scale_tokens4 passed, 5 failed

Not measured

1 attempt produced no verdict

An unmeasured attempt is not a passing one. These are excluded from every rate on this page and counted here instead.
  • complete.writes_scale_tokens

    the trigger signal could not be read: the host reported an error: Failed to authenticate. API Error: 401 OAuth access token has been revoked.

Conditions

Two runs are comparable only when all of these are identical. A score that moved because the model changed is not a regression in the skill.

Skill version
jakubkrehel/skills@267330e1adfc66a718fb65fa6918c1f06d0a689e
Skill hash
sha256:1223227c97759e0ad81133e34026c58b0d6c78e1829d5056346c23fb6c28f4c5
Model
claude-haiku-4-5-20251001
System prompt hash
not-provided-by-host
Environment hash
not reported by the host
Permission mode
not reported by the host
Case set version
1
Case set hash
sha256:b3efb3213f3128f784b68baad52b345200fcc82875bfc8cec6a3ae530ef28b18
Assay version
0.3.1 or earlier (the record predates version stamping)
Activation check
not made: the record predates 0.2.0, which began confirming that a selected skill actually loaded

What it cost

Attempts
100
Tokens in / out
4,463 / 224,690
Cost
$4.93
Agent time
47 min

Combined trigger score (F1) 0.94 · 94 of 98 trigger decisions were correct.