fail

ui-ux-pro-max

It missed 20 of 40 requests it should have handled — the skill stayed silent and the model answered on its own.

2026-09-03claude-haiku-4-5-2025100110 attempts per caseclaude-coderun-2026-09-03T16-34-04-660Z-2a900c03

Fired when it should have

fired in 20 of 40 attempts

50%

95% confidence between 35% and 65%. 40 attempts narrow it this far; more would narrow it further.

Was right when it fired

was right in 20 of 20 attempts

100%

95% confidence between 84% and 100%. 20 attempts settle it to a narrow range.

Every case

Instability

3 cases produced both outcomes

The same prompt, the same model, different results. A single run of any of these would have been a coin flip reported as a fact.
  • trigger.positive.pricing_page7 passed, 3 failed
  • trigger.positive.review_component3 passed, 7 failed
  • complete.persists_design_system2 passed, 8 failed

Conditions

Two runs are comparable only when all of these are identical. A score that moved because the model changed is not a regression in the skill.

Skill version
nextlevelbuilder/ui-ux-pro-max-skill@f3ac195224eac1eb0dfe1a3059c2a6add78ffbe3
Skill hash
sha256:dbffccf82b28bc19e7d448225f334dac60e5bdd6f7c73525df4dff224937239c
Model
claude-haiku-4-5-20251001
System prompt hash
not-provided-by-host
Environment hash
sha256:b989cbb10ba6ccf3457c875a6095782be9199102d00d631c26dbe80f412a6d72
Permission mode
not reported by the host
Case set version
1
Case set hash
sha256:3fbfb4a8980e80638f5ad91c9ae4fff717373e16ab496bbf99e865e7d2054e56
Assay version
0.3.1 or earlier (the record predates version stamping)
Activation check
not made: the record predates 0.2.0, which began confirming that a selected skill actually loaded

What it cost

Attempts
100
Tokens in / out
4,515 / 343,413
Cost
$6.00
Agent time
70 min

Combined trigger score (F1) 0.67 · 80 of 100 trigger decisions were correct.