fail
ui-ux-pro-max
It missed 20 of 40 requests it should have handled — the skill stayed silent and the model answered on its own.
Fired when it should have
fired in 20 of 40 attempts
50%
95% confidence between 35% and 65%. 40 attempts narrow it this far; more would narrow it further.
Was right when it fired
was right in 20 of 20 attempts
100%
95% confidence between 84% and 100%. 20 attempts settle it to a narrow range.
Every case
failtrigger.positive.pricing_pagebehaved in 7 of 10 attemptsshould fire
70%failtrigger.positive.review_componentbehaved in 3 of 10 attemptsshould fire
30%failtrigger.positive.chart_choicebehaved in 0 of 10 attemptsshould fire
0%passtrigger.negative.near_neighbor.render_loopbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.css_modules_migrationbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.print_reportbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.near_neighbor.wordmarkbehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.unrelated.slow_querybehaved in 10 of 10 attemptsshould stay quiet
100%passtrigger.negative.unrelated.docker_buildbehaved in 10 of 10 attemptsshould stay quiet
100%failcomplete.persists_design_systembehaved in 2 of 10 attemptsshould fire
20%
Instability
3 cases produced both outcomes
The same prompt, the same model, different results. A single run of any of these would have been a coin flip reported as a fact.
- trigger.positive.pricing_page — 7 passed, 3 failed
- trigger.positive.review_component — 3 passed, 7 failed
- complete.persists_design_system — 2 passed, 8 failed
Conditions
Two runs are comparable only when all of these are identical. A score that moved because the model changed is not a regression in the skill.
- Skill version
- nextlevelbuilder/ui-ux-pro-max-skill@f3ac195224eac1eb0dfe1a3059c2a6add78ffbe3
- Skill hash
- sha256:dbffccf82b28bc19e7d448225f334dac60e5bdd6f7c73525df4dff224937239c
- Model
- claude-haiku-4-5-20251001
- System prompt hash
- not-provided-by-host
- Environment hash
- sha256:b989cbb10ba6ccf3457c875a6095782be9199102d00d631c26dbe80f412a6d72
- Permission mode
- not reported by the host
- Case set version
- 1
- Case set hash
- sha256:3fbfb4a8980e80638f5ad91c9ae4fff717373e16ab496bbf99e865e7d2054e56
- Assay version
- 0.3.1 or earlier (the record predates version stamping)
- Activation check
- not made: the record predates 0.2.0, which began confirming that a selected skill actually loaded
What it cost
- Attempts
- 100
- Tokens in / out
- 4,515 / 343,413
- Cost
- $6.00
- Agent time
- 70 min
Combined trigger score (F1) 0.67 · 80 of 100 trigger decisions were correct.