fail

marketing-skills:product-marketing

It missed 6 of 6 requests it should have handled — the skill stayed silent and the model answered on its own.

2026-09-11claude-haiku-4-5-202510013 attempts per caseclaude-coderun-2026-09-11T14-54-16-671Z-912ad216

Fast mode — an early warning, not evidence

Only the trigger layer was measured. Declared assertions were not evaluated, and they are not counted as unknown.

At 3 attempts per case the intervals are wide by construction. Run without --fast before trusting a green result.

Collision matrix

Rows are the skill each case expects to win, columns the skill that fired first. Winning means firing first; a skill that fired later is counted under “also fired”. Names are shown without their common prefix marketing-skills:; not under it: run.

expected · fired firstwonnonesignupcropopupspaywallsonboardingcopywritingcopy-editingemailscold-emailseo-auditai-seoprogrammatic-seoschemaproduct-marketingrun
signup1 case0% 0/395% CI 0%–56%2··············1
cro2 cases33% 2/695% CI 10%–70%4·2·············
popups2 cases0% 0/695% CI 0%–39%6···············
paywalls1 case0% 0/395% CI 0%–56%3···············
onboarding1 case0% 0/395% CI 0%–56%3···············
copywriting / copy-editing1 case0% 0/395% CI 0%–56%3···············
copy-editing1 case0% 0/395% CI 0%–56%3···············
emails1 case67% 2/395% CI 21%–94%1·······2·······
cold-email1 case67% 2/395% CI 21%–94%1········2······
seo-audit1 case100% 3/395% CI 44%–100%··········3·····
ai-seo1 case100% 3/395% CI 44%–100%···········3····
programmatic-seo1 case0% 0/395% CI 0%–56%3···············
schema1 case100% 3/395% CI 44%–100%·············3··
product-marketing2 cases0% 0/695% CI 0%–39%6···············
none3 cases100% 9/995% CI 70%–100%9···············

Fired when it should have · target skill only

fired in 0 of 6 attempts

0%

95% confidence between 0% and 39%. 6 attempts narrow it this far; more would narrow it further.

Was right when it fired · target skill only

never fired

The skill did not fire in any attempt that was read, so there is nothing it could have been right or wrong about.

Every case

Instability

3 cases produced both outcomes

The same prompt, the same model, different results. A single run of any of these would have been a coin flip reported as a fact.
  • collide.cro.lead_form2 passed, 1 failed
  • collide.emails.welcome_sequence2 passed, 1 failed
  • collide.cold_email.outreach2 passed, 1 failed

Conditions

Two runs are comparable only when all of these are identical. A score that moved because the model changed is not a regression in the skill.

Skill version
coreyhaines31/marketingskills@5b2c0007766c6a1cf1d53fd8fc73e979e0821022
Skill hash
sha256:aff03848f835c2d0eb755346e1c1b13d6cbad75d63f21b2cfdb673bee910fb6e
Model
claude-haiku-4-5-20251001
System prompt hash
not-provided-by-host
Environment hash
sha256:cb6c8490e46dd898f7643f0668035ed1bbcb5380247d6cfec192a1a5e6519f05
Permission mode
acceptEdits
Case set version
3
Case set hash
sha256:ee4ae6435aa016c3823c65dd63a084c5bc7e56b3d26e411ea788dd3dac1a0d59
Assay version
0.4.3

What it cost

Attempts
60
Tokens in / out
2,799 / 135,395
Cost
$3.04
Agent time
28 min

Combined trigger score (F1) not measurable · 9 of 15 trigger decisions were correct.