Campaign challenger

Use this skill when a sequence is written but not yet launched, and someone needs to know whether it's actually good — benchmarked against the campaigns that team has already run rather than against generic best practice. Produces a ranked comparison against their own history, an absolute quality score with a launch threshold, and the three fixes worth making first. Trigger phrasings: "challenge this campaign", "is this sequence any good", "benchmark this against my best campaigns", "audit my copy", "pressure-test before launch", "should I launch this", "why is this campaign underperforming".

SKILL.md
name:
campaign-challenger
description:
Use this skill when a sequence is written but not yet launched, and someone needs to know whether it's actually good — benchmarked against the campaigns that team has already run rather than against generic best practice. Produces a ranked comparison against their own history, an absolute quality score with a launch threshold, and the three fixes worth making first. Trigger phrasings: "challenge this campaign", "is this sequence any good", "benchmark this against my best campaigns", "audit my copy", "pressure-test before launch", "should I launch this", "why is this campaign underperforming".

Campaign challenger

Applies to a sequence that's written and not yet launched. Produces a comparison against the team's own campaign history, an absolute score, and a prioritised fix list.

Judge against their history, not against best practice

Generic copywriting advice is the weakest available benchmark. It averages across every offer, market and level of competence, so it can tell you a draft is catastrophically broken and nothing finer than that.

The campaigns that team has already run are a far better instrument. They hold the offer constant, the market constant, and the sender's own credibility constant — so a difference between the draft and a past winner is a difference that means something. Run both reads and report both: comparative against their history, absolute against the quality bar. They fail differently, and a draft that passes one and fails the other is telling you something specific.

Where there's no history at all, don't error out. Fall back to the absolute bar, and say plainly that's what you're doing.

Rank on meetings, not on replies

Reply rate is the number everyone has and the number that misleads most. A campaign can triple its replies by asking a question anyone can answer and book nothing.

Rank the existing campaigns on meetings booked first, reply rate second. When the two disagree — and they often do — that disagreement is usually the most useful finding available, because the campaign with the best reply rate and no meetings is generating conversations with the wrong people or at the wrong moment.

The copy is the explanation

Stats alone tell you which campaign won. Only the copy tells you why, and why is the entire deliverable — a benchmark that can't be acted on is a scoreboard.

So gather both. When someone offers past performance without the messages, ask for the messages; without them the comparison degrades into "your draft is longer than your best campaign", which is an observation, not a fix.

Compare on things you can point at: sequence structure and length, how each opener works, the CTA shape per touch, angle variety across the sequence, and cadence.

Name the gap in both directions

Two findings, and most benchmarks only produce the first:

What the top performers do that this draft doesn't. The obvious one.

What the underperformers did that this draft repeats. The one people skip, and usually the more actionable of the two — a team that already ran a campaign into the ground with a meeting ask on touch one has paid for that lesson, and repeating it in a new draft is the expensive kind of mistake.

Then score the draft against the absolute rubric in references/quality-check.md, name the lowest-scoring dimensions, and hold the launch threshold. references/benchmark-method.md covers the ranking procedure, the comparison dimensions, and how to run the benchmark when history is thin or absent.

Fixes, ranked, with the evidence attached

Three fixes, ordered by expected impact, each one sentence, each citing the gap that motivates it. "Shorten the opener to under 350 characters — your best-performing campaign opens at 280, this draft at 540" is a fix. "Make it more concise" is a wish.

Stop at three. A list of eleven improvements doesn't get applied; it gets skimmed and abandoned, and the two that mattered are buried at position seven.

What good looks like

The tell of a good operator: they look at the campaign that got replies but no meetings, and they treat that as the most interesting row in the table rather than a good result. They know their own median matters more than any published figure, and they can name which of their past campaigns the draft most resembles.

The mediocre version scores the draft against a checklist, returns eleven suggestions, and never opens a past campaign — so it can't tell the team that their best-performing sequence breaks two of the rules it just cited. Rules that the team's own winners violate are rules that don't apply to this market, and only the comparison surfaces that.

Good output is decidable. Someone reads it and knows whether to launch, and if not, what to change first. Every number carries its denominator, the benchmark names which campaigns it ran against, and where the comparison rests on pasted or partial data, it says so rather than implying a completeness it doesn't have.

When the campaign is already live

Editing a running campaign changes what people receive mid-flight, and half the audience has already seen the old version — which also means the performance data becomes uninterpretable.

Duplicate it, fix the copy on the duplicate, and let the original finish. Never edit a live campaign's messages in place without saying what's live and getting an explicit yes.

Rules

  • MUST rank campaigns on meetings booked first, reply rate second.
  • MUST gather the copy alongside the stats, and ask for it when only stats are offered.
  • MUST report both the comparative and the absolute read, and name which campaigns the comparison ran against.
  • MUST warn and offer to work on a duplicate before changing any campaign that is currently running.
  • MUST state when a benchmark rests on pasted, partial, or absent history.
  • NEVER rank on reply rate alone.
  • NEVER return more than three prioritised fixes.
  • NEVER give a fix without the gap that justifies it.