Reference file

Benchmark method

benchmark-method.md

Running the benchmark

1. Get the draft

Take the sequence to be challenged — every touch, not just the opener. A benchmark on message one misses the failure mode that kills most sequences, which is four follow-ups that repeat it.

If no draft was supplied, ask for it and stop. There is nothing to do without one.

2. Gather the comparison set

Pull the team's past campaigns with both their performance and their copy. Detect what's available yourself rather than asking someone to describe their own tooling.

Three cases, in descending order of quality:

Full history — stats and copy. The real benchmark. Use it.

Stats only. Ask for the copy of at least the top and bottom performers. Without it the comparison can only produce shape observations, not causes. If it can't be had, say the benchmark is running on structure alone.

No history. Don't error. Fall back to the absolute rubric and the typical performance range for that campaign type, and state plainly that the comparison layer is missing.

A practical note: campaign copy is sometimes not retrievable as clean templates — stored per-variant, per-sender, or only ever materialised as sent messages. Reconstructing the structure from a handful of real sent messages is a valid fallback. The personalisation is already resolved, but structure, angle, length and CTA all survive, and those are what the benchmark reads. Say which path you're on so nobody mistakes a reconstruction for the source.

3. Rank

Meetings booked first. Reply rate second. Where the two disagree, that's a finding — surface it rather than smoothing it.

Always carry the sample size. A 14% reply rate on 40 sends is noise wearing a percentage sign, and a campaign ranked on it will send someone in the wrong direction with confidence.

4. Compare on nameable dimensions

Put the draft next to the ranked campaigns and compare on things that can be pointed at:

Dimension What to read
Sequence length Number of touches, and days end to end
Message length Characters per touch, especially the opener
Opening pattern Question, observation, statement, signal reference
CTA shape per touch And whether any shape repeats consecutively
Angle variety Whether follow-ups add an argument or restate the first
Cadence Channel order and gaps between touches

For each dimension, two readings:

  • What the top performers do that this draft doesn't.
  • What the underperformers did that this draft repeats.

The second is the one that gets skipped and usually the more valuable. A team that already burned a list with a meeting ask on touch one has paid for that lesson; a benchmark that doesn't catch the draft repeating it wasted the history it had access to.

5. Score against the absolute bar

Independently of the comparison, score the draft on the rubric in quality-check.md and hold the launch threshold. Name the lowest-scoring dimensions specifically — a bare number tells nobody what to fix.

Report both reads. A draft that beats the team's history but fails the absolute bar means their history is weak, not that the draft is ready. A draft that passes the absolute bar but sits below their own median means the bar is generic and their market is harder than average. Both are useful; neither is visible from one read alone.

6. Three fixes

Ordered by expected impact. One sentence each. Each cites the gap that motivates it.

Shorten the opener to under 350 characters — your best campaign opens at 280, this draft at 540.

Move the meeting ask off touch one to touch four — the two campaigns that asked on touch one sit bottom of your book on meetings.

Give touch three a new angle — it currently restates touch two, which is the pattern in the campaign that stalled at 2% replies.

Stop at three. Longer lists don't get applied.

A note on rules the winners break

When a team's best-performing campaign violates a rule from the absolute rubric, the rule loses in that market. Say so, rather than flagging the draft for the same thing.

This is the main reason to run the comparative read at all. Generic copywriting rules encode an average, and any team with real history has evidence about their specific market that outranks it.