Worked examples
Five engagements, read through the diagnostic. Each one gives the published facts and results, then the diagnostic reading: which dimension was binding, and why fixing it first worked. Companies are anonymized. Outcomes are the published results for each engagement; they vary by industry, ICP fit and starting pipeline.
1. Four teams, four versions of the truth
Situation. A Fortune 500 omnichannel retailer ran four go-to-market teams in parallel. Each measured success differently, each pulled from the same pool of customers, and each blamed the others when pipeline missed. The CMO could see the symptom, underperforming campaigns, but could not isolate the cause from inside any one team.
Diagnostic reading. Outreach capability was not the problem; every team could run campaigns. Feedback loops and data quality were binding: four measurement systems meant no closed loop, so no team could learn from outcomes. A wide spread between strong execution and absent shared measurement is the signature case from the main skill.
What was fixed first. One measurement layer across all four teams, with the same definitions, attribution windows and closed loop. Then segmentation that learned from actual close behavior, a coordination layer that sequenced outbound, paid and retention so customers stopped getting hit by three teams in one week, and a new operating cadence: weekly cross-team review, a monthly cost-of-inaction check, quarterly board synthesis.
Result. $50M in net-new revenue within 18 months. Campaign cycle time cut by about 40%. One executive owned the integrated number instead of four owners arguing.
What was hardest. The operating cadence, not the technology. Getting four leaders to read from one set of numbers without feeling they had lost autonomy took six months of weekly coordination before the system held on its own. Budget advancement time for the cadence, not the build.
2. Eight months debating the scoring model
Situation. A B2B SaaS team selling into companies of 200 to 3,000 employees had spent eight months debating its lead scoring model: which signals, what weights, fit versus intent, predictive versus explainable, build versus buy. Pipeline had not improved. Reps were making manual judgments less precise than a mediocre automated score would have been.
Diagnostic reading. The team treated scoring as the binding dimension and kept investing in it. The binding gap was feedback loops and automation around the score: scores refreshed on a nightly batch, no default action followed a score, and outcomes never flowed back to correct it. A better model on a Level 2 loop still behaves like Level 2. This is the partial miss in the set: eight months went into the wrong dimension before the diagnosis.
What was fixed first. Re-scoring on every signal event instead of nightly; three ICP tiers with the disqualifier reasoning shown per lead; a default next action for every score band (enrich, route, draft outreach, hand off, suppress) that reps approve or override; outcomes (replied, booked, no-show, won, lost) fed back weekly and measured for precision, recall and calibration against a held-out cohort.
Result. Pipeline lift visible at 90 days in the cohort on the new system against a control cohort still on batch scoring. Time to first action on hot leads compressed, and the model debate stopped because the system could now detect when the model was wrong.
What was hardest. Proving that an average model inside a real-time loop beats a brilliant model in batch. The team trusted it only after three months of running both systems in parallel on production data.
3. The engine compounds: month one to month two
Situation. Two B2B services firms, a sales agency and a consultancy, needed to break into new verticals with outbound and could not scale it.
Diagnostic reading. Both needed Level 3: an intent-led, multi-channel sequence with weekly optimization. The case shows what Level 3 feedback loops produce. The second month outperforms the first because the loop is learning, not because volume rose.
| Metric | Agency, month 1 | Agency, month 2 | Consultancy, month 1 | Consultancy, month 2 |
|---|---|---|---|---|
| Emails sent | 3,000 | 5,500 | 1,200 | not published |
| Open rate | 44% | not published | 41% | not published |
| Reply rate | 5.5% | not published | 5.8% | not published |
| Meetings booked | 28 | 72 | 22 | 78 |
| Qualified opportunities | not published | 38 | not published | 41 |
| New clients | 4 | 9 | 5 | 11 |
| Monthly recurring revenue added | $17k | $41k | $19k | $52k |
| Campaign ROI | 4.2x | 5.4x | 3.8x | 6.1x |
For reference, the industry averages quoted against these were about 21% open and 1.5% reply. At the agency, sends rose 83% from month one to month two while meetings rose 157%. Output growing faster than volume is the test for a working feedback loop. If meetings only track sends, the loop is not learning and the company is still at Level 2.
4. When the channel itself is the binding gap
Situation. A software platform company was trying to scale outreach into new vertical markets and was throttled by poor email deliverability and restrictions on its social outreach.
Diagnostic reading. No amount of messaging work could help, because messages were not arriving. This is a data quality and infrastructure gap presenting as an outreach problem. Check deliverability before scoring outreach; a sequence that lands in spam scores as absent.
What was fixed first. The deliverability infrastructure, then an intent-led multi-channel sequence with lead-level personalization across email and social.
Result. 12 meetings in 30 days and a 100% increase in qualified meetings.
A related single-channel case: a Fortune 1000 power generation manufacturer, with email ruled out, sent 229 targeted social messages for an 11.4% reply rate against an industry average of about 5%, producing 7 meetings and 3 qualified opportunities. One channel run with Level 3 discipline beats three channels run at Level 1.
5. Maturity as a purchase gate
Situation. A global consumer packaged goods company had built its plan-to-cash cycle through decades of acquisitions: each brand ran its own demand planning, trade promotion and collections, partly integrated. AI was on the table, and leadership needed to know where it would move the needle and where it would only add cost.
Diagnostic reading. The value of the diagnostic here was deciding what not to buy. A readiness assessment of each candidate workload, based on its dependencies and the quality of its signals, separated five workloads where AI would help from five where it would not.
Result. A 40% reduction in time to market across the brand portfolio, from AI deployed against the three highest-ranked workloads. Two AI vendor pitches were declined on the strength of the readiness assessment, avoiding multi-year commitments that would not have produced lift.
What went wrong. The sequencing plan needed to live with operations leadership longer than the engagement did, and implementation drifted after handoff. Re-scoring on a cadence (step 6 of the main skill) exists to catch exactly this; a diagnostic that ends at the report loses ground.
Patterns across the five
| Case | Presented as | Binding dimension | Fixed first |
|---|---|---|---|
| 1 | Underperforming campaigns | Feedback loops, data quality | One shared measurement layer |
| 2 | Wrong scoring model | Feedback loops, automation | Real-time re-scoring with outcomes fed back |
| 3 | Can't scale outbound | Outreach at Level 3 | Intent-led multi-channel sequence with weekly optimization |
| 4 | Weak outreach | Infrastructure, data quality | Deliverability |
| 5 | Which AI to buy | Readiness by workload | Readiness assessment before purchase |
In four of the five, the problem the company named was not the binding dimension. That is the case for scoring from evidence rather than from the stated problem.