Reference file

Evaluation benchmarks

evaluation-benchmarks.md

Evaluation Benchmarks: Provenance and How the Skill Uses Them

Buyer expectations (enterprise AI, 2026)

  • 70% of enterprise buyers prioritize speed of deployment in vendor selection; 57% expect POC ROI within 3 months and 11% expect it immediately; vendor evaluation is increasingly one-shot with little tolerance for a failed first test (a16z Enterprise Survey, 2026). The skill uses these to justify the hard clock and the midpoint readout: evaluation velocity is now table stakes, not a differentiator.

Trial-to-paid conversion by motion (2025-2026 aggregations)

  • Opt-in self-serve trials: 8-22%, median ~14%. Opt-out (card required): 35-55%, median ~44%. Sales-assisted trials/POCs: 35-70%, median ~55% (industry trial-benchmark aggregations published 2025-2026, compiling self-reported SaaS cohort data). These are blog-tier aggregations, not audited surveys; the skill uses them only for two robust conclusions that hold across every source: the ranges differ by motion far more than by execution quality, and cross-motion comparison is meaningless. Calibrate targets against your own motion's cohort within two quarters.

Activation as the dominant variable

  • Activation explains an estimated 60-75% of trial conversion variance; activated trials convert at 35-65% versus 2-8% for un-activated (same 2025-2026 aggregations). Directionally consistent across sources even where exact splits differ; the skill treats the direction (activation dominates) as reliable and the exact percentages as indicative.

Practice-based rules (no external study)

  • Default clocks (14 days self-serve, 30 days assisted) and the earned-extension rule.
  • The stall thresholds (no activation by day 3 self-serve, or by the first working session assisted).
  • The 30-day expired-trial cold marker and warm-outreach suppression.
  • The rule that a successful POC's commercial terms are agreed before the POC starts.

Replace with your own cohort numbers once two quarters of instrumented evaluations exist.