evaluation-benchmarks.md
Evaluation Benchmarks: Provenance and How the Skill Uses Them
Buyer expectations (enterprise AI, 2026)
- 70% of enterprise buyers prioritize speed of deployment in vendor selection; 57% expect POC ROI within 3 months and 11% expect it immediately; vendor evaluation is increasingly one-shot with little tolerance for a failed first test (a16z Enterprise Survey, 2026). The skill uses these to justify the hard clock and the midpoint readout: evaluation velocity is now table stakes, not a differentiator.
Trial-to-paid conversion by motion (2025-2026 aggregations)
- Opt-in self-serve trials: 8-22%, median ~14%. Opt-out (card required): 35-55%, median ~44%. Sales-assisted trials/POCs: 35-70%, median ~55% (industry trial-benchmark aggregations published 2025-2026, compiling self-reported SaaS cohort data). These are blog-tier aggregations, not audited surveys; the skill uses them only for two robust conclusions that hold across every source: the ranges differ by motion far more than by execution quality, and cross-motion comparison is meaningless. Calibrate targets against your own motion's cohort within two quarters.
Activation as the dominant variable
- Activation explains an estimated 60-75% of trial conversion variance; activated trials convert at 35-65% versus 2-8% for un-activated (same 2025-2026 aggregations). Directionally consistent across sources even where exact splits differ; the skill treats the direction (activation dominates) as reliable and the exact percentages as indicative.
Practice-based rules (no external study)
- Default clocks (14 days self-serve, 30 days assisted) and the earned-extension rule.
- The stall thresholds (no activation by day 3 self-serve, or by the first working session assisted).
- The 30-day expired-trial cold marker and warm-outreach suppression.
- The rule that a successful POC's commercial terms are agreed before the POC starts.
Replace with your own cohort numbers once two quarters of instrumented evaluations exist.