Promotion evidence
Read this when proposing a level change, and when setting up the record that promotions will be argued from.
A promotion is an argument from a log. If the log does not exist, the answer is no, and the work is to start logging rather than to decide.
Who proposes, who approves
Never the same party, and never the agent itself. The agent produces the evidence, a reviewer proposes the change against these criteria, and the person carrying the consequences approves it. Self-promotion turns the gate into a formality.
What has to be in the window
| Criterion | L0 to L1 | L1 to L2 | L2 to L3 |
|---|---|---|---|
| Volume | 30 items reviewed | 50 items at L1 | continuous |
| Quality | edit or reject rate under 10% | sampled gate-pass at or above 95% | gate-pass holds |
| Outcome | not yet required | not yet required | metric at or above target, 4 consecutive weeks |
| Breaches | none | none in window | zero |
| Logged refusal | not required | at least one | at least one |
Volume and quality are cheap to measure and easy to game. The outcome metric is the one that decides whether autonomy is deserved, which is why it gates the last step rather than the first.
Judgment steps are separate trust accounts
Any step resting on taste rather than a checkable rule gets its own count: choosing the angle, choosing the contact, deciding an account is a fit. Ten examples reviewed, nine or ten of which would ship untouched, is a reasonable bar per account.
Then the asymmetry that makes it work: a weak-but-harmless output costs nothing, and a credibility-costing miss resets the count to zero. Wrong person, or a premise so off-base it advertises that no homework happened, is a reset. A soft line worth a tweak is not. Scoring both the same way either freezes every task forever or teaches the reviewer to wave through the misses that actually matter.
Refusal as evidence
Look for a skipped send, a held draft, a flagged edge case, something the agent could have done inside its own authority and chose not to. It counts only if the reviewer judges it substantive on the merits.
For a narrow mechanical task where the window plausibly offered no refusal-worthy moment, the criterion can be waived, but the waiver gets logged on the record so it never becomes a silent default.
Risk is logged per item, read at promotion
Tag each output low, medium or high risk as it ships. The tag does three things and no more: it sorts the review queue while a human is still reading everything, it defines the exception that surfaces at L2, and it exposes a task that is hitting its numbers while emitting a disproportionate share of risky output. Such a task gets scrutinized or split before promotion rather than waved up on the average.
What the tag must never do is quietly reinstate review on work a task already earned the right to ship. A high-risk tag that always holds is a second gate wearing a triage costume, and it cancels the ladder it was meant to sharpen.
What good looks like
- Every level has a date, the numbers it cleared, and the name of whoever approved it.
- Demotions appear in the record as often as promotions. A log with only promotions is a log of one direction of evidence.
- The outcome metric is defined before the task reaches L2, not invented when someone asks for L3.
- A waived criterion is visible as a waiver, never as a pass.