References and the decision package
References surface what interviews miss. The decision package turns everything gathered into a logged decision.
Reference checks
References are where deal-breakers surface. Strong candidates are not exempt.
How many
| Hire type | Minimum | Must include |
|---|---|---|
| Type 1 | 3 | One off-list reference the candidate did not provide, and one former direct report |
| Type 2 | 2 | One former manager |
Tell the candidate off-list calls will happen. Never contact a current employer without explicit permission.
The script
Open: "I am considering {candidate} for {role}. The role has to deliver {top outcome}. I would value 15 minutes of your honest read."
- What did they do best? Give me a specific example.
- Where did they need to grow? What happened when that came up?
- How would you compare them to the best person you have seen in a similar role?
- Would you hire them again, for this role? (Listen for the pause.)
- If I hire them, what is the one thing I should know to set them up to succeed?
- Who else should I talk to?
Ask one question aimed at each evidence gap from the interviews.
Reading the call
| Signal | What it usually means |
|---|---|
| Specific examples, unprompted | Genuine endorsement |
| Adjectives with no examples | Polite or distant reference |
| A pause before "would you hire again" | A real reservation; ask about it directly |
| "In the right environment, they would thrive" | Reservation about the last environment |
| Strong praise for skills, silence on people | Possible team or culture issue |
Write each call up within an hour: the quote, the example, and any pause or qualification.
Decision package
The package answers the decision. It does not restate the question as pros and cons.
Per candidate
| Section | Contents |
|---|---|
| Outcome fit | Score per outcome, weighted, with the evidence behind each |
| Competency scores | Each critical and supporting competency, 1 to 4, with evidence |
| Deal-breakers | Clear, or which one is flagged |
| Behaviors observed | Which of the 4 to 6 behaviors were seen, and where |
| Mindset | Score 1 to 4 on the two mindset questions, the combination read (builder, talker, executor, drifter), and the evidence |
| Live exercise | Baseline score, learning delta, and the reading (capable and coachable, steep curve, plateaued, or no-hire) |
| Failure answer | What the candidate said they got wrong and changed, or the evidence gap if they could not |
| Referral or sponsor | Who referred or is sponsoring the candidate, and which conclusions rest on that trust rather than evidence |
| References | Summary of each call with quotes |
| Evidence gaps | What is still unknown, and what information would change the decision |
| Risks | What could go wrong, and the earliest signal it is going wrong |
| Recommendation | Hire, no-hire, or hire with named conditions |
Weighted outcome score
For each outcome, rate likelihood the candidate delivers it (1 to 4) based on evidence. Multiply by the weight and sum. The maximum is 400.
| Band | Reading |
|---|---|
| 320 and above | Strong evidence across the heaviest outcomes |
| 260 to 319 | Hireable if the gaps sit in supporting outcomes and have a plan |
| Below 260 | Do not hire for a Type 1 role |
A high total with a weak score on the heaviest outcome is still a no-hire.
Before deciding: the interviewer's operating rule
Do not ask whether you like the candidate. Answer these ten questions in writing.
- What outcomes must this person produce?
- What evidence says they have produced comparable outcomes before?
- Which critical behaviors did I observe, rather than infer?
- What did the live exercise reveal about how they think?
- What changed after I coached them?
- How do they explain failure?
- What evidence shows self-direction?
- What evidence contradicts my current opinion?
- What am I believing only because someone I trust recommended them?
- If this hire fails in twelve months, what signal from this process will I wish I had taken more seriously?
The last question matters most. It turns the process from confirming enthusiasm into surfacing the evidence that could make the decision wrong. Write that signal into the decision log as one of the "would change my mind" signals.
When authority overrides the evidence
Sometimes a more senior leader overrules the hiring manager. That is their right. Record it:
- the recommendation and the evidence behind it
- who overrode it, and the stated reason
- the specific risk signals to watch, and the checkpoint date
An override that resolves the evidence is good judgment. An override that ignores it is a risk the organization should see in writing.
The decision log
Hire: "Decided to hire {candidate} as {role} on {date}. Rationale: {one paragraph}. In the first 90 days, these signals would tell us we were wrong: {2 or 3 specific signals}. Checkpoint: {date}."
No-hire: "Decided not to hire {candidate} on {date}. Rationale: {one paragraph}. Would reconsider if: {specific conditions}."
Reversibility in the decision
| Type 1 | Type 2 | |
|---|---|---|
| Process time | Roughly twice the default | Fast |
| References | 3 or more, one off-list | 2 |
| Package | Full | Scorecard plus short log |
| Checkpoint | Day 90 | Day 30 |
| Ship bar on the kit rubric | 90, every criterion at 9 or above | 75, scorecard and kit at 8 or above |
The checkpoint audit
At the checkpoint, reopen the log. For each predicted signal: did it appear? Which outcomes are on track? Which interview question predicted actual behavior best? Record the answers. After five hires into the same role type, the scorecard and question bank are calibrated by evidence rather than opinion.