The Evidence Packet: 8 Fields to Prepare Before an Idea Review17 min read

A review board sits down at ten with five items on the agenda. Here’s what it actually received:
- A forty-slide business case with a rendered mock and a five-year revenue curve
- Two paragraphs pasted into an email
- A spreadsheet whose only tab was labeled “final v3”
- A deck from a sponsor nobody in the room wanted to challenge
- One packet that answered the obvious questions before anyone asked them
By eleven, one is funded and four are dead.
Nobody in that room compared ideas. They compared preparation, then called the result a decision.
Experienced reviewers, real criteria, honest intent. What reached them was missing the same eight things, which is a template problem and cheap to fix.
The Review Is Usually Lost Before Anyone Sits Down
Those submissions were written to persuade rather than to be evaluated, and not one said how much confidence the numbers deserved.
The general state of corporate decision making isn’t encouraging, and the numbers are worth sitting with:
- Only 20% of respondents say their organization excels at decision making
- 61% say most of their decision-making time is used ineffectively, and 57% of C-level respondents say the same
- Just 37% say their decisions are both high quality and fast
- A typical Fortune 500 company loses roughly 530,000 days of manager time a year to poor decision processes, about $250 million in wages
So the outcome is largely set before anyone sits down. And the input is the part you can actually fix.
In practice, innovation reviews compress all of that into a single hour, with real criteria and real authority in the room. What’s missing is comparable input.
Gut feel fills the space where evidence wasn’t supplied. Ill-prepared participants are often people who received nothing worth preparing from.
The Failure Modes Are Mostly Input Problems
Stage-Gate International, Robert Cooper’s organization, keeps its own list of gate meeting failure modes:
- Hit-or-miss attendance, with key people missing
- No clarity on who plays which role
- Poor informational decision support
- No agenda or flow that builds toward a decision
- Ineffective decision criteria
- Participants who arrive ill-prepared
- Lack of integrity, meaning people don’t do what they agreed to do
Read that list again: attendance, role clarity and integrity are governance problems. Most of what remains traces back to the input.
Poor decision support is a description of what was circulated, or wasn’t.
Here’s the thing a lot of innovation programs get backwards. They keep optimizing for volume, chasing more submissions into the funnel.
The constraint sits where submissions get judged, and decision quality beats idea quantity. More ideas through a review that can’t tell strength from formatting just adds coin flips.
A Standard Input Set Is a Documented Lever
This has been tested in the field with measured results.
A large natural-resources company imposed exacting preparation requirements on its meetings:
- Templates for meeting preparation
- Standardized pre-read protocols
- Deadlines
Bain reports that the change improved decision effectiveness and held the company at industry-leading performance.
Preparation was one of four changes there. The others were clearer purpose, redesigned agendas, and tighter meeting conduct.
Amazon runs the most quoted version of the same idea. Its employees don’t do PowerPoint, they write six-page narrative memos.
Meetings open with a silent study hall where everyone reads together before anyone speaks.
Truth-Seeking Versus Selling
Bezos’s argument was about cognition. Slides let you hide a weak argument behind bullet fragments, while a narrative forces claims to connect or visibly fail.
The PR/FAQ format tightens this further with a fixed field structure:
- Heading and subheading
- Summary paragraph
- Problem paragraph
- Solution paragraph
- Quotes and a getting-started section
- An external FAQ, then an internal FAQ covering finance, marketing, support, operations, HR and more
Reviewers read and annotate silently for 15 to 20 minutes before discussion begins. Its design goal is the sharpest three words here: truth-seeking versus selling.
That’s the distinction a packet exists to enforce. A pitch foregrounds upside and buries the load-bearing assumption on slide 34.
A packet is engineered to be evaluated, and it says on page one what would disprove the idea. Structure enables the creative work.
The Eight Fields of an Evidence Packet
Every submission fills the same eight fields, and the tiering model further down sets how many apply at each gate.
| Field | The Question It Answers | What Disqualifies It |
|---|---|---|
| 1. The decision requested | What are you asking this board to approve, and what happens if they say nothing? | A submission with no named decision and no named requester |
| 2. Problem and who has it | Whose problem is this, how do we know, and what do they do today instead? | A problem stated only from the inside of the company |
| 3. The proposed solution in one paragraph | What are we actually building or changing? | A description that needs a demo to be understood |
| 4. The value case with assumptions exposed | What’s the expected return, and which three assumptions is that number resting on? | A single number with no visible arithmetic |
| 5. Evidence log with strength ratings | What do we know, how do we know it, and how strong is each piece? | Evidence cited without a source or a strength rating |
| 6. The riskiest assumption and its test | Which assumption, if wrong, kills this, and what would settle it? | A risk register with no test attached and no owner |
| 7. Kill criteria | What would have to be true for us to stop, and who checks? | Criteria written so loosely they can never trigger |
| 8. Resource ask and reversibility | What’s being committed, and how expensive is it to reverse? | A resource ask with no exit cost stated |
Fixed fields do two jobs at once. They make submissions comparable, and they make omissions visible.
No exceptions for senior sponsors, and none for ideas that are “obviously” good. The tier sets the count, never the sponsor.
Of the two jobs, visibility matters more, since an empty cell is a question the board never had to think to ask. Every field feeds a real decision.
The First Four Fields Set Up the Decision
Field 1 is the one most submissions skip entirely. A packet that doesn’t name the decision leaves the board guessing whether it’s being informed or asked.
Name it, name who’s asking, and say what happens if the board stays silent. Fields 2 through 4 fail in their own ways:
- Field 2 gets written from inside the company instead of from the person with the problem
- Field 3 gets written in language only its author can parse
- Field 4 arrives as a single number with no visible arithmetic
A useful test for field 2: state the problem in the words the affected person uses. If you can’t, you’re describing an internal preference.
One paragraph is the constraint on field 3, and anything needing a demo isn’t ready for a decision. Field 4 is where the gaming happens.
Ask for the expected return and the three assumptions it rests on, because an unshown calculation can’t be challenged.
Fields Five and Seven Change Behavior Fastest
Submissions routinely lack these two, and they shift behavior quickest. Both work by making the submitter confront something first:
| Field | What It Forces | Why Teams Resist It |
|---|---|---|
| 5. Evidence log | Every claim sits in a row next to the word “source” | The empty rows become impossible to miss |
| 7. Kill criteria | Naming the conditions under which you’d stop your own project | Writing them before funding takes nerve |
Kill criteria are the closest thing to a sincerity test a review process has.
Teams that resist field 7 hardest usually need it most, and a board that never kills anything is just a queue. The evidence log changes how submissions get written.
Design the Packet Backwards From Your Scoring Criteria
The packet is an input to whatever criteria your gatekeepers already score on.
Stage-Gate’s six standard criteria are the common baseline:
- Strategic fit
- Product and competitive advantage
- Market attractiveness
- Technical feasibility
- Synergies with core competencies
- Financial reward versus risk
If your board scores on six things, the packet should make all six scoreable without a follow-up email.
Anything in the template that doesn’t feed a criterion is decoration. So write your criteria first, then map each field to the one it serves.
Any orphan field is a candidate for deletion, and any criterion with no supporting field explains a recurring argument.
One more structural rule, borrowed from Bain’s RAPID model: the packet has a single named author.
That Recommender has the visibility to assemble it honestly. Committee packets tend to average away the disagreement the board most needs to see.
Rate the Evidence Behind Every Claim
In 1951, a National Intelligence Estimate said an attack on Yugoslavia should be considered a “serious possibility.”
Sherman Kent, later chairman of the Board of National Estimates, asked his colleagues what they understood that phrase to mean.
Their answers ranged from 20% to 80% likely. Not one of them had realized they disagreed.
Kent’s fix was a calibrated vocabulary, each phrase carrying a center value and a margin:
- Almost Certain, around 93%
- Probable, around 75%
- Chances About Even, around 50%
- Probably Not, around 30%
- Almost Certainly Not, around 5%
- Possible, used only when no odds could responsibly be assigned
That’s your idea review. Everyone nods at “strong market interest” while privately holding numbers that differ fourfold.
The meeting then ends in an agreement that nobody actually shares.
The IPCC took the idea further and runs two parallel scales. One rates likelihood, from exceptionally unlikely to virtually certain.
The other rates confidence, derived from evidence quality and the degree of agreement.
Evidence counts as most robust when multiple independent, high-quality lines agree. Probability and evidence quality are different questions.
A Four-Level Evidence Scale
A survey-backed 70% and a pilot-backed 70% aren’t the same object.
Here’s a four-level scale you can paste into your template tomorrow:
| Level | What It Means | Example | How Much Weight It Carries |
|---|---|---|---|
| E1 Assertion | Internal opinion, analogy to another market, or a vendor claim | “Our sales team says customers keep asking for this” | Enough to justify investigating, never enough to justify investing |
| E2 Stated preference | Someone said they would | Survey results, interview quotes, a workshop vote | Directional. Enough to fund a small, reversible step |
| E3 Revealed behavior | Someone did something that cost them a little | Prototype usage, a signed letter of intent, a completed pilot | Strong enough to gate a funded stage |
| E4 Committed behavior | Someone did something that cost them real money or real switching effort | A paid pilot, a signed contract, live production data | Strong enough to gate an irreversible commitment |
One rule makes the scale work: a claim’s rating is set by the weakest evidence supporting it. Six E1 assertions and one E3 pilot is still E1.
A claim with no evidence still gets a row, rating left blank so the gap shows. The submitter assigns ratings, and any reviewer can challenge them.
An argument about whether something is E2 or E3 beats an argument about whether the idea “feels right.”
What Software Can and Can’t Do Here
The scale formalizes what discovery practice already does. Teresa Torres describes David Bland’s assumption mapping as evidence strength against importance.
This is where software earns its place:
- Enrich a submission with market context
- Cluster near-duplicate ideas
- Flag which required fields are still empty
- Surface an E1 claim that’s quietly carrying a decision needing E4 evidence
Whatever sits in the critical-and-weakly-evidenced corner gets tested first, and good tooling makes that corner obvious.
What it can’t do is manufacture evidence nobody gathered. Tools amplify the clarity you already have, and gathering the data stays your job.
Teams that can’t tell weak evidence from strong treat “we ran a survey” as settling what only a prototype could. No tool fixes that.
Scale the Packet to the Decision
The axis that should govern packet weight is reversibility, and budget correlates only loosely. Bezos’s 2015 shareholder letter splits decisions in two:
- Type 1, irreversible one-way doors that warrant heavy deliberation
- Type 2, reversible two-way doors that high-judgment individuals or small groups should make quickly
A cheap announcement can be a one-way door, an expensive cancellable pilot a two-way door. He set the threshold at roughly 70% of the information you want.
Innovation accounting adds a caution. Toma and Gons warn that very early ventures are too volatile to score on the mature scale.
Three Tiers, Assigned by the Process Owner
The tier decides how much packet the decision has to carry:
| Tier | Decision Type | Required Packet | Reviewer Time |
|---|---|---|---|
| Tier 1 Screening | Reversible, small, a two-way door | Fields 1, 2, 3 and 8 only, one page | 10 minutes, async, no meeting |
| Tier 2 Investment | A funded stage, recoverable but not free | All eight fields, E3 on the riskiest assumption, E2 on the value case | 20-minute read plus a 30-minute discussion |
| Tier 3 Commitment | Hard to reverse: market entry, platform build, public commitment | All eight fields, E4 on the riskiest assumption, plus a pre-mortem and a reference class check | 30-minute read plus a 60-minute discussion |
Anyone calling this heavyweight should look at where Cooper himself landed.
The creator of Stage-Gate now advocates “lean gates with teeth” and a Triple A system: adaptive and flexible, agile, accelerated.
Gates that demand little and refuse nothing are just meetings with a nicer name. Tier 3 should be rare, so people accept it when it appears.
The tier is assigned by the process owner when the submission lands. Every submitter alive will argue their idea is a two-way door.
The Pre-Read Is the Part People Skip
Circulating the packet and hoping people read it produces a room where almost nobody has. Amazon’s fix is to put the reading inside the meeting.
Filing early exists for tiering and scheduling. For anything that reaches a meeting, reading together buys debate time instead of exposition:
- The packet is filed 72 hours ahead for tiering, or it drops off that week’s agenda
- Tier 1 items never reach the room, since they’re read and decided async
- For Tier 2 and 3, the reading happens in the room, in silence, before discussion
- The recommender doesn’t present, they answer
- Every item is labeled Inform, Discuss or Decide
- The decision, the rationale and the kill criteria are logged before anyone leaves
Bain reports a smaller device that carries as much weight: label every agenda item Inform, Discuss, or Decide. In practice, that label gets skipped first.
Then there’s the attendance question, which many boards handle by inviting everyone. At one consumer-goods company, 40% of those involved added no value.
Removing them made decisions faster and better, and the same research names debate quality as the top predictor for big bets.
Add a Pre-Mortem for Tier 3
For Tier 3, add a scheduled pre-mortem rather than an optional workshop. Optional means it won’t happen.
Gary Klein’s method runs in four moves:
- Tell the team the project has already failed spectacularly
- Have each person privately write down every reason they can think of
- Go round-robin, one reason at a time, until all are recorded
- Afterward the owner reviews the list and strengthens the plan
The research Klein cites found prospective hindsight improves people’s ability to identify reasons for future outcomes by roughly a third.
In Klein’s own example, a pre-mortem surfaced that a team lacked time to prepare its business case.
That’s exactly the kind of thing nobody volunteers in a status update. Klein’s framing is psychological rather than analytical.
People won’t voice reservations during planning unless the format gives them permission, so build it into the calendar.
Log the Decision Before Anyone Leaves
Whatever gets decided, the outcome set has to include real options.
Stage-Gate’s four are Go, Kill, Hold and Recycle, and every Go carries an action plan and next gate date. Then close the way Amazon does: disagree and commit.
Recorded dissent is healthy, but the debate ends when the decision is logged. The agreement that dissolves in the corridor is the worst outcome.
Where This Goes Wrong
The first failure is that the packet turns into a beauty contest.
Give people a template and a word count and the strongest writer starts winning more than their share. That’s a different bias than the one you meant to remove.
The E-rating column is the counterweight. Reviewers score the evidence itself and say the rating out loud before discussing the claim.
A beautifully written E1 is still an E1.
Business Case Theater
The second failure is more corrosive, and most organizations are already living in it.
A CEO survey by MIT Sloan found average real hurdle rates of 12.2% overall and 11.6% for manufacturers.
The reported reason is the interesting part. Managers may set rates above their required returns to correct for optimistic sponsor projections.
Read that as an admission:
- The organization already knows its business cases are inflated
- It corrects downstream instead of fixing the input
- The correction taxes every honest submitter and barely touches the ones who padded
The optimism is a measurable class effect.
Bent Flyvbjerg’s World Bank analysis puts average cost overrun at 44.7% for rail, 33.8% for bridges and tunnels, and 20.4% for roads.
Demand forecasts are worse. Rail passenger numbers come in an average of 51.4% below forecast, with 84% missing by more than 20%.
The mature response already exists in government practice.
The UK Department for Transport publishes recommended optimism-bias uplifts built from reference class data, and requires promoters to apply them.
Publish the Uplift Instead of Hiding It
At full business case stage the uplifts run 20% for roads, 30% for rail, and 28% for fixed links. Earlier stages carry far larger numbers.
Set a house uplift per class, apply it openly, and publish the number. Then nobody has to inflate a forecast to stay competitive.
When the Template Becomes the Problem
The third failure is that the whole thing calcifies into documentation. McKinsey warns that complex process diagrams are rarely read beyond their authors.
A template that has crept from eight fields to twelve is how good ideas die of administration.
Run your template through three submissions, then cut every field that didn’t change a decision. Do it annually and the packet stays a tool.
Standardize the Input Before You Blame the Meeting
The biggest lever available to your idea review is eight fixed fields and an evidence rating on every material claim.
If you change three things this quarter, change these:
- The evidence log, with an E1 to E4 rating on every claim that carries weight
- Kill criteria written by the submitter before funding, not renegotiated after
- A pre-read deadline you hold even when a senior sponsor misses it
These three shift the conversation on their own, so new scoring models can wait. Watch whether a claim carries more weight than its evidence can bear.
That’s a review board doing its actual job, and it costs you one template revision rather than a change program.
Accept Mission’s own figures put the reduction in resource waste near 30%. Most of it comes from killing weak work sooner.
Download our ebook on idea evaluation and governance, or book a demo to see how Accept Mission structures scorecards, evidence ratings, and gate decisions.











