Founder Performance Research guide

What is a confidence grade? Why "high confidence" is not 90%

"High confidence" tells you about the evidence behind a claim. It does not tell you the odds, or what to do next.

Two hands hold an egg against a candling lamp in a dark shed so its inner shadow shows, sorted trays blurred behind
The mark is only as good as the reasons written beside it.

Direct answer

A confidence grade is a short label, such as high, medium or low, that records how well the evidence you have checked supports one specific claim, with the reasons written beside it. It is not a probability: the IPCC tells its authors that confidence should not be read as a probability, and US intelligence standards forbid mixing the two in one sentence. It is not permission to act either. That also depends on what a mistake would cost.

A confidence grade is a short label, such as high, medium or low, that records how well the evidence you have actually checked supports one specific claim, with the reasons written next to it. It answers “how solid is the support?” It does not answer “how likely is it?”, and it does not answer “should we act?”

That sounds like hair-splitting until you notice who insists on it. The IPCC tells its authors that confidence “should not be interpreted probabilistically”. US intelligence rules forbid putting a confidence level and a likelihood in the same sentence of an analytic report. Medical guideline panels rate the evidence and the recommendation separately. Three very different fields, one rule: “high confidence” is a statement about the evidence, not about the odds.

Where do confidence grades come from?

You will meet confidence grades in research summaries, risk registers, competitor reports, due-diligence notes and, increasingly, in the output of AI tools. The serious versions all separate the grade from something else.

Who uses itWhat gets gradedThe scaleWhat it is kept apart from
IPCC (climate assessments)How valid a finding is, judged from the type, amount, quality and consistency of the evidence and the degree of agreementvery low, low, medium, high, very highLikelihood, a separate probability scale where “likely” means 66 to 100%
GRADE (medical guidelines)The quality of a body of evidencehigh, moderate, low, very lowThe strength of the recommendation: strong or weak
US intelligence (ICD 203)The analyst’s confidence in the basis for a judgment: logic, evidence, quantity and quality of sourcese.g. “high confidence”The likelihood of the event, which may not share a sentence with it
The Agentic Author (Len’s book)How much you trust a claim before it has been checkedA strong, B good, C mixed, D speculative, plus an X flag for “remove or verify”Risk: what being wrong would cost, in its own field

The IPCC’s rules are in its guidance note for authors of the Fifth Assessment Report. The intelligence rule is in Intelligence Community Directive 203, signed in 2015. GRADE is described in the BMJ. The last row comes from The Agentic Author, where the grade sits in an evidence ledger next to a separate risk column.

What does a confidence grade look like in practice?

Here is a hypothetical example. A colleague forwards a headline: “Our main supplier is raising prices 10% in January.” It looks like one fact. It is three claims, and each deserves its own grade.

ClaimEvidence you openedGradeWhy
The supplier has announced a price increase for JanuaryThe supplier’s own letter to customers, dated, read in fullHighPrimary source, read directly, says exactly this
Prices rise 10% on everything we buy from themThe same letter says “up to 10%” and lists product groups; two of our five items are not on the listLowThe letter supports “up to 10% on the listed groups”, not “10% on everything”
Our other suppliers will followNothing yetUnratedThis is a forecast about the future, not a finding about evidence

Notice the second row. The fix is not to argue about whether it deserves medium. The fix is to rewrite the claim to what the evidence supports: “Prices on the listed product groups rise by up to 10% from January.” That narrower claim is graded high.

The third row needs a different tool. A forecast gets a probability and a date on which it resolves. Checking such numbers against outcomes over time is what calibration means.

What should sit next to the grade?

A bare “high” tells the next reader almost nothing. It could mean the source says exactly this, or that three colleagues agree, or that the writer simply expects it to be true. Keep four things beside the mark.

  1. The exact claim. Scope, place, date. “The supplier announced” and “prices rise on everything” are different claims with different grades.
  2. The evidence you actually opened. Document, passage, date or version. A link you have not opened is a lead, not evidence. Twenty reposts of one letter are one source, not twenty.
  3. The reason for the grade. What the evidence supports and where it stops. The IPCC asks its authors for “a traceable account” of how they judged the evidence. Agree in your team what high, medium and low mean before you compare grades across rows.
  4. The recheck trigger. Who looks again, and on which date or event. ICD 203 asks analytic reports, where appropriate, to name the indicators that would change their level of uncertainty. For the price claim, that might be the supplier’s new price list.

Some evidence ledgers grade the kind of source instead (primary, secondary, internal data, testimony), as in this explainer on evidence ledgers. Source type is a useful input to a confidence grade, but it is not the same thing. Label which one your column means.

What a confidence grade is not

It is not a probability. “High” does not secretly mean 90%. The IPCC keeps confidence and likelihood on two separate scales for exactly this reason, and ICD 203 bans mixing them in one sentence “to avoid confusion”. If you need odds, write a forecast with a number and a resolution date.

It is not permission. The GRADE authors put it plainly: “High quality evidence doesn’t necessarily imply strong recommendations, and strong recommendations can arise from low quality evidence.” Whether to act also depends on what a mistake would cost.

The LEVEL check shows where a grade fits. It is a five-step method from The Evidence-Aware Life, a book by Len that is available now, and you do not need the book to use it:

  1. Locate the claim: one sentence someone could check.
  2. Estimate the stakes: what acting wrongly, ignoring it wrongly, or waiting would cost.
  3. Value the evidence: its kind, quality, independence and relevance. The confidence grade lives here.
  4. Expose the gap: what is missing for this particular decision.
  5. Lean accordingly: act only as far as evidence and stakes allow.

A grade answers step three. Steps two, four and five are still yours. A medium-graded claim can justify a cheap test and still be too weak for a large, hard-to-reverse commitment.

Unrated is not low. Unrated means nobody has inspected the evidence yet, or could not get access. Low means someone looked and the support is weak. Contradicted means the evidence points the other way. Those are three different findings. Keep them apart, and write down why a claim is unrated.

An AI tool’s “confidence” is not a grade either. AI Agents for Startup Strategy, another of Len’s books, advises treating a system’s self-reported confidence as a cue for what to check first, not as a calibrated probability. The trail back to the source is what you review. For competitor intelligence, the same book adds a simple corroboration rule: no claim graded above low confidence without two independent sources.

Why keep a grade at all?

Because it makes the weak claims visible early, while they are cheap to fix. The Agentic Author calls an honest D (speculative) the most valuable grade on its ladder, because D is where the claims you want to be true are hiding. It also makes a useful point about what grading does: it does not verify a claim. It tells you which claims still need checking, and how far to trust them until then.

A grade card you can copy

Claim (exact wording, scope, date):
Evidence opened (what, where, which version):
Grade:   high / medium / low / unrated
Why this grade (what it supports, where it stops):
Evidence against, if any:
Recheck (who, and which date or event):

One-line rule: no grade without its reasons, and no probability hiding inside a grade.

Try this today (10 minutes)

Find one “high confidence”, “definitely” or “we are confident that” in a report, a deck or a chat thread from this week. Ask three questions: which claim exactly, what evidence was opened, and what would change the grade?

If nobody can answer, relabel it “unrated” until someone does. That is not a downgrade. It is the first honest entry in the record.

Cite this:What is a confidence grade? Why "high confidence" is not 90%.Len P. van der Hof. https://lenvanderhof.com/en/blog/what-is-a-confidence-grade/ ·

Terminology

Sources

  1. Guidance Note for Lead Authors of the IPCC Fifth Assessment Report on Consistent Treatment of Uncertainties (2010) · Intergovernmental Panel on Climate Change (IPCC)
  2. Intelligence Community Directive 203: Analytic Standards · Office of the Director of National Intelligence
  3. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations (BMJ, 2008) · BMJ (via PubMed Central)
  4. The Agentic Author
  5. AI Agents for Startup Strategy
  6. LEVEL framework

Further reading

Markdown for LLMs