You ask the model to size the market, check a rule, or summarise a paper. Back comes a clean paragraph. You paste it into the deck. Three weeks later a number was invented, a “leading study” was a blog post, and a deal moved on evidence that was never there. The output was not researched. It was fluent. Fluency is now the cheapest thing a model produces.
Score the trace
Evaluating reasoning is not asking “does this sound right?” That question grades the prose. The prose is the last mile.
A scored trace answers four cheaper questions.
- What was gathered? Named sources, or a vibe from pretraining.
- What was retrieved? A passage you can point at, or extra context stuffed in a prompt.
- What was assumed? The load-bearing premise. If it is false, the paragraph collapses.
- How sure are you entitled to be? A number you would defend, not a tone of voice.
If you cannot replay those four, you cannot grade the answer. You can only like it.
This is the same split as an agent versus an automation. A script fails on a node you named. A fluent guess fails in a sentence you admired.
Do not confuse this Evaluate with the Evaluate move in SENSE. SENSE Evaluate turns a market signal into a sourced claim for Monday. This page scores a model’s chain before you let that claim into the register.
A fluent miss you can replay
The paragraph said: “Enterprise adoption is already above 70 percent, per a recent leading study.” It scanned. It even had a comma.
The trace, if anyone had asked:
- Gathered: nothing you can name. No URL. No PDF. No table.
- Retrieved: none. There was no chunk.
- Assumed: that a round number plus “leading study” is a citation.
- Entitled confidence: zero, once the second line is blank.
The repair is not a better adjective. Write the claim, the evidence, the percentage you would bet. If line two is empty, the sentence does not ship. That is CALIBRATE applied to a paste, not a personality lecture.
CALIBRATE: confidence is a report
CALIBRATE is the practice of matching the strength of a judgment to the strength of the evidence, then updating when the facts move. Confidence is a report, not a personality trait.
Apply it to the model and to yourself.
The model will sound certain. That is a rendering choice. It is not a probability.
You will feel certain because the paragraph is tidy. That feeling is the certainty trap: a binary yes wearing a complete sentence. The Calibrated Mind is the book that trains the human side of this. English and Dutch editions are available now. The test on this page is already usable.
Write the claim in one line. Write the evidence in one line. Write a percentage you would actually bet. If you cannot write the second line, the first line is not a judgment. It is a mood.
A prediction journal is the scorecard that does not lie. Memory will defend the feeling. The journal will not.
GRAIN: name the chunk
If the answer depended on your files, your web, or a corpus, fluency is the wrong instrument.
RAG means the model must fetch named passages before it answers. The retrieval is the product. The generation is the last mile. If you cannot inspect which chunk was used, you do not have RAG. You have a language model with extra context.
GRAIN is the loop: Gather, Rank, Assemble, Inspect, Navigate. An answer stays tied to a chunk you can name. The RAG Engineer is the book. It is in production. The loop is already the test.
Ask the person who “added RAG” to walk one bad answer.
- Which passage was gathered?
- Why did it rank?
- What was assembled around it?
- Did anyone inspect a failure, or only a demo?
- What would you navigate to next?
If those lines are blank, stop arguing about embeddings. You are tuning a symptom.
What the other two books are for
ML for Strategic Founders is the adjacent literacy: how to scope and audit a model that touches revenue or reputation without becoming a data scientist. SIGNAL is that loop. It is in production. Use it when the question is “is this model allowed to sit on this decision,” not “did the paragraph scan.”
The Judgement Gym is the training side. REPS is a protocol for keeping your own conclusion first when analysis got cheaper. It is in design. Use it when the temptation is to delegate the judgment because the prose arrived in four seconds.
None of this is a ReasonKit benchmark. This site will not invent one.
A ten-minute audit
Take the last model answer you trusted.
| Underline | Source or “unknown” | Chunk in eight words | Entitled confidence | What would move it 20 points |
|---|---|---|---|---|
| first number | ||||
| first causal verb | ||||
| first proper name |
Unknowns that still shipped are the defect. Fluency hid them. The next answer you keep is the one that survives this list.
CALIBRATE and GRAIN are the names. The audit is the work. The books will catch up.