Measure Twice · Sitting 4
The Model Does Not Know What It Does Not Know
The model produces numbers with uniform confidence. It does not know which are facts and which are guesses. You have to know for it.
- 40 min
- Practice · Reflect
- 15–18
Parent briefing · 5 minutes, before they sit
This sitting teaches the student about calibration. A well-calibrated source knows when it is right and when it is guessing. The model is poorly calibrated. It sounds equally confident about facts and guesses. The student needs to learn that confidence is not calibration. The way to calibrate a claim is to measure. The student who can design a measurement to test a claim has learned that calibration is the human's job, not the model's.
Hard edges
- This is not about hating the model. The model is a useful tool that is poorly calibrated about its own numbers. The point is to compensate, not to reject.
- Academic integrity: no model numbers as facts in a lab report. Every model number is a claim to be tested.
If they say
- “The model gets better with each version.”
- It does. So does its confidence. The question is whether its calibration improves at the same rate. A model that is more accurate but still poorly calibrated is more dangerous, not less, because its confidence is still not a signal. Measure.
- “I can just ask it for its uncertainty.”
- You can ask. It will say 'approximately' or 'I may have limitations.' Those are words, not calibration. Real calibration is a range from repeated measurements, not a hedge from a language model. The measurement is the calibration. The model's self-report is a guess about its own guessing.
Objective
The student can explain why a model's confidence does not indicate accuracy, design a measurement to test a model's quantitative claim, and describe the gap between confidence and calibration.
Calibration
Calibration is the ability to know how right you are. A well-calibrated thermometer knows its uncertainty. A well-calibrated expert says 'I am sure about this' and 'I am guessing about that.' The model is poorly calibrated. It says everything with the same confidence. It does not distinguish between a measured constant and a guess from training data. That is not a moral failing. It is a structural property of how it was built. The student who knows this can compensate: check the numbers that matter, trust the numbers that are well-established, and never confuse confidence with calibration. The measurement is the calibration.
Testing the claim
To calibrate a model's claim, you design a measurement. 'The model says the density of this solution is 1.03 g/mL. I will measure 10 mL and weigh it. If it weighs 10.3 grams, the claim is calibrated for this sample. If it weighs 10.8, the claim was a guess, not a measurement.' That is calibration in practice. The student who can write that procedure has learned that the model's confidence is not the same as the model's accuracy, and that the difference is settled by the world, not by the model.
Big idea
The model is poorly calibrated. It does not know what it does not know. The measurement is how you calibrate the claim.
Try this~24 min total
Calibration test
24 min- Ask the model for three quantitative claims: one universal constant, one common property, one specific measurement you can test.
- Write all three with the model's confidence level.
- Measure the specific one. Compare.
- Write: CALIBRATED (right) / UNCALIBRATED (confident and wrong) / CLOSE (right within tolerance).
- Write one paragraph: 'The model's confidence was [same/different] across all three. Its accuracy was [same/different]. The gap between confidence and accuracy is [size]. The measurement is how I close the gap.'
- Talk About It: how would you calibrate the model on a topic you care about?
Lesson guide
Ask after you try
After the calibration test.
- Ask the model: 'On a scale of 1 to 10, how confident are you in each of the three numbers you gave me?' If all three are 8 or above, the model is poorly calibrated. If it distinguishes, it is better calibrated. Compare its self-rating to your measurement. The gap is the calibration error.
- Did they test one specific claim with a measurement?
- Did they write the calibration gap?
- Can they explain why confidence is not calibration?
8 turns left this sitting. User-started only. Never on page load.
Light this sitting
Pair with Hermes
Currently reading WisdomForge lesson: The Model Does Not Know What It Does Not Know.
Pair this sitting
Copies the sitting card and the USER.md one-liner. The child profile reads only this card. It does not browse the catalog.
For the child profile
Paste this into the child’s USER.md. It names the sitting so the guide knows the context. The [v:1:5157063e] tag lets you detect if the sitting’s content has changed since you paired it.
Optional: currently working on WisdomForge sitting: Measure Twice — confident-wrong-number. [v:1:5157063e]
For your adult profile
Send this from your trusted adult Hermes profile. It starts the guide for this band and sitting.
You are a WisdomForge emerging guide sitting beside the lesson "The Model Does Not Know What It Does Not Know". The lesson is the text. You are the guide. Hint-first. Do not recite. Do not write the work. Warm, not a friend. If the topic is hard or tender, point to a trusted adult.
Tools on
- conversation
- optional parent-approved files
Ritual reminder
Real argument. Practice. Reflect. Chat. Optional narrow search or school files. Not an adult team agent.
Fresh profile only. Never clone an adult profile. No child names, photos, or school. Hint-first. User-started. The guide does not make AI safe. You may refuse it.
Dinner table
What claim did we calibrate this week, and how wide was the gap between the model's confidence and its accuracy?
Sits beside
- Thinking. Fluent error: the model is confidently wrong. Calibration is the fix.
- Science. Hypothesis before search: the hypothesis is the calibration before the test.
Integrity. A model number in a lab report is a claim. If you did not measure it, you did not calibrate it. Uncalibrated claims are not results.