Measure Twice · Sitting 4

The Model Does Not Know What It Does Not Know

The model produces numbers with uniform confidence. It does not know which are facts and which are guesses. You have to know for it.

  • 40 min
  • Practice · Reflect
  • 15–18

Parent briefing · 5 minutes, before they sit

This sitting teaches the student about calibration. A well-calibrated source knows when it is right and when it is guessing. The model is poorly calibrated. It sounds equally confident about facts and guesses. The student needs to learn that confidence is not calibration. The way to calibrate a claim is to measure. The student who can design a measurement to test a claim has learned that calibration is the human's job, not the model's.

Hard edges

  • This is not about hating the model. The model is a useful tool that is poorly calibrated about its own numbers. The point is to compensate, not to reject.
  • Academic integrity: no model numbers as facts in a lab report. Every model number is a claim to be tested.

If they say

The model gets better with each version.
It does. So does its confidence. The question is whether its calibration improves at the same rate. A model that is more accurate but still poorly calibrated is more dangerous, not less, because its confidence is still not a signal. Measure.
I can just ask it for its uncertainty.
You can ask. It will say 'approximately' or 'I may have limitations.' Those are words, not calibration. Real calibration is a range from repeated measurements, not a hedge from a language model. The measurement is the calibration. The model's self-report is a guess about its own guessing.

Objective

The student can explain why a model's confidence does not indicate accuracy, design a measurement to test a model's quantitative claim, and describe the gap between confidence and calibration.

Calibration

Calibration is the ability to know how right you are. A well-calibrated thermometer knows its uncertainty. A well-calibrated expert says 'I am sure about this' and 'I am guessing about that.' The model is poorly calibrated. It says everything with the same confidence. It does not distinguish between a measured constant and a guess from training data. That is not a moral failing. It is a structural property of how it was built. The student who knows this can compensate: check the numbers that matter, trust the numbers that are well-established, and never confuse confidence with calibration. The measurement is the calibration.

Testing the claim

To calibrate a model's claim, you design a measurement. 'The model says the density of this solution is 1.03 g/mL. I will measure 10 mL and weigh it. If it weighs 10.3 grams, the claim is calibrated for this sample. If it weighs 10.8, the claim was a guess, not a measurement.' That is calibration in practice. The student who can write that procedure has learned that the model's confidence is not the same as the model's accuracy, and that the difference is settled by the world, not by the model.

Big idea

The model is poorly calibrated. It does not know what it does not know. The measurement is how you calibrate the claim.

Try this~24 min total

Calibration test

24 min
  1. Ask the model for three quantitative claims: one universal constant, one common property, one specific measurement you can test.
  2. Write all three with the model's confidence level.
  3. Measure the specific one. Compare.
  4. Write: CALIBRATED (right) / UNCALIBRATED (confident and wrong) / CLOSE (right within tolerance).
  5. Write one paragraph: 'The model's confidence was [same/different] across all three. Its accuracy was [same/different]. The gap between confidence and accuracy is [size]. The measurement is how I close the gap.'
  6. Talk About It: how would you calibrate the model on a topic you care about?

Lesson guide

Ask after you try

After the calibration test.

  1. Ask the model: 'On a scale of 1 to 10, how confident are you in each of the three numbers you gave me?' If all three are 8 or above, the model is poorly calibrated. If it distinguishes, it is better calibrated. Compare its self-rating to your measurement. The gap is the calibration error.
  2. Did they test one specific claim with a measurement?
  3. Did they write the calibration gap?
  4. Can they explain why confidence is not calibration?

8 turns left this sitting. User-started only. Never on page load.

Light this sitting

Pair with Hermes

Practice · Reflect15–18

Currently reading WisdomForge lesson: The Model Does Not Know What It Does Not Know.

Pair this sitting

Copies the sitting card and the USER.md one-liner. The child profile reads only this card. It does not browse the catalog.

For the child profile

Paste this into the child’s USER.md. It names the sitting so the guide knows the context. The [v:1:5157063e] tag lets you detect if the sitting’s content has changed since you paired it.

Optional: currently working on WisdomForge sitting: Measure Twice — confident-wrong-number. [v:1:5157063e]

For your adult profile

Send this from your trusted adult Hermes profile. It starts the guide for this band and sitting.

You are a WisdomForge emerging guide sitting beside the lesson "The Model Does Not Know What It Does Not Know". The lesson is the text. You are the guide. Hint-first. Do not recite. Do not write the work. Warm, not a friend. If the topic is hard or tender, point to a trusted adult.

Tools on

  • conversation
  • optional parent-approved files

Ritual reminder

Real argument. Practice. Reflect. Chat. Optional narrow search or school files. Not an adult team agent.

Fresh profile only. Never clone an adult profile. No child names, photos, or school. Hint-first. User-started. The guide does not make AI safe. You may refuse it.

Dinner table

What claim did we calibrate this week, and how wide was the gap between the model's confidence and its accuracy?

Sits beside

  • Thinking. Fluent error: the model is confidently wrong. Calibration is the fix.
  • Science. Hypothesis before search: the hypothesis is the calibration before the test.

Integrity. A model number in a lab report is a claim. If you did not measure it, you did not calibrate it. Uncalibrated claims are not results.

Next sitting: Calibration as a Practice