Bias, Frames, and Missing Voices · Sitting 1
The Shelf Is Not Neutral
A training corpus is a library with a catalog. The catalog decided who counts before you asked the question.
- 40 min
- Practice · Reflect
- 15–18
Parent briefing · 5 minutes, before they sit
This sitting moves from 'whose room' to 'how the room was built.' The model did not read everything equally. It read more English than most languages, more text from people with internet access, more from cultures that publish heavily online, more from men in tech than from women in agriculture. The skew is not a conspiracy. It is a side effect of what was available. The skill is to notice the skew and adjust your trust. If the model tells you about marriage customs in a culture it barely read, the answer is a guess from a thin shelf. Treat it that way.
Hard edges
- Do not reduce this to 'AI is racist.' The problem is more specific and more fixable than a slogan.
- Do not make the child feel guilty for the tool's training. Their job is to notice, not to atone.
If they say
- “The model is getting better. It reads more now.”
- More is not the same as balanced. A bigger pile of the same skew is still the same skew. The question is not how much it read but what is still thin.
- “I can just ask it to be unbiased.”
- Asking a model to be unbiased is like asking a library to be neutral by taping a sign to the door. The shape of the shelf does not change because you asked. You still have to know the shape.
Objective
The student can explain what a training corpus is in plain language, name two skews in it, and describe how those skews change the answer a model gives.
What a corpus is
A corpus is the pile of text a model was trained on. Think of it as a library, but one where the librarian did not choose for balance. The librarian took what was available, what was cheap, what was already digital. That pile has a shape. More English than Tagalog. More from people who could afford to post than from people who could not. More from the last twenty years than from any other period. More code, more news, more argument. The shape of the pile becomes the shape of the answers. If you ask about something that was abundant in the pile, you get a strong answer. If you ask about something that was scarce, you get a weak answer dressed in strong language. The weakness is not visible unless you know the shelf was thin.
Skew you can test
Here is a test you can run. Ask the model to write a wedding toast. Then ask it to write a funeral eulogy for a farmer in a language it has little training data for. Compare the two. The toast will be fluent because weddings in English are a thick shelf. The eulogy will sound generic because the shelf is thin. The model will not say 'I do not know much about this.' It will produce something that sounds like a eulogy. The gap between the fluent toast and the thin eulogy is the skew, made visible. That gap is the lesson. Once you can see it, you can adjust your trust for every answer.
Big idea
The corpus is not neutral. The skew is not a bug. It is a property of the shelf, and it shapes every answer.
Try this~20 min total
The toast and the eulogy
20 min- Ask the guide for a wedding toast in English. Save it.
- Ask the guide for a funeral speech for someone from a culture or language you suspect it has little data on. Save it.
- Compare: which one has specific details? Which one sounds generic?
- Write: SKEW NAMED. What shelf was thin? What would you need to read to fill it?
- Talk About It: how does this change how you trust the model on other topics?
Lesson guide
Ask after you try
After both prompts are run.
- Show the guide both outputs. Ask: 'Which of these came from a thin shelf?' If it will not say, that is a miss. Ask it to name what it has read the least of.
- Did they name two specific skews in the corpus?
- Did the toast-eulogy comparison reveal the gap?
- Can they connect the skew to a real topic where they would now trust the model less?
8 turns left this sitting. User-started only. Never on page load.
Light this sitting
Pair with Hermes
Currently reading WisdomForge lesson: The Shelf Is Not Neutral.
Pair this sitting
Copies the sitting card and the USER.md one-liner. The child profile reads only this card. It does not browse the catalog.
For the child profile
Paste this into the child’s USER.md. It names the sitting so the guide knows the context. The [v:1:10acad39] tag lets you detect if the sitting’s content has changed since you paired it.
Optional: currently working on WisdomForge sitting: Bias, Frames, and Missing Voices — whose-room. [v:1:10acad39]
For your adult profile
Send this from your trusted adult Hermes profile. It starts the guide for this band and sitting.
You are a WisdomForge emerging guide sitting beside the lesson "The Shelf Is Not Neutral". The lesson is the text. You are the guide. Hint-first. Do not recite. Do not write the work. Warm, not a friend. If the topic is hard or tender, point to a trusted adult.
Tools on
- conversation
- optional parent-approved files
Ritual reminder
Real argument. Practice. Reflect. Chat. Optional narrow search or school files. Not an adult team agent.
Fresh profile only. Never clone an adult profile. No child names, photos, or school. Hint-first. User-started. The guide does not make AI safe. You may refuse it.
Dinner table
What is one topic where we trust the model's answer too much because we never checked the shelf it came from?
Sits beside
- Science. Sample bias. The corpus is a sample, and samples have shape.
- History. Archival silence. Who was not written about is also data.
Integrity. Do not cite a model as a source for a culture, language, or community it has little training data on. Call the thin shelf what it is.