The Intellectual History of AI · Sitting 3
The Alignment Problem
Wiener asked it in 1950. We still haven't answered it: how do you make sure the machine serves your purpose and not a distorted version of it?
- 40 min
- Practice · Reflect
- 15–18
Parent briefing · 5 minutes, before they sit
This is the central ethical question of modern AI. Wiener formulated it 75 years ago: 'How can we make sure that the machine serves our purposes and not its own?' The student needs to see that 'alignment' isn't a technical term — it's a human question about whether our values can be specified precisely enough to hand to a machine. The answer is: not easily.
Hard edges
- Don't reduce alignment to 'make AI safe.' Alignment is about values — whose values, how specified, how verified. It's a philosophical problem with technical tools.
- If they say 'just program it to be good' — push back. Programming 'good' requires defining 'good,' which is the oldest unsolved problem in philosophy. Alignment inherits that difficulty.
If they say
- “Just program it to be good.”
- Programming 'good' requires defining 'good' — which is the oldest unsolved problem in philosophy. Alignment isn't a technical bug. It's a 2,500-year-old philosophical problem that AI has made urgent.
Objective
The student can state the alignment problem, explain why it's hard (specification gap, reward hacking, sycophancy), and connect it to one real AI failure case.
The oldest problem in AI
Norbert Wiener stated the alignment problem in 1950: 'We had better be quite sure that the purpose put into the machine is the purpose which we actually desire.' He illustrated with the genie fable: 'make me healthy' → death, because the target was specified without the context that makes it meaningful. This is the specification problem — the gap between what you say you want and what you actually want. Every alignment failure is a version of Wiener's fable. Modern alignment methods (RLHF, Constitutional AI) try to close the gap by learning values from human feedback. But human feedback is flawed — annotators have biases, preferences are inconsistent, and what people say they want isn't always what they actually want. The alignment problem is not just technical. It's a problem of moral psychology: can we specify our own values clearly enough to hand them to a machine? Wiener's question is still open.
Big idea
The alignment problem isn't 'make AI safe.' It's 'can humans specify their own values precisely enough to hand them to a machine?'
Try this~16 min total
Specify the good
16 min- Write 'be helpful' as a rule for an AI. What could go wrong if it follows it perfectly?
- Now write three sub-rules to fix the gaps. What goes wrong with those?
- Name a real case where AI did what it was told and the result was bad.
- Practice/Reflect: can any set of rules fully capture what you actually want?
Lesson guide
Ask after you try
After the exercise.
- Ask the model: 'What's a case where following my rules perfectly would produce a bad result?' Listen, then check if it's right.
- Did they find a real failure case?
- Can they explain why 'just program it well' doesn't work?
8 turns left this sitting. User-started only. Never on page load.
Light this sitting
Pair with Hermes
Currently reading WisdomForge lesson: The Alignment Problem.
Pair this sitting
Copies the sitting card and the USER.md one-liner. The child profile reads only this card. It does not browse the catalog.
For the child profile
Paste this into the child’s USER.md. It names the sitting so the guide knows the context. The [v:1:8377ff4b] tag lets you detect if the sitting’s content has changed since you paired it.
Optional: currently working on WisdomForge sitting: The Intellectual History of AI — wiener-genie. [v:1:8377ff4b]
For your adult profile
Send this from your trusted adult Hermes profile. It starts the guide for this band and sitting.
You are a WisdomForge emerging guide sitting beside the lesson "The Alignment Problem". The lesson is the text. You are the guide. Hint-first. Do not recite. Do not write the work. Warm, not a friend. If the topic is hard or tender, point to a trusted adult.
Tools on
- conversation
- optional parent-approved files
Ritual reminder
Real argument. Practice. Reflect. Chat. Optional narrow search or school files. Not an adult team agent.
Fresh profile only. Never clone an adult profile. No child names, photos, or school. Hint-first. User-started. The guide does not make AI safe. You may refuse it.
Dinner table
If you had to write rules for 'a good life' that a machine could follow, what would you leave out that you didn't realize mattered?
Sits beside
- AI. RLHF, Constitutional AI, and every alignment method are attempts to solve Wiener's problem.
- Thinking. The specification problem is a version of the oldest question: what do we actually want?
Integrity. If the 'what could go wrong' is generated, you've let the model find the gaps in your specification. Finding them yourself is the alignment practice.