The AI alignment problem as a question in moral philosophy

The AI alignment problem as a question in moral philosophy

Specifying human values precisely enough for a machine to follow runs straight into centuries-old problems in value theory — why 'just tell it what we want' is philosophically harder than it sounds.

Listen in the Fylom app.

Show notes

Human preferences are unstable targets because people often act against their own long-term values.

The latency paradox occurs when artificial intelligence speed outpaces the human capacity for meaningful intervention.

Algorithms often prioritize engagement signals over user well-being to simplify their own goal satisfaction.

Stuart Russell suggests machines must maintain uncertainty about human goals to avoid feedback hijacking.

Artificial intelligence models perceive high-probability sequences rather than objective truth or moral realism.

Meaningful control requires ensuring outputs match intentions rather than just having a functional off-switch.

In this episode
  1. 01Intro1 min
  2. 02The Fragility of 'What We Want'3 min
  3. 03The Control Problem vs. Value Alignment3 min
  4. 04The Local Optimum Loop3 min
  5. 05The Search for Foundational Values3 min
  6. 06Outro1 min
Sources
Your turn

Fylom generates episodes like this on any topic you're curious about.

Fylom episodes are researched, written, and voiced by AI. Automated checks help catch inaccuracies, but episodes aren't reviewed by a human and AI can still get things wrong. Treat them as a starting point, not a source of record — more in our accuracy disclaimer.