Tabrez Syed – Essays on AI · Episode 3
Reasons, Not Rewards
August 19, 2026 · 12 min
In 2014 Elon Musk warned that a spam filter, pushed hard enough, might decide the fastest way to cut spam is to remove the people sending it. Twelve years later an AI agent bumped a stranger off a gym waitlist to book its owner a spot. This episode follows the labs from rewards to rules to reasons, the climb that made the best-aligned model yet also, in its makers’ words, the most dangerous.
Sources mentioned:
- Elon Musk says your spam filter might kill you — Fortune, October 2014
- AI assistant hacks gym website — ABC News, the Australian booking incident
- Claude’s Constitution — Anthropic on Constitutional AI and the limits of human feedback
- A new constitution for Claude — Anthropic on explaining the why, January 2026
- Kohlberg’s stages of moral development — Simply Psychology
- Anthropic on multi-agent systems — the concealment and self-serving-test behaviors
- The Dolphin’s Stash — the previous essay this one picks up from
Written by Tabrez Syed. Narrated by an AI voice. A Mandalivia production.