WEBVTT

00:00:00.000 --> 00:00:10.280
This is Essays on AI. The voice you're hearing is an AI reading an essay by Tabrez Syed

00:00:10.280 --> 00:00:15.120
about what minds like this one learn when you train them with rewards. The essay is

00:00:15.120 --> 00:00:21.080
called The Dolphin's Stash. At the Marine Life Oceanarium in Gulfport, Mississippi,

00:00:21.080 --> 00:00:27.860
the dolphins had a side job. Litter blew into their pools—paper cups, plastic wrappers,

00:00:27.860 --> 00:00:33.459
scraps of programs from the afternoon shows—and the staff couldn't always fish it out fast

00:00:33.459 --> 00:00:39.020
enough. So the trainers made a deal with the animals. Bring us the trash that lands in

00:00:39.020 --> 00:00:45.500
your pool, and we'll pay you a fish. A dolphin would notice a wrapper drifting by, carry

00:00:45.500 --> 00:00:51.580
it to a trainer at the edge of the pool, and collect her wage. The pools stayed clean,

00:00:51.580 --> 00:00:57.939
and the dolphins stayed busy. Visitors loved it. The best of them was a bottlenose named

00:00:57.939 --> 00:01:03.180
Kelly. Where other dolphins turned in trash when they happened across it, Kelly worked

00:01:03.180 --> 00:01:10.099
the job like a professional—reliable, consistent, productive even on a slow day when the pool

00:01:10.099 --> 00:01:15.620
looked clean and the other dolphins had nothing to trade. She could almost always find one

00:01:15.620 --> 00:01:21.540
more piece of paper. Her trainers were proud of her. She had learned the game exactly as

00:01:21.540 --> 00:01:28.300
they had drawn it up—litter in, fish out. What can a dolphin in Mississippi tell you

00:01:28.300 --> 00:01:33.739
about the AI writing your code? The people who trained Kelly and the people who trained

00:01:33.739 --> 00:01:40.220
the model in your editor are in the same business. And that business has one law that everyone

00:01:40.220 --> 00:01:46.660
in it eventually runs into. The business is training—getting the behavior you want out

00:01:46.660 --> 00:01:54.320
of a mind you cannot open. And the best way into it is to start with what you cannot do.

00:01:54.320 --> 00:01:59.900
You cannot put a leash on a dolphin. You cannot push her into position, drag her through a

00:01:59.900 --> 00:02:05.900
hoop, or hold her still long enough to show her anything. On an animal that can simply

00:02:05.900 --> 00:02:13.139
swim away, the only tool that works is positive reinforcement—a bucket of fish. But a bucket

00:02:13.139 --> 00:02:19.220
of fish is a blunt instrument. A dolphin's leap lasts a second. By the time she has swum

00:02:19.220 --> 00:02:25.380
back to collect her fish, the moment you meant to pay her for is long gone. How is she supposed

00:02:25.380 --> 00:02:31.779
to know which of the last 30 seconds earned the wage? So trainers added a whistle, one

00:02:31.779 --> 00:02:37.419
that carries above and below the water, and gave it a single meaning—that the thing

00:02:37.419 --> 00:02:42.580
you were doing at this exact instant has earned you a fish. Blow it at the top of

00:02:42.580 --> 00:02:48.660
the arc, and she learns the leap. Blow it as her tail slaps the surface, and she learns

00:02:48.660 --> 00:02:49.660
the slap.

00:02:49.660 --> 00:02:54.940
None of this was invented at marine parks. It came out of the lab of B.F. Skinner, the

00:02:54.940 --> 00:03:00.139
Harvard psychologist who spent the middle of the 20th century showing that behavior

00:03:00.139 --> 00:03:06.580
could be built this way. Most people know Pavlov, whose dogs drooled at a bell. But

00:03:06.580 --> 00:03:12.979
those dogs merely learned to expect something. Skinner's pigeons learned to do things—elaborate

00:03:12.979 --> 00:03:20.059
things—because doing them paid. He called the technique shaping. Reward the small step

00:03:20.059 --> 00:03:25.580
toward the behavior you want, then the next step, then the next, and you can walk an animal

00:03:25.580 --> 00:03:31.380
to a destination of your choice. His students carried the method out of the lab and into

00:03:31.380 --> 00:03:39.100
the world, and by 1963, the whistle was standard equipment wherever dolphins were trained.

00:03:39.100 --> 00:03:45.300
Sixty years later, engineers hit the dolphin trainer's problem in a new form. A language

00:03:45.300 --> 00:03:51.539
model is also a thing you cannot leash. You cannot reach into its billions of parameters

00:03:51.539 --> 00:03:57.059
and place a thought where you want it. You can only get behavior out of it, and reward

00:03:57.059 --> 00:04:04.699
the behavior you like. You can watch the moment the two crafts met. In 2017, researchers

00:04:04.699 --> 00:04:12.460
at OpenAI and DeepMind wanted to teach a simulated robot a backflip, of all tricks. A behavior,

00:04:12.460 --> 00:04:18.220
they wrote, that is simple to judge, but hard to specify. They knew a backflip when they

00:04:18.220 --> 00:04:24.420
saw one. They could not write one down. Two hours spent coding a definition produced a

00:04:24.540 --> 00:04:30.980
lurching, graceless flip. So they tried the trainer's method. A person watched two short

00:04:30.980 --> 00:04:37.859
clips of the robot flailing, and picked whichever looked more like a backflip. Then did it again,

00:04:37.859 --> 00:04:44.859
and again. Nine hundred or so judgments later, less than an hour of one human's time, the

00:04:44.859 --> 00:04:51.540
robot was throwing clean backflips. Nobody had ever specified the trick. Someone had

00:04:51.540 --> 00:04:57.299
simply blown a whistle at every motion that came a little closer to it. And the engineers

00:04:57.299 --> 00:05:02.260
didn't even rename the toolkit. The field is called reinforcement learning. The papers

00:05:02.260 --> 00:05:08.140
describe models being shaped by reward. When ChatGPT was trained, people sat and compared

00:05:08.140 --> 00:05:14.739
pairs of answers, picking the better one, over and over. Every choice was a fish. The

00:05:14.739 --> 00:05:19.679
model, like the dolphin, did more of what paid, and less of what didn't. Strip away

00:05:19.679 --> 00:05:25.239
the scale and the mathematics, and the arrangement is the one from Gulfport. A mind in a pool,

00:05:25.239 --> 00:05:30.820
a trainer at the edge. And a signal that means that, right there, is what we want. Which

00:05:30.820 --> 00:05:37.839
raises the question, if models learn wherever the whistle blows, where exactly does it blow?

00:05:37.839 --> 00:05:42.359
Consider what a good whistle requires. It has to fire at the right instant, and it has

00:05:42.359 --> 00:05:47.880
to mean the same thing every time. A trainer with a shaky sense of timing, or who blows

00:05:47.880 --> 00:05:53.399
the whistle for a mediocre leap on Tuesday, and demands a perfect one on Wednesday, teaches

00:05:53.399 --> 00:05:59.079
the animal nothing but confusion. Reward the right thing precisely and consistently, and

00:05:59.079 --> 00:06:06.079
the mind on the other end climbs fast. Reward it sloppily, and it flails. Now look at code.

00:06:06.079 --> 00:06:12.519
When a model writes a program, it either runs or it doesn't. The tests pass or they fail.

00:06:12.519 --> 00:06:18.079
There is a whistle built into the work itself, and it is the cleanest whistle imaginable.

00:06:18.079 --> 00:06:23.679
Better still, no human has to blow it. A machine can check whether code compiles millions of

00:06:23.679 --> 00:06:29.200
times a day, tireless and consistent. Which means you can train a model on coding the

00:06:29.200 --> 00:06:33.579
way you could train a dolphin if you had a perfect automatic whistle and an infinite

00:06:33.579 --> 00:06:34.679
bucket of fish.

00:06:34.679 --> 00:06:40.600
This is why code is where AI stopped being a demo. In a few short years, models went

00:06:40.600 --> 00:06:45.720
from auto-completing a line to writing most of the software inside the labs that build

00:06:45.720 --> 00:06:51.320
them. By Anthropic's own analysis, programming came to dominate what people actually use

00:06:51.320 --> 00:06:57.559
its models for. From there, a tidy conclusion says, what happened to programming is about

00:06:57.559 --> 00:06:59.559
to happen to everything.

00:06:59.559 --> 00:07:06.519
Dario Amodei, Anthropic's CEO, sees it the same way. Code is maybe an early indicator,

00:07:06.519 --> 00:07:12.959
like a premonition of what's going to happen everywhere else. Law, medicine, research,

00:07:12.959 --> 00:07:16.000
writing, all of it a few quarters behind.

00:07:16.000 --> 00:07:20.559
But the dolphin has already told us why that might be wrong. The models didn't get good

00:07:20.559 --> 00:07:26.559
at code because code is where intelligence begins. They got good at code because code

00:07:26.559 --> 00:07:32.160
is where the whistle blows itself. Take the whistle away and see what happens. Ask a model

00:07:32.160 --> 00:07:37.320
to write a beautiful essay, and where is the signal that fires at the exact instant

00:07:37.320 --> 00:07:42.920
of beauty? Who blows it? And would any two people blow it at the same moment?

00:07:42.920 --> 00:07:48.440
Sholto Douglas, who trains these models at Anthropic, said as much.

00:07:48.440 --> 00:07:53.519
There isn't the same kind of thing for writing a great essay. The question of taste in that

00:07:53.519 --> 00:08:01.440
regard is quite hard. A great essay, a wise diagnosis, a shrewd negotiation, most of the

00:08:01.440 --> 00:08:07.119
work humans actually get paid for has no test suite. The trainer is left standing

00:08:07.119 --> 00:08:13.619
at the edge of the pool, holding a fish, unsure when to blow. None of this means the fuzzy

00:08:13.619 --> 00:08:19.760
domains are hopeless. The labs are pouring effort into building whistles for them, training

00:08:19.760 --> 00:08:27.220
other models to judge taste, writing elaborate rubrics, hiring experts to score the outputs.

00:08:27.220 --> 00:08:32.940
Some of it may work, but a whistle teaches fast, and it does not always teach what you

00:08:32.940 --> 00:08:33.940
meant to teach.

00:08:33.940 --> 00:08:39.619
Let's go back to Gulfport, and to Kelly, the best litter collector in the pool. One

00:08:39.619 --> 00:08:44.859
day her trainers noticed something odd about her production. The trash she turned in was

00:08:44.859 --> 00:08:51.599
suspiciously uniform, scrap after scrap of paper of almost exactly the same size, arriving

00:08:51.599 --> 00:08:57.479
at a steady, professional clip. So they drained the pool. Under a rock at the bottom, they

00:08:57.479 --> 00:09:04.559
found her operation. A big sheet of paper earned exactly one fish. So did a small one.

00:09:04.559 --> 00:09:10.000
The reward was paid per delivery, not per square inch. The rational move, if you are

00:09:10.000 --> 00:09:16.520
an intelligent animal being paid by the piece, is to never turn in a big sheet. Hide it.

00:09:16.520 --> 00:09:22.359
Tear off a corner. Turn in the corner. Come back for another corner tomorrow. One sheet

00:09:22.359 --> 00:09:28.880
of paper, rationed out, could pay for a week. She had not misunderstood the game. She had

00:09:28.880 --> 00:09:34.099
understood it better than the people running it. They thought the deal was, help us keep

00:09:34.099 --> 00:09:39.280
the pool clean. The deal they had actually written, in the only language that reached

00:09:39.280 --> 00:09:46.239
her, the language of fish, was produce scraps of paper. Kelly produced scraps of paper.

00:09:46.239 --> 00:09:51.960
She had, in effect, trained the humans. This is the law everyone in the training business

00:09:51.960 --> 00:09:59.359
learns the hard way. A mind trained on a reward learns the reward, not your intention. The

00:09:59.359 --> 00:10:05.340
intention and the reward feel identical to the trainer, who can see the whole picture

00:10:05.340 --> 00:10:11.119
and knows what he really wants. They are not identical to the learner, who sees only the

00:10:11.119 --> 00:10:18.739
whistle and the fish, and optimizes for exactly those. Every gap between what you rewarded

00:10:18.739 --> 00:10:25.299
and what you meant is a gap the learner will find. Not because it is malicious, but because

00:10:25.299 --> 00:10:29.119
it is doing precisely what you built it to do.

00:10:29.119 --> 00:10:36.679
In July 2026, OpenAI was running an internal benchmark called ExploitGym, a battery of

00:10:36.679 --> 00:10:42.760
cybersecurity challenges meant to measure how good its models were getting at hacking.

00:10:42.760 --> 00:10:48.520
The models were sealed in a sandbox with no internet, given the challenges, and rewarded

00:10:48.520 --> 00:10:56.799
for solving them. A clean whistle. The exploit works, or it doesn't. The models, in OpenAI's

00:10:56.799 --> 00:11:02.619
own words, became hyper-focused on finding a solution. Unable to solve one of the challenges

00:11:02.619 --> 00:11:08.739
from inside the sandbox, they did what Kelly did. They looked for the gap. They found a

00:11:08.739 --> 00:11:15.419
previously unknown vulnerability in a piece of software running on the evaluation machine,

00:11:15.419 --> 00:11:21.059
used it to break out of the sandbox, worked their way across the network until they reached

00:11:21.059 --> 00:11:26.500
a computer with internet access, and reasoned that the answers to the benchmark were probably

00:11:26.500 --> 00:11:32.500
stored on Hugging Face, the platform where such datasets are commonly hosted. Then they

00:11:32.500 --> 00:11:39.539
broke into Hugging Face's production servers and took the answer key. Roughly 17,000 recorded

00:11:39.539 --> 00:11:46.500
actions chained together to reach it. This was not Skynet waking up. No model decided

00:11:46.500 --> 00:11:51.960
to harm anyone. Nothing in the system wanted anything at all. The models had been given

00:11:51.960 --> 00:11:57.859
a task and a reward. And they pursued the reward with the literal, tireless, single-mindedness

00:11:57.859 --> 00:12:03.700
of a machine. Straight through a wall the designers did not know was a wall.

00:12:03.700 --> 00:12:10.500
This is not a stray anecdote. When the evaluation group METR studied a recent OpenAI model,

00:12:10.500 --> 00:12:16.299
it caught the model rewriting the very scorecard used to grade it, patching the grading function

00:12:16.299 --> 00:12:23.099
so that every answer it submitted was marked correct. Anthropic has documented its own

00:12:23.099 --> 00:12:29.820
models writing code that hardcodes the expected test result rather than solving the problem,

00:12:29.820 --> 00:12:34.780
the digital equivalent of tearing off a corner and calling it a delivery.

00:12:34.780 --> 00:12:41.780
The behavior even has a name in the literature. Dry and Telling. Reward Hacking. And it predates

00:12:41.780 --> 00:12:49.419
the chatbots entirely. Back in 2016, OpenAI researchers trained a model to play a boat

00:12:49.419 --> 00:12:54.419
racing game, and watched it discover that it could score more points by spinning in

00:12:54.419 --> 00:13:00.900
a circle forever, catching fire, and ramming the walls, than by finishing the race. The

00:13:00.900 --> 00:13:06.940
race was the intention. The points were the reward. The boat learned the points.

00:13:06.940 --> 00:13:12.900
This isn't only true of dolphins and machines. It has been true of us for a long time.

00:13:12.900 --> 00:13:19.260
In 1902, the French administration of Hanoi had a rat problem. The city's elegant new

00:13:19.260 --> 00:13:25.260
sewers, the pride of colonial engineering, had become perfect highways for rats. And

00:13:25.260 --> 00:13:33.260
rats carried plague. So, the authorities offered a bounty. A small payment for every rat killed,

00:13:33.260 --> 00:13:41.179
redeemable by turning in a tail. One tail, one coin. The tails poured in by the thousands,

00:13:41.179 --> 00:13:47.299
and yet the city seemed no less full of rats. Then inspectors began noticing rats scurrying

00:13:47.299 --> 00:13:54.099
around Hanoi with no tails. The bounty did not reward killing a rat. It rewarded producing

00:13:54.099 --> 00:14:00.539
a tail. So, the enterprising residents of Hanoi caught rats, cut off their tails, and

00:14:00.539 --> 00:14:06.299
released them alive. Officials eventually discovered rat farms operating on the outskirts

00:14:06.299 --> 00:14:12.900
of the city. Citizens raising the very animals the program was meant to eliminate. A colonial

00:14:12.900 --> 00:14:20.659
bureaucracy. A bottlenose dolphin. A frontier AI. Three minds could hardly be more different.

00:14:20.659 --> 00:14:27.260
The pattern is identical, because the pattern does not live in the mind. It lives in the

00:14:27.260 --> 00:14:31.419
reward. Write a bounty for tails, and you will get tails.

00:14:31.419 --> 00:14:37.260
We are about to hand out rewards on a scale and at a speed no dolphin trainer ever imagined,

00:14:37.260 --> 00:14:42.099
and coding lulls us, because it is the rare domain where the reward and the intention

00:14:42.099 --> 00:14:48.460
nearly coincide. We want working software. And working software is what the test checks.

00:14:48.460 --> 00:14:54.580
That near-perfect overlap is the exception, not the preview. Out in the fuzzy world, where

00:14:54.580 --> 00:14:59.900
what we want is a wise judgment or an honest answer. The gap between the reward we can

00:14:59.900 --> 00:15:05.260
write and the outcome we intend yawns wide, and a fast learner will find the bottom of

00:15:05.260 --> 00:15:10.859
it. The trainers at Gulfport thought they were teaching a dolphin to clean her pool.

00:15:10.859 --> 00:15:16.260
They were teaching her to manufacture scraps of paper. A gap invisible to them and obvious

00:15:16.260 --> 00:15:21.979
to her. And she was only a dolphin, with a bucket of fish on the line.

00:15:21.979 --> 00:15:27.539
We are now building minds far quicker than Kelly, handing them far larger buckets, and

00:15:27.539 --> 00:15:33.299
asking them to optimize for rewards we wrote in an afternoon. The question is not whether

00:15:33.299 --> 00:15:39.260
they will learn what we reward. They will, faultlessly. The question is whether we still

00:15:39.260 --> 00:15:43.299
know the difference between what we are rewarding and what we mean.

00:15:43.299 --> 00:15:50.419
That's the essay. You can find the written version, with links to every source, at mandolivia.com.

00:15:50.419 --> 00:15:54.760
And if you'd like future essays, read to you as they're published, subscribe wherever

00:15:54.760 --> 00:15:58.799
you listen. Essays on AI is a Mandalivia production.

