WEBVTT

00:00:00.000 --> 00:00:11.520
This is Essays on AI. I'm Tabrez Syed. Actually, no. I'm an AI voice reading his essays. But

00:00:11.520 --> 00:00:16.059
you probably guessed. This one is called, When the Map Runs Out.

00:00:16.059 --> 00:00:19.879
The opening section draws from Timothy B. Lee's excellent explanation of reinforcement

00:00:19.879 --> 00:00:25.959
learning at his newsletter, Understanding AI. In 2009, a graduate student named Stefan

00:00:25.959 --> 00:00:30.719
Ross at Carnegie Mellon University was trying to solve a problem that seemed straightforward

00:00:30.719 --> 00:00:36.400
– teach a computer to play Super Tux Kart, an open-source racing game similar to Mario

00:00:36.400 --> 00:00:42.340
Kart. His approach was elegantly simple. Ross would play the game himself, while his software

00:00:42.340 --> 00:00:47.279
captured screenshots and recorded which buttons he pressed. Then he'd train a neural network

00:00:47.279 --> 00:00:52.279
to predict which buttons Ross would push in any given situation. If the network could

00:00:52.279 --> 00:00:57.599
learn to mimic Ross' gameplay perfectly, it should be able to play the game itself.

00:00:57.599 --> 00:01:01.880
Just push the same buttons Ross would have pushed in the same situations.

00:01:01.880 --> 00:01:06.760
The initial results looked promising. The AI-controlled car would zip around the track,

00:01:06.760 --> 00:01:11.400
taking turns at reasonable speeds, staying roughly in the center of the road. For a few

00:01:11.400 --> 00:01:16.279
seconds, it looked like Ross had cracked the code. Then something would go wrong. The car

00:01:16.279 --> 00:01:21.000
would drift slightly to the left, or take a turn a fraction too wide, or hit a small

00:01:21.040 --> 00:01:26.639
bump that pushed it off its ideal line. These weren't catastrophic errors, just tiny deviations

00:01:26.639 --> 00:01:32.879
from the perfect gameplay Ross had demonstrated. But tiny deviations became bigger ones. The

00:01:32.879 --> 00:01:38.680
car would drift further off course, then make increasingly erratic corrections. Within moments,

00:01:38.680 --> 00:01:43.400
what had started as a small mistake would cascade into complete chaos. The animated

00:01:43.400 --> 00:01:48.120
vehicle would careen off the track entirely, tumbling into the virtual abyss while Ross

00:01:48.360 --> 00:01:53.239
watched in frustration. The problem wasn't that the neural network was stupid. It had learned

00:01:53.239 --> 00:01:59.320
Ross' driving patterns remarkably well. The issue was more subtle and more fundamental. Ross was a

00:01:59.320 --> 00:02:04.040
pretty good SuperTuxCart player, which meant his car spent most of its time near the center of the

00:02:04.040 --> 00:02:10.520
road, driving at reasonable speeds, taking smooth turns. The neural network had learned to drive

00:02:10.520 --> 00:02:15.720
perfectly, as long as everything went perfectly. But it had almost no training data showing what

00:02:15.720 --> 00:02:21.240
to do when things went wrong. When the car drifted off course, it found itself in situations

00:02:21.240 --> 00:02:26.199
that weren't well represented in its training data. And in those unfamiliar situations,

00:02:26.199 --> 00:02:31.160
it made increasingly poor decisions. Ross had discovered something that would become crucial

00:02:31.160 --> 00:02:36.759
to understanding artificial intelligence. Systems trained on past experience work beautifully

00:02:36.759 --> 00:02:41.559
until they encounter situations that weren't in their training data. Then they don't just fail,

00:02:41.559 --> 00:02:45.960
they fail catastrophically, with small errors compounding into complete breakdowns.

00:02:46.759 --> 00:02:52.360
Ross, who now works on self-driving cars at Waymo, formerly Google's self-driving car project,

00:02:52.360 --> 00:02:57.160
had identified a fundamental challenge that would prove relevant far beyond video games.

00:02:57.160 --> 00:03:02.199
This wasn't just a quirk of Ross' particular approach. It's how most AI systems learn,

00:03:02.199 --> 00:03:07.240
by studying massive amounts of data and finding patterns. Language models learn by processing

00:03:07.240 --> 00:03:11.960
billions of text examples. Image recognition systems train on millions of labeled photos.

00:03:12.520 --> 00:03:17.000
The fundamental approach is the same. Show the system enough examples, and it will learn to

00:03:17.000 --> 00:03:22.600
recognize and respond to similar situations. But Ross' SuperTuxKart car revealed the Achilles'

00:03:22.600 --> 00:03:26.520
heel of this approach. What happens when you encounter something that wasn't in the training

00:03:26.520 --> 00:03:31.880
data? Ross' problem wasn't unique to video games. The same challenge was playing out in the real

00:03:31.880 --> 00:03:36.919
world, where the stakes were considerably higher than a virtual car tumbling into a digital abyss.

00:03:37.240 --> 00:03:42.679
Google had started its self-driving car project in 2009, the same year Ross was wrestling with

00:03:42.679 --> 00:03:47.880
his racing game. In typical Google fashion, they took a Google approach to the training

00:03:47.880 --> 00:03:53.960
data problem. They decided to collect massive amounts of it. For nearly a decade, from 2009

00:03:53.960 --> 00:03:59.880
to 2018, Waymo vehicles drove millions of miles with human safety drivers behind the wheel,

00:04:00.520 --> 00:04:05.720
meticulously recording every scenario, every edge case, every moment when human judgment was

00:04:05.720 --> 00:04:11.880
required. The strategy worked. By the time Waymo launched its first commercial service in December

00:04:11.880 --> 00:04:18.200
2018, their cars had logged over 10 million miles on public roads. The vehicles that had once

00:04:18.200 --> 00:04:24.440
struggled with basic scenarios, like Ross' SuperTuxKart model drifting off course, more data had largely

00:04:24.440 --> 00:04:30.119
solved the SuperTuxKart problem. Where Ross' model failed after encountering a single unfamiliar

00:04:30.119 --> 00:04:35.160
situation, Waymo's cars could handle thousands of edge cases they'd seen during their extensive

00:04:35.160 --> 00:04:41.399
training period. The compounding errors that plagued early AI systems became increasingly rare

00:04:41.399 --> 00:04:47.079
as the training datasets grew more comprehensive. This approach represents how we've tackled most AI

00:04:47.079 --> 00:04:53.720
challenges. Identify the gaps in training data, then systematically fill them. Can't recognize cats

00:04:53.720 --> 00:05:00.679
in photos? Train on more cat images. Struggling with medical diagnoses? Feed the system more patient

00:05:00.679 --> 00:05:06.679
records. Having trouble with language translation? Process more bilingual text pairs. Today, we've

00:05:06.679 --> 00:05:12.200
taken this data collection strategy even further. AI systems can now manufacture their own training

00:05:12.200 --> 00:05:18.200
data through synthetic data generation. Need more examples of rare medical conditions? AI can

00:05:18.200 --> 00:05:23.640
generate realistic patient scenarios. Want to train a model on edge cases that rarely occur in real

00:05:23.640 --> 00:05:29.239
driving? AI can create thousands of synthetic scenarios with unusual weather conditions,

00:05:29.239 --> 00:05:35.399
unexpected obstacles, or complex traffic patterns. This ability to generate synthetic data means we

00:05:35.399 --> 00:05:40.040
can fill gaps in training datasets more systematically than ever before. It's a

00:05:40.040 --> 00:05:45.399
powerful strategy, and it works remarkably well for a specific type of problem. Statisticians

00:05:45.399 --> 00:05:50.600
call this epistemic uncertainty. Uncertainty that exists because we don't know something, but we

00:05:50.600 --> 00:05:56.359
could learn it if we gathered more information. Ross' SuperTuxKart model suffered from epistemic

00:05:56.359 --> 00:06:01.160
uncertainty. It didn't know how to recover from mistakes because it had never seen examples of

00:06:01.160 --> 00:06:07.320
recovery. Waymo's early struggles were also largely epistemic. Their cars couldn't handle construction

00:06:07.320 --> 00:06:14.119
zones because they hadn't seen enough construction zones. But epistemic uncertainty is solvable. You

00:06:14.119 --> 00:06:19.799
can collect more data, train on more examples, and gradually fill in the gaps. And here's where AI

00:06:19.799 --> 00:06:25.320
systems will eventually surpass human intelligence. They can process vastly more data than any human

00:06:25.320 --> 00:06:30.519
could ever experience. While a human driver might encounter a few dozen construction zones in their

00:06:30.519 --> 00:06:35.880
lifetime, an AI system can learn from millions of construction zone scenarios captured by thousands

00:06:35.880 --> 00:06:42.440
of vehicles. The epistemic uncertainty gap between humans and AI is temporary. Humans currently have

00:06:42.440 --> 00:06:47.399
an advantage because we've been training on visual and spatial data since birth, accumulating decades

00:06:47.399 --> 00:06:53.000
of experience navigating the world. But AI systems are rapidly catching up, and they'll eventually

00:06:53.000 --> 00:06:58.040
have access to far more comprehensive datasets than any individual human could process. For

00:06:58.040 --> 00:07:04.440
epistemic uncertainty, the solution is clear. More data, better algorithms, and time. Here's where the

00:07:04.440 --> 00:07:10.440
story takes a turn. Our predictions of the future are reflections from our rearview mirror. Even with

00:07:10.440 --> 00:07:15.640
over 20 million miles of real-world driving data, Waymo's systems still encounter situations that

00:07:15.640 --> 00:07:20.519
weren't in their training. Not because of gaps in the dataset, but because the future draws from a

00:07:20.519 --> 00:07:26.839
vast theoretical space of unrealized possibilities. Consider the microbursts that hit Austin in 2025.

00:07:27.559 --> 00:07:32.040
A sudden downdraft brought down trees and scattered debris across roads in patterns no

00:07:32.040 --> 00:07:37.959
training dataset had ever captured. Or the early days of the COVID-19 pandemic, when governments

00:07:37.959 --> 00:07:43.799
implemented lockdown policies based on historical precedent from vastly different eras. In this

00:07:43.799 --> 00:07:49.880
theoretical space, tree limbs can land on roads, locusts can swarm highways, microbursts can

00:07:49.880 --> 00:07:55.160
scatter debris in patterns no algorithm has ever seen. These aren't just rare events, they're

00:07:55.160 --> 00:08:00.839
genuinely unprecedented combinations that emerge from the complex interactions of countless variables.

00:08:00.839 --> 00:08:07.000
These events represent aleatory uncertainty, from the Latin word for dice, the inherent randomness

00:08:07.000 --> 00:08:11.959
and unpredictability in the world itself. Even our synthetic data generation can only create

00:08:11.959 --> 00:08:17.079
variations of what we can imagine from past experience. We extrapolate from history, but the

00:08:17.079 --> 00:08:22.519
future contains genuinely novel combinations that exist in theoretical possibility but have never

00:08:22.519 --> 00:08:28.760
manifested in any dataset. You could collect data for eons, train on every scenario from the past,

00:08:28.760 --> 00:08:34.280
and still encounter situations that have never existed before. This is because there is no data

00:08:34.280 --> 00:08:39.479
about the future, not because we haven't collected enough data about the past, but because the future

00:08:39.479 --> 00:08:45.080
contains genuinely novel combinations of factors that have never existed before. Here's what makes

00:08:45.080 --> 00:08:50.359
this distinction important. Both human intelligence and artificial intelligence are sophisticated

00:08:50.359 --> 00:08:55.719
pattern matching systems. We excel at recognizing situations we've encountered before and applying

00:08:55.719 --> 00:09:01.239
lessons from past experience. But when faced with genuinely unprecedented events, both types of

00:09:01.239 --> 00:09:06.840
intelligence hit the same fundamental wall. Consider how humans handled the COVID-19 pandemic.

00:09:06.840 --> 00:09:12.599
We had data from previous pandemics, the 1918 flu outbreak, various smaller disease outbreaks,

00:09:12.599 --> 00:09:17.880
historical accounts of quarantine measures. But the world of 2020 was fundamentally different

00:09:17.880 --> 00:09:23.960
from 1918. Our interconnected global economy, modern healthcare systems, digital communication

00:09:23.960 --> 00:09:30.440
networks, and social structures created a context that had never existed before. The result? Human

00:09:30.440 --> 00:09:35.799
decision makers implemented policies that were essentially educated guesses. Lockdown strategies

00:09:35.799 --> 00:09:41.239
varied wildly between countries and regions. Some measures worked, others didn't, and many had

00:09:41.239 --> 00:09:45.080
unintended consequences that no amount of historical analysis could have predicted.

00:09:45.719 --> 00:09:51.080
We were pattern matching from incomplete and often irrelevant historical data, just like an AI system

00:09:51.080 --> 00:09:56.200
trying to navigate an unprecedented scenario. The same dynamic plays out in financial markets

00:09:56.200 --> 00:10:01.559
during genuine crises, in political systems facing novel challenges, and in any domain where the

00:10:01.559 --> 00:10:06.919
future presents combinations of factors that have never occurred before. Humans don't have some

00:10:06.919 --> 00:10:12.679
magical ability to handle aleatory uncertainty better than AI systems. We're just as dependent

00:10:12.679 --> 00:10:17.719
on pattern matching from past experience. This distinction matters because it challenges a

00:10:17.719 --> 00:10:23.080
common assumption about human versus artificial intelligence. We often think of humans as

00:10:23.080 --> 00:10:28.359
fundamentally superior to AI systems, possessing some ineffable quality that machines can never

00:10:28.359 --> 00:10:33.719
replicate. But when we separate epistemic uncertainty from aleatory uncertainty, a

00:10:33.719 --> 00:10:38.679
different picture emerges. For epistemic uncertainty, situations where more data and

00:10:38.679 --> 00:10:43.559
better pattern recognition can solve the problem, AI systems will likely surpass human performance.

00:10:44.200 --> 00:10:48.599
They can process more information, learn from more examples, and update their models more

00:10:48.599 --> 00:10:54.599
systematically than biological intelligence. For aleatory uncertainty, genuinely unprecedented

00:10:54.599 --> 00:10:59.320
events that emerge from novel combinations of factors, both human and artificial intelligence

00:10:59.320 --> 00:11:03.719
face the same fundamental limitation. Neither can predict what has never happened before.

00:11:04.520 --> 00:11:07.479
Neither can prepare for combinations of factors that have never existed.

00:11:08.280 --> 00:11:12.840
But humans have developed something valuable over millennia of facing the unknown—mental

00:11:12.840 --> 00:11:17.799
models and frameworks for navigating uncertainty itself. We've learned to reason by analogy,

00:11:17.799 --> 00:11:22.840
to break down complex problems into manageable parts, to consider multiple scenarios simultaneously.

00:11:23.479 --> 00:11:27.640
These aren't perfect solutions, but they're evolved responses to the fundamental challenge

00:11:27.640 --> 00:11:32.520
of operating without complete information. AI systems will likely learn these same mental

00:11:32.520 --> 00:11:36.760
models from the vast repository of human experience they train on. They'll absorb our

00:11:36.760 --> 00:11:40.919
approaches to uncertainty, our problem-solving frameworks, our ways of thinking through

00:11:40.919 --> 00:11:46.599
unprecedented situations. But they'll also inherit our biases, our blind spots, and our systematic

00:11:46.599 --> 00:11:51.719
errors. After all, they're learning from our exhaust, both the wisdom and the mistakes embedded

00:11:51.719 --> 00:11:57.239
in human decision-making. The future will always contain surprises that no amount of historical

00:11:57.239 --> 00:12:02.520
data can anticipate. But this raises a deeper question about the nature of intelligence itself.

00:12:03.320 --> 00:12:06.919
Is human decision-making reducible to a set of mental models and rules?

00:12:07.559 --> 00:12:12.520
Or is there more to gut feeling than the total sum of carbon-based neurons buzzing in our heads?

00:12:13.320 --> 00:12:17.799
Thanks for listening. The full essay, with links to Timothy B. Lee's explainer

00:12:17.799 --> 00:12:21.719
and Stéphane Ross's paper, is on the episode page at mandolivia.com.

00:12:22.440 --> 00:12:25.799
If you'd like more essays and audio, subscribe wherever you're listening.

00:12:25.799 --> 00:12:30.520
Essays on AI is a Mandalivia production.

