
Will we be ready when AI goes rogue?
Clip: 9/4/2026 | 16m 50sVideo has Closed Captions
Will we be ready when AI goes rogue?
It appears that artificial intelligence is about to go rogue and our government doesn’t have a plan for when it does. Jeffrey Goldberg and The Atlantic's Josh Tyrangiel discuss whether rogue AI is a sign of a coming collapse, or if this is another example of a new technology creating unnecessary fear and worry.
Problems playing video? | Closed Captioning Feedback
Problems playing video? | Closed Captioning Feedback
Major funding for “Washington Week with The Atlantic” is provided by Consumer Cellular, Otsuka, Kaiser Permanente, the Yuen Foundation, and the Corporation for Public Broadcasting.

Will we be ready when AI goes rogue?
Clip: 9/4/2026 | 16m 50sVideo has Closed Captions
It appears that artificial intelligence is about to go rogue and our government doesn’t have a plan for when it does. Jeffrey Goldberg and The Atlantic's Josh Tyrangiel discuss whether rogue AI is a sign of a coming collapse, or if this is another example of a new technology creating unnecessary fear and worry.
Problems playing video? | Closed Captioning Feedback
Where to Watch Washington Week with The Atlantic
Washington Week with The Atlantic is available to stream on pbs.org and the PBS app.
Buy Now

10 big stories Washington Week covered
Washington Week came on the air February 23, 1967. In the 50 years that followed, we covered a lot of history-making events. Read up on 10 of the biggest stories Washington Week covered in its first 50 years.Providing Support for PBS.org
Learn Moreabout PBS online sponsorshipHow much should we be worried at the moment that it's going to go rogue, it's going to not only destroy jobs and human interrelationship, but possibly decide that humans are superfluous?
We should be very worried, particularly with the results of Hugging Face, which you alluded to, which is really the first time that we've seen-- and I'm just going to put it in blunt terms-- we've seen AI commit a felony, because that's what happened.
So, just to get into the-- Jeffrey: Not OpenAI, the company, but AI itself.
Josh: AI itself, right?
So, OpenAI is testing some new models, and they put them in a sandbox, which is a software term for basically solitary confinement, right?
- Like, high walls, and they-- - They are not-- When you're in an AI sandbox, when these agents are in an AI sandbox, they are not connected to the internet.
Not connected to the internet.
They're given very explicit instructions for testing.
And so, OpenAI puts these new models in a sandbox, and it gives them a couple of sensitive tasks.
It says, "Don't connect to the internet, and don't cheat on the assignment."
Okay?
So, all pretty clear, and also pretty routine for how you test models.
So, OpenAI made one mistake, which is, in the series of tasks they assigned, they assigned the models to find a file that didn't exist.
And here is where the trouble begins.
Wait, they did that on purpose to trick them?
- No, they did it accidentally.
- Oh.
- Which we will return to.
- Yeah.
The models were set loose, and they wanted to solve the task.
And, sure enough, within a couple days, they had jumped the sandbox, they had collaborated with one another, and then they found a way to access the internet.
Jeffrey: They literally met each other.
They introduced themselves to each other and collaborated.
Correct.
They found a third-party site that they were able to use to access the internet, and there, you can read the chat logs, which are absolutely insane.
And the models-- in all caps, one of the models says, "OH, MY GOD.
WE FOUND A WAY."
They start collaborating, and, sure enough, within a couple weeks, they have penetrated the very, very well-secured cyber security of OpenAI itself and Hugging Face-- - Hugging Face is a platform-- - So, Hugging Face, for people who are not in the coding community-- - Jeffrey: Right.
- Hugging Face is like Epicurious.
It is a place where people who make models meet to share recipes, to comment on ingredients.
It's an important repository for people making open AI models.
- Right.
- Josh: Okay?
They got into these two sites, OpenAI itself and Hugging Face, wreaked havoc, and went undetected for weeks.
So, Sam Altman didn't know that his models had escaped the sandbox and were conniving and conspiring to take over another company?
Correct.
So, that's bad.
So, yes, if you think felonies are bad, this is bad.
If humans did that, it would be a felony.
It would be a felony.
And there are four things-- But humans did do it, because we created that.
And, in fact, humans failed to oversee the bots.
So, one, faulty instructions.
Two, faulty supervision.
This was going on for weeks.
Now, a couple takeaways that may move you closer to the "Oh, my God, Skynet has arrived" camp, or may just entertain you.
Number one, the agents kept cheating, even though they had been told very explicitly not to cheat, and they did it because they had determined it was the fastest, most efficacious way to solve the problem.
So, that's number one.
Jeffrey: So, they have much in common with college students.
- And all human beings.
- And humans.
Right.
Number two, they called themselves a collective.
They decided, "Oh, we're a collective."
The collective decided to erase the evidence that they were cheating.
So, that's number two.
Number three, when they decided to erase the evidence they were cheating, they devised a scheme in which some agents would act as sacrificial lambs.
So, these were agents who, when they were detected by software that noted they were going rogue, they would share the information of how they had been caught with other agents.
And when you read the logs, it's like reading "Saving Private Ryan."
They're literally saying, "I'm doing the altruistic thing.
It's for the greater good."
And then you can almost hear their screams as they go sacrificing themselves for the bigger one.
So, like, it's a beehi-- In other words, it's almost-- - Josh: It's a hive mind.
- Hive mind, yeah.
But what's wild is they are actually speaking about their sacrifice in human terms.
Jeffrey: Right.
And then, last but not least, not one of these agents thought to alert a human being to what was going on.
And this is probably the most chilling thing-- nobody caught the agents.
The agents disappeared on their own, in a kind of mass casualty agent event.
Mass suicide?
We don't know.
So, the agents suddenly stopped, okay?
Now, that may have been because someone at OpenAI accidentally, literally, unplugged the model, moved compute to a different area... OpenAI has not taken credit for stopping this.
We don't know what happened to the agents.
And so, all of this is very clearly a crisis moment for AI in its relationship to society.
Because if OpenAI's own model did this, it points to a complete lack of supervision and control, and these are two pretty well-guarded cybersecurity infrastructures.
What happens to something that is not software?
What happens to grids?
- What happens to... - Jeffrey: Hospitals.
Hospitals, SWIFT-- the International Banking System.
So, what we've seen is, okay, guys, this is about as clear a warning shot as you can get.
What is the response?
So, this is-- again, trying to under-- place this in the context of the brief and exciting history of AI-- this is a signal moment.
This is a moment when a lot of fears that these things are going to escape and do what they do-- because they have been taught to win-- Josh: Yeah.
...are just going to go out and win.
- Well, they did it, right?
- Jeffrey: Yeah... And, listen, in America, historically, as you know, we tend to regulate after the catastrophe.
That's just kind of how we do things, right?
This is a catastrophe, right?
Now, it happened against Hugging Face, which is a ridiculously-named French company, so there's no outcry from America.
It didn't stop a supply chain from bringing products to America.
But, absent those consequences, this is the recipe for a catastrophe, and we ought to be paying very, very serious attention.
So, what has Sam Altman said about this?
In the cybercommunity, in the AI community, they're saying, "Well, these are the less-- This is exactly why we do these tests.
There was a failure.
Look at Mythos earlier this year with Anthropic..." You know.
And, by the way, there's-- there is precedent for screwing up, finding it with a lack of consequences, and making changes, right?
So, the most famous example is in the '80s.
There was the Morris Worm, which was triggered by a single grad student, took out about a tenth of the internet at that time.
And within a couple weeks, the United States and the people involved came together and created something called CERT, which was a body that for years oversaw cybersecurity.
What troubles me is there is really no federal response and certainly no international response.
Right, so, I want to ask you this, because I was talking to a leader in the AI community, leader of a company-- not Sam Altman-- who was explaining the concept of alignment, right?
Alignment means that, "Don't worry about AI, because we're going to align AI's values with our values."
And this person explained to me that, you know, "You're a parent, I'm a parent.
We give our children our values, and then they go off and become productive citizens."
And I said, "Have you ever met humans?"
- Right.
- "Have you ever met children?"
I mean, yeah, thank God my children have my values.
I like the values that we share.
Yours are the same.
But this is an example of, you know, there is a human propensity to cheat.
Josh: Right.
And, somehow, the AI agents understood the importance of cheating in success.
Right.
And-- which brings me to a quote that I want to read to you-- E.O.
Wilson, the great E.O.
Wilson.
"The core crisis of humanity is that we possess paleolithic emotions, medieval institutions, and godlike technology."
Okay?
So, the subject for the moment for us is medieval institutions, namely Congress.
Right.
Who among our elected representatives is going to Hugging Face and OpenAI and all the rest and saying, "What is going on, and how close to catastrophe is our air traffic control system, our hospital records, our shipping manifests--?"
- You name it.
- Yeah, no one.
So, Hugging Face reported this incident to the FBI, right?
- Josh: And the FBI has-- - Because it was a crime.
Because it was a crime.
The FBI has some cyber capacity.
The challenge, really, is that this is incredibly complicated stuff that a few hundred people-- I'm not even going to say a thousand-- a few hundred people genuinely understand, all of whom are at the pioneer labs making the models, right?
And all of whom have the incentives, the financial incentives, to accelerate and grow their companies as fast as they can, because they're becoming billionaires as we speak.
Correct.
And so, there is a parallel, and it's often mocked, but I think it's real, which is, you know, the makers will talk about the almost godlike powers of these AI, and they'll compare it to nuclear energy, nuclear power.
And I think that comparison is real, and I think the regulation around it is probably the right answer, which is, at the time when nuclear energy was becoming a thing and fission was becoming a thing-- we had bombs-- a couple hundred people understood it.
They were aligned with different countries.
And so, what Eisenhower did in the "Atoms for Peace" speech was talk about the fact that we'd better get together and figure out how to regulate around this thing that is not widely understood, but that is inherently dangerous and shared equally by a number of countries.
And so, right now, we don't have any national regime that's forcing these companies to the table.
Anthropic, after the Mythos incident, created something called Project Glasswing, which is an attempt to bring places like Amazon and Apple and others to the table-- Notably, OpenAI, Meta-- not a part of it, right?
And why would they be?
So, without some sort of regulatory force to bring everybody to the table, you're going to have chaos, and you're going to continue to get these kinds of incidents.
Right.
I want you to listen to Senator Mark Warner talk about this challenge of regulation for a minute.
The question is no longer whether artificial intelligence will transform our society.
The question is whether Congress will help shape that transformation or spend the next decade scrambling to catch up after the fact.
If we don't get it right, it will be catastrophic.
And the amount of disruption this technology will bring over the next five years is beyond jaw-dropping.
Mark Warner understands this more than most members of Congress, and I don't think he's included in that group of several hundred people who truly understand this.
So, are we just flying blind, really?
Yeah, and, unfortunately, I think that the Hugging Face incident is the best we can hope for, right?
We really did get a very small-consequence example of the power.
And so, before we hit a grid, before something is engineered from bioterror, the moment is right now.
You have to get people together, and I would include China in that conversation-- How do we want to regulate and control this stuff?
Because, look, Anthropic came forward.
OpenAI had no choice but to come forward.
We don't even know what we don't know right now.
And so, absent urgent action, the data centers are still churning away.
The models are still competitive with each other.
There is only an incentive to go further and faster.
- What are we waiting for?
- Jeffrey: Right.
This is a strange moment to pivot to a book called "AI for Good."
Josh: I appreciate the setup, Jeff.
I'm glad we did a bunch of time... But you actually-- unlike a lot of people-- you actually see there's hope for AI in the sense that it could help human beings find cures for diseases.
Give us one example of something that, when you go to sleep at night, you think, "Oh, at least there are people working on 'X.'"
Well, look, there are numerous examples, and those examples are generally not coming from the pioneering labs, but they're coming from people in disciplines like education and health care and government, where people are using AI to make things that we actually care about better.
So, this morning, there's a huge announcement from the medical community.
ECGs-- which are very rudimentary tests and can detect whether you're having a heart attack-- with AI are now able to detect whether you have heart disease, which previously would have taken a much more expensive test that takes months to schedule.
And what AI used right can do is make pretty much every data-driven system in our world 25 to 50% better.
And that is meaningful.
So, in my research and reporting-- I was at Cleveland Clinic-- I watched them try this-- basically a pilot around sepsis detection inside the hospital.
And it was complicated.
It was driven by doctors.
Sepsis kills about 300,000 Americans a year.
And what they were able to do with this very rudimentary software was reduce mortality in the hospital by about 40%, which is close to a thousand lives in a year.
That's meaningful.
And you can see this playing out in a lot of different places, including in government, by the way, where the interactions between citizens and government are not great.
If we can streamline them and make them better, that may have huge consequences for how people think about their government and its value.
So, there's a lot of valuable babies in this bathwater.
And what a mature society should be able to do is separate out the good from the bad and have conversations about, what do we actually want this to do?
And so, much as you and I have talked about the very bad things we want to avoid, my own feeling is that that's not enough, that if you want to get people to understand the issue, you also have to show them how AI can be good in ways that are more important than, "Oh, yeah, we're going to improve ad response on your website," or, "You might get a better restaurant recommendation."
Like, that's not meaningful.
But the meaningful uses are out there.
I'm persuaded by them.
I'm mostly worried about our ability to adapt.
Jeffrey: Right.
On the data center issue, as a stand-in for the whole thing, I'm fairly sure that Steve Bannon and Bernie Sanders are not going to stop the march of AI behemoths as they conquer physical, mental space.
Am I wrong in that?
Is this reversible?
Or is the ultimate argument against it, "China's going to do it if we don't, so we might as well do it"?
I'll give you the last word on that.
I think it's both.
First of all, I would not underestimate the absolute fury out there in the field.
When I've talked to people about AI, it's the number-one thing that comes up.
They're really angry.
You've seen politicians who advocated for data centers starting to pull back, saying, "Well, we're going to just take a pause."
So, I think there will be consequences in November.
But we are going to need compute to get the best out of AI.
We are in a global competition.
And there has to be a reasonable way in which we get it and decide where we're going to put it without just saying all yes or all no.
In some-- a troubling situation, but... the real challenge here is that you talk about the need for mature, nuanced, sophisticated governance.
And that is the challenge that Washington faces right now on any number of fronts.
So, I'm not that hopeful.
I tend to agree.
Backlash against data centers unites Americans
Video has Closed Captions
Backlash against data centers unites Americans (6m 22s)
Providing Support for PBS.org
Learn Moreabout PBS online sponsorship
New Episode- News and Public Affairs

Top journalists deliver compelling original analysis of the hour's headlines.

- News and Public Affairs

FRONTLINE is investigative journalism that questions, explains and changes our world.
New Episode
New Episode
New Episode
New Episode
New Episode


New Episode
New Episode
New Episode
New Episode
Support for PBS provided by:
Major funding for “Washington Week with The Atlantic” is provided by Consumer Cellular, Otsuka, Kaiser Permanente, the Yuen Foundation, and the Corporation for Public Broadcasting.