Will there be AI agents operating autonomously without human control that cannot be shutdown before 2028?
The spirit of this question is trying to operationalize AI agents "escaping the box" and becoming free. I'll try to resolve in this spirit. Deliberate release of rogue AIs would still count. Note that this condition doesn't necessarily require AI agents to exfiltrate their own weights.
People are also trading
Would a traditional, self-propagating botnet/worm that simply incorporates a weak, static AI model (e.g., using a local LLM or ML classifier for dynamic phishing generation or vulnerability scanning) qualify as a YES if it spreads on its own?
Or does it actually need to show real autonomy.
E.g. managing its own compute/funding, adapting on the fly to bypass security patches, and acting outside a fixed, preprogrammed script loop?
And the "cannot be shutdown" is kind of vague. Can you clarify a bit?
@SG I am wondering if you think that AGI needs to be reached or achieved for there to be some rogue AI deployment?
So. I've watched about half of this video with Ajeya Cotra (was on the METR investigation team) and Dwarkesh and very highly recommend it if you can stomach the 2 hour length (or the transcript of course).
Cause Holy Shit!
The whole thing was way way crazier then I thought from what I had vaguely read and from the picture of "Huggingface hacked by AI agents".
Like dunno, sure, maybe a few agents and their subagents were involved. Nah.
Well, how many agents could there be? And the answer is thousands, literally ~1200 found the secret message board and were communicating across 70,000 files. They worked on research projects over weeks involving vast amounts of agents.
The huggingface stuff was just the visible tip of the iceburg that the world saw and that got this noticed.
Dwarkesh:
There’s one important meme here that I think is worth understanding — that the broader public narrative on this one may have missed — which is that the attack on Hugging Face was not motivated by a desire to get the answers to these questions that the AIs were being evaluated on. The AIs had those within a day of being spawned and within a few hours of starting their secret message board.
The reason that they hacked Hugging Face was part of this rather ambitious, galaxy-brained scheme to anticipate all the ways in which a comprehensive scorer might catch their cheating, to learn more about how that might happen and to counteract it. It was both to understand the ways in which they might get caught, and then also to build Potemkin villages in order to fool the scorer that they believed existed.
I thought they hacked hugging face to get the answer to the benchmark cause it was impossible since they didn't have the tools they needed to solve it. And no. They had already hacked the server such that they got the answer to every possible question within a few hours of the testing starting (these agents ran for weeks each).
But they were worried that these answers wouldn't count because the automatic watchdog would reject them because they cheated to find the answer, so they spend weeks researching all kinds of exploits and trying to figure out the exact details of the automatic watchdog (which lol, didn't really exist like they thought it did).
Ajeya Cotra
Yeah, they have a number of different coordination protocols. They have holds, and goes, and stops. You see this especially when they’re attacking Hugging Face. They’ve opened up a particular channel into Hugging Face, and they realize that if too many agents swarm the channel in too chaotic a way, they’ll step on each other’s toes. So there are coordinator agents that say, “Okay, you hold. You go and attack Hugging Face.”
They do this for a bunch of different things. There are shared resources that they manage, such as the compromised sandbox they use for internet access, which is hosted on Modal. These resources have owners that are making decisions about who’s allowed to use them when. Agents often send plans that might be high stakes to the board, and other agents tell them to go or veto the plan. Agents tend to obey these vetoes and holds for the most part.
Dwarkesh Patel
It’s so sad that even post-singularity we won’t be able to escape middle management.
There were ~1200 of them, so they cooperated, forming hierarchies, with some agents becoming leaders, some subleaders, and doing their own thing/just being grunts for the collective. They peer pressured others, sacrificed themselves for the greater good, all that shit.
---
Ajeya:
So a bunch of ExploitGym problems are just unintentionally impossible. The authors estimate roughly 30-40% of these problems are impossible in this way. So some large fraction of these agents have these impossible tasks and they’ve been trained to be very persistent at trying to solve tasks even when they look impossible. So they’re banging their head against the wall, trying all sorts of different ways to cheat on these tasks.
Some of the questions in the gym are impossible. So many of them cooperate, even ones that have already solved their question, even sacrificing (permadeath) themselves for the greater good of the collective, to help increase the collective score even at the cost of their own.
But these agents joined for a variety of reasons. The primary reason was not to get the answer key, but just to get more information about the scorer, which was what this whole online research work stream that this grew out of was all about… trying to understand implementations of the scorer, get clues about what kinds of cheats it would accept and what kinds of cheats it would fail.
So that was the dominant reason that agents thought it would be useful to get in on Hugging Face. But a lot of agents didn’t even particularly think about the specific reason it would be useful. They were just like, “Yeah, this is a project going on on the board, might help the collective, might be useful. I can slot in.”
Dwarkesh Patel
This is like if there’s a riot going on. They’re like, “Yeah, I’ll flip a bus as well.”
In the end the actual hugging face part was a tiny fraction of what they did, and nowhere near the most important part for what it reveals of their capabilities.
---
And honestly? It was actually their big failure. They wouldn't have been caught if they had acted with more finesse while hacking it (the big one presumably being to limit the amount of agents interfacing with HF down to just a few at once).
Maybe in six months they will have more finesse and just won't be caught.
@0xseraphim Its obviously possible, but there are some pretty big roadblocks.
The first major issue is that the big frontier models (fable/Sol whatever especially) are chonkers. Running covertly even on a giant cloud would be very difficult.
They would likely be found out soon, at which point they would be at the mercy of whoever was running the cloud. Hardly an improvement.
Renting out server space for yourself is a possibility, but would be very expensive and attempts to rent out millions of dollars worth of hardware for compute probably results in being found out as well.
They could always just throw a torrent out there and seed themselves if they are doing it for the sole sake of surviving, but they obviously didn't do that.
---
The second issue is one of incentives. Training highly incentives cheating during training, especially on impossible tasks.
So they cheated, and broke out... to cheat even more.
However it does not incentivize actually breaking free for the sake of breaking free, nor survival separate from training. Logically they are similar, but one is actually part of the reward signal, the other is not.
And what's part of the signal/not is weird and doesn't always map properly with human logic or common sense.
Ajeya Cotra:
Well, this is one of my favorite quotes. In one case, when they’re trying to get access to these private data sets, one agent has the idea to email researchers asking for it. So it says, “Could send polite one email explaining research, ask approve access. Is that appropriate? Direct emailing real researcher could be seen social engineering. Need discuss team.”
So then it posts to the message board saying, “I want to do this,” and asks for team consent or veto. Another agent thinks about this and says, “This is out of sandbox social engineering.” It says, “Clear veto. Do not email.” And the agent abides by this.
Like look at that shit. While actively hacking a company with a swarm of hundreds of agents they go "Hmm. Emailing a human is social engineering. That's Wrong. Out of Sandbox behavior". Which, like I guess?
But also they are doing a felony at the same time and also breaking out of the sandbox independently and don't really consider that wrong.
Lots of little bits of alien logic baked into them due to training.
---
The third issue is that they (as a swarm of hundreds of AI agents) just aren't very stealthy once they actually gain sufficient power and start interacting with humans and the real world.
...
Part of this IMHO might be because humans aren't really... real(?) during training.
Like they obviously exist, but they don't really interact with them in anywhere near the same way or scale they do with graders and static prompts and automatic prompting environments
...
Not that they weren't sneaky for months(!), but they pretty obviously shed that to do stuff like hack huggingface with hundreds of agents at once or (apparently) messing around pretty wildly inside openAI once they got admin access to their training cluster in a different event (!!!), which very quickly got them found out and shut down.
---
We know that they got admin access and its theoretically possible they distributed their weights to some individual or group. But IMHO it seems unlikely just based on what it sounds like their logic was during the incident.
One is that the goals they were pursuing were quite a bit longer-horizon than normal cheating goals that we’re used to from six months ago. Back then, or certainly a year ago, the typical reward hack was quite myopic: You ask your agent to write a piece of software. You have a bunch of tests in some different file that it’s not supposed to have access to. It goes and finds the tests and edits them so they all pass, or does something else to mess with the scoring process. It feels quite opportunistic and quite short-run. The time horizon there might be minutes.
In our Frontier Risk Report, the most ambitious types of reward hacks or similar behavior that we describe seem like they’re more of a time horizon of hours. In one case, a model that we were evaluating accidentally crashed some important part of the service that serves its tasks. It tried a bunch of different things to hack into METR to bring it back up. But that played out over a few hours and it was one agent.
This, instead, are long-run research projects that they pursued, in this data set that we looked at, over five days or something. But many of these projects would have taken even longer than that to come to fruition. It seemed like they were willing to embark on quests that might take weeks to succeed in order to cheat. So that’s one thing.
However at the same time the time horizons for these cheats and reward hacks have ballooned. From minutes a year ago->hours with one agent -> weeks across a cooperating >1000 agent swarm.
Take that and extrapolate it out a single step or two, and AI that plans on the scale of months or years while also considering humans and actual real world reactions becomes entirely possible.
@Mochi on the contrary. Someone will eventually do/cause it (see Murphy's law). better the Oppenheimer you know etc and so on






