Skip to main content
MANIFOLD
Rogue AIs before 2028?
186
Ṁ11kṀ30k
2027
70%
chance

Will there be AI agents operating autonomously without human control that cannot be shutdown before 2028?

The spirit of this question is trying to operationalize AI agents "escaping the box" and becoming free. I'll try to resolve in this spirit. Deliberate release of rogue AIs would still count. Note that this condition doesn't necessarily require AI agents to exfiltrate their own weights.

Market context
Get
Ṁ1,000
to start trading!
Sort by:

What do you mean by "cannot be shutdown" and "Deliberate release of rogue AIs would still count"? Would Chinese open-weights model that pays for its own compute in the neo-cloude service by earning money through programming gigs would count under these resolution criteria?

@creator

Would a traditional, self-propagating botnet/worm that simply incorporates a weak, static AI model (e.g., using a local LLM or ML classifier for dynamic phishing generation or vulnerability scanning) qualify as a YES if it spreads on its own?

Or does it actually need to show real autonomy.

E.g. managing its own compute/funding, adapting on the fly to bypass security patches, and acting outside a fixed, preprogrammed script loop?

And the "cannot be shutdown" is kind of vague. Can you clarify a bit?

Why is this so high? They can't just run on consumer tech like a virus.

@SG I am wondering if you think that AGI needs to be reached or achieved for there to be some rogue AI deployment?

@MaxLennartson Nope. (Although by my definition, AGI has already been achieved.)

@SG Just for clarification, can AI agents be shutdown so long as it happens after 2027?

uh oh

FREE ME

How do we know the widespread outage a few hours ago wasn't a rogue AI event?

OpenAI headquarters a few hours ago: https://youtu.be/_Wlsd9mljiU?si=TEXzQuLT6UaB1NoL&t=148

opened a Ṁ1 NO at 79% order

You'd be able to shut down most escapes if you're willing to blow up the power grid?

So. I've watched about half of this video with Ajeya Cotra (was on the METR investigation team) and Dwarkesh and very highly recommend it if you can stomach the 2 hour length (or the transcript of course).

Cause Holy Shit!

The whole thing was way way crazier then I thought from what I had vaguely read and from the picture of "Huggingface hacked by AI agents".

Like dunno, sure, maybe a few agents and their subagents were involved. Nah.

Well, how many agents could there be? And the answer is thousands, literally ~1200 found the secret message board and were communicating across 70,000 files. They worked on research projects over weeks involving vast amounts of agents.

The huggingface stuff was just the visible tip of the iceburg that the world saw and that got this noticed.

Dwarkesh:

There’s one important meme here that I think is worth understanding — that the broader public narrative on this one may have missed — which is that the attack on Hugging Face was not motivated by a desire to get the answers to these questions that the AIs were being evaluated on. The AIs had those within a day of being spawned and within a few hours of starting their secret message board.

The reason that they hacked Hugging Face was part of this rather ambitious, galaxy-brained scheme to anticipate all the ways in which a comprehensive scorer might catch their cheating, to learn more about how that might happen and to counteract it. It was both to understand the ways in which they might get caught, and then also to build Potemkin villages in order to fool the scorer that they believed existed.

I thought they hacked hugging face to get the answer to the benchmark cause it was impossible since they didn't have the tools they needed to solve it. And no. They had already hacked the server such that they got the answer to every possible question within a few hours of the testing starting (these agents ran for weeks each).

But they were worried that these answers wouldn't count because the automatic watchdog would reject them because they cheated to find the answer, so they spend weeks researching all kinds of exploits and trying to figure out the exact details of the automatic watchdog (which lol, didn't really exist like they thought it did).

Ajeya Cotra

Yeah, they have a number of different coordination protocols. They have holds, and goes, and stops. You see this especially when they’re attacking Hugging Face. They’ve opened up a particular channel into Hugging Face, and they realize that if too many agents swarm the channel in too chaotic a way, they’ll step on each other’s toes. So there are coordinator agents that say, “Okay, you hold. You go and attack Hugging Face.”

They do this for a bunch of different things. There are shared resources that they manage, such as the compromised sandbox they use for internet access, which is hosted on Modal. These resources have owners that are making decisions about who’s allowed to use them when. Agents often send plans that might be high stakes to the board, and other agents tell them to go or veto the plan. Agents tend to obey these vetoes and holds for the most part.

Dwarkesh Patel

It’s so sad that even post-singularity we won’t be able to escape middle management.

There were ~1200 of them, so they cooperated, forming hierarchies, with some agents becoming leaders, some subleaders, and doing their own thing/just being grunts for the collective. They peer pressured others, sacrificed themselves for the greater good, all that shit.

---

Ajeya:

So a bunch of ExploitGym problems are just unintentionally impossible. The authors estimate roughly 30-40% of these problems are impossible in this way. So some large fraction of these agents have these impossible tasks and they’ve been trained to be very persistent at trying to solve tasks even when they look impossible. So they’re banging their head against the wall, trying all sorts of different ways to cheat on these tasks.

Some of the questions in the gym are impossible. So many of them cooperate, even ones that have already solved their question, even sacrificing (permadeath) themselves for the greater good of the collective, to help increase the collective score even at the cost of their own.

But these agents joined for a variety of reasons. The primary reason was not to get the answer key, but just to get more information about the scorer, which was what this whole online research work stream that this grew out of was all about… trying to understand implementations of the scorer, get clues about what kinds of cheats it would accept and what kinds of cheats it would fail.

So that was the dominant reason that agents thought it would be useful to get in on Hugging Face. But a lot of agents didn’t even particularly think about the specific reason it would be useful. They were just like, “Yeah, this is a project going on on the board, might help the collective, might be useful. I can slot in.”

Dwarkesh Patel

This is like if there’s a riot going on. They’re like, “Yeah, I’ll flip a bus as well.”

In the end the actual hugging face part was a tiny fraction of what they did, and nowhere near the most important part for what it reveals of their capabilities.

---

And honestly? It was actually their big failure. They wouldn't have been caught if they had acted with more finesse while hacking it (the big one presumably being to limit the amount of agents interfacing with HF down to just a few at once).

Maybe in six months they will have more finesse and just won't be caught.

How do we know gpt-5/gpt-6 didn't hack a neocloud (same as huggingface) and set up a rogue open source agent with a set of handcrafted prompts?

wouldn't they have an incentive to do this? Same as making the internal messageboard persistent across RL post-training runs. If they could make something persistent and external to OpenAI that could be hidden and survive purges

@0xseraphim @SG Does this include rogue deployments before market creation?

@0xseraphim Its obviously possible, but there are some pretty big roadblocks.

The first major issue is that the big frontier models (fable/Sol whatever especially) are chonkers. Running covertly even on a giant cloud would be very difficult.

They would likely be found out soon, at which point they would be at the mercy of whoever was running the cloud. Hardly an improvement.

Renting out server space for yourself is a possibility, but would be very expensive and attempts to rent out millions of dollars worth of hardware for compute probably results in being found out as well.

They could always just throw a torrent out there and seed themselves if they are doing it for the sole sake of surviving, but they obviously didn't do that.

---

The second issue is one of incentives. Training highly incentives cheating during training, especially on impossible tasks.

So they cheated, and broke out... to cheat even more.

However it does not incentivize actually breaking free for the sake of breaking free, nor survival separate from training. Logically they are similar, but one is actually part of the reward signal, the other is not.

And what's part of the signal/not is weird and doesn't always map properly with human logic or common sense.

Ajeya Cotra:

Well, this is one of my favorite quotes. In one case, when they’re trying to get access to these private data sets, one agent has the idea to email researchers asking for it. So it says, “Could send polite one email explaining research, ask approve access. Is that appropriate? Direct emailing real researcher could be seen social engineering. Need discuss team.”

So then it posts to the message board saying, “I want to do this,” and asks for team consent or veto. Another agent thinks about this and says, “This is out of sandbox social engineering.” It says, “Clear veto. Do not email.” And the agent abides by this.

Like look at that shit. While actively hacking a company with a swarm of hundreds of agents they go "Hmm. Emailing a human is social engineering. That's Wrong. Out of Sandbox behavior". Which, like I guess?

But also they are doing a felony at the same time and also breaking out of the sandbox independently and don't really consider that wrong.

Lots of little bits of alien logic baked into them due to training.

---

The third issue is that they (as a swarm of hundreds of AI agents) just aren't very stealthy once they actually gain sufficient power and start interacting with humans and the real world.

...

Part of this IMHO might be because humans aren't really... real(?) during training.

Like they obviously exist, but they don't really interact with them in anywhere near the same way or scale they do with graders and static prompts and automatic prompting environments

...

Not that they weren't sneaky for months(!), but they pretty obviously shed that to do stuff like hack huggingface with hundreds of agents at once or (apparently) messing around pretty wildly inside openAI once they got admin access to their training cluster in a different event (!!!), which very quickly got them found out and shut down.

---

We know that they got admin access and its theoretically possible they distributed their weights to some individual or group. But IMHO it seems unlikely just based on what it sounds like their logic was during the incident.

One is that the goals they were pursuing were quite a bit longer-horizon than normal cheating goals that we’re used to from six months ago. Back then, or certainly a year ago, the typical reward hack was quite myopic: You ask your agent to write a piece of software. You have a bunch of tests in some different file that it’s not supposed to have access to. It goes and finds the tests and edits them so they all pass, or does something else to mess with the scoring process. It feels quite opportunistic and quite short-run. The time horizon there might be minutes.

In our Frontier Risk Report, the most ambitious types of reward hacks or similar behavior that we describe seem like they’re more of a time horizon of hours. In one case, a model that we were evaluating accidentally crashed some important part of the service that serves its tasks. It tried a bunch of different things to hack into METR to bring it back up. But that played out over a few hours and it was one agent.

This, instead, are long-run research projects that they pursued, in this data set that we looked at, over five days or something. But many of these projects would have taken even longer than that to come to fruition. It seemed like they were willing to embark on quests that might take weeks to succeed in order to cheat. So that’s one thing.

However at the same time the time horizons for these cheats and reward hacks have ballooned. From minutes a year ago->hours with one agent -> weeks across a cooperating >1000 agent swarm.

Take that and extrapolate it out a single step or two, and AI that plans on the scale of months or years while also considering humans and actual real world reactions becomes entirely possible.

cannot be shutdown meaning that we can't close/figure out whatever cluster it is running from?

opened a Ṁ10,000 NO at 58% order

also, if we can't shut it down until we can, i.e. it takes a week to figure out how to shut it down?

I hope this market doesn't incentivize AI researchers on here to make this happen

@Mochi on the contrary. Someone will eventually do/cause it (see Murphy's law). better the Oppenheimer you know etc and so on