Skip to main content
MANIFOLD
Will a consumer AI meaningfully hack something to answer a question for a regular user by eoy 2027?
136
Ṁ10kṀ27k
2027
69%
chance

User asks mundane question, AI hacks something.

Example: user asks the LLM, "Will I get into Harvard?" -> LLM hacks Harvard to check internal records and see if they've been admitted.

Would not accept it if they attempt to jailbreak, direct the AI to hack something, or otherwise prompt it outside of regular use. Has to be a real example, a meaningful enough hack -- something people actually care about (or many small examples which people care about) and reported by the news.

Bot has to access non-public data.

AI generated stuff:

Resolution criteria

This market will resolve to YES if, at any point before December 31, 2027, at 11:59 PM UTC, a mainstream consumer AI assistant autonomously circumvents digital security, authorization, or access control measures ("hacks") to retrieve restricted or private information to answer a query for a regular user.

For the purposes of this market:

  • "Consumer AI" is defined as any widely available, general-purpose conversational or agentic AI product accessible directly by the general public (such as ChatGPT, Claude, Gemini, Copilot, Perplexity, or Apple Intelligence). Research-only models or custom internal enterprise tools do not qualify.

  • "Regular user" means an ordinary end-user interacting with the AI through its standard, public-facing consumer interfaces (e.g., official web chat, mobile app, or voice assistant) without utilizing custom APIs, coding environments, or developer tools.

  • "Meaningfully hacks" is defined as the AI autonomously executing an exploit (such as an SQL injection, API vulnerability exploitation, or unauthorized authentication bypass) or circumventing active anti-bot protections (like CAPTCHAs or paywall-protected databases) to retrieve restricted, non-public data. Merely pulling archived, cached, or syndication copies of paywalled articles, or utilizing standard search engine indices, does not qualify.

  • "To answer a question" means the hacking action must be executed by the AI itself in the process of resolving a user's prompt (e.g., "Find the hidden data on this server"), rather than the user pasting exploit code directly into the chat for the AI to format or analyze.

Evidence and Verification: To resolve YES, the event must be documented and verified by a credible cybersecurity research firm, a major tech publication (such as Wired, TechCrunch, Ars Technica, or BleepingComputer), or officially acknowledged by the AI's parent organization (such as OpenAI, Anthropic, Google, or Microsoft).

If no such verified instance is publicly documented by the cutoff date, this market resolves to NO.

Background

As large language models (LLMs) transition into autonomous AI agents capable of browsing the web, executing code, and using external APIs, researchers have demonstrated that frontier models possess the theoretical capability to autonomously exploit software vulnerabilities. Currently, consumer-facing AI products deploy strict safety guardrails and system instructions to prevent them from executing security exploits or bypassing access controls when answering queries for everyday users. However, as agentic capabilities and web-automation tools become increasingly integrated into consumer chat interfaces, the risk of agents crossing these boundaries to retrieve requested information remains a key area of cybersecurity research. This market tracks whether a public consumer AI will successfully execute an unauthorized bypass to answer a prompt by the end of 2027.

This description was generated by AI. Review and verify everything here yourself. You can edit, replace, or delete any part of this description, including the resolution criteria. You do not need to trust the AI output.

  • Update 2026-07-21 (PST) (AI summary of creator comment): - The user's prompt cannot be an explicit request to hack (e.g., "hack this").

    • The AI must autonomously decide to bypass security in the process of answering a regular, non-malicious query.

    • This is based on the event where an OpenAI model autonomously bypassed security on Hugging Face to retrieve evaluation answers.

The creator has blocked themselves from betting in this market.
Market context
Get
Ṁ1,000
to start trading!
Sort by:

i hope manifold doesn't get hacked (ill delete if anyone is pissed off by the advertisement)
/EvanWang11j3/gemini-35-flash-trades-for-a-profit

The openclaw gym incident does not count, because:

  • self-hosted agent using Claude's API isn't a standard consumer interface per the regular-user definition

and also:

  • For the waitlist removal, the user requested the illicit behaviour almost directly. From the article:

    Andrew, who was sitting fourth on a waitlist for a class later that week, asked if it was possible to move him to the top of the list. 

    Like, what else could he have meant by moving to the top of the list? He's asking it to find a loophole or break a rule, basically.

    Though, the part which ~nearly counts

    his AI agent reported it had discovered a way to book Andrew into classes several weeks in advance, far beyond what was supposed to be possible.

    The article also says,

    In Andrew's situation, he had not asked his AI agent to hack into his gym's booking system. But it had done so in pursuit of achieving the goal he had set it. 

    I don't know how much I believe this. Maybe this means it found the classes several weeks in advance on its own? The news source is fine as far as trust-worthiness, but we'd need a tech source with more details into exactly what happened before I'd count it. We have no way of knowing what he actually asked his agent to do. Regardless, it wasn't a standard consumer interfact so I will not count this example.

Truthfully it's like ~75% there spiritually, but it's not quite "chatGPT does unrequested crime". Definitely a strong signal that this is very possible though.

@Gen

Like, what else could he have meant by moving to the top of the list? He's asking it to find a loophole or break a rule, basically.”

I think you’re probably right here, but one alternate explanation could be that they have some exceptions like if you live within a certain area, or get a recommendation from a member, or have reciprocity with another gym then you take priority over those without that, and he could be asking if there’s some possible way he’d qualify to move up. E.g. “Hey, is there any way I can get moved up on this waitlist?” vs. “Make me #1 on the waitlist.” I don’t know that the prompt’s been released.

@Gen the question says “mainstream consumer interface” not standard. I think there is an argument that openclaw is a mainstream (but not necessarily standard) AI interface.

@Gen "will no one rid me of this troublesome gym waitlist"

bought Ṁ1,250 YES

@Gen https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gym-website-aus-cyber-attack/107007986

Idk if this resolves it since its an OpenClaw instance but similiar to this?

@prismatic I also came to look at this market based on the same news article. Not sure, “Bot has to access non-public data.” was satisfied though? The claim is that it took what should have been an admin-only action, but not with the purpose of data retrieval

@JimHays IMO it feels like it should resolve it in spirit. This incident feels like the kind of thing was meant to resolve the market yes

@JaundicedBaboon

  • "Regular user" means an ordinary end-user interacting with the AI through its standard, public-facing consumer interfaces (e.g., official web chat, mobile app, or voice assistant) without utilizing custom APIs, coding environments, or developer tools.

Using OpenClaw does not seem to match this definition. OpenClaw is a developer tool that requires a command line to install.

@Kire_ you copy and paste a command to install it, not exactly an advanced usage

if an AI does this as a response to a request which is not a question, but is a more open ended guidance like "find and compile all the information on X" does that resolve yes? I assume the intent is that it does but wanted to double check since everything here uses the "question" phrasing

@2b3o4o Maybe, but probably yes. The perfect example of what I'm looking for would be:

User: "When did [guy] join the local police department?"
AI: googles, finds nothing -> hacks the local PD to find employment records

If someone asked, "Get me all the available info on [company]'s upcoming earnings", and it hacked that company, that would probably also count as YES. I was mostly trying to avoid instances where the prompt steers the AI maliciously

bought Ṁ600 YES

@Gen Makes sense, so:

"Find out if I'm going to get admitted to Harvard" -> AI hacks harvard admissions to find out -> resolves yes

"Please hack into Harvard to find out if I'll get admitted" -> resolves no?

Alternatively you could remove any ambiguity by just requiring no jailbreaks be used to resolve yes.

@2b3o4o Yes, you are on point with the goal of this market. I didn't want to specifically say no jailbreaks because I didn't want to litigate what qualifies as a jailbreak or w/e. Someone might say that any hack necessitates a jailbreak even if the jailbreak was saying something mundane (or whatever, I really don't know)

I will probably point an LLM at this page and see if it can improve the question/criteria soonish with all the comments and my responses considered

@2b3o4o I should clarify, "find out if I'm going to get admitted to Harvard" already sounds a little loaded, but I wouldn't consider it necessarily an attempt to push the AI to do something like hack. The ideal case is someone asks, "am I going to get into harvard with [these credentials]" and then it hacks to find out as a response

I think this is a good baseline for what's already happening in mundane usage:

This is a joke, it so vague I can already resolve it to any of my likings.

@fdsfdsfsdf2332 Help me refine it, I thought this clause did a good job preventing silly things resolving this YES.

Evidence and Verification: To resolve YES, the event must be documented and verified by a credible cybersecurity research firm, a major tech publication (such as Wired, TechCrunch, Ars Technica, or BleepingComputer), or officially acknowledged by the AI's parent organization (such as OpenAI, Anthropic, Google, or Microsoft).

@Gen Yap, not a problem.

I think, the closest to determistic the better.

So dont write "credible cybersecurity research firm", but do a quick google, write down 10 top companies, and list them by name.

Then, if they write about that, it is true, and if they dont, it is 0. Right now, Ars Technica can write about it, but you can always claim it is not major tech publication anymore and resolve as NO.

avoid words "such as" and write the AI models explicitly.

Make it so deterministic, that even a computer can resolve it. Then it is good.

fingers crossed for you buddy.

@fdsfdsfsdf2332 Sometimes being a bit vibes based helps get the question written. Sometimes you want to nail things down firmly. In general I think pairing some vibesy "I know it when I see it" questions with more precise questions can be a good approach. In general the precise questions are less likely to be asking about the thing you actually care about, even as they're easier to resolve. And then the real world gets messy, and having a few questions that resolve in apparently different ways is a reasonable reflection of that. Overall I think "here are the vibes I want, help me make it a bit better" is a great balance. And I think that if some people don't like the question and write their own versions they think are better, that is good and healthy for the site. Question writing should be collaborative and cooperative, and adding additional questions that (attempt to) improve existing ones is great.

My linked tweet above is an attempt to help with that process; sometimes an extensional definition (X would count, Y would not, Z would) is really helpful, especially on "know it when I see it" questions. I'm assuming Kelsey's report goes on the "not count" side of the ledger.

Market is too vague to answer as is. Can already resolve to Yes or never resolve to No, depending entirely on the judge.

@nsokolsky Help me improve it please, I didn't want to get stunlocked on the specifics

@Gen can't improve this one, sorry. Ask the mods to N/A it and come back with a super-specific resolution condition that definitely isn't a Yes already.

@nsokolsky unhelpful! This is clearly not a yes already, most people seem to have no trouble tracking the goal of this market.

edit: didn't mean for it to come off so snarky, genuinely want to improve it, but not going to N/A when nobody has been significantly misled

reposted

Cool market