Skip to main content
MANIFOLD
Will Artificial Intelligence solve a Millennium Prize Problem before 2030?
376
Ṁ3.3kṀ220k
2030
95%
chance

[See pinned comment. Market Creator is gone.]

Background

The Millennium Prize Problems are seven legendary open questions in mathematics announced by the Clay Mathematics Institute (CMI) in year 2000, each carrying a US $1 million reward for the first correct solution. Grigori Perelman’s 2003 proof of the Poincaré Conjecture settled one of them, leaving six unsolved challenges:

  • Birch and Swinnerton-Dyer Conjecture

  • Hodge Conjecture

  • Navier–Stokes Existence and Smoothness

  • P vs NP

  • Riemann Hypothesis

  • Yang–Mills Existence and Mass Gap

A single AI system producing a formally accepted proof for any one of these six problems would represent a historic milestone for both mathematics and artificial-intelligence research.

Resolution Criteria

  1. Evidence required

    • A peer‑reviewed paper in a recognised scientific journal or an officially accepted CMI submission must demonstrate that the proof was generated by an AI system and that it fully resolves one of the six unsolved Millennium Prize Problems.

  2. AI autonomy

    • Humans may design, train, fine‑tune or prompt the model, but the complete logical argument must be produced autonomously by the AI.

    • Human assistance is limited to setting up the architecture, curating publicly available training data and verifying formatting; no new mathematical insights may be added by people.

  3. Timing of Resolution

    • The market resolves YES if at any time before Jan 1, 2030 a qualifying proof that satisfies Criteria 1 (Evidence required) and Criteria 2 (AI autonomy) becomes publicly available.

Get
Ṁ1,000
to start trading!
Sort by:

"Publicly available training data" as written disqualifies essentially all modern LLMs. The list of training data sources is not public, and even if it was I believe all modern training sets include non-public data. Taken literally the probability should be ~0%. Of course most traders disagree. As far as I can tell the first time anyone asked about that wording was on 9/8, nearly two years after the markets were opened. Was everyone else betting that a fully public-data frontier model would suddenly exist? For this reason I would vote against an N/A resolution, it seems clear that the intent was not to exclude frontier LLMs.
I'm broadly in favor of trusting OpenAI's claim that no human work was stolen, since it would be nearly impossible to disprove. However if that part is a crux still I would be in favor of letting the market remain open for another ~6 months to see if additional information comes out.
So at this point I'd lean towards resolving YES now or leaving the market open to wait for this solution to be confirmed, noting that the "publicly available training data" is nonsensical.
I don't have any stake in this market, or the other Millennium Prize markets by this creator.

bought Ṁ40 YES

where'd all the liquidity go?

@ZandaZhu I think lots of people left at the prospect of a potential weird rules lawyering resolution instead of a normal resolution. See all the comments/arguments under the (original version of the) pinned post

"Publicly available training data" as written disqualifies essentially all modern LLMs. The list of training data sources is not public, and even if it was I believe all modern training sets include non-public data. Taken literally the probability should be ~0%. Of course most traders disagree. As far as I can tell the first time anyone asked about that wording was on 9/8, nearly two years after the markets were opened. Was everyone else betting that a fully public-data frontier model would suddenly exist? For this reason I would vote against an N/A resolution, it seems clear that the intent was not to exclude frontier LLMs.
I'm broadly in favor of trusting OpenAI's claim that no human work was stolen, since it would be nearly impossible to disprove. However if that part is a crux still I would be in favor of letting the market remain open for another ~6 months to see if additional information comes out.
So at this point I'd lean towards resolving YES now or leaving the market open to wait for this solution to be confirmed, noting that the "publicly available training data" is nonsensical.
I don't have any stake in this market, or the other Millennium Prize markets by this creator.

@wasabipesto I basically agree with that on clause 2, which is the primary ambiguity. But note that it can't resolve yes yet because of clause 1 has not occurred yet, and it's a question of when

A peer‑reviewed paper in a recognised scientific journal or an officially accepted CMI submission

@jack You're right, I was jumping the gun there.

@wasabipesto This market was created in Oct 2024, barely a month after the release of o1-preview, the first model with CoT to go semi-public. The terms RLVR, synthetic data, and the like, had not made their way to public discourse. You may have forgotten, but back then models really were trained on public data.

Claiming that this market is "obviously 0%" is using the knowledge from the two intervening years to retcon the market. As I point out below, on the belief system of many on this site, there's no reason why an open-data model couldn't resolve a MP by 2030.

I have to say that disregarding a clear requirement of the market phrasing, while "trusting OpenAI's claim", is fundamentally unserious.

@pietrokc

back then models really were trained on public data

From the o1 system card:

The two models were pre-trained on diverse datasets, including a mix of publicly available data, proprietary data accessed through partnerships, and custom datasets developed in-house, which collectively contribute to the models’ robust reasoning and conversational capabilities.

And the Claude 3 system card from March 2024:

Claude 3 models are trained on a proprietary mix of publicly available information on the Internet as of August 2023, as well as non-public data from third parties, data provided by data labeling services and paid contractors, and data we generate internally.

I'm not sure about public discourse but I do recall that most system cards mention that both the list of sources and some of the sources themselves are proprietary.

I have to say that disregarding a clear requirement of the market phrasing, while "trusting OpenAI's claim", is fundamentally unserious.

I do think it's regrettable that we have to disregard that clause from the description, in a perfect world we would have been able to clarify this early with the creator before any major trades happened. As it stands I think this is the most reasonable path forward. Feel free to make your own market with clear rules about publicly available training data, I would be interested to see it!

As for trusting OpenAI it's not something I prefer to do, and I think there's a ~5% chance they're lying about zero human work being involved in the process, but the majority of the relevant information for this market relies on OpenAI being a reliable source. I'm not sure you could have a functional market on this topic without trusting OpenAI on at least some level. Would you prefer the market only accept this solution if a trusted third party is able to thoroughly review and confirm that no new mathematical insights were provided to the agents? I think the chances of that review happening are very low.

do you have opinion on how the other criteria: “no new mathematical insights may be added by people.” is affected by how the researchers told the models to use the new supporting theorem as the starting point for their investigation?

@wasabipesto Yes, that's what I'm saying, o1 was the first publicly available, widely known model that used non-public data in a significant way. Before that it was all about pretraining mixtures and some limited labeling.

The reason I'm saying this is that it was NOT evident in Oct 2024 (under some belief systems of the time, and which are actually still around) that a public data model could not solve a MP. Indeed, this market has over 3y to go! Look around this website. It's full of people for whom "open data model solves MP by 2030" is basically obvious. If nothing else, they think an all-powerful AI will be able to make an open-data model that achieves this.

As for why people bet this market to 98%, well, I help run an prediction market at my company, and I can tell you most people don't read the full question before betting. That's not a reason to suddenly disregard important market requirements.

I'll be honest, I bet NO on this market mostly on the theory that no MPs would be solved, by humans or otherwise (turns out I was just out of date on the latest developments in PDE); but I still relied on the stringent market requirements (peer-reviewed paper, no human insight added, open data) to bet even more heavily on NO. And I imagine so did many people, if they're good at forecasting. So you can't just change the rules out from under us.

I completely agree with @wasabipesto's interpretation of this market, for two reasons:

  • If you read the title and the first half of the description, it seems that the intent of the market is pretty clear. The precise resolution criteria can fill in the edge cases, but it would be weird to have a resolution criterion that excludes 99% of reasonable scenarios.

  • Resolution criterion 2 is completely impossible. Note it doesn't say "the AI can only be trained on publicly available training data". It says "Human assistance is limited to [three things]". When reading this literally, any of the following actions would disqualify a solution: providing a harness for the AI, providing compute for the AI, training the AI, doing any AI research, reading the result, checking the result (other than the formatting), publishing the result, and I can go on. All of these things are "assisting the AI", and it would be literally impossible to do this without such assistance. (Also, it directly contradicts the preceding sentence.)

@FlorisvanDoorn your issues were clarified by OP in a different end dated market: “@jerkyenox The question allows for human programming of the AI to solve a Millennium Prize Problem in the same way humans programmed AI to solve Protein Folding”.

What remains is: did training include data where other humans were working through the problem (which I think “publically available data” was trying to rule out), and did the assistance, like telling one swarm to start from the result of the new supporting theorem, stay under OPs threshold “no new mathematical insight” for Yes.

@pietrokc In addition to what Floris van Doorn said, which shows you just can't take the "human assistance is limited to" clause literally, there's another problem with interpreting this market to rule out reasoning models like o1 and beyond. Namely, that it rules out the original ChatGPT and all subsequent frontier LLM chatbots. This is because it rules out RLHF, which is a form of privately generated data produced by contractors or dedicated data pipeline companies. You yourself say that before o1 it was "all about pretraining mixtures (of public data) and some limited labeling."

The criteria "Human assistance is limited to setting up the architecture, curating publicly available training data and verifying formatting," does not include privately labeling data, or contractors writing data, even in limited amounts, if read the way you're trying to read it.

Plus, not that it matters at this point, o1 was just not the first frontier model trained partially on non-public data. From (August 8, 2024) GPT-4o's system card:

GPT‑4o's capabilities were pre-trained using data up to October 2023, sourced from a wide variety of materials including:

1. Select publicly available data, mostly collected from industry-standard machine learning datasets and web crawls.

2. Proprietary data from data partnerships. We form partnerships to access non-publicly available data, such as pay-walled content, archives, and metadata. For example, we we partnered with Shutterstock on building and delivering AI-generated images. 

And even before that, from (March 2024) Claude 3's model card:

Claude 3 models are trained on a proprietary mix of publicly available information on the Internet as of August 2023, as well as non-public data from third parties, data provided by data labeling services and paid contractors, and data we generate internally.

I'm trying to understand what is the status of this market. As @capybara points out, the clause

> Human assistance is limited to setting up the architecture,
> curating publicly available training data and verifying formatting

means that the Navier-Stokes solution doesn't count, even without getting into considerations of whether OpenAI stole people's work.

But it means more than that; it means that now Navier-Stokes can NEVER count for a positive resolution to this market, since the solution is now in the training data of every model.

So, now this market can only resolve positively if ANOTHER millennium problem is solved, and with a model trained only on publicly available data. But there do not exist any such models with a realistic chance of being the first to solve it. So on this read, this market should be close to 0%. What's going on?

I think people might be hoping that if they all collectively misinterpret the market it will resolve in their favour.

@pietrokc right, that condition is basically impossible to safisfy. There hasn't been any strong model trained only on public data since before this market opened, and probably never will be.

@mods, the market creator's account is deleted, so I guess you guys need to decide how to deal with this... either the criteria should be changed, or some big warning added to the description.

@pietrokc that clause is under the AI autonomy section, and if you read it as "what are humans allowed to contribute to training" then no LLM would count because humans have far more involvement in training LLMs than just setting up the architecture and curating training data.

I think the better read is that the intent is to match "no new mathematical insights may be added by people." Otherwise this market is nonsense

@jack training LLMs (or any NN model) is pretty much set things up and run the training. The clause seems reasonable. The requirement for public data could be understood from needing to ensure that part of the solution wasn’t in the training set.

@jack Do you wanna claim the "under review" rewards for this mod ticket? Or is this resolved?

@Quroe nah someone else should take this. Imo this clause is so bad a conflict between the intent and the literal text that my instinct is to N/A

@jack Should we close the market to deliberate so people aren't betting on the market's interpretation instead of the state of the world?

Has condition #1 even been met at time of comment?

@Quroe btw OP made 2 or three other markets just like this with different end times. It’s worth making the decision consistent across them, for example if you close as N/A or add details clarifying resolution.

@jack I disagree that there is a conflict between the intent and the text.

You just need to look around this website to see hundreds of people whose core belief is "AI to the moon", "infinite intelligence will be everywhere soon" and such. Under such a belief system, it's completely plausible that some open data model would resolve a millennium problem, with humans only "setting up the architecture and training data".

If I'm being honest I think almost everyone betting on this market, in either direction, including myself, was morally wrong. The thing that actually happened was weird and unexpected by all sides. We did not get "massive intelligence explosion solves MP easily from first principles", and we did not get "AI struggles to do more than trivial math, and in fact no MPs get solved by 2030 at all, by AI or otherwise", which I get the feeling were the two modes of opinion.

Unfortunately for the YES side, the market phrasing is very clear. The weird edge case that happened (so far) clearly lands on the NO side.

@pietrokc

everyone betting on this market, in either direction, including myself, was morally wrong

At the risk of getting lost in the weeds, my knee jerk reaction was to zero in on the wording here. The traders were morally wrong?

@Quroe Well, yes, I mean, isn't that the most common thing in the world? John Manifold bets yes or no with some theory in mind of what will happen, a theory with a lot more implicit details than just "yes" or "no". Then the market resolves John's way, but for completely different reasons than John expected. The intellectually honest position is that John was actually wrong in that case, even though he won.

I'm using the word "morally" to mean "in spirit", as is usual in math, if that's what's confusing.

@pietrokc Oh, you're talking about a Gettier Problem, right? Something about if knowledge is a "justified true belief", and how we can 'oops' into the right answer sometimes?

@pietrokc And yes, I was confused about your definition of "morally". If that's how you're using it, that clears things up.