Imagine that any math problem you can write down on a piece of paper that a team of Fields medalists can solve, AI can as well. Until recently, I would've predicted that that was an AGI-complete problem. Of course people used to think grandmaster-level chess would require AGI. Until 2022 I was sure that commonsense reasoning and being able to explain jokes would require AGI.
If subhuman general intelligence can be a superhuman mathematical intelligence, that will be another big update for me.
FAQ
1. What if AGI happens first?
This is a conditional prediction market. If AGI, as defined in my other market, happens first, this resolves N/A.
2. Does the AI need to max out the FrontierMath benchmark for this to resolve YES?
Yes, and every math benchmark, plus gold-medal performance on the International Math Olympiad [Update: now achieved, as of summer 2025]. Even acing the Putnam.
3. What if it's essentially true but there are rare exceptions?
The spirit of the question is that we'd only consider an AI failure to be an exception if it failed for a reason other than being insufficiently brilliant at math. Like tricksy wording, or any trick question. The posing of the question has to be non-adversarial.
4. What about a book-length question?
Tentative answer so far: The problem has to be posed on a single human-readable sheet of paper or equivalent. But a question can cite any peer-reviewed math paper as background. (Dumping an impenetrable tome on the arXiv doesn't count.) If you have an example where this feels limiting, let me know. My suspicion is that all interesting math problems can be posed on a single page and in any case it won't harm the spirit of this question to limit ourselves to such.
4. What about research taste?
That's a big part of being a mathematician and isn't required for this market. The AI just has to be superhuman at answering questions, not asking them.
5. What about cost and speed?
The AI has to dominate the best humans on all metrics. We'll find an authoritative source for the market value of mathematicians' time if it comes down to that.
6. What about availability to the public?
Not required. If there's any doubt about the veracity of claims that this has been achieved, we'll discuss and delay resolution as needed.
7. What if the AI is sometimes super- and sometimes sub-human at math?
In some senses that's already the case but there may be ambiguous edge cases. As an extreme example, imagine that the AI is so blatantly superhuman that it cracks a famous open problem, yet it's routinely stumped or wrong on problems amateur human mathematicians can do. For the spirit of the question for this market, we'll try to assess whether we'd consider a human with the AI's math abilities to be the greatest mathematician of all time. (Or the greatest raw math prodigy of all time -- see FAQ 4 on the distinction between problem solving and knowing what questions to ask. The latter is a key part of being a successful mathematician and is explicitly not part of this prediction.)
Related markets
https://manifold.markets/dreev/in-what-year-will-we-have-agi
https://manifold.markets/jack/will-an-ai-outcompete-the-best-huma-cj3ul7a2g2
https://manifold.markets/MatthewBarnett/will-an-ai-achieve-85-performance-o
https://manifold.markets/Manifold/what-will-be-the-best-performance-o-nzPCsqZgPc
[ignore the subhuman clarifications that keep automatically appearing below this line]
Update 2026-09-23 (PST) (AI summary of creator comment): - Clarified the definition of superhuman as either solving everything humans can faster and cheaper, OR solving major open problems without otherwise being too blatantly subhuman at math.
Proposed defining the human/superhuman threshold as what a team of 4 Fields Medalists can solve in 4 days.
People are also trading
What do @traders think of pinning down the human/superhuman threshold as "what a team of 4 Fields Medalists can solve in 4 days"? Up to now I had only referred to what a "team of Fields Medalists" could solve in "a day or a week". Since 4 is the average of 1 day and 7 days, might as well go with 4?
I'm trying to be extra careful and transparent since I have a large position myself. Currently for YES but when this market was new I put thousands of mana on NO.
Mostly my thinking now is that the unit distance problem and the Jacobian conjecture and especially Navier-Stokes already put us far beyond what was required for YES. But as @pietrokc is arguing, we also need clarity on jaggedness. AI is far beyond the threshold on some problems but below it on others. FAQ#7 actually tried to anticipate this, posing what was meant as a wildly hypothetical edge case: AI that's "so blatantly superhuman that it cracks a famous open problem, yet it's routinely stumped or wrong on problems amateur human mathematicians can do." I went on to suggest that we'd resolve such ambiguity by asking if a human with the AI's abilities would be considered the greatest math problem-solving prodigy of all time.
Which is annoyingly fuzzy, now that we're actually in something like that wild hypothetical. But I'm not sure we need to adjudicate the prodigy question because in the hypothetical the ambiguity arose from the combination of cracking a famous open problem and routinely being shown up by amateur mathematicians. Can anyone (@pietrokc?) make the case that it's routine for amateur mathematicians to show up Astra or Fable (let alone the latest internal models)? In the example of the SAIR competition the human winners were professionals.
Further clarification on FAQ items 2 vs 7: FAQ#2 says we need to saturate every benchmark. FAQ#7 says the AI can potentially be routinely outclassed by amateurs and still count as superhuman if it also cracks a major open problem. The seeming contradiction is because most of the FAQ just presumed AI would not solve a major open problem. So it's basically saying "superhuman means solving everything humans can faster and cheaper OR solving major open problems without otherwise being too blatantly subhuman at math".
@dreev I would not consider current AI results to make someone the best maths problem solved of all time as per FAQ7. Although it's getting close and I expect it to be unambiguously the case by the end of the year.
@dreev Well... The market's deadline is "before 2030" which is approximately a lifetime and a half in this cursed singularity. If we already find ourselves debating precise definitions because of capability jaggedness late-2026 (more than three years to go still to smooth out the edges!), the industry will be able to clear any resolution criteria you throw at it within a year at most, and it looks to me like we might even get there by the end of this year. So I wouldn't spend too much time and effort endlessly refining criteria that will fall within the time boundary as an already foregone conclusion; my biggest concern here is actually "assuming no AGI yet"; i.e. we only need to clearly and unambiguously separate between these two major conditions, and the rest will be self-evident enough.
@dreev Some disorganized thoughts.
I think there is a contradiction between "considering AI the greatest raw math prodigy of all time" and the fact that AI can be scaled up indefinitely by just spending more money.
Here is the fact / intuition that I find it helpful to come back to. We have known that a computer could prove every provable theorem / solve every solvable problem since the 1930s. The technical result is "the set of theorems of ZFC is recursively enumerable" if you wish to google or ask models (in case you're not familiar). It means you can set off a computer with very simple initial configuration, and if you give it enough time and memory, it will eventually spit out every provable theorem, while spitting out no unprovable statements.
Of course there is nothing special about computers. A human could carry all this out by hand. Or you could parallelize it across a billion humans to speed it up. But this clashes with the colloquial meaning of "raw math prodigy". If Ramanujan had obtained his amazing results by just blindly listing all provable theorems for billions of years, then occasionally using a time machine to bring the theorems back to our time, would we consider him a prodigy? Or if he got the results by cloning himself a billion times and spreading this blind work?
So if we're going to put a 4-day cap on the "team of Fields medalists", we need to put a cap on AI as well, which is your FAQ#5, but I feel is being forgotten in your latest post.
This concern has already manifested in practice. OpenAI solved Navier-Stokes by spending more than $20M. Given that their result was just the final brick in an edifice that humans had been constructing for decades, and that $20M can pay the salary of ~10 top mathematicians for ~10 years, it is not at all clear that solving N-S with AI was cheaper.
the unit distance problem and the Jacobian conjecture and especially Navier-Stokes already put us far beyond what was required for YES
I don't think this follows at all. The market is explicitly about the AI dominating humans in ALL problem solving. If a few examples were enough, then this market would have resolved positively in the 1990s if not before: Robbins algebra - Wikipedia
@pietrokc Re: recursive enumerability: true fact! It's a little like how infinite monkeys on infinite typewriters eventually reproduce Shakespeare. And you're right about FAQ#5 (cost and speed). It's meant to plug that loophole.
I take your point that for Navier-Stokes in particular it's not obvious whether humans could've gotten there with a similar budget. Here are Claude's estimates on human-vs-AI costs for the biggest open problems AI has solved:

@pietrokc Re: the Robbins conjecture: the above open problems seem pretty different to me, but I'm not a mathematician. As for how to interpret this market, it was originally something like "presumably AI won't solve major open problems pre-AGI but will it mog humans?" and then FAQ#7 considered the case that it actually were to solve major open problems, implying that that would suffice if it weren't otherwise too dumb.
So how dumb is it otherwise, could be the key question.
If we don't have good examples of problems that amateurs can solve that AI can't, then I think that'll mean we're already at YES.
(I know this isn't likely to matter for the actual resolution since there are 3.3 years for these i's/t's to get dotted/crossed but I'm actually very interested in the question of when exactly the superhuman math problem-solving milestone should be considered to have been crossed. Probably sometime within 2026?)
@dreev What are the human costs in those estimates based on? Wild ranges aside, they kind of treat these problems as something that can be solved by throwing money at human mathematicians which, if it were actually the case, would've been done decades ago by big businesses who'd directly benefit from having those solutions on hand. But mathematical ability isn't a commodity; there's no guarantee that any number of people would be able to figure many of these problems out regardless of the available time and motivation. And if we're talking the best of them, money is still not the bottleneck.
@moozooh Yeah, those are just wild guesses at the market rates of Fields Medalists and how long those problems would've taken them. I agree we can't put much stock in those numbers and for the biggest open problems maybe the answer should be like a quadrillion dollars if the whole worldwide math community was stumped for decades. There's some ambiguity with Navier-Stokes but it seems like for everything else, the AI solutions cost much less than if you had to hire humans to solve them.
Mostly the point of talking cost is just to close the loophole where the AI solves math problems by quasi brute force with absurd amounts of compute. The point of automation is to do it faster/cheaper than humans.
@dreev Yeah, being cheaper is a foregone conclusion imo, we still have three years to drop intelligence costs and they're already more than competitive in this particular domain.
@dreev I don't have time for a longer response now but those Claude estimates are totally off. As is to be expected; what data would Claude be drawing on, to make these calculations?
I fear a lot of folks in these threads have a caricature of math in mind. It is absolutely not the case that "AI mogs humans" even in the case of Navier-Stokes, as I alluded to in my "last brick in an edifice" comment.
To give just one example: $200k/year is a very respectable tier-1 US research university salary. The salary is actually much lower in Europe! (These are the people who actually solve most problems, Fields medal publicity and idol worship notwithstanding.) But that salary doesn't buy you an Erdos unit distance solver; it buys a solver of whatever problems the prof feels like solving, following whatever strategy they feel is most likely to succeed, most fruitful even if unsucessful, etc. With AI, on the contrary, you can tell it exactly what to do, and you can spawn one copy for each strategy you can conceive of.
I claim, and I think this is general community consensus, that if you could have paid a combinatorial geometer to specifically work on disproving Erdos unit distances, there is a very good chance they would have been done in 6mo. You may read Gowers's remarks about this.
@pietrokc Yeah, I guess Claude imagined Fields Medalists would demand sky-high wages. But note that some rows in that table had AI as 1000x cheaper, so even paying grad student rates wouldn't change much.
As for Gowers and the unit distance problem, I think the cope around that was reasonable at the time, back in May. What does Gowers say now? Again, we don't have to fully adjudicate this. Any remaining ambiguity about cost/speed should resolve soon enough.
Interestingly, Toby Ord concluded the other day that the insane amount OpenAI spent on Navier-Stokes was purely to compress it into a week so they could get the scoop. He estimates that they could've spent 10% as much doing it in a month or so. A single agent, at 1% of the cost, could've done it in a year. https://www.lesswrong.com/posts/6cb7qd3RSkgnviCpf/swarm-scaling
I think the best way to argue we're not at YES yet is to point to the biggest failures of frontier models we can find. Is there any math problem you can solve that a frontier model can't?
This feels like it's there in spirit, right? And within epsilon for the most persnickety possible reading of the market criteria. FrontierMath is close to saturated and AI aces the IMO and is close on the Putnam.
Reviewing each of the FAQ items:
No AGI yet: N/A
Saturating all benchmarks: ✅ - 𝜀
No Adversarial posing of questions: N/A
No multi-page problem statements: N/A [4b. Research taste not required: N/A]
Cost and speed: ✅
Public availability not required: N/A
Jaggedness: ✅
I don't see how any but FAQ#2 could militate for NO and we've still got over 3 years to go to close that epsilon gap. Even if we don't quite close that gap, FAQ#7 suggests that we could still get to a YES despite AI not completely dominating humans at all math problem-solving, as long as the peaks of its achievements are high enough. As of the Navier-Stokes solution, we've got that in spades.
PS: Navier-Stokes did cost millions in compute, but humans couldn't solve that one at all. Even if you think a team of Fields Medalists could've solved it, I'm pretty sure the cost, paying them market rates, would've been even higher. For the hardest contest and benchmark problems, the compute cost and speed seems to blow humans away. And, again, still 3 years for it all to get even better/cheaper/faster.
Btw, if I end up too biased, I'm happy to defer to a neutral arbiter for resolution. Just let me know what you think I'm missing above.
(And note that when this market was new I dumped thousands of mana into NO. Then reality played out.)
@dreev I would say it's not there yet based on public information. The open problems that have been solved are a small fraction of the most important. We can also see frontier maths open problems to have a more precise denominator. Further all the solutions were within human reach as far as I can tell and have been less clever than you would expect based on the problems prestige. In particular the most impressive results are all counter examples so far. Based on this I don't think it outperforms a team of fields medalists for a week on > 50% of problems. However if Open AIs internal results are as good as they suggest this might well be moot.
@dreev I don't think your assessment is correct.
As I had warned below, the claim in this market is difficult to establish either way. However, recently there was a SAIR competition on the inverse Galois problem. Basically, for each one of the 25,000 transitive permutation groups contained in S_24, find a polynomial whose Galois group is that group. What's interesting about this competition is that everyone was highly encouraged to use AI, and (I believe) got infinite free credits to do so. So here is a scenario in which dozens (hundreds?) of people were trying to combine their efforts with all that AI could offer, to solve the same 25,000 problems. Everyone had two months to do it.
Nevertheless, the winning team did not use AI at all. And you can see that they were the only team to solve this problem for hundreds upon hundreds of specific groups.
More generally, mathematics is incredibly vast, and I worry that this market will be resolved based on vibes from announcements like "OpenAI solved 100 open problems". There is a big PR firehose spraying us with the (very impressive!) positive examples, but what is the mechanism for us to learn of the negative examples?
There's thousands of papers posted on arxiv every day, of which definitely hundreds solve some problem which hadn't been solved before, and of which quite likely dozens had AI tried on it and it didn't work.
I've already exited this market with substantial profit but I'm a bit worried how this will be adjudicated at expiry. How does one tell what "a team of Fields medalists can solve" without, well, giving it to them?
One way this market could resolve NO is if there are some clearly easy problems that AI cannot solve. But these are getting harder to find.
Since getting "a team of Fields medalists" together almost certainly won't happen (they are busy people with diverse specializations), the other plausible way this could resolve NO is if AI cannot solve the problems that led to people winning Fields medals.
It's pretty clear to professionals that this is currently the case; AI today could not have solved, say, the three-dimensional Kakeya conjecture, if starting from where Hong Wang did. But how would we prove this? If you ask an AI today, it will know the solution from the internet. Even if you asked an AI a week before a solution was posted, it would have known all the partial results that Wang, Zahl &co got over the years leading up to it.
So how is it proposed to resolve this market?
@pietrokc Great clarification. I had in mind a day or a week for the hypothetical team of Fields medalists -- not long enough to do something career-making. Does that sound fair to people before I update the market's FAQ?
@dreev My concern is that "what a hypothetical team of Fields medalists can accomplish in a week", which is not even is well-defined in the first place, will be defined for the purposes of this market by people who are not professional mathematicians -- taking it even farther from its already nebulous meaning.
@pietrokc I'm very open to ideas for operationalizing this better. And I'll do my best to be fair regardless. The consensus of the math community should be the ground truth, and there's a decent chance there'll be no ambiguity about that.
https://x.com/dmitryrybin1/status/2079904005652893709?s=46
So this is the quality of prompt required to get the AI to solve a major unsolved problem:
https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de063
Post 1: Construct a counterexample to general (non-planar) case of Dinitz Garg Goemans conjecture. You should do a breakthrough and find a structured counterexample.
[thought for a long time and didn't solve it]
Post 2: please continue research and find a complete unconditional counterexample
[thought for a long time and didn't solve it]
Post 3: Continue the search. Have a clear strategy obtained from deeper understanding of the problem structure.
[thought for a long time and didn't solve it]
Post 4: it's enough of partial results. let's finish with a complete unconditional counterexample
[thought for a long time and did solve it!]
The twitter thread had like 5 other people post other counterexamples that disproved other open mathematical questions lol.
The fact that these are all counterexamples are telling, but that's still majorly superhuman in large parts of math.
IMHO its abilities are doubtlessly still spikey (for instance you could probably still trick it with some vision based problems), but this resolving yes is far more likely then it was even a month ago.
Big jumps in priors this year, not just the recent Jacobian Conjecture news: https://www.reddit.com/r/singularity/comments/1v1aie6/apparently_the_jacobian_conjecture_was_just/