Skip to main content
MANIFOLD
Which AI lab will get a perfect score on the IMO 2026?
134
Ṁ1.4kṀ24k
Aug 15
53%
Google
51%
OpenAI
47%
Aristotle
20%
Meta
17%
Z.ai / Zhipu
16%
Moonshot
12%
DeepSeek
9%
xAI
Resolved
YES
Anthropic

Regular AI IMO rules apply

Market context
Get
Ṁ1,000
to start trading!
Sort by:

anthropic is so tuffington and auraful

@Bayesian Was the YES for Anthropic because of Deedy's result that I posted, or because of the Opus 5 system card?

Basically, what decisions have you made about the criteria?

@AhronMaline the system card. Didnt make a decision against deedy, but system card is reliable and doesn’t seem controversial

bought Ṁ20 NO

@AhronMaline This is not an official attempt. It's also graded by AI. Also, it depends whether the market creator will accept such attempts.

@DottedCalculator Didn't he already say it doesn't need to be "official"? There's no "statistical trick" here, Fable did it on the first try. Sol "needed two tries" for one of the questions, so that's more questionable.

As for who does the grading - yeah, that's a known problem with these markets. But if the results seem perfect superficially and we don't have experts offering to check, then what criterion could we demand that's better than AI grading?

@AhronMaline I think such situations should have been clarified earlier. I don't see a clear response from the market creator.

I might try to read the Fable solutions if I have time, but I do expect that it should be a 42.

bought Ṁ50 YES

@Bayesian I think Aristotle already reported a perfect score

@Chimpy oh nevermind

@Bayesian can we have a clarification on what it means for a lab to get perfect? Does someone running 5.6 Pro and getting perfect count or does it need to be some sort of “official” attempt?

@DottedCalculator Hmmmmmmm it needs to be a real success, so there would be more scrutiny if only an independent 5.6 pro user reported this success, but if they bring the receipts it would count.

@DottedCalculator is this not a 41/42? I don’t want to answer a definite yes or no bc ive not gone into the weeds enough to be confident but leaning no

@Bayesian suppose theoretically I take these solutions and grade it as a 42/42 (I actually believe the problem 2 in the linked paper should be a 7 rather than a 6 under IMO rubrics), would that count?

EDIT: let me explain a little more about problem 2. i'm just looking at the section where the author explains why it got a 6. with the exception of point 4, which is just a minor typo and should never be a dock, the rest of the complaints are configuration issues. the IMO rubric explicitly states to not dock for configuration issues (the last time configurations were docked was 2023/6).

@DottedCalculator hmmm if the majority of uninvolved mathematicians with relevant experience agree with u i kinda think it should count yeah, but from you only or from one offs i would prolly say that’s not sufficient

@Bayesian how would we find this "majority of uninvolved mathematicians with relevant experience" to judge such cases

@DottedCalculator we hope it doesn’t come to that but uh yeah it’s tough

@Bayesian what about in the case where 50% of 5.6 pro runs result in a 42 and the others don't, would that still count? what if it's 1% and I run it 100 times?

@DottedCalculator if you run it 100 times and only report the success it doesnt count, but if u run it 100 times and have the ai judge which one is correct and u grade it once that’s fine. Basically no statistical tricks

@Bayesian what would be the difference between me testing it 100 times and 100 different people testing it once and only the one success happens to be posted online?

@DottedCalculator sorry, both arent allowed, the thing allowed is you having 100 parallel instances testing it once and then chatting with one another about what solution is right, and if they are good enough at grading themselves then they solved the problem as a team, individually from human help and statistical tricks. so to speak

@Bayesian how would you know whether there were statistical tricks involved (intentionally or unintentionally)

@DottedCalculator mortal human judgement

@Bayesian this is a strange and unexpected rule. Presumably there are many individuals attempting to test various models on the IMO, and of course there is a heavy selection effect about what gets reported. How could we ever decide whether a success is one out of some large number of tries?

@AhronMaline i want to exclude obvious gotchas but count real successes. do you suggest some improvement/alternative?

@Bayesian If we do have multiple known tries with the same conditions, then it makes sense to set a threshhold of what percent of the time it gets a perfect score. Maybe 33%?

But in practice, I guess most individuals will only run it once, because of the expense. And then if one succeeds, we never know whether it's luck or some factor like a better harness, which should be valid. So I would want to just treat all such single runs as OK.

@AhronMaline hmm I would be pretty put off if they weren't either official runs or consistently succeeding tbh