Skip to main content
MANIFOLD
What will be the best score on First Proof - Second Batch?
21
Ṁ1.1kṀ4.2k
resolved Jul 1
100%15%
6
0.9%
0
3%
1
4%
2
8%
3
17%
4
18%
5
8%
7
12%
8
7%
9
6%
10

https://1stproof.org/

Resolves to the best score that an AI model yields on the 1st Proof benchmark's second batch.

They will release ten (more) math questions, based on open research questions from top mathematicians.

Previous round resolved to my best judgment, since there were some slight disputes as to the best score, but miraculously this was without controversy! This round should be easier to adjudicate as I think the 1stproof org will grade their own benchmark and publish the results in June.

If the questions are graded with partial scores, I will resolve with PROB to the two nearest integers. So, for example, if the answer is 7.8, that would resolve 20% to 7 and 80% to 8. I promise this makes sense and is normal.

I will not bet on this market, so I can be an unbiased judge. I may consult with leading mathematicians (if I can find them). I will probably extend the market until just before the results are released.

Market context
Get
Ṁ1,000
to start trading!

🏅 Top traders

#TraderTotal profit
1Ṁ254
2Ṁ126
3Ṁ67
4Ṁ54
5Ṁ37
Sort by:
reposted

It's not super easy to tell from the document, but I'm giving the score as the number of questions that the best AI system got either "flawless" or "minor revisions" (which were generally not math-related, but related to citation errors!), which came out to 6 for System A: Problems 1,2,3,5,7,9.

This should very likely have been 7, but uhhhh:

I think these results are impressive. Batch 2 was designed to be slightly more challenging and have less trivial questions than Batch 1, if I understand the methodology correctly (I probably don't).

Extending this; should be announced on June 10th!!!

reposted

Re-upping this in light of recent events

A little surprised the consensus is so high on 5+. For one, the submission process is tighter on the second batch, with no room for human selection amid multiple solutions, AFAICT, and it's also possible given the publicity, the problems will be slightly more challenging.

@bens i am much more bullish on model progress so i bet 8+