Skip to main content
MANIFOLD
Will Anthropic’s next Sonnet model exceed 65% on terminal bench?
11
Ṁ100Ṁ5.4k
resolved Jul 27
Resolved
NO

Will be looking toward https://www.tbench.ai/ for evals, using the terminus 2 scaffolding.

Only counts if the number in the model’s name increments, so a new Claude Sonnet 4.5 checkpoint does not count.

If a new Sonnet model is not released by 2027 this will resolve NA

Market context
Get
Ṁ1,000
to start trading!

🏅 Top traders

#TraderTotal profit
1Ṁ230
2Ṁ170
3Ṁ55
4Ṁ48
5Ṁ0
Sort by:

I resolved this no. Tbench.ai never posted results, but it seemed reasonable to resolve it this way because evidence suggested Sonnet 4.6 could not achieve this score, and my description never promised to only consider tbench.ai results. Sorry for any confusion.

@JaundicedBaboon I don't think anybody is going to test Sonnet 4.6 on Termial Bench, Anthropic had Sonnet 4.6's Terminal-Bench 2.0 score at 59.1%, but nobody has submitted its results to the leaderboard yet. I don't know if you think Anthropic's results are good enough of if you want to continue waiting for somebody to submit results to the leaderboard.