In AI 2027, at the end of 2026, the leading AI lab released a model dubbed Agent 1 mini, which was 10x cheaper than the previous state of the art Agent 1. Claude Fable's benchmark scores are mostly above the level predicted for Agent 1 in AI 2027 (outscoring the projections for OS World and Cybench, with Mythos Preview scoring similarly to the projected METR 80% time horizon, though somewhat below). This market seeks to track if a model of approximately Agent 1 mini capabilities is released around the time AI 2027 projected.
This market will resolve YES if any AI is released that scores at least as well as Claude Fable 5 on at least 8 out of 10 of the following benchmarks, and whose output tokens are at least 10x cheaper than Claude Fable 5's, by the end of March 2027. The model must be released by the end of March to the public, but resolution may happen later due to the benchmarks being run later.
Resolution sources are included with each benchmark, though these are subject to change for now
SWE-bench Pro: 80.0%
https://llm-stats.com/benchmarks/swe-bench-pro
Frontier Code 1.1: 53.5%
https://cognition.com/frontiercode
Terminal bench: 44.5%
https://www.tbench.ai/
Automation bench: 17.4%
https://llm-stats.com/benchmarks/automationbench
OS-World Verified: 86%
https://osworld-v1.xlang.ai/
ARC AGI 2: 89.2%
https://arcprize.org/leaderboard
Epoch Capabilities Index: 162
https://epoch.ai/eci?subset-view=graph&subset-tab=Software+engineering&view=graph&tab=release-date
FrontierMath Tier 4: 88%
https://epoch.ai/frontiermath/tiers-1-4?view=graph&tab=release-date&tier=Tier+4+%28v2%29
Artificial Analysis Intelligence Index: 62
https://artificialanalysis.ai/#intelligence
Humanity's Last Exam: 55.5%
https://artificialanalysis.ai/evaluations/humanitys-last-exam
If any of these benchmarks are no longer updated by this date, it will be removed from the resolution criteria, with a threshold of greater than or equal than 75% of the remaining benchmarks being used for resolution.
If no such model is released by the end of March 2027, this market resolves no.
I will not trade in this market.
See the version for the end of 2026 here:
https://manifold.markets/BenAybar/will-any-model-at-least-as-good-as-qzn2sLyP2C?r=QmVuQXliYXI
People are also trading
@Bayesian I wanted a higher bar than just tying the eci but if others feel similarly that this would be better for the spirit of the question I'm open to changing it
[edited comment] sorry missed a part of your description!
@Bayesian if they update the benchmark and don't put old models on it, I'll remove the benchmark from the scoring and use a threshold of at least 75% of the remaining benchmarks instead.