Resolution criteria
This market evaluates whether an open-weight AI model with a memory requirement or size footprint of under 32GB VRAM (at ok quant, usually q4, and also it is about VRAM, it may need eg up to 96gb total, if MoE, 32gb is a counter for the part which is really expensive to run at ok speeds) matches or exceeds the benchmark performance of Anthropic's Claude Opus 5.5 by February 2027.
Each answer corresponds to a specific benchmark or evaluation metric:
SWE-bench: Resolves to YES if an open-weight model with a model/weights size or standard runtime footprint under 32GB achieves a score greater than or equal to Claude Opus 5.5 on SWE-bench (Verified, Pro, or standard test set, evaluated under comparable agent scaffolding) by February 28, 2027. Otherwise, resolves to NO.
[Reserved options]: New benchmark or capability criteria may be added by users. Each added option will resolve YES if a qualifying <32GB open-weight model matches or outperforms Claude Opus 5.5 on that benchmark by February 28, 2027, according to official leaderboards or peer-reviewed/verified evaluation reports.
Official leaderboards such as SWE-bench or official technical reports published by Anthropic and open-model providers will serve as primary sources of truth. If Claude Opus 5.5 is not released or evaluated on a given benchmark by the resolution date, or if no qualifying open model meets the criteria, options without comparable data resolve based on available public evaluations (same benches), N/A if no eval at that bench, for .subjective judgements is my subjective judgement category.
Background
Qwen3.8 27B is said by some benches to equal or outpace Opus 4.6 which came out in February of 2026, that is, ~6 months before qwen3.8, crude (linear) estimate expects median to be models of similar size be outpacing such models as opus 5.5, fable 5.1, gpt 6 astra ~six months from their release. Qwen 3.8 27b fits into 32 vram at high speed, good quant and big context. Hence 50% on that month.
Main topic is autonomous coding, like all those "clone of Minecraft in 7 minutes challenge".
I want to know what markets lean from that.
I noticed that it seems I often won when bet while feeling markets are mispriced, so I may bet here on objective benches (won't bet on my subjective judgement obviously).
Related (not my):
https://manifold.markets/Laplace/an-open-weight-model-beats-opus-47?r=RW5pU2Np