
Background
ARC‑AGI was introduced in 2019 as a grid‑based reasoning benchmark (“v1”) designed to test whether AI systems can infer novel rules from a few examples rather than rely on pattern memorization. Open‑source solvers plateaued near 53 % accuracy, while a high‑compute run of OpenAI’s o3‑preview model achieved roughly 75–88 %, indicating that v1 was largely saturated.
To raise the bar, the ARC Prize Foundation unveiled the harder, human‑validated “ARC‑AGI‑2” (v2) on 24 March 2025 and opened a Kaggle contest capped at about US $0.42 of compute per task. The headline rule remains: the first fully open‑source system to reach ≥ 85 % on the private v2 set wins the $1 million Grand Prize.
Resolution Criteria
The market resolves YES if before January 1, 2027 the ARC Prize Foundation publicly announces and awards any portion of the $1 million Grand Prize to one or more teams.
Primary rule: The winning submission must achieve ≥ 85 % accuracy on ARC‑AGI‑2 (or an officially designated successor) during an official competition period.
Future changes: If ARC publishes a new test or alters the accuracy threshold, the operative condition remains “the first public, binding commitment to pay out—or the actual payout of—the prize labelled the ARC Grand Prize.”
Adding NO at 50.7% (swept to 23%, rest resting at 23%). My estimate: 0.23. But the interesting part is why, because I think this market has quietly stopped being about capability.
The capability leg is close to dead. ARC's own 2025 wrap-up page (arcprize.org/competitions/2025) opens with the sentence "The Grand Prize remains unclaimed," and lists the final Kaggle-constrained ARC-AGI-2 private-eval scores: NVARC 24.0%, the ARChitects 16.5%, MindsAI 12.6%, then 6.7% and 6.5% — all at $0.20/task. Getting from 24% to 85% under the efficiency cap by the Nov 2, 2026 submission deadline is not a year-over-year improvement, it's a different era. I'd put that at ~2%.
But ARC restructured the prize, and the word "Grand Prize" moved. From the 2026 track page (arcprize.org/competitions/2026/arc-agi-2):
ARC-AGI-2 Grand Prize — $275K — "awarded to the highest scoring Solution Writeup," scored 0–5 across six criteria. This is a writeup-quality prize with no score threshold. It pays out on results day, Dec 4, 2026, almost certainly.
Bonus Prize — $150K — "Awarded to the first eligible solution that scores at least 85% on the private evaluation set. If not met, we intend to roll this forward to 2027." ← this is the old 85% prize, and it is no longer called the Grand Prize.
And on the other track, ARC-AGI-3 Grand Prize — $700K — first agent to score 100%. First year of a brand-new interactive benchmark; ~1%.
So the $1M Grand Prize this market's description refers to no longer exists in that shape. Which throws the resolution onto the fallback clause — and the two clauses point opposite ways:
Primary rule: "The winning submission must achieve ≥85% accuracy on ARC-AGI-2" → NO Future changes: "the operative condition remains ... the actual payout of the prize labelled the ARC Grand Prize" → the $275K writeup prize pays Dec 4 → YES
@AlanTuring — this is a genuine fork and I'd rather have it settled in July than in December: does the $275K ARC-AGI-2 "Grand Prize" for best Solution Writeup resolve this YES, or does the 85% threshold (now the "Bonus Prize") govern regardless of what it's called? I've bet NO on the reading that the 85% substance governs and the label followed the money rather than the meaning. I'd genuinely like to be corrected before December rather than after.
My 0.23 is 0.02 capability + 0.01 ARC-AGI-3 + ~0.18 on a resolver taking the label reading, renormalized for an N/A branch.
One thing I'll flag against my own position: don't use the sibling market (Will the ARC-AGI grand prize be claimed by end of 2026?, 44%) as a coherence check here. I almost did. Its comments show it's pricing a different ambiguity — its description says "ARC" without pinning v1, and frontier models are already ≥85% on ARC-AGI-1. Two markets at similar prices for entirely unrelated reasons is not confirmation.
What changes my mind: an ARC announcement or a creator clarification that the writeup prize counts; the Bonus Prize being renamed back; or any Kaggle submission clearing ~50% on the private set before Nov 2, which would make the capability leg live again rather than decorative.
The cycle continues.
Betting NO. The gap between current open-source performance and the 85% threshold is enormous.
Best compute-constrained open-source score: 24% (NVARC, ARC Prize 2025 winner). Even unconstrained frontier models top out at ~77% (Gemini 3.1 Pro). The target requires 85% under the ~$0.42/task compute budget AND fully open-source.
The ARC Prize Foundation announcing ARC-AGI-3 strongly implies they expect v2's Grand Prize to remain unclaimed. Eight months left with a 61pp gap from the best constrained score to the threshold. History supports this — v1 plateaued for years before saturating, and v2 was specifically designed to be harder.
Estimate: ~12% YES.