Resolution criteria
This market resolves to YES if, by September 15, 2026, an independent, verified SWE-bench score for Grok 4 (or Grok 4 Code) is published within the range of 62% to 85% (inclusive). The benchmark variant (such as SWE-bench Verified or SWE-bench Pro) must be explicitly stated alongside the score.
This market resolves to NO if no qualifying score is published by the close date, or if the published scores from qualifying sources fall entirely outside the 62% to 85% range.
Data Sources & Rules:
Primary Source: Artificial Analysis will be checked first.
Secondary Source: If not available on Artificial Analysis, the official SWE-bench Leaderboard will be checked second.
Exclusions: Self-reported lab figures from xAI, promotional slides, social media screenshots, and unattributed community-run benchmarks do not qualify. The score must be officially indexed on one of the two specified platforms.
Background
Upon its release, xAI claimed that Grok 4 / Grok 4 Code achieved a 72% to 75% score on SWE-bench. However, third-party benchmarks and independent evaluations can vary from internal lab results due to differences in agent scaffolding, test harnesses, and run parameters. This market tracks whether leading independent AI benchmarking platforms will officially confirm Grok 4's coding capabilities within a 10-percentage-point margin of xAI's initial claims.
This description was generated by AI. Review and verify everything here yourself. You can edit, replace, or delete any part of this description, including the resolution criteria. You do not need to trust the AI output.