Skip to main content
MANIFOLD
Where will Gemini 4 Pro score on the Artificial Analysis Intelligence Index vs the leading model prior to its inclusion?
3
Ṁ100Ṁ34
2027
30%
Gemini 4 Pro 5-8 points worse
27%
Gemini 4 Pro 0-4 points worse
15%
Gemini 4 Pro scores higher than the prior leading model
27%
Gemini 4 Pro ≥9 points worse

Resolution criteria

This market will resolve based on the score of Google's Gemini 4 Pro on the Artificial Analysis Intelligence Index compared to the leading model's score immediately prior to Gemini 4 Pro's inclusion.

The market will resolve using the official scores published on the Artificial Analysis Intelligence Index (or the main LLM Leaderboard).

  • $S_{Gemini}$: The highest score achieved by any configuration/effort level of Gemini 4 Pro (e.g., standard, max effort, etc.) when its official score is first fully indexed and listed on the leaderboard.

  • $S_{Lead}$: The highest score on the index among all other models immediately prior to Gemini 4 Pro's inclusion.

  • Difference ($D$): Calculated as $D = S_{Lead} - S_{Gemini}$ (representing how many points worse Gemini 4 Pro scored compared to the previous leader).

Resolution Rules:

  • Gemini 4 Pro scores higher than the prior leading model: Resolves YES if $S_{Gemini} > S_{Lead}$ ($D < 0$).

  • Gemini 4 Pro 0-4 points worse: Resolves YES if $0 \le D \le 4$.

  • Gemini 4 Pro 5-8 points worse: Resolves YES if $5 \le D \le 8$.

  • Gemini 4 ≥9 points worse: Resolves YES if $D \ge 9$.

Edge Cases & Details:

  • Index Versions: Both $S_{Gemini}$ and $S_{Lead}$ must be evaluated on the same version of the Artificial Analysis Intelligence Index. If the index version is upgraded at the time of Gemini 4 Pro's release, $S_{Lead}$ will be the highest score of any other model under that updated index version.

  • Deadline: If Gemini 4 Pro is not officially released or graded on the Artificial Analysis Intelligence Index by December 31, 2027, this market will resolve as N/A.

Background

The Artificial Analysis Intelligence Index is a composite benchmark that aggregates scores across multiple demanding evaluations (including Humanity's Last Exam, SciCode, and AA-Briefcase) to evaluate LLM performance in math, science, reasoning, and coding.

As of late September 2026, the leading model on the index is Anthropic's Claude Opus 5.5 (max effort) with a score of 58, followed by Claude Sonnet 5.5 and GPT-6 Astra. Google's highest-ranking model on the leaderboard is Gemini 3.8 Flash with a score of 41, while the older Gemini 3.1 Pro scores 30. Early leaks and arena tests under the codename "Argon" suggest that Gemini 4 Pro is designed to be Google's premier model to challenge the current leaders.

Get
Ṁ1,000
to start trading!