Recent releases from leading labs have intensified competition on MathArena benchmarks, with OpenAI’s GPT-6 Astra (max) topping expected performance at 90.7% shortly after its early September 2026 launch, ahead of Anthropic’s Claude-Opus-5 and Claude-Fable-5.1 variants. Moonshot AI’s Kimi K-series models show strong results on AIME-style and research-level problems, while open-weight entries from Z.ai and Alibaba remain competitive on cost-adjusted metrics. Frontier models now routinely exceed 95% on saturated olympiad tests yet face steeper challenges on proof generation and arXiv-derived tasks, creating uncertainty around any specific year-end threshold. Traders are watching for additional model drops or capability jumps before December 31 that could shift the leaderboard.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated$119,649 Vol.
1575
74%
1600
29%
$119,649 Vol.
1575
74%
1600
29%
Results from the "Score" column under the "Text Arena | Math" Leaderboard tab at https://arena.ai/leaderboard/text/math-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Market Opened: Apr 2, 2026, 6:07 PM ET
Resolver
0x65070BE91...Results from the "Score" column under the "Text Arena | Math" Leaderboard tab at https://arena.ai/leaderboard/text/math-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Resolver
0x65070BE91...Recent releases from leading labs have intensified competition on MathArena benchmarks, with OpenAI’s GPT-6 Astra (max) topping expected performance at 90.7% shortly after its early September 2026 launch, ahead of Anthropic’s Claude-Opus-5 and Claude-Fable-5.1 variants. Moonshot AI’s Kimi K-series models show strong results on AIME-style and research-level problems, while open-weight entries from Z.ai and Alibaba remain competitive on cost-adjusted metrics. Frontier models now routinely exceed 95% on saturated olympiad tests yet face steeper challenges on proof generation and arXiv-derived tasks, creating uncertainty around any specific year-end threshold. Traders are watching for additional model drops or capability jumps before December 31 that could shift the leaderboard.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated



Beware of external links.
Beware of external links.
Frequently Asked Questions