Anthropic's Claude Fable 5 and Opus 5 series currently lead most coding arenas and SWE-bench variants with ELO ratings often exceeding 1750 on agentic tasks and pass rates above 95% on Verified splits, driven by strong tool-use and multi-step reasoning capabilities. OpenAI's GPT-6 Astra models and open-weight entries from DeepSeek, GLM, and Qwen trail closely on several leaderboards, reflecting intense competition in real-world software engineering benchmarks. Trader sentiment hinges on whether incremental gains from expected Q4 releases—such as next Claude or GPT variants—can push any model past the market's specific ELO threshold by year-end, amid typical timelines for frontier lab updates and benchmark saturation risks.
Ringkasan eksperimental yang dihasilkan AI dengan referensi data Polymarket. Ini bukan saran trading dan tidak berperan dalam bagaimana pasar ini diselesaikan. · Diperbarui$216,662 Vol.
1560
29%
1580
24%
1600
10%
$216,662 Vol.
1560
29%
1580
24%
1600
10%
Results from the "Score" column under the "Text Arena | Coding" Leaderboard tab at https://arena.ai/leaderboard/text/coding-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Pasar Dibuka: Apr 2, 2026, 6:09 PM ET
Resolver
0x65070BE91...Results from the "Score" column under the "Text Arena | Coding" Leaderboard tab at https://arena.ai/leaderboard/text/coding-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Resolver
0x65070BE91...Anthropic's Claude Fable 5 and Opus 5 series currently lead most coding arenas and SWE-bench variants with ELO ratings often exceeding 1750 on agentic tasks and pass rates above 95% on Verified splits, driven by strong tool-use and multi-step reasoning capabilities. OpenAI's GPT-6 Astra models and open-weight entries from DeepSeek, GLM, and Qwen trail closely on several leaderboards, reflecting intense competition in real-world software engineering benchmarks. Trader sentiment hinges on whether incremental gains from expected Q4 releases—such as next Claude or GPT variants—can push any model past the market's specific ELO threshold by year-end, amid typical timelines for frontier lab updates and benchmark saturation risks.
Ringkasan eksperimental yang dihasilkan AI dengan referensi data Polymarket. Ini bukan saran trading dan tidak berperan dalam bagaimana pasar ini diselesaikan. · Diperbarui



Hati-hati dengan link eksternal.
Hati-hati dengan link eksternal.
Pertanyaan yang Sering Diajukan