OpenAI’s GPT-5.6 Sol and GPT-6 Astra models lead FrontierMath v2 leaderboards with 83–97% scores on Tier 4 research-level problems and have already produced verified solutions to multiple Erdős and open research questions, including novel proofs and counterexamples accepted under Epoch AI’s framework. This dominance stems from scaled reasoning, tool use, and iterative verification pipelines that outpace Anthropic’s Claude Opus 4.x and Google DeepMind’s Aletheia on private benchmarks. Epoch AI’s ongoing evaluations of FrontierMath: Open Problems create near-term resolution catalysts through December 2026, with trader consensus reflecting OpenAI’s consistent recent breakthroughs and the low likelihood of rivals closing the gap before the next model releases or benchmark updates.
Résumé expérimental généré par IA à partir des données Polymarket. Ceci n'est pas un conseil de trading et ne joue aucun rôle dans la résolution de ce marché. · Mis à jourView resolved

Méfiez-vous des liens externes.
Méfiez-vous des liens externes.
Questions fréquentes