Google's Gemini series trails the HLE leaderboard, where Anthropic's Claude Fable 5.1 and Opus 5 lead at 59–65% on expert-authored questions spanning math, physics, and other domains. Recent Gemini 3.1 Pro and 3.7–3.8 Flash variants post 45–48% in text-only or closed-book evaluations, reflecting gains in multi-step reasoning but persistent gaps versus Claude on specialized knowledge tasks. Traders price in continued Google release cadence—including Gemini 4 pre-training and Flash iterations targeting agentic capabilities—against the benchmark's remaining headroom before year-end saturation. Key catalysts include any verified model updates or leaderboard shifts by December 31 that could push a Gemini variant past 50–55% implied probability thresholds.
Polymarketデータを参照したAI生成の実験的な要約。これは取引アドバイスではなく、このマーケットの解決方法には一切関係ありません。 · 更新日$92,724 Vol.
50%以上
90%
55%以上
42%
60%以上
21%
65%以上
12%
70%以上
4%
$92,724 Vol.
50%以上
90%
55%以上
42%
60%以上
21%
65%以上
12%
70%以上
4%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
マーケット開始日: Jul 23, 2026, 6:56 PM ET
リゾルバー
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
リゾルバー
0x65070BE91...Google's Gemini series trails the HLE leaderboard, where Anthropic's Claude Fable 5.1 and Opus 5 lead at 59–65% on expert-authored questions spanning math, physics, and other domains. Recent Gemini 3.1 Pro and 3.7–3.8 Flash variants post 45–48% in text-only or closed-book evaluations, reflecting gains in multi-step reasoning but persistent gaps versus Claude on specialized knowledge tasks. Traders price in continued Google release cadence—including Gemini 4 pre-training and Flash iterations targeting agentic capabilities—against the benchmark's remaining headroom before year-end saturation. Key catalysts include any verified model updates or leaderboard shifts by December 31 that could push a Gemini variant past 50–55% implied probability thresholds.
Polymarketデータを参照したAI生成の実験的な要約。これは取引アドバイスではなく、このマーケットの解決方法には一切関係ありません。 · 更新日



外部リンクに注意してください。
外部リンクに注意してください。
よくある質問