Google's Gemini series trails the HLE leaderboard, where Anthropic's Claude Fable 5.1 and Opus 5 lead at 59–65% on expert-authored questions spanning math, physics, and other domains. Recent Gemini 3.1 Pro and 3.7–3.8 Flash variants post 45–48% in text-only or closed-book evaluations, reflecting gains in multi-step reasoning but persistent gaps versus Claude on specialized knowledge tasks. Traders price in continued Google release cadence—including Gemini 4 pre-training and Flash iterations targeting agentic capabilities—against the benchmark's remaining headroom before year-end saturation. Key catalysts include any verified model updates or leaderboard shifts by December 31 that could push a Gemini variant past 50–55% implied probability thresholds.
Експериментальне резюме, згенероване ШІ з посиланням на дані Polymarket. Це не торгова порада і не впливає на вирішення цього ринку. · ОновленоHighest Google Gemini score on Humanity’s Last Exam in 2026?
$92,724 Обс.
50%+
90%
55%+
42%
60%+
21%
65%+
12%
70%+
4%
$92,724 Обс.
50%+
90%
55%+
42%
60%+
21%
65%+
12%
70%+
4%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Ринок відкрито: Jul 23, 2026, 6:56 PM ET
Вирішувач
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Вирішувач
0x65070BE91...Google's Gemini series trails the HLE leaderboard, where Anthropic's Claude Fable 5.1 and Opus 5 lead at 59–65% on expert-authored questions spanning math, physics, and other domains. Recent Gemini 3.1 Pro and 3.7–3.8 Flash variants post 45–48% in text-only or closed-book evaluations, reflecting gains in multi-step reasoning but persistent gaps versus Claude on specialized knowledge tasks. Traders price in continued Google release cadence—including Gemini 4 pre-training and Flash iterations targeting agentic capabilities—against the benchmark's remaining headroom before year-end saturation. Key catalysts include any verified model updates or leaderboard shifts by December 31 that could push a Gemini variant past 50–55% implied probability thresholds.
Експериментальне резюме, згенероване ШІ з посиланням на дані Polymarket. Це не торгова порада і не впливає на вирішення цього ринку. · Оновлено



Обережно з зовнішніми посиланнями.
Обережно з зовнішніми посиланнями.
Часті запитання