Google's Gemini models currently trail the Humanity's Last Exam (HLE) frontier, with Gemini 3.1 Pro and 3.8 Flash variants posting text-only scores near 47% on Artificial Analysis and similar leaderboards as of mid-September 2026, while Anthropic's Claude Fable 5.1 and Opus 5 lead at 55-65% depending on thinking modes or tools. Rapid benchmark progress—driven by extended chain-of-thought reasoning and specialized knowledge—has lifted the frontier from single digits at the benchmark's 2025 launch to the mid-50s percent range, outpacing original projections. Trader focus centers on Google's aggressive 3.x release cadence, including recent Flash iterations, alongside pre-training for Gemini 4 expected later in 2026. Key catalysts include further reasoning improvements and potential agentic enhancements that could narrow the gap before year-end resolution.
Polymarket ডেটা রেফারেন্স করে পরীক্ষামূলক AI-জেনারেটেড সারাংশ। এটি ট্রেডিং পরামর্শ নয় এবং এই মার্কেট কীভাবে রেজলভ হয় তাতে কোনো ভূমিকা রাখে না। · আপডেটেড$80,437 Vol.
50%+
79%
55%+
25%
60%+
17%
65%+
12%
70%+
4%
$80,437 Vol.
50%+
79%
55%+
25%
60%+
17%
65%+
12%
70%+
4%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
মার্কেট ওপেন হয়েছে: Jul 23, 2026, 6:56 PM ET
রেজলভার
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
রেজলভার
0x65070BE91...Google's Gemini models currently trail the Humanity's Last Exam (HLE) frontier, with Gemini 3.1 Pro and 3.8 Flash variants posting text-only scores near 47% on Artificial Analysis and similar leaderboards as of mid-September 2026, while Anthropic's Claude Fable 5.1 and Opus 5 lead at 55-65% depending on thinking modes or tools. Rapid benchmark progress—driven by extended chain-of-thought reasoning and specialized knowledge—has lifted the frontier from single digits at the benchmark's 2025 launch to the mid-50s percent range, outpacing original projections. Trader focus centers on Google's aggressive 3.x release cadence, including recent Flash iterations, alongside pre-training for Gemini 4 expected later in 2026. Key catalysts include further reasoning improvements and potential agentic enhancements that could narrow the gap before year-end resolution.
Polymarket ডেটা রেফারেন্স করে পরীক্ষামূলক AI-জেনারেটেড সারাংশ। এটি ট্রেডিং পরামর্শ নয় এবং এই মার্কেট কীভাবে রেজলভ হয় তাতে কোনো ভূমিকা রাখে না। · আপডেটেড



বাহ্যিক লিংক থেকে সাবধান।
বাহ্যিক লিংক থেকে সাবধান।
সচরাচর জিজ্ঞাসা