OpenAI models currently trail Anthropic’s Claude Fable 5.1 and Opus 5 on Humanity’s Last Exam, with the strongest reported OpenAI scores (GPT-5.4 Pro and GPT-6 Astra) in the mid-to-high 50s percent range versus Claude’s 59–65 percent as of early September 2026. Trader sentiment reflects OpenAI’s rapid release cadence of reasoning-enhanced GPT variants and tool-augmented evaluations that have closed much of the gap since the benchmark’s January 2026 launch, yet Anthropic maintains a narrow lead through superior performance on graduate-level math, physics, and multi-step reasoning questions. Key catalysts through year-end include any new OpenAI model launches or inference optimizations that could push scores above 60 percent before December 31 resolution, alongside continued benchmark revisions via HLE-Rolling that may alter relative standings.
基於Polymarket數據的AI實驗性摘要。這不是交易建議,也不影響該市場的結算方式。 · 更新於$76,741 交易量
50%以上
90%
55%以上
78%
60%+
38%
65% 以上
19%
70%以上
11%
$76,741 交易量
50%以上
90%
55%以上
78%
60%+
38%
65% 以上
19%
70%以上
11%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
市場開放時間: Jul 23, 2026, 6:53 PM ET
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
OpenAI models currently trail Anthropic’s Claude Fable 5.1 and Opus 5 on Humanity’s Last Exam, with the strongest reported OpenAI scores (GPT-5.4 Pro and GPT-6 Astra) in the mid-to-high 50s percent range versus Claude’s 59–65 percent as of early September 2026. Trader sentiment reflects OpenAI’s rapid release cadence of reasoning-enhanced GPT variants and tool-augmented evaluations that have closed much of the gap since the benchmark’s January 2026 launch, yet Anthropic maintains a narrow lead through superior performance on graduate-level math, physics, and multi-step reasoning questions. Key catalysts through year-end include any new OpenAI model launches or inference optimizations that could push scores above 60 percent before December 31 resolution, alongside continued benchmark revisions via HLE-Rolling that may alter relative standings.
基於Polymarket數據的AI實驗性摘要。這不是交易建議,也不影響該市場的結算方式。 · 更新於



警惕外部連結哦。
警惕外部連結哦。
Frequently Asked Questions