Recent releases from Anthropic and OpenAI, including Claude Opus 5 (max) at 84.4% and GPT-5.6 variants near 80% on MathArena's competition leaderboard, have driven trader sentiment toward yes. These models demonstrate continued gains on uncontaminated olympiad problems through scaled reasoning and post-training, though proof-heavy tasks lag below 40%. Moonshot's Kimi series and other open-weight entries trail at 60-70%, underscoring closed-lab advantages in math-specific capabilities. With four-plus months until year-end, labs are expected to iterate on agentic workflows and larger context windows; however, saturation on existing benchmarks and potential release delays introduce uncertainty around crossing higher thresholds like 90%.
Polymarketデータを参照したAI生成の実験的な要約。これは取引アドバイスではなく、このマーケットの解決方法には一切関係ありません。 · 更新日$112,943 Vol.
1575
81%
1600
26%
$112,943 Vol.
1575
81%
1600
26%
Results from the "Score" column under the "Text Arena | Math" Leaderboard tab at https://arena.ai/leaderboard/text/math-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
マーケット開始日: Apr 2, 2026, 6:07 PM ET
Resolver
0x65070BE91...Results from the "Score" column under the "Text Arena | Math" Leaderboard tab at https://arena.ai/leaderboard/text/math-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Resolver
0x65070BE91...Recent releases from Anthropic and OpenAI, including Claude Opus 5 (max) at 84.4% and GPT-5.6 variants near 80% on MathArena's competition leaderboard, have driven trader sentiment toward yes. These models demonstrate continued gains on uncontaminated olympiad problems through scaled reasoning and post-training, though proof-heavy tasks lag below 40%. Moonshot's Kimi series and other open-weight entries trail at 60-70%, underscoring closed-lab advantages in math-specific capabilities. With four-plus months until year-end, labs are expected to iterate on agentic workflows and larger context windows; however, saturation on existing benchmarks and potential release delays introduce uncertainty around crossing higher thresholds like 90%.
Polymarketデータを参照したAI生成の実験的な要約。これは取引アドバイスではなく、このマーケットの解決方法には一切関係ありません。 · 更新日



外部リンクに注意してください。
外部リンクに注意してください。
よくある質問