Recent LiveBench mathematics results reflect a tight race among frontier large language models, with market odds clustering near 45-50% for leading contenders as traders weigh incremental gains from reasoning-focused releases. OpenAI's GPT-5 series and o-series variants currently post the strongest scores on the contamination-resistant math tasks, driven by scaled test-time compute and targeted fine-tuning, while Anthropic's Claude 5 Opus and related reasoning modes trail but remain competitive on proof-style and competition problems. Chinese labs like DeepSeek and Moonshot trail further due to narrower gaps in advanced contest math, though rapid iteration keeps them in play. With resolution at month-end, any new model drops, efficiency updates, or benchmark refreshes in August could shift the narrow margins, underscoring how live evaluation rewards timely capability jumps over static leaderboards.
Polymarket ডেটা রেফারেন্স করে পরীক্ষামূলক AI-জেনারেটেড সারাংশ। এটি ট্রেডিং পরামর্শ নয় এবং এই মার্কেট কীভাবে রেজলভ হয় তাতে কোনো ভূমিকা রাখে না। · আপডেটেডOpenAI 45%
Anthropic 38%
Alibaba 25%
Xiaomi 23%

OpenAI
42%

Anthropic
33%

Alibaba
25%

Xiaomi
23%

DeepSeek
22%

SpaceXAI
22%

22%

Moonshot
21%

Meta
20%

ByteDance
18%

Baidu
9%

Z.ai
9%

Thinky
6%

Nvidia
7%

Mistral
3%

Tencent
3%

Amazon
3%

Meituan
1%

StepFun
1%

Microsoft
1%

MiniMax
-
OpenAI 45%
Anthropic 38%
Alibaba 25%
Xiaomi 23%

OpenAI
42%

Anthropic
33%

Alibaba
25%

Xiaomi
23%

DeepSeek
22%

SpaceXAI
22%

22%

Moonshot
21%

Meta
20%

ByteDance
18%

Baidu
9%

Z.ai
9%

Thinky
6%

Nvidia
7%

Mistral
3%

Tencent
3%

Amazon
3%

Meituan
1%

StepFun
1%

Microsoft
1%

MiniMax
-
Results from the “Mathematics” column of the leaderboard at https://livebench.ai/#/?cats=Mathematics, with the latest available LiveBench release selected and the category set to “Mathematics,” will be used to resolve this market.
Models will be ranked according to the specified score, with higher scores ranked ahead of lower scores. If two or more models have exactly the same score as displayed on the leaderboard, the model with the lower listed "cost per successful task" will be ranked ahead. If a tie still remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., if the two models are tied by exact score and cost per successful task, “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies first place under this ranking.
The resolution source for this market is the LiveBench leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to “Other.”
মার্কেট ওপেন হয়েছে: Jul 29, 2026, 6:27 PM ET
রেজোলিউশন সোর্স
https://livebench.ai/#/?cats=MathematicsResolver
0x69c47De9D...Results from the “Mathematics” column of the leaderboard at https://livebench.ai/#/?cats=Mathematics, with the latest available LiveBench release selected and the category set to “Mathematics,” will be used to resolve this market.
Models will be ranked according to the specified score, with higher scores ranked ahead of lower scores. If two or more models have exactly the same score as displayed on the leaderboard, the model with the lower listed "cost per successful task" will be ranked ahead. If a tie still remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., if the two models are tied by exact score and cost per successful task, “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies first place under this ranking.
The resolution source for this market is the LiveBench leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to “Other.”
রেজোলিউশন সোর্স
https://livebench.ai/#/?cats=MathematicsResolver
0x69c47De9D...Recent LiveBench mathematics results reflect a tight race among frontier large language models, with market odds clustering near 45-50% for leading contenders as traders weigh incremental gains from reasoning-focused releases. OpenAI's GPT-5 series and o-series variants currently post the strongest scores on the contamination-resistant math tasks, driven by scaled test-time compute and targeted fine-tuning, while Anthropic's Claude 5 Opus and related reasoning modes trail but remain competitive on proof-style and competition problems. Chinese labs like DeepSeek and Moonshot trail further due to narrower gaps in advanced contest math, though rapid iteration keeps them in play. With resolution at month-end, any new model drops, efficiency updates, or benchmark refreshes in August could shift the narrow margins, underscoring how live evaluation rewards timely capability jumps over static leaderboards.
Polymarket ডেটা রেফারেন্স করে পরীক্ষামূলক AI-জেনারেটেড সারাংশ। এটি ট্রেডিং পরামর্শ নয় এবং এই মার্কেট কীভাবে রেজলভ হয় তাতে কোনো ভূমিকা রাখে না। · আপডেটেড
বাহ্যিক লিংক থেকে সাবধান।
বাহ্যিক লিংক থেকে সাবধান।
সচরাচর জিজ্ঞাসা