Recent LiveBench coding snapshots show Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol variants trading narrow leads through agentic tasks, tool use, and code generation, with scores clustered within a few points of each other. July releases of these models, alongside Claude Opus 5 and incremental updates from Google, Moonshot, and open-weight labs like DeepSeek, have tightened the competitive field on contamination-resistant benchmarks. Trader sentiment reflects this parity plus the likelihood of further October refreshes or new model drops shifting relative rankings before resolution, while lower odds for Mistral, ByteDance, and others align with their trailing positions on current leaderboards.
Polymarket ডেটা রেফারেন্স করে পরীক্ষামূলক AI-জেনারেটেড সারাংশ। এটি ট্রেডিং পরামর্শ নয় এবং এই মার্কেট কীভাবে রেজলভ হয় তাতে কোনো ভূমিকা রাখে না। · আপডেটেডZ.ai 47%
Alibaba 46%
OpenAI 45%
SpaceXAI 45%

Z.ai
47%

Alibaba
46%

OpenAI
45%

SpaceXAI
45%

Anthropic
45%

Moonshot
45%

Baidu
45%

Microsoft
44%

Meta
44%

StepFun
43%

Thinky
43%

Nvidia
43%

Xiaomi
28%

Mistral
28%

MiniMax
28%

ByteDance
27%

Tencent
26%

26%

Meituan
25%

DeepSeek
24%

Amazon
19%
Z.ai 47%
Alibaba 46%
OpenAI 45%
SpaceXAI 45%

Z.ai
47%

Alibaba
46%

OpenAI
45%

SpaceXAI
45%

Anthropic
45%

Moonshot
45%

Baidu
45%

Microsoft
44%

Meta
44%

StepFun
43%

Thinky
43%

Nvidia
43%

Xiaomi
28%

Mistral
28%

MiniMax
28%

ByteDance
27%

Tencent
26%

26%

Meituan
25%

DeepSeek
24%

Amazon
19%
Results from the “Coding” column of the leaderboard at https://livebench.ai/#/?cats=Coding, with the latest available LiveBench release selected and the category set to “Coding,” will be used to resolve this market.
Models will be ranked according to the specified score, with higher scores ranked ahead of lower scores. If two or more models have exactly the same score as displayed on the leaderboard, the model with the lower listed "cost per successful task" will be ranked ahead. If a tie still remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., if the two models are tied by exact score and cost per successful task, “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies first place under this ranking.
The resolution source for this market is the LiveBench leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to “Other.”
মার্কেট ওপেন হয়েছে: Aug 12, 2026, 8:05 PM ET
Resolver
0x69c47De9D...Results from the “Coding” column of the leaderboard at https://livebench.ai/#/?cats=Coding, with the latest available LiveBench release selected and the category set to “Coding,” will be used to resolve this market.
Models will be ranked according to the specified score, with higher scores ranked ahead of lower scores. If two or more models have exactly the same score as displayed on the leaderboard, the model with the lower listed "cost per successful task" will be ranked ahead. If a tie still remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., if the two models are tied by exact score and cost per successful task, “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies first place under this ranking.
The resolution source for this market is the LiveBench leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to “Other.”
Resolver
0x69c47De9D...Recent LiveBench coding snapshots show Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol variants trading narrow leads through agentic tasks, tool use, and code generation, with scores clustered within a few points of each other. July releases of these models, alongside Claude Opus 5 and incremental updates from Google, Moonshot, and open-weight labs like DeepSeek, have tightened the competitive field on contamination-resistant benchmarks. Trader sentiment reflects this parity plus the likelihood of further October refreshes or new model drops shifting relative rankings before resolution, while lower odds for Mistral, ByteDance, and others align with their trailing positions on current leaderboards.
Polymarket ডেটা রেফারেন্স করে পরীক্ষামূলক AI-জেনারেটেড সারাংশ। এটি ট্রেডিং পরামর্শ নয় এবং এই মার্কেট কীভাবে রেজলভ হয় তাতে কোনো ভূমিকা রাখে না। · আপডেটেড
বাহ্যিক লিংক থেকে সাবধান।
বাহ্যিক লিংক থেকে সাবধান।
সচরাচর জিজ্ঞাসা