Recent benchmark leadership by Shanghai Artificial Intelligence Laboratory's Atria Dawn Preview on verified agentic evals like Terminal-Bench and OSWorld has highlighted Chinese labs' rapid gains in multi-step tool use and long-horizon workflows, keeping Alibaba, Moonshot, and DeepSeek in contention alongside established Western players. OpenAI's GPT-6 Astra and Agents API beta, Anthropic's Claude Opus 5 with expanded context and effort controls, and Google's Gemini updates plus MCP protocol for device integration have driven incremental improvements in demonstrated agent reliability, yet persistent gaps below human baselines on complex tasks sustain market uncertainty. With probabilities clustered tightly among top labs and elevated odds on "Other," traders appear focused on near-term releases, enterprise adoption metrics, and new benchmarks through November as the key swing factors in this closely contested ranking.
Ringkasan eksperimental yang dihasilkan AI dengan referensi data Polymarket. Ini bukan saran trading dan tidak berperan dalam bagaimana pasar ini diselesaikan. · DiperbaruiMicrosoft 46%
Anthropic 26%
SpaceXAI 26%
DeepSeek 26%

Microsoft
46%

Anthropic
26%

SpaceXAI
26%

DeepSeek
26%

Nvidia
26%

ByteDance
26%

Baidu
26%

Amazon
26%

OpenAI
25%

Meta
25%

Meituan
25%

Tencent
25%

Xiaomi
25%

MiniMax
25%

14%

Alibaba
14%

Mistral
14%

Z.ai
13%

Moonshot
13%
Microsoft 46%
Anthropic 26%
SpaceXAI 26%
DeepSeek 26%

Microsoft
46%

Anthropic
26%

SpaceXAI
26%

DeepSeek
26%

Nvidia
26%

ByteDance
26%

Baidu
26%

Amazon
26%

OpenAI
25%

Meta
25%

Meituan
25%

Tencent
25%

Xiaomi
25%

MiniMax
25%

14%

Alibaba
14%

Mistral
14%

Z.ai
13%

Moonshot
13%
Results from the "Rank" column under the "Agent Arena" Leaderboard tab at https://arena.ai/leaderboard/agent filtered for "Labs" will be used to resolve this market.
Models marked “AutoEval” at the applicable check time will not be considered, regardless of whether they display a rank or score.
AI companies will be ordered primarily by their Lab Rank at the market’s check time. If the results based on the lab ranking are ambiguous or unavailable, the relevant AI companies will be ordered according to their highest-ranking AI model in the leaderboard’s “Models” view. If two or more models are tied on rank, they will be ordered by which model is listed higher on the leaderboard. If a tie still remains, alphabetical order of AI lab/company names as listed in this market group will be used as a final tiebreaker (e.g., “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies third place under this ranking.
The resolution source for this market is the arena.ai Agent Arena Leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to "Other".
Pasar Dibuka: Sep 17, 2026, 8:04 PM ET
Resolver
0x69c47De9D...Results from the "Rank" column under the "Agent Arena" Leaderboard tab at https://arena.ai/leaderboard/agent filtered for "Labs" will be used to resolve this market.
Models marked “AutoEval” at the applicable check time will not be considered, regardless of whether they display a rank or score.
AI companies will be ordered primarily by their Lab Rank at the market’s check time. If the results based on the lab ranking are ambiguous or unavailable, the relevant AI companies will be ordered according to their highest-ranking AI model in the leaderboard’s “Models” view. If two or more models are tied on rank, they will be ordered by which model is listed higher on the leaderboard. If a tie still remains, alphabetical order of AI lab/company names as listed in this market group will be used as a final tiebreaker (e.g., “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies third place under this ranking.
The resolution source for this market is the arena.ai Agent Arena Leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to "Other".
Resolver
0x69c47De9D...Recent benchmark leadership by Shanghai Artificial Intelligence Laboratory's Atria Dawn Preview on verified agentic evals like Terminal-Bench and OSWorld has highlighted Chinese labs' rapid gains in multi-step tool use and long-horizon workflows, keeping Alibaba, Moonshot, and DeepSeek in contention alongside established Western players. OpenAI's GPT-6 Astra and Agents API beta, Anthropic's Claude Opus 5 with expanded context and effort controls, and Google's Gemini updates plus MCP protocol for device integration have driven incremental improvements in demonstrated agent reliability, yet persistent gaps below human baselines on complex tasks sustain market uncertainty. With probabilities clustered tightly among top labs and elevated odds on "Other," traders appear focused on near-term releases, enterprise adoption metrics, and new benchmarks through November as the key swing factors in this closely contested ranking.
Ringkasan eksperimental yang dihasilkan AI dengan referensi data Polymarket. Ini bukan saran trading dan tidak berperan dalam bagaimana pasar ini diselesaikan. · Diperbarui
Hati-hati dengan link eksternal.
Hati-hati dengan link eksternal.
Pertanyaan yang Sering Diajukan