Recent September 2026 model releases have fragmented the AI agent landscape, with no single leader ahead of the November deadline. Anthropic’s Claude Fable 5.1 and Mythos 5.1 lead long-horizon autonomy and Terminal-Bench scores, while OpenAI’s GPT-6 Astra excels on ARC-AGI, ExploitBench, and computer-use tasks; its new Agents API further boosts practical deployment via managed multi-agent workflows and sandboxes. Open-weight contenders from Z.ai (GLM), DeepSeek, and Alibaba (Qwen) post competitive agentic benchmarks at lower cost, sustaining trader parity. Safety coordination among OpenAI, Anthropic, and Google adds regulatory uncertainty, keeping implied probabilities closely matched across labs.
基于Polymarket数据的AI实验性摘要。这不是交易建议,也不影响该市场的结算方式。 · 更新于Nvidia 51%
SpaceXAI 47%
DeepSeek 47%
Moonshot 46%

Nvidia
51%

SpaceXAI
47%

DeepSeek
47%

Moonshot
46%

百度
46%

美团
46%

微软
46%

Anthropic
61%

Mistral
41%

阿里巴巴
28%

OpenAI
26%

Z.ai
26%

MiniMax
25%

谷歌
23%

Meta
14%

亚马逊
14%

字节跳动
14%

小米
25%

腾讯
8%
Nvidia 51%
SpaceXAI 47%
DeepSeek 47%
Moonshot 46%

Nvidia
51%

SpaceXAI
47%

DeepSeek
47%

Moonshot
46%

百度
46%

美团
46%

微软
46%

Anthropic
61%

Mistral
41%

阿里巴巴
28%

OpenAI
26%

Z.ai
26%

MiniMax
25%

谷歌
23%

Meta
14%

亚马逊
14%

字节跳动
14%

小米
25%

腾讯
8%
Results from the "Rank" column under the "Agent Arena" Leaderboard tab at https://arena.ai/leaderboard/agent filtered for "Models" will be used to resolve this market.
Note: Models marked “AutoEval” at the applicable check time will not be considered, regardless of whether they display a rank or score.
Models will be ordered primarily by their leaderboard rank at the market’s check time. If two or more models are tied on rank, they will be ordered by which model is listed higher on the leaderboard. If a tie still remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies first place under this ranking.
The resolution source for this market is the Agent Arena Leaderboard found at https://arena.ai/leaderboard/agent. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve based on another resolution source.
市场开放时间: Sep 17, 2026, 8:02 PM ET
Results from the "Rank" column under the "Agent Arena" Leaderboard tab at https://arena.ai/leaderboard/agent filtered for "Models" will be used to resolve this market.
Note: Models marked “AutoEval” at the applicable check time will not be considered, regardless of whether they display a rank or score.
Models will be ordered primarily by their leaderboard rank at the market’s check time. If two or more models are tied on rank, they will be ordered by which model is listed higher on the leaderboard. If a tie still remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies first place under this ranking.
The resolution source for this market is the Agent Arena Leaderboard found at https://arena.ai/leaderboard/agent. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve based on another resolution source.
Recent September 2026 model releases have fragmented the AI agent landscape, with no single leader ahead of the November deadline. Anthropic’s Claude Fable 5.1 and Mythos 5.1 lead long-horizon autonomy and Terminal-Bench scores, while OpenAI’s GPT-6 Astra excels on ARC-AGI, ExploitBench, and computer-use tasks; its new Agents API further boosts practical deployment via managed multi-agent workflows and sandboxes. Open-weight contenders from Z.ai (GLM), DeepSeek, and Alibaba (Qwen) post competitive agentic benchmarks at lower cost, sustaining trader parity. Safety coordination among OpenAI, Anthropic, and Google adds regulatory uncertainty, keeping implied probabilities closely matched across labs.
基于Polymarket数据的AI实验性摘要。这不是交易建议,也不影响该市场的结算方式。 · 更新于
警惕外部链接哦。
警惕外部链接哦。
常见问题