Moonshot AI’s July 2026 launch of Kimi K3, a 2.8-trillion-parameter mixture-of-experts model, has been the main catalyst lifting Kimi scores on Humanity’s Last Exam. The release delivered strong gains in tool-augmented reasoning and agentic workflows, placing K3 in the mid-50s percent range on recent leaderboards and ahead of most open-weight rivals. Earlier K2-series variants had already shown notable tool-use improvements, with scores rising from the low 30s to over 44 percent when search and code interpreters were enabled. Anthropic’s Claude Fable and Opus variants continue to lead overall HLE performance, but Moonshot’s rapid iteration and focus on long-context agent capabilities keep Kimi competitive. Traders are watching for any late-2026 model updates or inference optimizations that could further close the gap before year-end.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated$54,989 Vol.
45%+
84%
50%+
37%
55%+
19%
60%+
11%
65%+
8%
$54,989 Vol.
45%+
84%
50%+
37%
55%+
19%
60%+
11%
65%+
8%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Market Opened: Jul 23, 2026, 6:43 PM ET
Resolution Source
https://agi.safe.ai/Resolver
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Resolution Source
https://agi.safe.ai/Resolver
0x65070BE91...Moonshot AI’s July 2026 launch of Kimi K3, a 2.8-trillion-parameter mixture-of-experts model, has been the main catalyst lifting Kimi scores on Humanity’s Last Exam. The release delivered strong gains in tool-augmented reasoning and agentic workflows, placing K3 in the mid-50s percent range on recent leaderboards and ahead of most open-weight rivals. Earlier K2-series variants had already shown notable tool-use improvements, with scores rising from the low 30s to over 44 percent when search and code interpreters were enabled. Anthropic’s Claude Fable and Opus variants continue to lead overall HLE performance, but Moonshot’s rapid iteration and focus on long-context agent capabilities keep Kimi competitive. Traders are watching for any late-2026 model updates or inference optimizations that could further close the gap before year-end.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated



Beware of external links.
Beware of external links.
Frequently Asked Questions