This market will resolve according to the company that owns the model that has the highest rank based on the Agent Arena Leaderboard (https://arena.ai/leaderboard/agent) when the table under "Agent Arena" filtered for "Models" is checked on September 30, 2026, 12:00 PM ET. Results from the "Rank" column under the "Agent Arena" Leaderboard tab at https://arena.ai/leaderboard/agent filtered for "Models" will be used to resolve this market. Models will be ordered primarily by their leaderboard rank at the market’s check time. If two or more models are tied on rank, they will be ordered by which model is listed higher on the leaderboard. If a tie still remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies first place under this ranking. The resolution source for this market is the Agent Arena Leaderboard found at https://arena.ai/leaderboard/agent. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve based on another resolution source.
Anthropic’s dominant 91% market-implied odds reflect its Claude models’ consistent leadership on agentic benchmarks, including top scores in long-horizon tool use, coding workflows, and autonomous task completion as of early September 2026. Recent releases such as Claude Fable 5.1 and Mythos 5.1, combined with Managed Agents infrastructure and the Model Hardware Standard for physical-device control, have strengthened its position in production-grade agent deployments across enterprise and research settings. OpenAI, Google, and others trail in current leaderboards, though a major capability leap or benchmark shift from competitors before month-end could narrow the gap. Traders view the current positioning as skin-in-the-game consensus on demonstrated reliability rather than guaranteed future performance.