This market will resolve according to the model that has the highest arena rank based on the arena.ai Text Arena (Overall) when the table under the "Leaderboard" tab is checked on the specified date, 12:00 PM ET. Results from the "Rank" column under the "Text Arena | Overall" Leaderboard tab at https://arena.ai/leaderboard/text/overall-no-style-control with style control off (Adjustments: None) and filtered for "Models" will be used to resolve this market. Note: Models marked “AutoEval” at the applicable check time will not be considered, regardless of whether they display a rank or score. No new model will be added to this market after market creation. Any model not explicitly listed in this market will be encompassed under the "Other" option. Models will be ordered primarily by their leaderboard rank at the market’s check time. If two or more models are tied on rank, they will be ordered by their Arena score, including any underlying, unrounded, granular values reflected in the data below the leaderboard. If a tie still remains, alphabetical order of model names as listed in this market group (full string, including suffixes such as “-high” and “-max”) will be used as a final tiebreaker (e.g., if two models remain tied, “claude-opus-5-high” would be ranked ahead of “claude-opus-5-max”). This market will resolve to the model that comes first according to this order. The resolution source for this market is the arena.ai Text Arena (Overall). If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to "Other".
Anthropic’s September 1 release of Claude Fable 5.1, which quickly topped independent leaderboards with the highest composite scores for agentic coding, reasoning, and long-context tasks, drives the overwhelming 96%+ market-implied probability that none of the listed Claude Opus variants will rank as the single best model on September 7. Recent GPT-6 Astra and Gemini 3.8 Flash launches further compress the gap at the frontier, while Opus 5 and earlier Opus 4.x releases trail on updated benchmarks such as Terminal-Bench and Frontier-Bench. Trader consensus treats these specific high/max-effort Opus variants as unlikely to overtake the newest flagships absent a major undisclosed capability jump or resolution-criteria shift before the snapshot date.