How the contract works
Probability
How the price has moved
Analysis
Context
What moves the probability
Timing of the next Claude flagship
A Claude release in the autumn that reaches the leaderboard with enough votes to stabilise is the single most direct path to Yes. A release in December risks not accumulating enough votes to overtake incumbents before the 31 December reading, and a release slipping into 2027 makes the question moot. This driver dominates every other consideration.
Google and OpenAI release cadence
Both labs have shipped frontier models at a pace that has repeatedly reclaimed the top Arena position within weeks of losing it. Each new Gemini or GPT flagship in the fourth quarter pushes this probability down, and does so quickly, because the leaderboard rewards whoever is newest at the top. This is the main force keeping the Claude line well below even money.
The Remove Style Controls view
Settlement reads the leaderboard without style adjustment, so verbosity and formatting are not corrected for. That view has tended to favour models producing longer, richly formatted answers, which historically has not been how Claude is tuned. It is a persistent structural drag of a few percentage points rather than a swing factor.
xAI, Alibaba and the long tail
Grok models have ranked at or near the top of the Arena, and Chinese labs including Alibaba and Moonshot have closed much of the gap on open-weight releases. Every additional credible contender at the top thins the probability available to Anthropic even if Claude improves in absolute terms. This matters at the margin, worth single digits.
Leaderboard methodology changes
LM Arena has revised its rating methodology and views before, including how style is handled and how new models are staged into the board. Any change to the Remove Style Controls view or to how ties are ordered would reshuffle the settlement reading without any model changing. Low probability, but a genuine source of uncertainty on a single-day snapshot.
The case for
- Anthropic has led this leaderboard before: Claude 3 Opus took the top spot in early 2024, the first non-OpenAI model to do so, which shows the outcome is achievable rather than theoretical.
- If Anthropic ships a flagship Claude in the September-to-November window, it would have several weeks of vote accumulation before the 31 December 2026 reading โ enough for a rating to stabilise at the top.
- The top of the LM Arena board has turned over repeatedly rather than settling with one lab, so incumbency at the number one row has proved fragile and short-lived.
- Settlement is a single-day snapshot, which means Anthropic does not need to lead for the year โ only on 31 December 2026.
The case against
- Anthropic's product emphasis on coding and enterprise agents is not what a general-audience preference vote rewards, and the Remove Style Controls view makes no correction for the verbose, heavily formatted answers that tend to win those votes.
- Google and OpenAI have each reclaimed the top position within weeks of losing it, so even a successful Claude launch can be displaced before the 31 December reading.
- With xAI, Alibaba, Moonshot and others clustered near the top, the probability mass at the number one row is split across more credible contenders than in 2024.
- The tracked price has moved down from its opening level rather than up, which indicates the market has not seen anything during the recorded period that strengthened the Anthropic case.
