Will Claude be the top-ranked AI model on LM Arena at the end of 2026?
chance the market gives this event โ not your chance of being right
- Yes โ The event happens
- 62%
- No โ The event does not happen
- 38%
Trade this contract
In short
The market treats this as unlikely. The core reason is structural rather than about any single model release: LM Arena's text leaderboard ranks models by aggregated human preference votes in blind pairwise comparisons, and Anthropic has rarely led that particular scoreboard, while Google, OpenAI and xAI have all pushed frontier models to the top of it in recent cycles. A Claude release that lands late in the year and holds the number one line on the Remove Style Controls view through 31 December 2026 would move this quickly; anything short of first place, however narrow, settles the contract at nothing.
How the contract works
Probability
How the price has moved
Context
Analysis
What moves the probability
Timing of the next Claude flagship
A Claude release in the autumn that reaches the leaderboard with enough votes to stabilise is the single most direct path to Yes. A release in December risks not accumulating enough votes to overtake incumbents before the 31 December reading, and a release slipping into 2027 makes the question moot. This driver dominates every other consideration.
Google and OpenAI release cadence
Both labs have shipped frontier models at a pace that has repeatedly reclaimed the top Arena position within weeks of losing it. Each new Gemini or GPT flagship in the fourth quarter pushes this probability down, and does so quickly, because the leaderboard rewards whoever is newest at the top. This is the main force keeping the Claude line well below even money.
The Remove Style Controls view
Settlement reads the leaderboard without style adjustment, so verbosity and formatting are not corrected for. That view has tended to favour models producing longer, richly formatted answers, which historically has not been how Claude is tuned. It is a persistent structural drag of a few percentage points rather than a swing factor.
xAI, Alibaba and the long tail
Grok models have ranked at or near the top of the Arena, and Chinese labs including Alibaba and Moonshot have closed much of the gap on open-weight releases. Every additional credible contender at the top thins the probability available to Anthropic even if Claude improves in absolute terms. This matters at the margin, worth single digits.
Leaderboard methodology changes
LM Arena has revised its rating methodology and views before, including how style is handled and how new models are staged into the board. Any change to the Remove Style Controls view or to how ties are ordered would reshuffle the settlement reading without any model changing. Low probability, but a genuine source of uncertainty on a single-day snapshot.
The case for
- Anthropic has led this leaderboard before: Claude 3 Opus took the top spot in early 2024, the first non-OpenAI model to do so, which shows the outcome is achievable rather than theoretical.
- If Anthropic ships a flagship Claude in the September-to-November window, it would have several weeks of vote accumulation before the 31 December 2026 reading โ enough for a rating to stabilise at the top.
- The top of the LM Arena board has turned over repeatedly rather than settling with one lab, so incumbency at the number one row has proved fragile and short-lived.
- Settlement is a single-day snapshot, which means Anthropic does not need to lead for the year โ only on 31 December 2026.
The case against
- Anthropic's product emphasis on coding and enterprise agents is not what a general-audience preference vote rewards, and the Remove Style Controls view makes no correction for the verbose, heavily formatted answers that tend to win those votes.
- Google and OpenAI have each reclaimed the top position within weeks of losing it, so even a successful Claude launch can be displaced before the 31 December reading.
- With xAI, Alibaba, Moonshot and others clustered near the top, the probability mass at the number one row is split across more credible contenders than in 2024.
- The tracked price has moved down from its opening level rather than up, which indicates the market has not seen anything during the recorded period that strengthened the Anthropic case.
Trade this contract
- No external wallet needed
- gas covered
Venues (1)
- KalshiRecommendedYes62%0.62
- Volume (24h)
- US$14.1k
- Fee
- 1.63%
Probability
- Claude62%
- ChatGPT14%
- Gemini11%
- Grok6%
- Muse Spark3%
- Kimi3%
- Qwen1%
- Ernie1%
Resolution rules
The outcome is read from the LM Arena text leaderboard, Remove Style Controls view, as published on 31 December 2026. It resolves Yes only if a model developed by Anthropic in the Claude family occupies the number one position on that date, and No if the top-ranked model belongs to any other developer, including OpenAI, Google, xAI, Meta, Alibaba, Moonshot or Baidu. Ties are broken by the ordering shown on the leaderboard itself. All venues listed here settle from that same source and view, so price differences between contracts reflect different developers within the same series rather than different settlement sources.
Calculation methodology โLocal context
What to watch
Common questions
- What exactly settles this market?
- The LM Arena text leaderboard, viewed with style controls removed, as published on 31 December 2026. If a model from Anthropic's Claude family holds the number one position on that view on that date, the contract settles Yes. Any other developer at number one โ OpenAI, Google, xAI, Meta, Alibaba, Moonshot, Baidu or anyone else โ settles it No.
- Why does the spread between venues look so wide?
- Every contract listed settles from the same leaderboard, the same view and the same date, so there is no source disagreement to explain a 62.6-point gap. The aggregate covers several contracts in one leaderboard series, one per developer, so the high and low readings are the market's ranking of different labs rather than two opinions about Anthropic.
- What does a price of 0.20 actually mean?
- It means buyers and sellers currently agree the outcome happens about two times in ten. A contract settles at $1 if the outcome occurs and at nothing if it does not, so the price is a direct read of the market-implied probability. It is not a forecast from any institution โ it is the level at which trade is happening.
- What if the leaderboard changes or is unavailable on the resolution date?
- Settlement is defined as the ranking published on the resolution date in the Remove Style Controls view, and ties are broken by the ordering the leaderboard displays. If LM Arena revises its methodology during the year, the reading still comes from whatever that view shows on 31 December 2026. That is why methodology changes are a genuine risk factor and not a footnote.
- Has Claude ever been number one on LM Arena?
- Yes. Claude 3 Opus briefly took the top spot in early 2024, the first time a model from outside OpenAI led the board. That episode is the strongest single argument that the outcome is achievable, and it also illustrates how quickly the lead has changed hands since.
- Does Anthropic need to lead all year?
- No. The market reads a single day. A model that leads from September and slips in the final week settles at nothing, while a model that reaches number one in December and holds it through the 31st settles Yes.