Menu
Tech

Will Claude be the top-ranked AI model on LM Arena at the end of 2026?

Resolution: Updated:
62%

market consensus

chance the market gives this event โ€” not your chance of being right

Yes โ€” The event happens
62%
No โ€” The event does not happen
38%

Trade this contract

Open Kalshi siteYes 0.62
  • No external wallet needed
  • gas covered
Buy the opposite sideNo 0.38

In short

The market treats this as unlikely. The core reason is structural rather than about any single model release: LM Arena's text leaderboard ranks models by aggregated human preference votes in blind pairwise comparisons, and Anthropic has rarely led that particular scoreboard, while Google, OpenAI and xAI have all pushed frontier models to the top of it in recent cycles. A Claude release that lands late in the year and holds the number one line on the Remove Style Controls view through 31 December 2026 would move this quickly; anything short of first place, however narrow, settles the contract at nothing.

How the contract works

Each contract is a claim on one outcome and settles at $1 if that outcome happens, or at nothing if it does not. The price is simply the level at which buyers and sellers currently agree on the chance: a contract trading at 0.30 implies the market thinks the event happens about three times in ten. Here the outcome settled is whether a model in Anthropic's Claude family occupies the number one position on the LM Arena text leaderboard, Remove Style Controls view, as published on 31 December 2026. Ties are resolved by the ordering the leaderboard itself displays. Settlement is a single reading on that date, so a model that led for most of the year but slips in December pays nothing. A position does not have to be held to the end โ€” it can usually be sold beforehand at whatever the price is at that moment, which is how holders react to a major launch without waiting for the year to close.
What the market thinks happens
$100
Yes62%

The event happens

Costs now
$0.62
If you put in $100
$161
No38%

The event does not happen

Costs now
$0.38
If you put in $100
$263
0%25%50%75%100%12:2617:5823:3105:0310:3516:07
ConsensusKalshi

How the price has moved

The recorded history for this series starts at 37% on 29 July 2026 and now sits at a consensus of 20% for the Claude line, with 768 price observations logged and a series-wide range from 36% to 100%. The drift down from the opening level is the whole story: the market has revised the Anthropic case lower during the tracked window, while a different leg of the same series has at times printed at or near certainty, which is what happens when one lab is visibly sitting at the top of the board. The record is too short to attribute the decline to any one publicly reported trigger, and no single release date accounts for it. What can be said is that the fall came alongside a widening of the series into a clear favourite plus a low-teens second tier, rather than a tightening race.

Context

LM Arena โ€” formerly the LMSYS Chatbot Arena โ€” ranks large language models by human preference. Users send a prompt, receive two anonymous responses, and vote for the better one. Those votes are aggregated into a rating and published as a public leaderboard, which has become the most widely cited informal scoreboard in the industry, quoted in launch blog posts and in earnings-call talking points. The leaderboard offers several views; this market reads the text leaderboard with style controls removed, meaning the ranking is not adjusted for response length or formatting. Anthropic is one of three labs generally treated as at the frontier, alongside OpenAI and Google DeepMind. Its Claude family has built its reputation largely on coding, long-context work and enterprise deployment rather than on conversational preference contests. Claude 3 Opus did briefly take the top spot on the Arena in early 2024, the first time a model from outside OpenAI led the board โ€” evidence that the outcome is possible, not merely hypothetical. Since then the top of the leaderboard has turned over repeatedly, with Google and OpenAI trading the lead and xAI's Grok models pressing close behind. The question resolves on a single day. Whoever holds the number one row on 31 December 2026 decides it. That makes release timing as important as model quality: a lab that ships a flagship in November has its rating locked in by year-end, while one that ships in January gets nothing from it.

Analysis

The consensus across venues sits at 20%, which places Claude clearly behind the field but not out of it. That is roughly the probability the market attaches to Anthropic both shipping a model that outperforms whatever Google, OpenAI and xAI have on the board in December, and doing so on a metric โ€” aggregated blind human preference โ€” that has historically not been Anthropic's strongest suit. Anthropic's public emphasis has been on coding agents, tool use and enterprise reliability. Arena voters reward answers that read well to a general user, and with style controls removed the ranking makes no adjustment for length or formatting, which has historically favoured models that produce longer, more elaborately structured replies. The headline spread of 62.6 percentage points between the highest and lowest venue needs reading carefully. Every contract listed settles by the same source, the same leaderboard view and the same date, so there is no source disagreement to explain a gap that size. What that spread actually reflects is that the aggregate covers several contracts in one leaderboard series โ€” a set of one-per-developer questions, of which Claude is one line. The leg trading at 63% and the leg at 0% are not two opinions about Anthropic; they are the market's ranking of different labs. Read that way, the distribution says something sharper than the spread suggests: the market has one clear favourite well above even money, a second tier in the low teens, and a tail of developers priced near zero. Total recorded volume across the series is $7,771,901, with the individual legs ranging from about $441,000 to roughly $1.79 million. That is real depth for a question about a voluntary, crowd-voted leaderboard, and it means the pricing is not a thin quote. It also shows attention is spread across several labs rather than concentrated on one โ€” the four largest legs each carry over $1 million, which is what a genuinely contested race looks like in volume terms. The price history is short. The first recorded observation is dated 29 July 2026 at 37%, and 768 observations have been logged since the series began being tracked, with a recorded range from 36% all the way to 100%. The 100% reading belongs to the series aggregate rather than the Claude line โ€” a leaderboard leg can print at or near certainty when a specific model is sitting at the top with weeks to go. For the Claude question itself, the meaningful comparison is the opening 37% against the current consensus in the 20s: the market has become less convinced during the tracked period, not more. The decisive mechanic is timing. LM Arena ratings move when a new model is added and accumulates votes, which takes days to weeks. That gives the market a fairly clear checkpoint structure: any flagship released in the autumn has time to settle into a stable rating before the 31 December reading, while anything announced in the final fortnight of the year may not have enough votes to move the ranking. For this to resolve Yes, Anthropic needs a top-of-board model on the leaderboard and holding through the final week โ€” not merely a well-reviewed release.

What moves the probability

  • Timing of the next Claude flagship

    A Claude release in the autumn that reaches the leaderboard with enough votes to stabilise is the single most direct path to Yes. A release in December risks not accumulating enough votes to overtake incumbents before the 31 December reading, and a release slipping into 2027 makes the question moot. This driver dominates every other consideration.

  • Google and OpenAI release cadence

    Both labs have shipped frontier models at a pace that has repeatedly reclaimed the top Arena position within weeks of losing it. Each new Gemini or GPT flagship in the fourth quarter pushes this probability down, and does so quickly, because the leaderboard rewards whoever is newest at the top. This is the main force keeping the Claude line well below even money.

  • The Remove Style Controls view

    Settlement reads the leaderboard without style adjustment, so verbosity and formatting are not corrected for. That view has tended to favour models producing longer, richly formatted answers, which historically has not been how Claude is tuned. It is a persistent structural drag of a few percentage points rather than a swing factor.

  • xAI, Alibaba and the long tail

    Grok models have ranked at or near the top of the Arena, and Chinese labs including Alibaba and Moonshot have closed much of the gap on open-weight releases. Every additional credible contender at the top thins the probability available to Anthropic even if Claude improves in absolute terms. This matters at the margin, worth single digits.

  • Leaderboard methodology changes

    LM Arena has revised its rating methodology and views before, including how style is handled and how new models are staged into the board. Any change to the Remove Style Controls view or to how ties are ordered would reshuffle the settlement reading without any model changing. Low probability, but a genuine source of uncertainty on a single-day snapshot.

The case for

  • Anthropic has led this leaderboard before: Claude 3 Opus took the top spot in early 2024, the first non-OpenAI model to do so, which shows the outcome is achievable rather than theoretical.
  • If Anthropic ships a flagship Claude in the September-to-November window, it would have several weeks of vote accumulation before the 31 December 2026 reading โ€” enough for a rating to stabilise at the top.
  • The top of the LM Arena board has turned over repeatedly rather than settling with one lab, so incumbency at the number one row has proved fragile and short-lived.
  • Settlement is a single-day snapshot, which means Anthropic does not need to lead for the year โ€” only on 31 December 2026.

The case against

  • Anthropic's product emphasis on coding and enterprise agents is not what a general-audience preference vote rewards, and the Remove Style Controls view makes no correction for the verbose, heavily formatted answers that tend to win those votes.
  • Google and OpenAI have each reclaimed the top position within weeks of losing it, so even a successful Claude launch can be displaced before the 31 December reading.
  • With xAI, Alibaba, Moonshot and others clustered near the top, the probability mass at the number one row is split across more credible contenders than in 2024.
  • The tracked price has moved down from its opening level rather than up, which indicates the market has not seen anything during the recorded period that strengthened the Anthropic case.

Trade this contract

Venues (1)

Open Kalshi siteYes 0.62
  • No external wallet needed
  • gas covered

Venues (1)

Probability

  • Claude62%
  • ChatGPT14%
  • Gemini11%
  • Grok6%
  • Muse Spark3%
  • Kimi3%
  • Qwen1%
  • Ernie1%

Resolution rules

Determined by
LM Arena leaderboard (Remove Style Controls view) as of December 31, 2026
Resolution date

The outcome is read from the LM Arena text leaderboard, Remove Style Controls view, as published on 31 December 2026. It resolves Yes only if a model developed by Anthropic in the Claude family occupies the number one position on that date, and No if the top-ranked model belongs to any other developer, including OpenAI, Google, xAI, Meta, Alibaba, Moonshot or Baidu. Ties are broken by the ordering shown on the leaderboard itself. All venues listed here settle from that same source and view, so price differences between contracts reflect different developers within the same series rather than different settlement sources.

Calculation methodology โ†’

Local context

The Anthropicโ€“OpenAIโ€“Google race is the axis around which a large slice of US equity value now turns, and English-language readers hold that exposure whether or not they follow the leaderboard. Amazon and Google are both investors in Anthropic; Microsoft's position is with OpenAI. Which lab is seen to be at the frontier feeds directly into cloud contract wins, capital-expenditure guidance at the hyperscalers, and the valuation multiple applied to the US megacap complex that sits inside almost every index fund, UK workplace pension and Canadian or Australian superannuation portfolio. There is a second channel through policy and procurement. Leaderboard position is used as shorthand in Washington, Brussels and Westminster debates about which labs are ahead and how they should be regulated, and it feeds enterprise buying decisions โ€” including at the Indian IT services firms that build delivery practices around whichever model families their clients standardise on. A leaderboard reading is a narrow measure of a narrow thing, but it is one of the few public, comparable numbers in the argument, which is why a single row on a public website moves conversations far from it.

What to watch

Three things. First, any Anthropic flagship announcement between now and roughly mid-November 2026 โ€” that is the last window in which a new model can plausibly accumulate enough Arena votes to hold the number one row on 31 December. Second, Google and OpenAI fourth-quarter releases, which have historically reclaimed the top position within weeks and would compress the Claude line fast. Third, the leaderboard itself: LM Arena has revised methodology and views before, and because settlement reads only the Remove Style Controls text leaderboard as published on the resolution date, any change to that view or to tie ordering matters as much as a model launch. The final week of December is where the price and the outcome converge, since a lead in November is worth nothing if it is lost by New Year's Eve.

Common questions

What exactly settles this market?
The LM Arena text leaderboard, viewed with style controls removed, as published on 31 December 2026. If a model from Anthropic's Claude family holds the number one position on that view on that date, the contract settles Yes. Any other developer at number one โ€” OpenAI, Google, xAI, Meta, Alibaba, Moonshot, Baidu or anyone else โ€” settles it No.
Why does the spread between venues look so wide?
Every contract listed settles from the same leaderboard, the same view and the same date, so there is no source disagreement to explain a 62.6-point gap. The aggregate covers several contracts in one leaderboard series, one per developer, so the high and low readings are the market's ranking of different labs rather than two opinions about Anthropic.
What does a price of 0.20 actually mean?
It means buyers and sellers currently agree the outcome happens about two times in ten. A contract settles at $1 if the outcome occurs and at nothing if it does not, so the price is a direct read of the market-implied probability. It is not a forecast from any institution โ€” it is the level at which trade is happening.
What if the leaderboard changes or is unavailable on the resolution date?
Settlement is defined as the ranking published on the resolution date in the Remove Style Controls view, and ties are broken by the ordering the leaderboard displays. If LM Arena revises its methodology during the year, the reading still comes from whatever that view shows on 31 December 2026. That is why methodology changes are a genuine risk factor and not a footnote.
Has Claude ever been number one on LM Arena?
Yes. Claude 3 Opus briefly took the top spot in early 2024, the first time a model from outside OpenAI led the board. That episode is the strongest single argument that the outcome is achievable, and it also illustrates how quickly the lead has changed hands since.
Does Anthropic need to lead all year?
No. The market reads a single day. A model that leads from September and slips in the final week settles at nothing, while a model that reaches number one in December and holds it through the 31st settles Yes.

Related events

62%/ 38%
Yes / No