Menu
Tech

Will Anthropic Have the Top AI Model on the Chatbot Arena Leaderboard by End of 2026?

Resolution: Updated:

In short

The market leans toward Anthropic holding the top spot but leaves real room for the opposite outcome. That lean reflects Anthropic's recent momentum in flagship model releases, offset by a long pattern of Google and OpenAI models trading the number-one position on this specific leaderboard. A new flagship release from any competitor before the December check would be the most likely thing to flip it.

Editorial illustration for: Will Anthropic Have the Top AI Model on the Chatbot Arena Leaderboard by End of 2026?

How the contract works

This contract settles at $1 if Anthropic holds the top rank on the Chatbot Arena text leaderboard, with style control off, at the specified check time on 31 December 2026, and at nothing otherwise. The price at any moment is simply the market's current estimate of that chance; a contract priced at 0.30, for example, would imply traders see roughly a three-in-ten chance of that outcome, not this specific market's price. Ties are broken first by unrounded Arena score and then alphabetically by company name, which matters if two labs finish statistically close. A position bought today can typically be sold before the 31 December check, at whatever price the market has moved to by then, rather than held to settlement.
What the market thinks happens
$100
Yes68%

The event happens

Costs now
$0.68
If you put in $100
$147
No32%

The event does not happen

Costs now
$0.32
If you put in $100
$313

Probability

History starts collecting once the event is tracked

How the price has moved

The contract currently sits at 67% on Polymarket, the only venue tracking this question, with $62,772 in volume behind that level. There is no second venue to compare against for a spread, and no day-over-day or week-over-week move is recorded in the available data, so the honest description is that the market has settled at a level treating Anthropic as clearly favored without a documented recent trigger for that stance. Given the thin volume relative to the roughly five months remaining until settlement, the current level should be read as an initial market view rather than one that has been heavily tested by new information.

Analysis

Context

Chatbot Arena, run by LMArena, ranks large language models by having anonymous users compare two model responses side by side and vote for the better one. The resulting Arena score, calculated without any style adjustment in the version this contract tracks, has become one of the most-cited public rankings of chatbot quality, even though critics note it can reward conversational polish as much as raw capability. Anthropic (Claude), OpenAI (GPT and o-series), Google DeepMind (Gemini), xAI (Grok), and Meta (Llama) all compete for the top slots, and the leaderboard has changed hands repeatedly since 2024 as each lab has shipped new flagship models. Anthropic's Claude models have generally sat in the top tier of the leaderboard through 2025 and into 2026, often ranked highly but not always first, with Google's Gemini series and OpenAI's GPT and o-series models frequently occupying the very top rank instead. The contract settles by checking who holds the top rank on 31 December 2026 at 12:00 PM ET, a single snapshot rather than a track record over the year.
The consensus figure across the one venue currently tracking this question, Polymarket, sits at 67%, backed by $62,772 in trading volume. That volume is modest for a contract with roughly five months left to run, which suggests a market that has taken a view but has not attracted heavy positioning either way. Because only one venue lists this contract, there is no cross-venue spread to check for disagreement between pools of traders, which is itself informative: a single-venue market with light volume can move more on a single new model release than a deeply traded one would. The 67% level implies the market sees Anthropic as clearly more likely than not to hold the top spot, but far from a settled matter, leaving roughly a one-in-three chance priced for a rival to be on top instead. Historically, the Arena's style-uncontrolled scoring has tended to reward longer, more polished-sounding answers, a dynamic that has sometimes worked against Claude's more concise default style even when Claude scores well on coding and reasoning benchmarks elsewhere. Google's Gemini and OpenAI's GPT and o-series models have repeatedly captured the top Arena rank at various points over 2024 and 2025, showing that leadership on this specific leaderboard has been genuinely contested rather than a one-lab story. What has likely shifted the market toward Anthropic is the pace of its recent flagship releases and apparent narrowing of the style gap, but the coming months carry real event risk: any one of the three main rivals shipping a strong new flagship model before the December snapshot could retake the top rank quickly, since Arena rankings can move within days of a major release entering voting rotation.

What moves the probability

  1. Style-uncontrolled scoring

    The leaderboard version used for settlement has style control off, which historically rewards longer or more polished-sounding answers. Claude's typically more concise default style has sometimes cost it ground against rivals on this exact metric, a modest but persistent drag on Anthropic's odds.

  2. Release timing before the check date

    Arena rankings can shift within days of a new flagship model entering the voting pool. Whichever lab ships its strongest model closest to 31 December 2026 gets a fresh boost right when it matters most for settlement.

  3. Competing flagship releases from Google and OpenAI

    Gemini and GPT or o-series models have repeatedly held the top Arena rank over 2024 and 2025. A major release from either lab in the second half of 2026 is the single biggest threat to an Anthropic finish.

  4. xAI as a third contender

    Grok's rapid release cadence adds another lab capable of briefly topping the leaderboard. Its presence widens the field beyond a two-way Anthropic-versus-incumbents contest.

  5. Tie-break rule favoring Anthropic alphabetically

    If Anthropic and a rival finish with statistically indistinguishable unrounded scores, the alphabetical tiebreak favors a company starting with A over Google, OpenAI, or xAI. This is a small structural edge that only matters in a near-exact tie.

The case for

  • Anthropic ships a flagship Claude update before the December snapshot that outperforms rivals in blind pairwise voting on the Arena's style-uncontrolled scoring.
  • No rival lab releases a stronger model in the weeks immediately before 31 December 2026, when a late release could otherwise flip the top rank quickly.
  • In a near-exact tie on unrounded Arena score, the alphabetical tiebreak rule works in Anthropic's favor.
  • The market's current 67% level already reflects a view that Anthropic's recent releases have closed much of the historical gap to Google and OpenAI on this specific leaderboard.

The case against

  • Google's Gemini and OpenAI's GPT or o-series models have repeatedly held the top Arena rank through 2024 and 2025, showing the top spot is genuinely contested rather than settled.
  • The style-uncontrolled scoring format has historically been less favorable to Claude's typically concise responses than to more verbose competitor outputs.
  • xAI's Grok adds a third serious contender capable of topping the leaderboard on short notice after a strong release.
  • Trading volume of $62,772 on a single venue is thin, meaning the current level may not have absorbed much information about labs' late-2026 release plans.

What to watch

Watch for flagship model releases from Anthropic, OpenAI, Google DeepMind, and xAI between now and December 2026, since a strong new release from any of them can move the Arena rank within days of entering the voting pool. The exact check happens on 31 December 2026 at 12:00 PM ET, so releases timed for the fourth quarter of 2026 carry outsized weight. Also watch whether LMArena is functioning normally at the check time, since the rules delay resolution rather than default to either outcome if the leaderboard is unavailable.

Trade this contract

Venues (1)

More about this event

Venues (1)

Probability

  • Will Anthropic have the best AI model at the end of December 2026?68%
  • Will OpenAI have the best AI model at the end of December 2026?11%
  • Will xAI have the best AI model at the end of December 2026?3%
  • Will Meta have the best AI model at the end of December 2026?2%
  • Will Moonshot have the best AI model at the end of December 2026?2%
  • Will DeepSeek have the best AI model at the end of December 2026?1%
  • Will Z.ai have the best AI model at the end of December 2026?1%
  • Will ByteDance have the best AI model at the end of December 2026?0%
  • Will Amazon have the best AI model at the end of December 2026?0%
  • Will Mistral have the best AI model at the end of December 2026?0%
  • Will Microsoft have the best AI model at the end of December 2026?0%

Resolution rules

Determined by
https://lmarena.ai/leaderboard/text
Resolution date

This settles using the Chatbot Arena LLM Leaderboard on lmarena.ai, specifically the Rank section of the text leaderboard with style control switched off. The check happens on 31 December 2026 at 12:00 PM ET; whichever model sits at the top at that moment determines the outcome, with ties resolved first by unrounded Arena score and then alphabetically by company name. If the leaderboard cannot be accessed at that exact time, resolution is delayed until it comes back online rather than defaulting to either side.

Calculation methodology โ†’

Local context

For English-language tech and finance readers, this contract is a proxy for a story they already follow closely: the competitive standing of Anthropic against OpenAI, Google, and xAI, three companies whose valuations and product roadmaps feature heavily in US tech coverage and, by extension, in Nasdaq-listed technology stocks. A shift in perceived AI leadership among these labs has repeatedly moved headlines and, at times, equity prices for the public companies tied to them, even when the underlying leaderboard question is narrow and technical.

Common questions

What exactly settles this contract, and when
The Chatbot Arena text leaderboard, style control off, checked on 31 December 2026 at 12:00 PM ET via LMArena's public leaderboard page. If Anthropic holds the top rank at that moment, the contract settles at $1; otherwise it settles at nothing.
What does the current price actually mean
The price is the market's live estimate of the probability Anthropic finishes on top, not a guarantee. A price near 0.67 implies traders collectively see this as roughly a two-in-three chance, but that figure moves as new information, such as model releases, arrives.
What happens if the leaderboard is down or unclear at the check time
The rules state that if the leaderboard is unavailable at the specified check time, resolution is simply delayed until it becomes available again. There is no default outcome triggered by a delay.
How are ties between labs handled
Ties are broken first by the unrounded Arena score, which can differ even when rounded scores look identical, and then alphabetically by company name if scores are truly indistinguishable.
Why does Anthropic sometimes rank below Google or OpenAI on this leaderboard despite strong coding benchmarks elsewhere
The Arena leaderboard used here has style control turned off, which tends to reward longer or more polished-sounding responses. Claude's typically more concise style has, at times, scored lower on this specific measure than on benchmarks that judge substance over presentation.
Is this the only leaderboard this contract cares about
Yes, the contract resolves specifically on the LMArena text leaderboard's Rank section with style control off. Other benchmarks or leaderboards, however Anthropic performs on them, have no bearing on settlement.

Related prediction events

Tokenized stocks

Market-implied probabilities that provide context for this assetโ€™s catalysts.

Related events