Menu
Tech

Will Alibaba Have the Top-Ranked Chinese AI Model on LMArena at the End of September 2026?

Resolution: Updated:

In short

The market treats this as highly likely but not fully settled. Alibaba's Qwen models have held the top Chinese slot on LMArena for much of 2026, and the main risk is a fresh flagship release from a rival lab landing before the 30 September checkpoint. A new checkpoint from Moonshot, DeepSeek or Zhipu that gathers enough human votes before the deadline is the clearest way this could flip.

Editorial illustration for: Will Alibaba Have the Top-Ranked Chinese AI Model on LMArena at the End of September 2026?

How the contract works

A contract on this market settles at $1 if Alibaba's model holds the top rank among qualifying Chinese AI companies on the LMArena Text Arena Overall leaderboard (style control off) when checked on 30 September 2026 at 12:00 PM ET, and at $0 otherwise. The price at any given time simply reflects what buyers and sellers currently think the chance of that outcome is; a contract trading at 0.30, for example, would imply the market sees roughly a three-in-ten chance, though that is a hypothetical, not this market's actual level. Ties are broken first by the underlying numerical Arena score and then alphabetically by company name, so a genuine dead heat is unlikely to produce an ambiguous result. A position in this market can generally be sold before the 30 September checkpoint at whatever price the market has moved to by then.
What the market thinks happens
$100
Yes87%

The event happens

Costs now
$0.87
If you put in $100
$115
No13%

The event does not happen

Costs now
$0.13
If you put in $100
$769

Probability

History starts collecting once the event is tracked

How the price has moved

The available data shows a consensus of 87% on the only tracked venue, Polymarket, with total volume of $129,850. No breakdown of the move over the last day or week, or the range since the market opened, has been reported, so it is not possible to say whether this level represents a recent shift or a level that has held steady since trading began. What can be said plainly is that a level in the high-80s combined with modest single-venue volume points to a market that sees the outcome as likely without treating it as a formality, and without enough liquidity yet to call it fully settled.

Analysis

Context

LMArena, formerly known as Chatbot Arena, ranks AI language models by having people compare pairs of anonymous model responses and vote for the better one. The resulting Text Arena Overall leaderboard is widely cited in AI industry coverage as a real-world proxy for model quality, distinct from academic benchmarks that labs can optimize for directly. Because voting is continuous and crowdsourced, rankings can shift gradually as new votes accumulate, not just when a company ships a new model. The Chinese AI sector has multiple labs releasing competing large language models on an irregular schedule: Alibaba's Qwen series, Moonshot's Kimi, DeepSeek's models, Zhipu/Z.ai's GLM line, plus entries from Tencent, ByteDance, Baidu, Xiaomi, MiniMax, Meituan and StepFun. Alibaba's Qwen has occupied the top Chinese position on LMArena's overall leaderboard for extended stretches through 2025 and into 2026, making it the incumbent this market is effectively asking whether to bet against. This specific market checks a single snapshot: the leaderboard's state on 30 September 2026 at 12:00 PM ET, as published on arena.ai. It resolves Yes only if Alibaba's model outranks every other qualifying Chinese company's model at that exact moment.
The consensus figure of 87% comes from a single tracked venue, Polymarket, with total volume of $129,850. That volume is modest for a technology-sector market, which means the price reflects a relatively small pool of positions rather than broad, deep price discovery; it should be read as a directional signal, not a precise statistical forecast. With only one venue quoting this market, there is no cross-venue spread to check for disagreement, and no day-by-day or week-by-week move has been reported, so the honest read is that the current level is the best available snapshot rather than the endpoint of a documented trend. The structural reason the market sits close to certainty rather than at a coin flip is incumbency. Alibaba's Qwen family has repeatedly held the top Chinese-model rank on LMArena's overall leaderboard through 2025 and into 2026, which means the Yes outcome does not require Alibaba to do anything new before 30 September 2026 โ€” it only requires the status quo to hold for another ten days from the time this market was assessed. That is a fundamentally easier bar than requiring a challenger to actively overtake an incumbent. The remaining uncertainty, priced at roughly 13%, comes from the release cadence of Chinese rivals. Moonshot, DeepSeek, Zhipu/Z.ai, ByteDance, Tencent and Baidu all ship model updates on schedules that are not publicly predictable months in advance, and LMArena adds new checkpoints to its leaderboard as labs submit them. A new flagship model from any of these companies that draws enough human preference votes before the 30 September, 12:00 PM ET cutoff could displace Qwen, even temporarily. Because the leaderboard uses style control off, raw human preference โ€” including sensitivity to response length and formatting โ€” determines rank, which adds some day-to-day noise around any close contest. The settlement rules also narrow the field in Alibaba's favor by limiting eligible competitors to companies the market defines as primarily Chinese, and by using a strict score-then-alphabetical tie-break that removes most scope for an ambiguous result. Together, incumbency, a defined and narrow competitor list, and the short remaining window to the checkpoint explain why the market prices this outcome as likely rather than uncertain, while still leaving room for a late model release to matter.

What moves the probability

  1. Qwen's leaderboard incumbency

    Alibaba's Qwen models have held the top Chinese-company rank on LMArena's overall leaderboard for much of 2026, so the Yes outcome mainly requires that position to persist rather than requiring new progress. Incumbency lowers the bar considerably compared with a market asking whether a challenger will overtake the leader.

  2. Rival release cadence

    Moonshot's Kimi, DeepSeek, and Zhipu/Z.ai's GLM line, plus ByteDance, Tencent and Baidu, ship new flagship checkpoints on unpredictable schedules. Any one of them landing a strong new model before 30 September 2026 and drawing enough votes could flip the top Chinese slot away from Qwen.

  3. Crowdsourced voting dynamics

    LMArena's ranking updates continuously as human votes accumulate, with style control off, meaning raw preference including response length can shift rankings even without a new model release. This adds genuine day-to-day noise close to the 30 September checkpoint.

  4. Narrow, defined competitor list

    The settlement rules only count companies deemed primarily Chinese, explicitly naming Alibaba, Moonshot, DeepSeek, Z.ai, Tencent, ByteDance, Baidu, Xiaomi, MiniMax, Meituan and StepFun. This removes ambiguity about which entrants count and keeps the field of realistic challengers relatively small.

  5. Tie-break mechanics

    Ties are resolved first by the underlying Arena score, then alphabetically by company name, which sharply reduces the chance of an ambiguous or contested result at the 30 September, 12:00 PM ET check.

The case for

  • Alibaba's Qwen has held the top Chinese ranking on LMArena for much of 2026, so the Yes outcome only requires that position to hold through 30 September 2026.
  • No confirmed rival flagship release has been reported that is positioned to overtake Qwen before the 12:00 PM ET checkpoint on 30 September 2026.
  • Alibaba continues to iterate on the Qwen series, and continuous model refreshes reduce the risk that a rival's static advantage accumulates unopposed.
  • The settlement rules' tie-break process, based on underlying Arena score then alphabetical order, favors a clean resolution rather than a contested near-tie.

The case against

  • Moonshot, DeepSeek, or Zhipu/Z.ai could release a new flagship checkpoint before 30 September 2026 that draws enough human preference votes to outrank Qwen on LMArena.
  • LMArena rankings shift as votes accumulate even without a new release, so Qwen's position could erode purely from evolving voter sentiment in the days before the checkpoint.
  • With style control off, formatting and response-length effects can influence rankings in ways that are difficult to predict for any single company, including Alibaba.
  • The market's total volume of $129,850 on a single venue is relatively thin for a technology question, meaning the 87% consensus reflects a narrower set of positions than a deeper, multi-venue market would produce.

What to watch

The key date is 30 September 2026 at 12:00 PM ET, when the LMArena Text Arena Overall leaderboard (style control off) is checked for this market's resolution, with settlement following on 1 October 2026. Between now and then, any public release of a new flagship model from Moonshot, DeepSeek, Zhipu/Z.ai, ByteDance, Tencent, Baidu, Xiaomi, MiniMax, Meituan or StepFun is the main event that could move this market, since LMArena adds new checkpoints as labs submit them and rankings can shift once enough human votes accumulate. Any update to Alibaba's own Qwen lineup in the same window would work in the opposite direction.

Trade this contract

Venues (1)

More about this event

Venues (1)

Probability

  • Will Alibaba have the best Chinese AI model at the end of September 2026?87%
  • Will Moonshot have the best Chinese AI model at the end of September 2026?11%
  • Will Z.ai have the best Chinese AI model at the end of September 2026?2%
  • Will DeepSeek have the best Chinese AI model at the end of September 2026?0%
  • Will ByteDance have the best Chinese AI model at the end of September 2026?0%
  • Will Baidu have the best Chinese AI model at the end of September 2026?0%
  • Will Tencent have the best Chinese AI model at the end of September 2026?0%

Resolution rules

Determined by
https://arena.ai/leaderboard/text/overall-no-style-control
Resolution date

This market is determined by the LMArena Text Arena Overall leaderboard with style control off, published at arena.ai. It resolves Yes if, among models from the listed Chinese AI companies, Alibaba's model holds the highest rank when the leaderboard is checked on 30 September 2026 at 12:00 PM ET. Ties are broken first by underlying Arena score, then alphabetically by company name, and companies not primarily operating within the Chinese technology ecosystem are excluded regardless of where they compete.

Calculation methodology โ†’

Local context

English-language technology and AI coverage treats LMArena as one of the more closely watched public gauges of the US-China AI competition, alongside benchmark suites like MMLU or coding leaderboards, because it reflects human preference rather than a lab's own reported scores. Which Chinese company leads that leaderboard feeds directly into commentary read by this audience about which firm โ€” Alibaba, Moonshot, DeepSeek or others โ€” is setting the pace in China's AI sector, a storyline that in turn shapes how Alibaba's business, including its cloud and AI division, is discussed in markets these readers already follow.

Common questions

What exactly settles this market, and when?
The market resolves based on the LMArena Text Arena Overall leaderboard with style control off, checked on 30 September 2026 at 12:00 PM ET, with settlement following on 1 October 2026. The rank of Alibaba's model relative to other qualifying Chinese-company models at that exact moment determines the outcome.
What does the market price actually mean?
The price is the market's current estimate of the probability that Alibaba's model holds the top Chinese rank at the checkpoint, based on what buyers and sellers on Polymarket are currently willing to trade at. It is not a guarantee, and it can move between now and the 30 September checkpoint as new information, such as a model release, emerges.
Which companies count as 'Chinese AI companies' for this market?
The settlement rules name Alibaba, Moonshot, DeepSeek, Z.ai, Tencent, ByteDance, Baidu, Xiaomi, MiniMax, Meituan and StepFun as qualifying, and specify that companies not headquartered in and not primarily operating within the Chinese technology ecosystem do not count, even if they compete.
What happens if the leaderboard shows a tie or an unclear result?
Ties are broken first by the underlying numerical Arena score, and if that is still tied, alphabetically by company name. This makes a genuinely unresolved outcome unlikely under the stated rules.
Why does a crowdsourced leaderboard like LMArena matter for judging AI models?
LMArena ranks models by aggregating anonymous human preference votes between pairs of model outputs, which is widely treated in AI industry reporting as a check against benchmarks that labs can optimize for directly. It has limitations, including sensitivity to formatting and response length, which is part of why this market's rules specify style control off explicitly.
Has Alibaba's Qwen usually led the Chinese pack on LMArena?
Qwen has held the top Chinese-company position on LMArena's overall leaderboard for extended periods through 2025 and into 2026, which is the main reason the market prices continuation of that position as the more likely outcome.

Related events