How the contract works
Probability
How the price has moved
Analysis
Context
What moves the probability
Anthropic's release cadence
Anthropic's ability to keep a Claude model at the top of LiveBench's coding average depends on how often and how strongly it ships updates between now and 31 December 2026. A track record of leading coding-specific benchmarks is the main reason the market sits above even odds. Any gap in Anthropic's release schedule while a rival ships gives the No side an opening.
Competing lab releases
OpenAI, Google and xAI each have the scale to release a frontier coding-focused model before year-end, and any one of them topping LiveBench would flip this market toward No. This is the single largest source of downside risk to the Yes price.
Moonshot AI and open-weight competition
Moonshot AI's Kimi models represent a lower-cost entrant that has already shown it can compete on coding tasks, and its own market on this same LiveBench ranking suggests traders take that threat seriously. This pulls some probability away from Anthropic even if it does not directly lift OpenAI or Google.
LiveBench's rotating question sets
LiveBench periodically refreshes its coding problems to limit contamination, which means rankings can shift for reasons unrelated to a genuine capability gap. This adds a layer of measurement noise that keeps the market from converging quickly.
Market immaturity
With only 143 observations and roughly a day of trading history, the price itself is not yet a reliable signal. The 42%-to-99% range recorded so far reflects thin liquidity more than a settled consensus, and the price should be expected to keep moving as volume builds.
The case for
- Anthropic's Claude models have repeatedly ranked at or near the top of coding-specific benchmarks over the past two years, which is the strongest evidence supporting a Yes outcome.
- For Yes to resolve, Anthropic needs to either maintain its current position or regain the top LiveBench Coding Average spot specifically on 31 December 2026, when the market checks the ranking.
- Anthropic's commercial focus on developer and agentic coding tools gives it a direct incentive to prioritize coding performance in every model release between now and year-end.
The case against
- OpenAI, Google or xAI could release a frontier model before 31 December 2026 that overtakes Claude on LiveBench's coding average, any of which would resolve this market No.
- Moonshot AI's Kimi models have already demonstrated competitive coding performance, and a further release from that lab could take the top spot at lower cost than the US labs.
- The market's own history, a swing from 42% to as high as 99% within roughly a day of trading, shows that current pricing is unstable and could look very different by the time more volume and information accumulate.
