Back to Blog
News

Qwen 3.8 Max: 2.4 Trillion Parameters, Benchmarks, and Pricing (2026)

Unifie TeamJuly 20, 20269 min read

Qwen 3.8 Max explained. Everything Alibaba has confirmed about its 2.4 trillion parameter model, what remains unverified, how it compares to Kimi K3 and Fable 5, and what the well-documented Qwen 3.7-Max tells us about what to expect.

Qwen 3.8 Max: 2.4 Trillion Parameters, Benchmarks, and Pricing (2026)

On July 19, 2026, during the World AI Conference in Shanghai, Alibaba's Qwen team previewed its next flagship model, called it one of the most powerful systems available today, and said it trails only Claude Fable 5. Then it published almost nothing to back that up. This is the complete picture of what's confirmed, what's missing, and why the timing of this announcement matters as much as the model itself.

What Alibaba Actually Said

Qwen 3.8 Max Preview is a 2.4 trillion parameter multimodal model, positioned by Alibaba as the successor to Qwen 3.7-Max. The company described it as "second only to Fable 5" among the systems it internally benchmarked, and the preview is live today through Alibaba's Token Plan endpoint, the Qoder and QoderWork coding platforms, and a public chat interface. Alibaba says full open weights will follow "soon," without naming a date or license.

That's it. That's the entire confirmed announcement.

What's Missing, and Why That Matters for a 2.4 Trillion Parameter Model

This is the part worth sitting with before treating any comparison involving Qwen 3.8 Max as settled. As of July 19, there was no Qwen 3.8 technical report, no published benchmark table, no model card, no Artificial Analysis entry, no OpenRouter listing, no Hugging Face checkpoint, no standard per-token pricing, and no announced weight-release date.

More specifically, Alibaba has not disclosed:

  • The active parameter count. Qwen 3.8 Max is described as using a Mixture-of-Experts design, following the pattern of earlier Qwen releases, but total parameter count tells you almost nothing about running cost or inference behavior without knowing how many of those 2.4 trillion parameters actually activate per token.
  • The context window. Community integrations are citing 1 million tokens, but Alibaba hasn't published an official spec sheet confirming that number.
  • Any independently verifiable benchmark score. No SWE-Bench, LMSYS Arena, GPQA, AIME, or agent-eval results exist for this specific model as of publication.
  • Standard per-token pricing. Access currently runs through subscription tiers (Lite, Standard, Pro) via Alibaba's Token Plan rather than transparent per-million-token rates, and the Token Plan's Credits conversion hasn't been published in a way that allows a direct dollar comparison against competing models.
  • Whether the specific 2.4T checkpoint being previewed today is even the version that eventually ships open-weight, assuming that happens at all.

This gap matters more than it would for a smaller or less consequential release. A 2.4 trillion parameter claim, without architecture details, is a number that sounds impressive and tells you almost nothing verifiable about how the model actually performs or what it costs to run.

The Timing Is Arguably the Real Story

Qwen 3.8 Max's preview landed exactly two days after Moonshot AI, a startup in which Alibaba holds a 36 percent stake, released Kimi K3, a 2.8 trillion parameter open-weight model whose strong benchmark showing against Anthropic and OpenAI's flagship models rattled global technology stocks. That sequencing is hard to read as coincidental. Alibaba has an established cadence of shipping a new Qwen Max tier roughly every four to six weeks through 2026, a July release following a May Qwen 3.7 launch fits that pattern on schedule, but the specific choice to preview a headline parameter count and a direct claim against Fable 5 within 48 hours of K3's launch reads as a competitive response to Moonshot's momentum, timed for maximum attention at a major industry conference, rather than a fully baked release Alibaba was ready to fully document.

Worth noting too: shipping closed, undocumented previews before publishing full specs isn't new for Alibaba specifically. Both Qwen 3.7-Max and the earlier Qwen 3.6-Max-Preview also launched behind Alibaba Cloud Model Studio without immediate open weights, so a promise of open weights "soon" for 3.8 is plausible given Alibaba's separate open releases like Qwen3-Coder-480B under Apache 2.0, but it isn't guaranteed, and it would represent a break from the pattern its two immediate predecessors followed.

What Qwen 3.7-Max Tells Us About What to Expect

Since Qwen 3.8 Max doesn't have its own published benchmark table yet, the most useful available signal is what its immediate predecessor actually delivered, since Qwen has historically published detailed launch documentation for each flagship tier.

Qwen 3.7-Max, released roughly two months earlier, was built specifically around long-horizon agent workflows rather than general chat, a framing Qwen itself called the "Agent Frontier." Its most notable published results:

Benchmark Qwen 3.7-Max Notable comparison
GPQA Diamond 92.4% Ahead of Opus 4.6 Max (91.3%), behind GPT-5.5 (93.6%)
SWE-Bench Verified 80.4% Narrowly behind Opus 4.6 Max (80.8%) and DS-V4-Pro Max (80.6%)
Terminal-Bench 2.0 69.7%
HMMT 2026 Feb (competition math) 97.1% Highest in its comparison table
Humanity's Last Exam 41.4 Ahead of Opus 4.6 Max (40.0)
MCP-Mark (agent tool use) 60.8 Ahead of GLM-5.1 (57.5) and Opus 4.6 (56.7)
Kernel Bench L3 (autonomous optimization) 1.98x speedup, 96% win rate Behind Opus 4.6 Max (2.63x/98%), well ahead of K2.6 Thinking and DS-V4-Pro Max

Qwen 3.7-Max also demonstrated genuinely unusual long-horizon behavior: during an 86-hour reinforcement learning training run, the model autonomously flagged 1,618 reward hacking cases and added 13 new heuristic rules to its own training loop without human intervention, and in a separate 35-hour autonomous kernel optimization run across 432 evaluations and 1,158 tool calls, it sustained meaningful progress well past the 30-hour mark, a point where most agent models stop improving. Whether that self-monitoring and long-horizon persistence carries forward into 3.8's larger architecture is a reasonable expectation given Alibaba's stated research direction, but it isn't confirmed for the new model specifically.

The consistent caveat across every independent writeup of Qwen 3.7-Max's numbers: these are vendor-published benchmarks from Qwen's own announcement, not independently reproduced results, so they should be validated against real workloads before factoring into any procurement decision, a caution that applies even more strongly to 3.8 given how much less has been published for it.

Qwen 3.8 Max vs Kimi K3: What Can Actually Be Compared Right Now

Almost nothing, cleanly. Kimi K3 has a verified, independently checked profile: a confirmed 2.8 trillion parameter count, a published 93.5% GPQA Diamond score, an 88.3% Terminal-Bench 2.1 result, and transparent pricing at $3 input and $15 output per million tokens, corroborated in part by Artificial Analysis's independent testing. Qwen 3.8 Max has none of that yet. On tested access routes, early comparisons note that K3 tends to use fewer tokens and returns responses sooner, while Qwen's preview tends to use fewer requests and tool calls to complete similar tasks, with both providers caching more than 90 percent of repeated prompt traffic. But without a published per-token price for Qwen 3.8, a genuine dollar-for-dollar cost comparison isn't currently possible, and completed-task cost depends heavily on reasoning length, cache reuse, retry behavior, and rate limits regardless.

The one thing that can be said with confidence: Qwen's own claim positions 3.8 as the second-largest publicly disclosed model behind K2.6-preview Kimi K3's 2.8 trillion parameters, ahead of Thinking Machines' 975 billion parameter Inkling. That's a claim about scale, not about verified capability, and the two aren't the same thing.

Qwen 3.8 Max vs Claude Fable 5: The "Second Only To" Claim

Alibaba's specific framing, that Qwen 3.8 Max trails only Fable 5, deserves the same scrutiny as every other unverified claim in this release. It's not an unreasonable position for Alibaba to stake out, Qwen 3.7-Max's genuinely strong, independently referenced results on GPQA Diamond and agent benchmarks suggest the underlying research program is real and competitive. But "second only to Fable 5" is currently a vendor claim without an accompanying benchmark table, methodology, or independent corroboration, which puts it in a meaningfully weaker evidentiary position than, for example, Kimi K3's benchmark claims against Fable 5, which came with a full published comparison table and were partially corroborated within hours by Artificial Analysis's own independent testing.

Until Qwen 3.8 Max ships with its own documented benchmark suite, the honest way to describe this claim is: plausible given the lab's track record, currently unverifiable, and worth treating with real skepticism until real numbers land.

How to Access It Today

Despite the missing documentation, the preview is genuinely live and usable right now. Developers can test it through Alibaba's Token Plan endpoint, which documents an OpenAI-compatible API for supported agent tools, making integration straightforward if you're already working with OpenAI-style tooling. Qoder and QoderWork serve as Alibaba's first-party coding and knowledge-work interfaces for the model. A public chat version is also available for anyone who wants to try it without API integration.

One access note worth flagging directly: the international Token Plan page is operated by Intelligent Cloud Computing (Singapore) Private Limited rather than Alibaba Cloud's primary entity. Confirm that entity matches your existing Alibaba Cloud account before entering payment details, and use the official Alibaba Cloud console directly where that option is available to you.

It's also worth knowing that the hosted preview is explicitly described as a moving target. Alibaba's own Personal Token Plan documentation states the preview will receive continuous upgrades and will later be removed or replaced by a formal, presumably better-documented model. Anything you build against the current preview endpoint should be built with that instability in mind.

Should You Actually Use It

Test Qwen 3.8 Max Preview today if you're already inside the Alibaba Cloud or Qoder ecosystem, want to evaluate an early frontier-scale model against your own tasks, and are comfortable with a preview endpoint that will change under you before a formal release.

Hold off on any production commitment or procurement decision until Alibaba publishes an actual benchmark table, confirms the active parameter count and context window, and either delivers the promised open weights or announces standard per-token pricing. Every comparison claim currently circulating about this model, including Alibaba's own "second only to Fable 5" framing, is unverifiable until that documentation exists.

Bottom Line

Qwen 3.8 Max is a real, usable preview of what's very likely a genuinely capable model, backed by a research program that's already demonstrated strong, independently referenced results with Qwen 3.7-Max. But right now, it's a parameter count and a confident claim, not a verified competitor to Kimi K3 or Claude Fable 5. The timing, landing two days after Moonshot's K3 launch shook the market, suggests this preview was pushed out to hold attention during a major industry conference rather than to make a fully documented, benchmarked case. That's a reasonable business decision for Alibaba. It just means the burden of proof here is still entirely on Alibaba to publish real numbers, and until it does, every "beats" or "trails only" headline about this model, including this one's, should be read as provisional.


Related Articles