1 CLICK = 30 BLOG POST + BACKLINKS
Qwen 3.8 just leaked past Kimi K3 with 2.4 trillion parameters. Alibaba claims it's #2 to Fable 5 — here's what's real, what's hype, and what's free.
Two days. That's how long it took Alibaba to answer Moonshot AI's Kimi K3. On July 19, 2026, in the middle of the World AI Conference in Shanghai, Alibaba's Qwen team pulled back the curtain on Qwen 3.8, a 2.4 trillion-parameter model the company is already calling one of the strongest AI systems available anywhere. Not the strongest. Alibaba is bold, but not delusional — the company says Qwen 3.8 lands "second only to Fable 5," a direct nod to Anthropic's newest frontier release.
I've watched enough of these announcements to know the drill. A lab drops a headline parameter count, a chart with no axis labels, and a promise that weights are "coming soon." Sometimes that pans out into something genuinely useful. Sometimes it evaporates into a rebranded API endpoint nobody remembers by September. Qwen 3.8 could go either way right now, and that uncertainty is exactly why it's worth digging into properly instead of just repeating the press release.
This isn't Alibaba's first rodeo. The Qwen family has quietly become one of the most-downloaded open model lineups on Hugging Face, and the company has a track record of eventually delivering on open-weight promises — even when it takes months longer than the announcement implies. So while the skepticism is warranted, so is the attention.
Qwen 3.8 is the newest flagship in Alibaba's Qwen series of large language models, and it's a significant jump from its predecessor. Where Qwen 3.7-Max was already a serious contender — scoring 92.4% on GPQA Diamond and 80.4% on SWE-bench Verified — Qwen 3.8 scales the architecture up dramatically.
Here's what's confirmed so far:
2.4 trillion total parameters, built on a sparse Mixture-of-Experts (MoE) design
The first Qwen model above 1 trillion parameters to go fully multimodal — it handles text, images, video, and documents in one system
Currently accessible only as a preview build called Qwen3.8-Max-Preview
Available through Alibaba's Token Plan, Qoder, and QoderWork platforms
An open-weight release is promised, but with no confirmed date, license, or model card yet
That last point matters more than it sounds like it should. Alibaba's Max-tier models have historically stayed closed-source — this would be the first time the company opens up a flagship at this scale, according to reporting from MarkTechPost, which broke down the announcement in detail on the day it happened.
2.4 trillion parameters is a genuinely massive figure — for context, most widely-used open models today sit in the tens-of-billions to low-hundreds-of-billions range. Crossing the 2-trillion mark puts Qwen 3.8 in rarefied territory.
But here's the asterisk nobody's putting in their headline: Alibaba hasn't disclosed how many of those 2.4 trillion parameters are actually active at inference time. For a sparse MoE model, that's not a footnote — it's the whole ballgame. Qwen's own lineup proves the point:
Qwen3-235B-A22B carries 235 billion total parameters but activates only 22 billion per token
Qwen3-30B-A3B activates roughly 3 billion of its total parameters
If Qwen 3.8 follows that same pattern, the "2.4 trillion" headline number could be doing a lot of heavy lifting for a much smaller effective compute cost per query. That's not a knock on the model — MoE architecture is precisely what makes massive models affordable to run. It's just a reason to treat the raw parameter count as marketing shorthand, not a performance guarantee.
Buried under the parameter-count fireworks is the detail I think matters more long-term: this is Qwen's first model above 1 trillion parameters that can genuinely see. Qwen developer Shuai Bai described it as the team's first multimodal system at this scale, capable of processing images, video, and documents alongside text — not bolted-on OCR, but a native part of the architecture.
That's the difference between "a bigger chatbot" and "a system a business can actually put in front of a document pipeline, a video archive, or a design workflow."
How multimodal AI models process images and video?
Let's be precise, because "release" is doing a lot of ambiguous work in most coverage of this story.
What's live today? (as of July 21, 2026):
✔️ A preview endpoint called Qwen3.8-Max-Preview
✔️ Access through Alibaba's Token Plan subscription
✔️ Access through Qoder and QoderWork
✔️ A documented context window reportedly near 1 million tokens, per Qwen Cloud integration notes (though this is not yet confirmed in an official model card)
✔️ Reasoning controls with low, high, and xhigh settings — xhigh is the default
What's NOT live yet?
❌ Open-weight download / self-hosting
❌ A published license
❌ A complete official benchmark table
❌ A standard, published per-token API price
❌ An official model card
In plain terms: you can test-drive Qwen 3.8 today through Alibaba's hosted platforms, but you can't yet download it, self-host it, or point to an independently verified benchmark chart and say "here's exactly how good this is." That gap between "announced" and "available" is the single most important thing to understand before you get excited about this model — a distinction the team at Coursiv's technical breakdown also flags as the key caveat right now.
Qwen 3.8 didn't appear in a vacuum. Two days before Alibaba's announcement, Moonshot AI — a company Alibaba actually holds a 36% stake in — released Kimi K3, an open-weight model at 2.8 trillion parameters, briefly claiming the title of largest open-source model available. Moonshot has reportedly hit $300 million in annual recurring revenue as of June, and Bloomberg has reported plans for the company to go public within roughly six months.
Read that again: Alibaba partly owns Moonshot, and Alibaba just announced a competing model days after its own portfolio company's big launch. That's not coincidence — that's a company hedging its bets across two horses in the same race, according to reporting from The Decoder. Whether that's clever portfolio strategy or a sign of internal competition at Alibaba is genuinely up for debate, and I don't think there's a clean answer yet.
This is the section everyone actually wants, so let's cut through the marketing language.
Alibaba's own claim is that Qwen 3.8 ranks "second only to Fable 5" — a direct reference to Anthropic's Claude Fable 5 model. That's a bold thing to say out loud, and to be fair to Alibaba, it's a specific enough claim that it can be checked once independent benchmarks land.
Here's the honest scorecard as it stands today:
Parameter count: Qwen 3.8 (2.4T) trails Kimi K3 (2.8T) in raw size
Modality: Qwen 3.8 wins here — it's multimodal from above 1 trillion parameters; Kimi K3's headline strength has been text and code
Availability: Kimi K3 is already open-weight; Qwen 3.8 is preview-only
Independent verification: Neither has a comprehensive third-party benchmark table published as of this writing
Predecessor comparison: Alibaba says Qwen 3.8 should beat Qwen3.7-Max — which itself scored 92.4% on GPQA Diamond and 80.4% on SWE-bench Verified — on coding, full-stack development, data analysis, and office workflows
That predecessor comparison is actually the most useful data point available right now, because it's the only one with real numbers attached to it. Everything about Qwen 3.8 vs Claude or Qwen 3.8 vs Kimi K3 is currently vibes-based until Alibaba (or a third party) publishes an actual benchmark table.
Cautiously, maybe. Alibaba doesn't have a habit of making wildly indefensible claims about Qwen models — the Qwen3.7-Max numbers cited above were independently reproducible. But "second only to Fable 5" is also exactly the kind of headline-grabbing framing a company reaches for when it wants press coverage during a major conference, and WAIC Shanghai is about as high-visibility a stage as it gets for a Chinese AI lab.
My honest read: treat this as a hypothesis worth testing, not a settled fact. The preview is live. If you have Token Plan access, running your own side-by-side comparison against a task you actually care about will tell you more than any headline will.
Nobody's giving away a 2.4 trillion-parameter model for nothing — but the entry price is genuinely low enough that the "practically $0" framing isn't pure clickbait. Let's get specific, because pricing is exactly where most coverage of this launch goes vague.
Alibaba bundles Qwen3.8-Max-Preview into its Token Plan, a credit-based subscription rather than a traditional per-token API rate. There's no standalone pay-as-you-go price published for Qwen 3.8 specifically — you're buying access to a plan that happens to include it, alongside other models like qwen3.7-max, GLM-5.2, DeepSeek-V4-Pro, and Wan2.7-Image-Pro.
Here's the current tier breakdown:
✔️ Lite Plan — $6/month: 2,500 credits every 7 days, 700 credits every 5 hours, room for 1-2 concurrent AI agents
✔️ Standard Plan — $18/month: 10,000 credits every 7 days, 3,000 credits every 5 hours, 3-4 concurrent agents
✔️ Pro Plan — $68/month: 40,000 credits every 7 days, 12,000 credits every 5 hours, 6-8 concurrent agents — built for teams
On top of those base prices, Alibaba is running the preview at 10% of standard pricing across the board. Stack a documented off-peak discount window on top of that — some reporting puts off-hours usage as low as roughly 1-2% of the standard rate — and you get why the "$0 to test" framing in this article's headline isn't an exaggeration so much as a rough summary of what early adopters are actually paying right now.
I'll be blunt: credit systems are a pain to budget against, and this is coming from someone who's tracked enough subscription-AI pricing to know the pattern. A few things worth flagging before you hand over a card number:
Credits, not tokens, means your actual cost-per-task is fuzzy until you've run real workloads through it
The 10% preview discount is, by definition, temporary — budget for the full-price number, not the promotional one
There's no confirmed date for when "preview" pricing rolls into permanent pricing
If your team needs predictable, forecastable per-token costs today, the older Qwen tiers remain the safer bet, a point eesel AI's hands-on review makes explicitly after testing the preview directly
If you want to try it yourself rather than take anyone's word for it (mine included), the path is straightforward:
Head to Alibaba Cloud's Token Plan pricing page and pick a region-appropriate tier
Generate your API key through the Qwen Cloud console
Point a compatible tool — Qwen Code, or any harness supporting the Qwen Cloud Coding Plan provider setting — at the qwen3.8-max-preview model ID
Alternatively, skip the raw API entirely and use it through Qoder or QoderWork, Alibaba's own agentic coding IDE and desktop assistant, where the preview model ships with aggressive credit multipliers baked in
Step-by-step guide to setting up Qwen Code
Strip away the marketing language and here's what's actually documented about Qwen 3.8's architecture and limits, based on Qwen Cloud's own integration metadata:
Total parameters: 2.4 trillion, sparse Mixture-of-Experts design
Active parameters per query: Not disclosed (see caveat below)
Context window: 983,616 tokens — close enough to round to "1 million" for practical purposes, though this figure comes from integration notes rather than a formal model card
Maximum output: 131,072 tokens
Reasoning modes: Low, high, and xhigh — with "thinking" always enabled by default and xhigh as the documented default setting
Modalities: Text, images, video, and documents
That context window is enormous by any standard — large enough to hold a lengthy codebase, a full legal contract, or hours of transcript in a single pass without chunking. If it holds up under real-world load rather than just synthetic testing, that alone makes Qwen 3.8 worth a serious look for anyone doing document-heavy or repo-scale work.
I flagged this in Part 1 and it's worth repeating here with more weight, because it's the single biggest gap in Alibaba's disclosure. Sparse MoE models like Qwen 3.8 don't run all their parameters on every query — a routing mechanism activates only a subset of "experts" per token. That's why trillion-parameter models are affordable to serve at all.
Alibaba has not published the activated-parameter count for Qwen 3.8. Compare that to the transparency around older Qwen releases, where Qwen3-235B-A22B's naming convention literally spells out the active count (22 billion) right in the model name. The silence here isn't necessarily suspicious — Max-tier models have historically had less public architecture disclosure than the open-source line — but it does mean any comparison of Qwen 3.8's "size" to Kimi K3's 2.8 trillion parameters is comparing two numbers that may not measure the same thing at all.
For context on real inference costs: a 2.4 trillion-parameter model at 4-bit quantization would need roughly 1.2 terabytes just to hold the weights, a figure independently calculated and cited in MarkTechPost's coverage of the launch. That's before you even get to activation memory during inference. This is precisely why the open-weight promise matters so much — self-hosting a model at this scale is a genuinely different infrastructure conversation than self-hosting something in the 30-70 billion parameter range most teams are used to.
Not every AI announcement deserves your attention. Here's my honest breakdown of who should be paying attention to this one today versus who should wait.
Worth testing the preview now if you are:
✔️ A developer already inside the Qwen or Alibaba Cloud ecosystem, where switching to the preview costs almost nothing at the discounted rate
✔️ A team evaluating multimodal document or video-processing pipelines who wants an early look at a trillion-parameter-class option
✔️ Someone doing comparative AI benchmarking who wants fresh data points before the model stabilizes
✔️ A cost-sensitive team currently priced out of frontier-model API rates elsewhere
Worth waiting on if you are:
❌ A business that needs a signed license and stable pricing before committing budget
❌ A team that requires self-hosting or on-premises deployment for compliance reasons — the open-weight release isn't here yet
❌ Anyone making a purchasing decision based purely on the "second only to Fable 5" claim — that's unverified marketing language until independent benchmarks exist
❌ A team without infrastructure to handle a model of this scale if and when open weights do land — remember that 1.2-terabyte weight-storage figure
Alibaba is explicitly pitching Qwen 3.8 at full-stack development, data analysis, and office-workflow tasks — not general chit-chat. The integration with Qoder and QoderWork reinforces that positioning; this is a model built to be dropped into an agentic coding harness and run against real repositories, not just chatted with in a browser tab.
If that's your use case, the practical question isn't "is it ranked #2 in the world" — it's "does it save you money or time on your actual repo's agent tasks at acceptable quality." One line from ExplainX's technical guide captures this well: if a model fails on multi-hour refactors, no ranking claim fixes that. That's the right lens for evaluating this launch, and it's the one most coverage skips entirely in favor of repeating Alibaba's self-assessment.
Best AI coding agents compared
Here's where I'll say something most coverage of this launch won't: the honest answer right now is "probably, but not for the reason Alibaba wants you to think."
The "second only to Fable 5" line is the hook that got this story into every AI newsletter this week, and it's also the least useful piece of information in the entire announcement. It's a self-graded claim with no accompanying benchmark table, no third-party eval, and no methodology disclosed. Treat it the way you'd treat a movie studio calling its own film "the best of the year" in the trailer — maybe true, but not evidence.
What's actually verifiable and actually interesting:
✔️ A massive, genuine architecture jump — 2.4 trillion parameters is real, confirmed, and not disputed by anyone covering this
✔️ The first Qwen flagship above 1 trillion parameters to be truly multimodal — that's a real capability shift, not a marketing spin
✔️ A near-1-million-token context window, which if it holds up under real workloads is genuinely competitive territory
✔️ Aggressive, low-cost preview access — you can be testing this today for the price of a couple of coffees a month
✔️ A stated open-weight commitment that, if honored, would be a first for Alibaba's Max tier and a meaningful shift in the open-source AI landscape
What's still fog:
❌ No independent benchmark table
❌ No confirmed activated-parameter count
❌ No open-weight release date, license, or Hugging Face repository
❌ No permanent pricing once the preview discount ends
My take: this is a legitimately significant technical release wrapped in a marketing claim that's outrunning its own evidence. Both things are true at once, and most of the coverage this week picked one and ignored the other.
I'd get more bullish on Qwen 3.8 if any of the following happen in the next few weeks:
Alibaba (or a credible third party like LMArena or an academic eval group) publishes an actual benchmark table
The open-weight release lands with a genuinely permissive license, matching or beating the openness of Kimi K3
Independent developers report real-world results on long-context, multi-hour agentic tasks — not synthetic benchmarks
I'd get more skeptical if:
The "open-weight soon" promise slips past a few months with no update, following a pattern Alibaba's Max tier has shown before
The preview quietly gets more expensive once the 10% discount phases out, with no corresponding jump in verified quality
The activated-parameter count, when it finally surfaces, turns out to be small enough that the 2.4T headline was mostly about compute efficiency marketing rather than raw capability
Qwen 3.8 is Alibaba's newest flagship large language model, previewed on July 19, 2026 during the World AI Conference in Shanghai. It's a 2.4 trillion-parameter, sparse Mixture-of-Experts model that's also Qwen's first system above 1 trillion parameters to handle images, video, and documents alongside text.
Not fully. What's live is a preview build called Qwen3.8-Max-Preview, accessible through Alibaba's Token Plan subscription and through its Qoder and QoderWork platforms. The complete open-weight model, along with its license and official benchmark data, has not been published.
Access runs through Alibaba's Token Plan, starting at $6/month for the Lite tier (2,500 weekly credits), $18/month for Standard, and $68/month for Pro. During the preview period, Qwen3.8-Max-Preview is offered at roughly 10% of standard pricing, with additional off-peak discounts on top of that.
Alibaba claims Qwen 3.8 is "second only to Fable 5" among the systems it benchmarked, but this claim has not been independently verified with a published benchmark table. Treat it as an unconfirmed self-assessment until third-party evaluations are available.
It depends on what you're measuring. Kimi K3 has more total parameters (2.8 trillion vs. 2.4 trillion) and is already open-weight, while Qwen 3.8 is ahead on modality — it's Qwen's first trillion-parameter-class model with native image and video processing. Neither has a complete independent benchmark comparison published yet.
Alibaba has committed to an open-weight release "soon," but has not published a date, license, or repository. This would mark a departure from Alibaba's historical practice of keeping its Max-tier models closed-source, so the promise is worth watching but not yet a certainty.
Qwen Cloud's integration metadata lists a 983,616-token context window with a maximum output of 131,072 tokens. This figure comes from developer integration notes rather than a formal, official model card, so treat it as reported-but-not-fully-confirmed.
Step back from Qwen 3.8 specifically for a second, because the bigger story here is the pace itself. Two major trillion-parameter-class model announcements — Kimi K3 and Qwen 3.8 — landed within 48 hours of each other, from companies that are financially intertwined with one another. That's not normal industry cadence. That's a sprint.
For businesses and developers watching from the sidelines, the practical lesson isn't "go bet everything on Qwen 3.8 today." It's that the gap between "largest model in the world" headlines is now measured in days, not quarters, and the models making those headlines are increasingly reaching the market as previews with real gaps in disclosure — pricing, benchmarks, active-parameter counts — that used to come standard at launch.
That's worth remembering the next time a headline says a model is "second only to" anything. Ask what it's second only to, who's grading the test, and whether you can actually download it yet.
I'll be updating this post as Alibaba fills in the blanks — the benchmark table, the license, the activated-parameter count. Until then, the preview is real, it's cheap to try, and it's genuinely worth testing on your own workload rather than taking anyone's ranking claim, including Alibaba's, at face value.
Also Read: