📊 Full opportunity report: Qwen3.8-Max’s Latest AI Results: A Step Closer To The Top? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced that its Qwen3.8-Max model, with 2.4 trillion parameters, has achieved top-tier benchmark scores, positioning it close to the leading models. The open weights are set to be released next week, marking a significant step in open AI model availability.

Alibaba has publicly released comprehensive benchmark results for its Qwen3.8-Max model, confirming its status as one of the most powerful AI models to date, with 2.4 trillion parameters. This marks a significant milestone in the company’s AI development and signals a closer step toward competing with top-tier models like GPT-5.6.

On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, revealing a model built on sparse mixture-of-experts architecture, with approximately 95 billion active parameters per query. The model is multimodal, capable of processing text, images, and videos, with text output. The benchmark scores place it near the top of several key evaluations, including a score of 86.6 on Terminal-Bench 2.1, surpassing models like Claude Fable 5 and only behind GPT-5.6 Sol at 88.8.

Alibaba also confirmed that the open weights for the model will be available next week, alongside a smaller 27-billion-parameter version, Qwen3.8-27B. The full benchmark results were withheld during the preview phase but are now publicly accessible, providing transparency about the model’s capabilities and limitations. The 2.4 trillion parameters are primarily a model size figure; the active parameters engaged during operation are about 95 billion.

At a glance
updateWhen: announced August 3, 2023; benchmarks pu…
The developmentAlibaba officially released benchmark results for its Qwen3.8-Max model, confirming its high performance and upcoming open-weight release.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba’s Benchmark Performance

The release of Qwen3.8-Max benchmark results demonstrates Alibaba's progress toward competing with leading AI models globally. Its high scores in multimodal and agentic tasks suggest it could influence AI deployment strategies, especially as the open weights become available next week. However, the model’s limitations on software engineering benchmarks highlight ongoing challenges in achieving comprehensive performance.

For the AI community and industry stakeholders, this development signals increased competition and innovation, especially as open models gain ground. The upcoming open-weight release could enable broader experimentation and deployment, potentially reshaping AI development dynamics.

Amazon

AI development and training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Launches

Alibaba’s AI development has been marked by strategic teasers and stealth previews, culminating in the July preview of Qwen3.8-Max. The model was initially identified through community detection on the Code Arena leaderboard and later confirmed at the World AI Conference in Shanghai. Prior to this, Alibaba’s models had been largely proprietary, with limited public benchmarks or open releases.

The company’s approach involved a staged reveal, starting with a stealth preview, followed by selective disclosures, and now full benchmark publication. This pattern aligns with industry practices of building anticipation and validating performance through independent testing.

Compared to competitors like Meta’s Kimi K3 and other large models, Alibaba’s model is notable for its size, multimodal capabilities, and focus on agentic tasks, with a clear emphasis on transparency through benchmark sharing.

"We are committed to transparency and will release the open weights next week, enabling broader access and experimentation."

— Alibaba spokesperson

Amazon

high-performance AI model training servers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Deployment

It is not yet clear how the Qwen3.8-Max model will perform in real-world applications beyond benchmarks, especially in software engineering tasks where it trails behind Fable 5. The licensing terms for the open weights remain unpublished, raising questions about usage restrictions and commercial deployment. Additionally, the impact of the active-parameter count on practical inference efficiency is still to be tested in diverse environments.

Amazon

GPU clusters for AI research

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Upcoming Release and Community Testing of Open Weights

The open weights for Qwen3.8-Max are scheduled for release next week, which will enable researchers and developers to evaluate its capabilities directly. The smaller 27-billion-parameter version, Qwen3.8-27B, is designed for deployment on high-memory single machines, making it accessible for broader experimentation. Industry observers will closely monitor how the model performs outside benchmark settings and how licensing terms evolve.

Further updates are expected as Alibaba publishes detailed documentation and user feedback begins to emerge from early adopters.

Scaling AI: The AI Governance and Security Playbook for Executives

Scaling AI: The AI Governance and Security Playbook for Executives

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the main capabilities of Qwen3.8-Max?

Qwen3.8-Max is a multimodal, 2.4 trillion-parameter model capable of processing text, images, and videos, with strong performance in various benchmark evaluations, especially in agentic and multimodal tasks.

When will the open weights for Qwen3.8-Max be available?

The open weights are scheduled to be released next week, enabling broader access for research and deployment.

How does Qwen3.8-Max compare to other large models like GPT-5.6?

In benchmark tests, Qwen3.8-Max scores close to GPT-5.6 in some areas, particularly in multimodal and agentic benchmarks, but trails behind in software engineering tasks.

What limitations does Qwen3.8-Max have?

While its benchmark results are impressive, the model still underperforms on certain deep software engineering benchmarks and its licensing terms remain unpublished, which could affect usage.

Will the smaller Qwen3.8-27B model be useful for deployment?

Yes, the 27-billion-parameter version is designed for deployment on high-memory single machines, making it accessible for practical applications and experimentation.

Source: ThorstenMeyerAI.com

You May Also Like

Webinar follow-up personalization tool for B2B consultants

A new webinar follow-up personalization tool for solo B2B consultants is entering testing, aiming to improve engagement and lead conversion.

SpaceX Owns Every Layer of AI Now. The Model Is Still the Weak Link.

SpaceX completes its acquisition of Cursor for $60 billion, owning every layer of AI infrastructure but still facing limitations in model performance.

The Hidden Reason Many Nanotech Startups Fail

Hidden challenges often derail nanotech startups before they succeed; uncover the crucial obstacles and how to overcome them to ensure your innovation thrives.

Forezai · TradingAgents: A Trading Firm Made of Agents

Forezai introduces TradingAgents, a novel multi-agent research framework mimicking a trading desk, emphasizing structured disagreement and oversight for market decisions.