📊 Full opportunity report: Qwen3.8-Max’s Latest AI Results: A Step Closer To The Top? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba announced that its Qwen3.8-Max model, with 2.4 trillion parameters, has achieved top-tier benchmark scores, positioning it close to the leading models. The open weights are set to be released next week, marking a significant step in open AI model availability.
Alibaba has publicly released comprehensive benchmark results for its Qwen3.8-Max model, confirming its status as one of the most powerful AI models to date, with 2.4 trillion parameters. This marks a significant milestone in the company’s AI development and signals a closer step toward competing with top-tier models like GPT-5.6.
On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, revealing a model built on sparse mixture-of-experts architecture, with approximately 95 billion active parameters per query. The model is multimodal, capable of processing text, images, and videos, with text output. The benchmark scores place it near the top of several key evaluations, including a score of 86.6 on Terminal-Bench 2.1, surpassing models like Claude Fable 5 and only behind GPT-5.6 Sol at 88.8.
Alibaba also confirmed that the open weights for the model will be available next week, alongside a smaller 27-billion-parameter version, Qwen3.8-27B. The full benchmark results were withheld during the preview phase but are now publicly accessible, providing transparency about the model’s capabilities and limitations. The 2.4 trillion parameters are primarily a model size figure; the active parameters engaged during operation are about 95 billion.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Implications of Alibaba’s Benchmark Performance
The release of Qwen3.8-Max benchmark results demonstrates Alibaba's progress toward competing with leading AI models globally. Its high scores in multimodal and agentic tasks suggest it could influence AI deployment strategies, especially as the open weights become available next week. However, the model’s limitations on software engineering benchmarks highlight ongoing challenges in achieving comprehensive performance.
For the AI community and industry stakeholders, this development signals increased competition and innovation, especially as open models gain ground. The upcoming open-weight release could enable broader experimentation and deployment, potentially reshaping AI development dynamics.
AI development and training hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on Alibaba’s AI Model Launches
Alibaba’s AI development has been marked by strategic teasers and stealth previews, culminating in the July preview of Qwen3.8-Max. The model was initially identified through community detection on the Code Arena leaderboard and later confirmed at the World AI Conference in Shanghai. Prior to this, Alibaba’s models had been largely proprietary, with limited public benchmarks or open releases.
The company’s approach involved a staged reveal, starting with a stealth preview, followed by selective disclosures, and now full benchmark publication. This pattern aligns with industry practices of building anticipation and validating performance through independent testing.
Compared to competitors like Meta’s Kimi K3 and other large models, Alibaba’s model is notable for its size, multimodal capabilities, and focus on agentic tasks, with a clear emphasis on transparency through benchmark sharing.
"We are committed to transparency and will release the open weights next week, enabling broader access and experimentation."
— Alibaba spokesperson
high-performance AI model training servers
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Remaining Questions About Model Deployment
It is not yet clear how the Qwen3.8-Max model will perform in real-world applications beyond benchmarks, especially in software engineering tasks where it trails behind Fable 5. The licensing terms for the open weights remain unpublished, raising questions about usage restrictions and commercial deployment. Additionally, the impact of the active-parameter count on practical inference efficiency is still to be tested in diverse environments.
As an affiliate, we earn on qualifying purchases.
Upcoming Release and Community Testing of Open Weights
The open weights for Qwen3.8-Max are scheduled for release next week, which will enable researchers and developers to evaluate its capabilities directly. The smaller 27-billion-parameter version, Qwen3.8-27B, is designed for deployment on high-memory single machines, making it accessible for broader experimentation. Industry observers will closely monitor how the model performs outside benchmark settings and how licensing terms evolve.
Further updates are expected as Alibaba publishes detailed documentation and user feedback begins to emerge from early adopters.

Scaling AI: The AI Governance and Security Playbook for Executives
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What are the main capabilities of Qwen3.8-Max?
Qwen3.8-Max is a multimodal, 2.4 trillion-parameter model capable of processing text, images, and videos, with strong performance in various benchmark evaluations, especially in agentic and multimodal tasks.
When will the open weights for Qwen3.8-Max be available?
The open weights are scheduled to be released next week, enabling broader access for research and deployment.
How does Qwen3.8-Max compare to other large models like GPT-5.6?
In benchmark tests, Qwen3.8-Max scores close to GPT-5.6 in some areas, particularly in multimodal and agentic benchmarks, but trails behind in software engineering tasks.
What limitations does Qwen3.8-Max have?
While its benchmark results are impressive, the model still underperforms on certain deep software engineering benchmarks and its licensing terms remain unpublished, which could affect usage.
Will the smaller Qwen3.8-27B model be useful for deployment?
Yes, the 27-billion-parameter version is designed for deployment on high-memory single machines, making it accessible for practical applications and experimentation.
Source: ThorstenMeyerAI.com