📊 Full opportunity report: Latest AI Performance Figures For Qwen3.8-Max: What’s Next? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba unveiled detailed performance figures for Qwen3.8-Max, confirming a 2.4 trillion-parameter model with strong benchmark results. Open weights are set to ship next week, signaling significant progress in open AI models.

Alibaba has officially released detailed performance figures for its Qwen3.8-Max model, confirming it as a 2.4 trillion-parameter, multimodal AI system with strong benchmark results. This announcement marks the transition from preview to full availability, with open weights scheduled to ship next week, making it a significant development in open AI models.

On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, revealing it as a 2.4 trillion-parameter, sparse mixture-of-experts model based on the Qwen3.5 architecture. The model demonstrates competitive performance across several benchmarks, notably achieving 86.6 on Terminal-Bench 2.1, surpassing Claude Fable 5 and only behind GPT-5.6 Sol at 88.8.

Alibaba confirmed that the active parameters per query are approximately 95 billion, with the total network size at 2.4 trillion parameters, marking it as the largest open-weight model announced to date. The model supports multimodal inputs—text, images, and videos—and outputs text, with a context window of 983,616 tokens. The benchmark results were obtained using Alibaba’s own testing environment, providing transparency about performance metrics.

The company also announced the upcoming release of open weights for the model next week. The open weights are expected to be a checkpoint that exceeds the capabilities of smaller models like Qwen3.8-27B, which will also be released. The 27B model is designed for deployment on single high-memory machines, making it accessible for local inference and development.

While the full benchmark table confirms strong performance, some limitations remain. Alibaba’s claims about the model’s agentic capabilities are supported by significant improvements in long-horizon tasks, but certain software engineering benchmarks still show notable gaps compared to competitors like Fable 5. The company emphasizes that the model’s agentic and long-horizon capabilities have improved dramatically through reinforcement learning environment scaling.

At a glance
updateWhen: announced August 3, 2023; benchmarks an…
The developmentAlibaba announced the full benchmark table and confirmed the upcoming release of open weights for Qwen3.8-Max, a 2.4 trillion-parameter AI model, after weeks of speculation.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba’s Open-Weight Release for AI Development

The announcement of Qwen3.8-Max’s detailed benchmark results and the upcoming release of open weights mark a key milestone in the democratization of large-scale AI models. With a 2.4 trillion-parameter size and competitive performance, Alibaba’s model challenges existing leaders and provides a new open-source option for developers and researchers. This development could accelerate innovation, foster more open research, and influence the future landscape of multimodal AI systems.

Moreover, the availability of a 95-billion-parameter model that can run on high-memory hardware makes advanced AI more accessible for local deployment, potentially reducing reliance on proprietary APIs. However, the model’s agentic capabilities and benchmark gaps highlight ongoing challenges and the need for further validation and testing in real-world scenarios.

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Launches and Benchmarking

Alibaba’s recent AI strategy has involved stealth previews and selective disclosures, building anticipation around its large models. In July, the company previewed Qwen3.8-Max as a slogan, followed by a stealthy appearance on the Code Arena leaderboard under the alias 'kaleb.' The official confirmation came during the World AI Conference in Shanghai, where Alibaba revealed the model’s parameters and capabilities, emphasizing its multimodal and agentic features.

The company’s approach has been strategic, releasing limited previews at discounted prices and withholding full benchmark data until now. Prior to Qwen3.8-Max, Alibaba’s models, such as Qwen3.7-Max, demonstrated incremental improvements, but the new release signifies a leap in both scale and performance, aligning with broader industry trends toward massive, multimodal models.

"Qwen3.8-Max sets a new standard for multimodal AI performance, and we are committed to making its weights accessible to foster innovation."

— Alibaba spokesperson

/Modern GPU Programming with Rust and CUDA 13: Mastering Parallel Computing, GPU Acceleration, Memory Optimization, AI Systems, and High-Performance Application Development (Learning Express Series)

/Modern GPU Programming with Rust and CUDA 13: Mastering Parallel Computing, GPU Acceleration, Memory Optimization, AI Systems, and High-Performance Application Development (Learning Express Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Open-Weight Licensing and Capabilities

It remains unclear what the licensing terms will be for the 2.4 trillion-parameter open weights, and whether they will be fully open-source or have restrictions. The impact of the open weights on real-world deployment and how agentic capabilities will hold up in practical applications are still to be tested. Additionally, the benchmark results, while strong, do not cover all use cases, leaving some performance aspects unverified.

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

  • Powerful Microcontroller: ESP32-S3 with 16MB Flash and 8MB PSRAM
  • AI Vision & Voice Capabilities: Camera and audio for AI interactions
  • Supports OpenCV & YOLO: Face tracking and human pose estimation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps: Open Weights Release and Broader Adoption

Alibaba is scheduled to release the open weights of Qwen3.8-Max next week, which will enable developers to test and deploy the model locally. The community will closely evaluate its agentic performance, especially on long-horizon tasks and software engineering benchmarks. Further updates may include refinements, licensing clarifications, and additional benchmark results as the model is adopted in diverse applications.

Acer Veriton AI Mini Workstation Personal Computer GN100-UD11 Series

Acer Veriton AI Mini Workstation Personal Computer GN100-UD11 Series

  • Powerful AI Performance: 1 PFLOPS FP4 AI with NVIDIA GB10 Superchip
  • Pre-installed NVIDIA DGX OS: Optimized for full NVIDIA AI stack
  • High-Performance GPU and CPU: Fifth-gen Tensor Cores with 20-core Arm CPU

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

When will Alibaba release the open weights for Qwen3.8-Max?

The open weights are scheduled to be released next week, with details to be confirmed by Alibaba.

How does Qwen3.8-Max compare to other large AI models?

It demonstrates competitive benchmark performance, notably surpassing some models on specific tasks like Terminal-Bench and PaperBench, and is only behind GPT-5.6 Sol at maximum effort.

What are the main limitations of Qwen3.8-Max?

While it excels in some areas, it still trails significantly on software engineering benchmarks like SWE-bench Pro and FrontierSWE, indicating room for improvement in certain technical tasks.

Will the model’s agentic capabilities be reliable in real-world use?

It is still uncertain how well the agentic improvements will translate outside controlled benchmarks, and further testing is needed to confirm its practical effectiveness.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

EMV 3‑DS 2.3: What the Latest Spec Adds for Merchants

Introducing EMV 3‑DS 2.3’s latest updates that empower merchants with enhanced security and seamless user experience—discover what these changes mean for your business.

Behind the Scenes of Real‑Time Payment Orchestration Engines

Unlock the secrets behind real-time payment orchestration engines that ensure speed, security, and compliance—discover how they stay ahead of evolving threats.