AI Power Rankings: Where Does Qwen3.8-Max Truly Stand?
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: AI Power Rankings: Where Does Qwen3.8-Max Truly Stand? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the broad availability of Qwen3.8-Max, confirming its 2.4 trillion parameters and benchmark results. The model shows strong performance in multimodal tasks but trails in software benchmarks, raising questions about its true ranking.

Alibaba has officially made Qwen3.8-Max broadly available, confirming it as a 2.4 trillion-parameter, multimodal AI model with strong benchmark results. This marks the first time the model’s full specifications and performance data are publicly released, ending weeks of speculation and stealth preview.

On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, revealing it as a sparse mixture-of-experts model built on the Qwen3.5 architecture, with approximately 95 billion active parameters per query. The model’s specifications include multimodal capabilities—text, image, and video input with text output—and a 983,616-token context window. The benchmark results show it surpasses several competitors in certain tasks, such as Terminal-Bench 2.1 (86.6), and dominates in multimodal and agentic benchmarks, like Parametric CAD Bench (91.5) and OmniDocBench (92.1). However, it trails significantly in deep software engineering benchmarks, such as SWE-bench Pro (67.7) and FrontierSWE (73.5), compared to Fable 5. The announcement also confirms the open weights will be available next week, with a 27B checkpoint targeted at local deployment. The company’s framing emphasizes the model’s agentic improvements, especially in long-horizon tasks, achieved through reinforcement learning environment scaling. The release strategy involved a stealth preview, followed by a staged disclosure of specifications and benchmark data, which has influenced market perception and share price movements.

At a glance
reportWhen: announced August 3, 2023; full release…
The developmentAlibaba officially released Qwen3.8-Max with benchmark data, confirming its specifications and open weights, after weeks of speculation.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Qwen3.8-Max’s Benchmark Performance

The release of Qwen3.8-Max and its benchmark data provides a clearer picture of Alibaba's position in the AI landscape, especially regarding large-scale, multimodal models. Its strong multimodal and agentic performance suggests advancements in practical AI applications, while its shortcomings in deep software benchmarks highlight ongoing challenges. The open weights' upcoming release signals a potential shift towards more accessible, high-capacity models, but the model's true capabilities and licensing details remain to be fully clarified. For developers and industry watchers, this development underscores the competitive race to build more versatile and scalable AI systems, with Alibaba positioning itself as a serious contender.

AI/ML Definitive Guide: Architecture, Models, Big Data, Deployment, Open-Source Tools, Cloud Services, MLOps, LLMs, Gen AI

AI/ML Definitive Guide: Architecture, Models, Big Data, Deployment, Open-Source Tools, Cloud Services, MLOps, LLMs, Gen AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Releases

Alibaba’s AI development has been marked by strategic stealth and selective disclosures, with notable models like Kimi K3 and the anonymous ‘kaleb’ model surfacing briefly before official confirmation. The company’s approach has involved staged previews, limited access, and a focus on benchmark performance, often emphasizing multimodal and agentic capabilities. The recent announcement follows a period of intense industry speculation, fueled by market reactions to competitor launches such as Meta’s Llama 3 and OpenAI’s GPT-5. Alibaba’s previous models have demonstrated incremental improvements, but the firm’s latest move with Qwen3.8-Max marks a significant shift toward transparency and open access, albeit with limitations related to licensing and deployment scale.

"We are committed to transparency and open access, and the upcoming release of weights will enable broader experimentation and deployment."

— Alibaba spokesperson

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

  • Powerful Microcontroller: ESP32-S3 with 16MB Flash and 8MB PSRAM
  • AI Vision & Voice Capabilities: Camera and audio for AI interactions
  • Supports OpenCV & YOLO: Face tracking and human pose estimation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Licensing and Deployment

Details about the licensing terms for the open weights remain unpublished, raising questions about potential restrictions on commercial use. It is also unclear whether the 2.4 trillion parameter checkpoint will be fully functional at scale or require further optimization. The long-term stability of agentic gains, especially in real-world applications, has yet to be validated through independent testing. Additionally, the full benchmark performance of the 27B checkpoint, designed for local deployment, has not been disclosed, leaving its comparative standing uncertain.

Evals for AI Engineers: Systematically Measuring and Improving AI Applications

Evals for AI Engineers: Systematically Measuring and Improving AI Applications

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps in Alibaba’s AI Model Rollout

Alibaba plans to release the open weights of Qwen3.8-Max next week, enabling wider testing and deployment. Industry analysts will closely monitor independent benchmark results, especially for the 27B checkpoint. Further disclosures on licensing, licensing restrictions, and real-world application performance are expected in the coming months. The company may also introduce updates or new variants to enhance agentic capabilities and software engineering performance, continuing its strategic push in the AI race.

Mastering Small Language Models: A Practical Guide to Building Lightweight NLP Systems with Python, Transformers, and Quantization Techniques

Mastering Small Language Models: A Practical Guide to Building Lightweight NLP Systems with Python, Transformers, and Quantization Techniques

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the main capabilities of Qwen3.8-Max?

Qwen3.8-Max is a multimodal model capable of processing text, images, and videos, with strong performance in agentic reasoning and complex task execution, especially in multimodal and long-horizon applications.

When will the open weights for Qwen3.8-Max be available?

The open weights are scheduled to be released next week, allowing developers to deploy and experiment with the model locally.

How does Qwen3.8-Max compare to competitors like GPT-5 or Fable 5?

In benchmark tests, Qwen3.8-Max outperforms some models like Claude Opus 4.8 and Fable 5 in certain tasks but trails behind GPT-5 at the highest effort levels, particularly in software engineering benchmarks.

What are the licensing implications of the open weights?

Licensing details remain unpublished, raising questions about restrictions on commercial use and whether the model can be freely deployed without attribution or revenue-sharing obligations.

What does this development mean for the AI industry?

This marks a significant step toward more open, high-capacity models, intensifying competition among AI developers and raising expectations for accessible, versatile AI systems in the near future.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
You May Also Like

How AI Is Transforming Ergonomic Office Chairs In 2026

In 2026, AI-driven ergonomic office chairs are transforming workplace comfort and support, with advanced features tailored to individual needs.

The Forecast Is the Plan.

Major AI labs publicly commit to automating AI R&D by 2026, signaling a shift from research goals to execution plans with significant implications for the industry.

Engineering Is Automated. Research Is the Residual.

Recent benchmarks show AI can automate core engineering tasks, leaving research as the remaining challenge, with implications for AI development timelines.

AI Oversight In Action: Amazon’s Talks With U.S. Authorities And The Anthropic Ban

Amazon’s recent discussions with U.S. officials have triggered a crackdown on Anthropic models, highlighting increased AI oversight efforts.