Alibaba released Qwen3.8-Max on August 3, 2026, the most capable model in its Qwen large language model (LLM) family and its first Max-class release to carry a committed open-weights publication date.
Architecture and Scale
<cite index="8-3,8-4">Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts (MoE) model that accepts text, image, and video as input and returns text.</cite> The MoE design means the full parameter count is not active during a given inference pass. <cite index="5-1">It activates only 95 billion parameters at a time, significantly reducing costs and latency while handling text, image, and video processing with a context window of up to 1 million tokens.</cite>
Availability and Pricing
<cite index="4-6">The model is available to global developers via Alibaba Cloud's Model Studio application programming interfaces (APIs), as well as through QwenWork, the company's workplace AI agent platform.</cite> <cite index="5-3">Alibaba priced the model at about 40% of Claude Opus 5 for input tokens and 24% for output tokens in international markets, while offering substantially lower costs through cache hits.</cite> <cite index="1-4,1-5">Alibaba said it will release the full model weights for public download in the coming week, and a smaller variant, Qwen3.8-27B, designed for more hardware-constrained environments, is also planned for open-source release.</cite> <cite index="4-3">The move marks Alibaba's return to open-sourcing its top-tier AI models after keeping several recent flagship releases proprietary earlier this year.</cite>
Benchmark Performance
<cite index="8-13,8-14">Alibaba published a full benchmark table with this release; Qwen3.8-Max scores 86.6 on Terminal-Bench 2.1, ahead of Claude Opus 4.8 and Claude Fable 5 at 84.6, but behind GPT-5.6 Sol (max) at 88.8.</cite> <cite index="8-19">The model tops most vision rows, including OSWorld-Verified at 86.1, Parametric CAD Bench at 91.5, and OmniDocBench 1.5 at 92.1.</cite> <cite index="6-6,6-7">On the crowdsourced comparison platform Arena.AI, Qwen3.8-Max immediately became the highest-ranking Chinese model for text tasks, though it still trails several Anthropic offerings including Claude Fable 5; for vision tasks, it ranked second globally, behind only a variant of Fable 5.</cite> <cite index="17-8,17-9">Most benchmark data is from Alibaba's own launch table, and the gap between vendor-reported and independently verified scores typically narrows over weeks.</cite>
Enterprise and Agentic Positioning
<cite index="12-3,12-4">Rather than emphasizing conversational intelligence, Alibaba is positioning the model as an autonomous coworker capable of executing projects that span days rather than minutes; according to the company, Qwen3.8-Max can autonomously complete software projects lasting more than 10 days, reproduce research papers involving thousands of lines of code, perform iterative chip-design optimization, and continuously revise plans using multimodal feedback loops.</cite> <cite index="1-8">Alibaba highlighted one autonomous coding demonstration in which the model completed a 16-day software engineering project without human intervention, producing an open-sourced framework called oh-my-cli.</cite> <cite index="12-5">Those demonstrations remain company-produced and have not yet been broadly replicated by independent evaluators.</cite>
Competitive Context
<cite index="6-4">Qwen3.8-Max sits close to the size of domestic rival Moonshot AI's Kimi K3, which launched last month with 2.8 trillion parameters.</cite> <cite index="4-4">The release signals Alibaba's entry into a recent round of powerful releases by Chinese developers aggressively narrowing the gap with leading United States labs.</cite> <cite index="9-8,9-9">The launch continues a fast release cadence from the Qwen team, which has shipped a steady run of models this year — a tempo that keeps Alibaba in the conversation each time a rival claims the lead.</cite>