On July 19, 2026, Alibaba's Qwen team teased its next flagship model, the 2.4-trillion-parameter "Qwen3.8," announcing that the weights would be released soon. However, what is currently available is only "Qwen3.8-Max-Preview," offered through the subscription-based Token Plan and integrated into Qoder-series products—not a downloadable model. While Qwen itself claims the model is "second only to Fable 5," there is still no performance table to verify that ranking. This is an announcement that opened a paid trial window first, while leaving both the performance table and the weight release for later.

AD

A Subscription Preview Moved Faster Than the 2.4 Trillion Figure

According to Qwen Cloud's guidance, Qwen3.8-Max-Preview is offered exclusively under the Token Plan. It supports a 1-million-token context, thinking mode, function calling, and built-in tools. On the other hand, structured output is listed as unsupported. Under the individual Token Plan specifications, users can access visual understanding in addition to text generation.

Qwen developer Shuai Bai described the model as the team's first multimodal model exceeding 1 trillion parameters, capable of processing images, video, and documents. Its scope is not limited to code generation. Qoder's announcement claims the model outperforms Qwen3.7-Max on tasks requiring long procedures, such as full-stack development, data analysis, and Office workflows.

However, no specific margin of improvement has been disclosed. Qwen3.7-Max, announced by Alibaba on May 20, was also a flagship model aimed at agentic coding, complex reasoning, and long-running task execution. What is confirmed with Qwen3.8 is its stated scale of 2.4 trillion parameters, the fact that it is the Qwen team's first multimodal model exceeding 1 trillion parameters, and the announcement of an upcoming open-weight release. Exactly which tasks improved, and by how much, must be measured separately.

No Numbers to Verify the Claim of Being "Second Only to Fable 5"

The figure of 2.4 trillion offers a clue for thinking about the capacity of knowledge and patterns a model can hold. But it does not necessarily mean all 2.4 trillion parameters are computed every time a single token is generated. With a Mixture of Experts (MoE) architecture, only the experts needed at that moment can be selectively activated. Whether Qwen3.8 is an MoE or a dense model, and how many parameters are used per token, has not been disclosed.

The quantization format, recommended inference configuration, and measured speed are also unknown. As a result, it is impossible to calculate the number of GPUs required, memory needs, or inference costs from the 2.4-trillion figure alone. Total parameter count is not a number that directly represents deployment weight.

The same problem applies to the claim of being "second only to Fable 5." Qwen has not disclosed which tasks were compared, which version of Fable 5 was used, the amount of thinking involved, the agent execution environment, or the number of trials. Without benchmark results or a model card, this remains Qwen's own self-assessment. While the 1-million-token context and tool support can be confirmed as product specifications, the model's overall ranking has not yet been established.

AD

The Difference from Kimi K3: How Much Was Disclosed Before Release

The teaser for Qwen3.8 came right after the debut of the 2.8-trillion-parameter Kimi K3. Neither model has distributed weights at this point; both are first accepting users through their respective apps and APIs. What differs is the amount of material disclosed ahead of the open-weight release.

Item Qwen3.8 Kimi K3
Stated parameter count 2.4 trillion 2.8 trillion
Current availability Token Plan and Qoder-series Max-Preview Kimi products and API
Weight release Announced as "soon," no date given Announced by July 27
Architecture Undisclosed KDA, AttnRes, Stable LatentMoE; selects 16 out of 896 experts
Context 1 million tokens 1 million tokens
Performance evaluation Results and conditions undisclosed Official evaluation table and execution conditions published
Pricing visibility Monthly Credit-based system; per-usage rate not listed $3 per million input tokens, $15 per million output tokens

It cannot be said that Kimi K3, being 400 billion parameters larger, is therefore higher-performing. Kimi has also not disclosed its active parameter count, and it cannot be estimated simply from the 16-out-of-896 expert selection ratio. Moreover, Kimi's evaluation table mixes different agent environments and Fable 5 fallback conditions. Even when more numbers are made public, they do not amount to a universal ranking measured under identical conditions.

Still, Kimi disclosed its architecture, API pricing, evaluation conditions, and weight release deadline first. Qwen3.8, by contrast, announced its delivery method and self-assessment before revealing the model's actual contents. The competition between the two is not simply a matter of comparing 2.4 trillion against 2.8 trillion in scale—it will be decided by whether third parties can reproduce the performance and operating costs after distribution.

The 90–98% Discount Is Not an API Price Cut

Alibaba is pushing the Preview's initial rollout through Credit discounts. In Qoder, the consumption multiplier during normal hours has been lowered from 0.5x to 0.05x—a 90% reduction. From 10 p.m. to 8 a.m. Singapore time (11 p.m. to 9 a.m. Japan time), the multiplier drops further to 0.01x, a 98% reduction. No end date has been set for the campaign.

This discount applies to the Credit consumption multiplier within Qoder. It does not mean that the per-million-token API price has dropped by 90% or 98%. The consumed amount is counted against Qoder's own contracted quota, and this is not a campaign that distributes additional Credits. Expert mode and some sub-agent calls may be excluded from the discount.

The limited-time pricing for the individual Token Plan is $6 per month for Lite, $18 for Standard, and $68 for Pro. Even the cheapest Lite plan has two caps: 700 Credits per 5 hours and 2,500 Credits per 7 days. Individual API keys are limited to interactive use within tools like Claude Code and Cursor, and cannot be used for automated scripts, application backends, or non-interactive batch processing.

In other words, Qwen3.8 at this stage has not been rolled out as a general-purpose, pay-as-you-go API. Instead, a pathway has been set up first through monthly plans and in-house agent products, allowing users to try out real long-running tasks. The steep Credit discounts, too, are applied specifically to this trial pathway.

AD

The Open-Weight Contest Begins Only After Distribution

It would be premature to call Qwen3.8 an open-weight model at this point. What Qwen has announced is a plan to release the weights—there is no distribution file or license yet. Neither a model card nor a technical report has been published. The stage has not yet been reached where users can bring the model into their own environments, choose quantization and inference infrastructure, and rerun the same tasks.

The materials that need to be confirmed at release time are clear: the number of active parameters used per token, the expert architecture, the quantization format, the required computing resources, a reproducible evaluation procedure, and the license governing commercial use and modification. The per-usage API pricing outside the Token Plan is also essential for comparing self-hosted operation against cloud usage.

If Alibaba announces a weight release date and provides this information alongside it, the 2.4-trillion-parameter scale will, for the first time, become useful material for deciding on self-hosted deployment. As for the claim of being "second only to Fable 5," it can be verified even before the weights arrive—simply by aligning the version compared, the amount of thinking involved, the execution environment, and the number of trials. What begins after the weights are distributed is the evaluation of whether the same results and operating costs can be reproduced in one's own environment. What's needed for the next judgment is not a new adjective, but numbers and distributed materials measured under consistent conditions.