ByteDance has begun pretraining an AI model that could reach as many as 10 trillion parameters. The Financial Times reported this on August 7, 2026, citing three people familiar with the matter. If the plan holds, the total parameter count would be about 3.6 times that of Moonshot AI's publicly disclosed Kimi K3. However, ByteDance's model has settled on neither its final scale nor its architecture, and Anthropic has not disclosed comparable figures. Ten trillion is not a finished performance benchmark—it is simply the upper bound of model scale currently under consideration.
10 Trillion Parameters Is Not a Finalized Spec
According to the Financial Times, ByteDance is pretraining a model with up to 10 trillion parameters. One source explained that this process typically takes three to six months, and if it proceeds smoothly, additional training will follow before the model is released. The final parameter count will also be determined at a later stage. As of August 7, 10 trillion is a candidate upper limit, not the specification of a completed model.
What the reporting confirms is only that development is in the early stages of pretraining. After completing this process, ByteDance is expected to apply additional training and release the model if things go well. No benchmark results or product-level availability terms have emerged yet. The total parameter count during training cannot be used to predict release timing or real-world performance.
ByteDance's officially announced Seed2.0 series consists of three lines: Pro, Lite, and Mini. Pro targets complex tasks requiring extended reasoning, Lite balances quality with response speed, and Mini aims at high-throughput reasoning and dense deployment. The company has not disclosed the parameter counts of these models on its public pages. Nor has it clarified the name of the model reported here or its relationship to Seed2.0.
Kimi K3's 2.8 Trillion Total, but Only 104 Billion Active Per Token
Among Chinese players with comparable official figures, Moonshot AI's Kimi K3 has disclosed a total parameter count of 2.8 trillion. ByteDance's candidate figure of up to 10 trillion would be about 3.6 times that. However, for Kimi K3, only 104 billion parameters are active when processing a single token. Translating total parameter counts directly into computational load or response performance risks misreading the actual situation.
This gap arises because Kimi K3 employs a Mixture-of-Experts (MoE) architecture. MoE selects a subset of "experts" depending on the input, giving the model large overall capacity without activating the entire model every time. Kimi K3 selects 16 out of 896 experts per token—a design that separates the model's overall capacity from the computational load at inference time.
Whether ByteDance's new model has a dense architecture or an MoE structure—and if MoE, how many experts are activated—remains unknown. The volume and quality of training data, the method of additional training, and the execution environment including tools are all undisclosed. Therefore, while 10 trillion versus 2.8 trillion may serve as a comparison of development scale, it does not translate into a claim of 3.6 times greater performance.
Missing Disclosures for Comparison With Anthropic
The Financial Times framed ByteDance's plan as a move approaching Anthropic's frontier models, and the paper also cited industry estimates comparing the two companies. However, Anthropic has not disclosed its parameter count. Furthermore, in its announcement on June 9, 2026, the company explained that Fable 5 and Mythos 5 use the same underlying model but differ in the scope of safety measures applied.
Given Anthropic's official explanation, it is difficult to place industry estimates on the same scale as ByteDance's candidate figures. What Anthropic publicly discloses are capability and safety evaluations, along with availability. It also provides pricing, but does not reveal model architecture or total parameter counts. Fable 5 is offered to the general public, while Mythos 5 is limited to a select group of vetted organizations. At least officially, the difference between the two lies not in the size of the underlying model but in safety measures and distribution methods.
Competitive gaps need to be measured through evaluations run under identical conditions. In agentic tasks, inference time and available tools—in addition to the model itself—drive the results. Scores can shift depending on the number of trials and the execution infrastructure used. Whether ByteDance has caught up to Anthropic can only be judged once the trained model is released and third-party evaluations under matched conditions are published.
A Long Game That Avoids Distillation, and What to Watch Next
This development approach carries a strategic bet that goes beyond mere scale. According to the Financial Times, ByteDance's foundation model division, Seed, has for over a year avoided distillation—the practice of using other companies' model outputs as training signals—and has instead pursued independent development. Founder Zhang Yiming reportedly told the team at an internal meeting in late July not to overly worry about short-term lags and to aim for "world-leading model capability" over the long term.
Distillation is a technique that uses the outputs of a strong model as training data to transfer capability to another model. ByteDance's leadership reportedly believes that independent development is necessary to surpass competitors.
What will matter for judgment is not the final total parameter count, but the scale active per token, the amount of compute used in training, and third-party evaluations. To assess competitiveness as a product, API pricing and regional availability will also be essential. After completing the pretraining process, said to take three to six months, how much will ByteDance disclose? The success or failure of the 10-trillion-parameter model will be determined not by the size of the number, but by whether it can outperform Anthropic and Kimi K3 under matched conditions.
