DeepSeek's new funding round is now expected to exceed 80 billion yuan, surpassing its initial target. On October 6, Bloomberg reported that CATL and Tencent have pledged some of the largest stakes, with negotiations in their final stage ahead of an initial public offering (IPO) planned for early 2027. To understand why so much money is flowing to a company known for low-priced AI, it helps to separate two things: lowering processing costs through model design, and building the infrastructure needed to deliver that processing at scale. The latest model's design and the release of software for Huawei's AI chips show a company trying to raise efficiency while also expanding its supply of computing resources.
Funding Could Top 80 Billion Yuan, but Talks Are Still Not Final
The company's initial funding target was about 50 billion yuan, and it is now reportedly in sight of securing more than 80 billion yuan. Based on signed term sheets, the final total could approach 100 billion yuan. Bloomberg also noted that terms could change during final negotiations. Investors expressing their intent to invest is a different stage from funds actually being paid in, and the two should be kept distinct.
The roughly 500 billion yuan valuation the company had initially envisioned is also a separate figure from the amount raised. It has not been confirmed as the post-funding valuation. CATL declined to comment, and neither Tencent nor DeepSeek responded. This is not an official corporate announcement but a report based on people familiar with the matter.
Input Processing and Cache Efficiency Underpin the Low Prices
DeepSeek's "V4.1-Flash," released on September 10, uses networks of different sizes for the stage that processes the input context and the stage that generates the response. This design difference matters for use cases where a large amount of text is read and a relatively short answer is returned. AI agents that read code and documents and keep working while receiving results from tools also tend to have a heavier processing load on the input side than on the output side.
V4.1-Flash enlarges the core model itself while reducing the number of parameters actually used for input processing. Comparing the official model card specifications with the earlier V4-Flash shows that the model's total size has grown while the number of active parameters per token is tuned by purpose.
| Official model card specification | V4-Flash | V4.1-Flash |
|---|---|---|
| Core model parameters | 284 billion | 552 billion |
| Active parameters for input processing | 13 billion | 8 billion |
| Active parameters for output generation | 13 billion | 16 billion |
The source is the "Base Model" comparison table in DeepSeek's V4.1-Flash model card, checked on October 6. The original notation of B (billion) has been standardized, and active parameter counts are per-token values. The core model figures do not include every component, such as separately provided memory mechanisms. Nor can processing speed or power consumption improvements be calculated directly from active parameter counts alone.
Another improvement is the reduction of the "KV cache," which stores intermediate computation results from the context that has been read. When continuing a long conversation or a series of tasks, recomputing the past context from scratch each time is costly. V4.1-Flash passes computation results obtained on the input side to the output side, and also shares and reuses the cache across layers. Compared with the earlier V4-Flash, DeepSeek says it has cut the high-bandwidth memory (HBM) capacity needed for the KV cache to one quarter and the SSD storage capacity to one eighth. However, this is a comparison of the KV cache portion only, and it does not mean that total server memory capacity or equipment costs fall by the same proportion.
Pricing also reflects the differences between input, output, and whether context can be reused. As of October 6, the official API prices per million tokens for V4.1-Flash are as follows.
| Billing item | Peak hours | Off-peak hours |
|---|---|---|
| Input with cache reuse | $0.006 | $0.003 |
| Input without cache reuse | $0.30 | $0.15 |
| Output | $1.20 | $0.60 |
Peak hours are 1:00–4:00 and 6:00–10:00 UTC on weekdays excluding Chinese public holidays, and off-peak rates apply all day on weekends and Chinese public holidays. Input whose past context can be reused from the cache becomes much cheaper, while newly read input and response generation carry separate unit prices. To compare the cost of running an AI agent, therefore, it is necessary to look not only at how much input can be reused but also at how long the responses are and how many times processing is repeated.
Huawei Support Links the Model to Computing Infrastructure
On September 30, DeepSeek added support for Huawei's Ascend to its compute library "TileKernels." It is designed so that NVIDIA GPUs and Huawei NPUs can be handled through the same Python interface and used selectively depending on the execution environment. In its public description, DeepSeek also thanks Huawei for technical support in implementing the Ascend version.
"DeepGEMM-Ascend," released the same day, is a matrix operation library developed and validated for the Ascend 950. In neural networks, matrix operations that multiply large numbers of values together are executed over and over, so the performance of this kind of foundational software is important for drawing out a chip's capabilities. Even if model weights are available, running them efficiently on different semiconductors requires compute implementations suited to that hardware.
Bloomberg also reported a plan to deploy at least 160,000 Huawei AI accelerators at a data center under construction in Inner Mongolia. However, this figure does not indicate the scale of facilities already in operation. The release of the compute libraries, by contrast, is a software-side development that can be verified separately from future facility plans.
That said, support for Huawei hardware in these libraries does not mean that every step from model training to inference has moved to Huawei chips. Securing semiconductors, building the necessary software, and running large-scale computing facilities reliably are each separate challenges. They cannot be solved with money alone; software and supply-chain readiness are required at the same time.
What Tencent and CATL Each Bring to the Relationship
Tencent's WorkBuddy and CodeBuddy are named as partner products adopting V4.1-Flash in DeepSeek's September 10 announcement. This means there were already touchpoints for delivering DeepSeek's models to actual users before the investment reports. For companies providing AI agents and development support services, a model that can be run repeatedly at low cost is also important when designing service pricing and usage volumes.
CATL has its own touchpoint in the power infrastructure that supports data centers. Its 2025 ESG report explicitly lists energy storage systems and energy management for data centers as uses of its batteries. Investment in AI data centers does not begin and end with the companies developing models. Equipment that supplies large amounts of electricity reliably is also required, and battery makers are part of that surrounding industry.
Given these existing businesses, Tencent could be involved in the growth of the AI market as a user of AI models as a service, and CATL from the power infrastructure side supporting data centers. However, it has not been confirmed that this is the purpose of the investment itself. Assuming arrangements such as priority supply in exchange for investment, or CATL supplying batteries to DeepSeek, would confuse publicly known business activities with undisclosed investment terms.
At the IPO, Balancing Open Models With Revenue Will Be Tested
The V4.1-Flash repository and model weights are released under the MIT license. Companies and developers can operate the model themselves, but running it reliably requires users to prepare considerable computing equipment and an operating environment of their own. The released model weights and DeepSeek's own paid API are different options for using the same model.
DeepSeek itself also invites inquiries from large-scale adopters able to prepare GPU and storage clusters on the order of 2,000 GPUs. This does not mean that 2,000 GPUs are necessarily required to deploy DeepSeek's models; it is guidance assuming large-scale operation. Even if the efficiency of the model itself improves, serving it to many users requires large-scale computing facilities.
According to Bloomberg, after the funding is completed the company is expected to move on to a corporate restructuring ahead of the IPO. The early-2027 listing timing also does not mean that a formal application or approval has been completed at this point. An important issue in evaluating the business going forward will be how it secures the revenue to sustain large-scale computing facilities while continuing to develop open models.
In assessing DeepSeek after the funding, what matters is not only who invested but also how much of the planned computing capacity is actually running, and whether current API prices and supply capacity can be maintained as users grow. If efficiency gains in input processing and caching can be tied to a stable large-scale service, developers will find it easier to use long-context work and AI agents that repeat processing many times at lower cost. Beyond a raise of more than 80 billion yuan, the question will be whether DeepSeek can supply enough capacity to support demand at that scale.
