How many of the latest-generation GPUs a company can secure has long been treated as the deciding factor in AI inference infrastructure investment—and many companies have poured budgets into chasing that belief. But it's the companies holding what are dismissed as outdated chips that are increasingly finding themselves in an advantageous position. The $1.6 billion acquisition talks with Intel had collapsed in December 2025. That same SambaNova announced on July 8, 2026, that it had completed a first close of $1 billion in a Series F round led by General Atlantic, reaching a valuation of $11 billion. Benchmarks released simultaneously showed a configuration incorporating previous-generation NVIDIA H200 chips hitting up to 850 tokens per second.

AD

Why Did the Valuation Balloon Sevenfold in Under Seven Months After the Deal Fell Through?

In its 2021 Series D round (led by SoftBank Vision Fund 2, raising $676 million), SambaNova's valuation surpassed $5 billion. But the acquisition price Intel offered in December 2025 came to only about $1.6 billion, and the negotiations collapsed. Compared to the earlier valuation, that represented a plunge to less than a third—a figure that reveals just how precarious SambaNova's position had become amid the overheated AI infrastructure investment boom.

Two and a half months after the deal fell through, in February 2026, SambaNova raised over $350 million in a Series E round, with Intel joining as an investor at that point as well. Nearly five months further on, on July 8, 2026, the company announced it had secured $1 billion (approximately ¥162.4 billion, at ¥162.4 to the dollar) in the first close of a Series F round led by General Atlantic, reaching a valuation of $11 billion. Measured against the $1.6 billion valuation at the time of the collapsed deal, that's a recovery of roughly 6.9x in under seven months. Few recent AI infrastructure companies have seen their valuations rebound this far in roughly half a year.

Alongside existing investors Intel Capital and Vista Equity Partners, Seligman Ventures, T. Rowe Price, Capital Group, BlackRock, and QIA newly joined this round. SambaNova has disclosed that it expects to bring in additional investors and complete a second close within a few weeks. The participation of major asset managers like T. Rowe Price and BlackRock represents a lineup that wasn't present back in the 2021 Series D round, signaling that the pool of capital flowing into inference chip companies is widening.

Why Choose the Outdated H200 Over the Latest B300?

The benchmark configuration unveiled at RAISE Summit 2026 wasn't a lineup of cutting-edge hardware. SambaNova adopted a heterogeneous setup: four NVIDIA H200 units handled prompt preprocessing (prefill), while response generation (decode) was assigned to 16 of the company's own RDUs mounted in its SambaRack SN50. Running this configuration with MiniMax M2.7, a Mixture-of-Experts (MoE) model developed by China's MiniMax, the company recorded up to 850 tokens per second on short-context inputs and more than 450 tokens per second even on long-context inputs—results it says were verified by the third-party firm Artificial Analysis.

aa-sambanova-benchmark-0708.webp

Prefill is the process of batch-processing input tokens in parallel, a task where raw compute power for handling massive matrix operations matters most. Decode, on the other hand, is a sequential process that generates one token at a time, where the bottleneck is memory bandwidth for repeatedly reading out the model's weights. The computational demands of these two stages are fundamentally different. The H200, a general-purpose GPU equipped with 141GB of HBM3e memory, is well-suited to the parallel computation prefill requires, and SambaNova has chosen to leave that portion to its existing H200 assets rather than deploying its own chips there.

By assigning decode alone to the RDUs, the company argues it can achieve fast inference without needing to procure large quantities of expensive latest-generation GPUs. That said, SambaNova has not disclosed a specific speed multiplier comparing this setup against a configuration using the same H200s without RDUs involved (an H200-only setup). While the absolute figure of 850 tokens per second is provided, the comparison baseline showing how much performance would drop without the heterogeneous configuration is not included in this announcement. What this setup aims for is operational efficiency—extracting performance while holding back investment in the latest chips.

The H200 itself is an NVIDIA product and still counts toward NVIDIA's sales, but if it can be proven that combining outdated-generation GPUs with the company's own reconfigurable chips achieves inference speeds on par with the latest generation, the cycle by which companies feel compelled to upgrade to the newest chips could slow. SambaNova's configuration—confining the H200 to the limited domain of prefill while concentrating the more profitable decode processing on its own RDUs—can also be read as a move chipping away at a corner of the inference market that NVIDIA has long dominated. What this configuration demonstrates is not a convenient fact for NVIDIA.

AD

What Distinguishes RDUs From GPUs?

GPUs have a fixed instruction set, and regardless of which model is being run, software executes instructions on top of the same circuit configuration. SambaNova's RDU (Reconfigurable Dataflow Unit), by contrast, has no predetermined instruction set. It reconfigures the tile-shaped compute units within the chip to match the computational graph of whatever model is being run, assembling circuitry anew each time to follow the actual flow of data. The two chips are fundamentally different in design philosophy from the ground up.

A GPU is like a general-purpose factory with a fixed production line, where every product must be adapted to fit that line. An RDU, in contrast, is a factory that reconfigures its production line itself each time an order comes in, which tends to boost processing efficiency by eliminating unnecessary steps. However, this reconfiguration requires a dedicated compiler and tuning, and the range of models it supports is more limited than what a GPU can handle.

SambaNova describes the advantages of this architecture as delivering 5x the performance and 4x the memory bandwidth compared to its previous-generation SN40, and states that a configuration combined with NVIDIA B300 achieves 10x the throughput. These multipliers come from the company itself, and independent third-party verification has not yet been confirmed. The dataflow-based design itself is not a new innovation unique to this announcement—SambaNova has consistently employed this approach since its founding. The $11 billion valuation now reached can be read as confirmation that this design philosophy is finally becoming economically viable in the inference market.

JPMorganChase's Adoption and Intel's Dual Role

JPMorganChase has adopted the SN40 and SN50 as its inference partners for on-premises environments, becoming one example of a financial institution running high-speed inference infrastructure within its own data centers. The financial industry, for regulatory reasons, tends strongly to avoid inference via cloud APIs and instead demands on-premises processing—an area that presents a natural business opportunity for dedicated chip makers like SambaNova. The same dynamic applies to Japanese financial institutions and data center operators considering on-premises inference infrastructure under domestic regulations. SambaNova's chips, alongside this announcement, have also been backed up by a concrete track record of real-world adoption.

After the acquisition talks collapsed in December 2025, the two companies signed a multi-year agreement for Xeon-based AI inference development, and Intel also invested in the Series E round. As a result, Intel's stake reportedly rose from 6.8% the previous year to 8.2% as of February 2026. In addition, Intel CEO Lip-Bu Tan also serves as SambaNova's chairman, creating a structure in which the profits from the surge in valuation of a company Intel failed to acquire are instead being recouped through its position as an investor. Even though the acquisition itself failed, the decision to deepen a technology partnership by increasing its equity stake has, as a result, also funneled the upside of this $11 billion valuation back to Intel.

AD

What Numbers Will Be Asked Next, Ahead of a 2027 IPO

Regarding this funding round, CEO Rodrigo Liang said, "An $11 billion valuation shows that fast inference has come to play a central role in enterprise AI stacks," and reportedly positions a US IPO in 2027 as a strong option. In the valuation race among inference-specialized chip companies, Cerebras—which raised $1 billion in a Series H round in February 2026—has reached a valuation of $23 billion, putting SambaNova's $11 billion in second place behind it. What both companies have in common is that they've stepped back from the performance race in training chips and are instead building up their valuations through efficiency in inference, the practical operational layer. While competition for training chip procurement remains dominated by NVIDIA and AMD, the valuations of these two companies confirm that room still remains for dedicated architectures to carve out a place in the downstream process of inference.

With its figure of 850 tokens per second, SambaNova has demonstrated the headroom for performance gains achievable simply through how existing chips are combined. Converted to Japanese yen, $1 billion corresponds to roughly ¥162.4 billion and $11 billion to roughly ¥1.7864 trillion, and SoftBank—which led the 2021 Series D round—is also reportedly one of the partners rolling out the SN50 this time. If the US IPO expected in 2027 materializes, it will serve as the market's verdict on how far an operational model that pairs older-generation GPUs with reconfigurable chips can be valued. Whether comparison data against an H200-only configuration is disclosed in the IPO filing documents should also form part of that verdict. What is starting to become an evaluation criterion for investors is the operational skill of how to combine and squeeze every drop out of the chips already on hand—and what this valuation recovery shows is the fact that specification-sheet cutting-edge status is not the deciding factor.