An unconfirmed report that Anthropic's custom AI chip will be manufactured on Samsung's 2nm process spread on October 7, 2026. According to the report, Broadcom is involved in the design, and the chip is aiming for roughly 4,000 TOPS of compute and a thermal design power of 1,300W per chip. The race to secure the computing resources that run Claude may be moving from buying existing chips to designing dedicated hardware. But the name of a manufacturing process alone does not tell us whether a chip will be fast, cheap, or available at the scale needed. How the memory is connected, and what power and cooling infrastructure the chip runs on, will also decide whether it pays off.

AD

An unconfirmed design: 4,000 TOPS and 1,300W

The report originates from posts by Jukan and SemiconductorsX. It says Broadcom would handle design and implementation of an ASIC for Anthropic, which could be manufactured on Samsung Foundry's SF2. An ASIC is an integrated circuit designed for a specific use. Unlike a GPU, which handles a wide range of workloads, it can optimize its circuits and data flow for a narrower set of tasks.

The specifications reported so far are below. None of them is a confirmed product specification.

Item Unconfirmed information
Design and implementation Broadcom
Manufacturing candidate Samsung Foundry, SF2 2nm process
Compute performance About 4,000 TOPS, reportedly INT8
Thermal design power (TDP) About 1,300W per chip
Memory 192GB, up to 288GB of HBM3E depending on design conditions
Initial production volume About 200,000 chips

The text of the original posts could not be checked directly, and the specifications in the table are limited to information cross-checked through articles that cited them. Articles that spread from the same posts do not count as independent corroboration. As of October 8, 2026, no official announcement confirming these specifications or a manufacturing contract has been found. The fab location and the timing of mass production or deployment are also unconfirmed.

The relationship between Samsung and Anthropic itself, however, has been made public. In its May 28 funding announcement, Anthropic named Samsung, along with Micron and SK hynix, as a strategic infrastructure partner. But a relationship supporting memory and computing infrastructure is a separate matter from a manufacturing order for a custom ASIC.

The same distinction applies to Broadcom. Anthropic's April 6 announcement describes securing multiple gigawatts of next-generation TPU compute capacity from Google and Broadcom, to come online starting in 2027. That is a contract to use Google's AI semiconductors, and it cannot be used to confirm a contract to manufacture an Anthropic-specific chip.

If Samsung is chosen, 2nm and packaging come together

Samsung's official process listing states that SF2 mass production began in 2025. SF2 is a GAA process using second-generation MBCFET. GAA is a transistor structure in which the gate surrounds the channel that carries current, and it carries forward technology Samsung has used since its 3nm generation. The move to 2nm is not the first time Samsung has adopted GAA.

For AI accelerators, after the compute circuitry is built, memory must be placed close to it and fed large volumes of data. Samsung's I-Cube is a 2.5D packaging technology that places chips side by side on an intermediate component called an interposer and connects them with short wiring. Because Samsung handles memory and advanced packaging in addition to contract manufacturing, it can coordinate design and delivery schedules between components, which is an advantage as a candidate.

This combination has a publicly announced case involving a Japanese company. The customer confirmed in Samsung's July 9, 2024 announcement as adopting 2nm and I-Cube S is Preferred Networks, and it is not an announcement supporting adoption for Anthropic. That announcement explains that Samsung would provide the manufacturing process and packaging together, with GAONCHIPS handling design.

Samsung has already pitched this arrangement of offering advanced process and packaging together to customers. But the announcement of the Preferred Networks project does not mean shipments were completed or performance was demonstrated. Still less does it show that Samsung can meet the yield and supply volumes required for another large ASIC. That has to be verified through the design and mass-production process of that particular product.

Even in the Anthropic case, adoption of I-Cube has not been decided. Even if Samsung were to take on manufacturing, the HBM would not necessarily be Samsung's as well. A single company having multiple technologies is not the same as the price and delivery schedule of a finished product built with them being settled, and verification remains to be done.

AD

288GB of memory, and the waiting time TOPS doesn't capture

In its February 2024 announcement of HBM3E development, Samsung presented a product stacking 12 layers of DRAM at 36GB per stack. Separately, an official explanation of packaging based on SFF 2022 describes an I-CubeS 8 configuration carrying eight HBM stacks and two logic dies.

Calculating capacity alone, 36GB × 8 stacks = 288GB. The 288GB figure in the report can be explained by a combination like this. However, this is an arithmetic example using a memory capacity and a package stack count published at different times, not a result confirming that the two are compatible or that Anthropic has adopted them. "12 layers" refers to the number of DRAM dies stacked inside one HBM stack, and "8 stacks" refers to the number of HBM units placed on the package.

Large memory capacity matters not only because a model can fit, but also because of the number of requests that can be processed at once. In large language models, memory is used by the model's weights and by the KV cache, which reuses information from generation in progress. The more long inputs and concurrent requests are handled, the more capacity the cache tends to occupy. Even so, a larger capacity does not make responses faster in the same proportion.

NVIDIA's explanation of inference optimization separates the stage that processes the input in bulk from the stage that generates output sequentially. The former is easy to parallelize. In the latter, weights and cache must be read out to produce each next token, so the speed of moving data from memory tends to constrain latency. A token is a small unit of text that the model processes.

For that reason, 4,000 TOPS cannot be converted directly into Claude's response speed. TOPS is a measure of trillions of operations per second, and INT8 refers to 8-bit integer calculation. Comparing peak values also requires conditions such as computational precision, whether operations on data containing zeros were skipped, and how operations were counted. Without HBM transfer speed and the conditions under which the model is run, it is not possible to determine how much of the compute circuitry can actually be kept busy.

With batch processing, which groups multiple requests, the loaded weights can be used to advance many calculations and raise utilization of the compute circuitry. However, large batches consume memory and are also limited by the latency users will tolerate. And if a model is split across multiple chips, time for exchanging data between chips is added. Drawing out the performance of a dedicated chip requires aligning not only the circuitry but also the software that allocates computation and the design of communication.

The power needed for 200,000 chips, and whether a custom chip pays off

If all 200,000 chips in the reported initial volume are estimated at 1,300W each at the same time, the total is 260MW. The calculation is 200,000 × 1,300W ÷ 1,000,000 = 260MW. Even eight chips alone would add up to 10.4kW of TDP, a scale that demands substantial power and heat-removal capacity on the server side.

This is a reference value obtained by summing an unconfirmed TDP, not actual power consumption or data center power-receiving capacity. TDP is a figure that guides thermal design, and it is not a measurement showing that a chip constantly consumes that power. It is also not certain that all 200,000 chips would run at once in one location, and the power for networking equipment and cooling facilities is not included in this calculation. Even so, it shows that deployment is not complete once the chip's manufacturer is decided.

Even if a single chip uses a lot of power, the power per unit of processing could be kept down if it can generate more answers of the same quality at the same latency. Conversely, even with high peak performance, equipment will be underused if it spends long periods waiting on memory or communication. What should be compared is token throughput and power consumption with the same model, quality and latency conditions. Calculations must also include power and cooling facilities and operating costs in addition to purchase price.

Custom design also carries the difficulty of continuing to match future models. If a model's processing methods change, software work is needed to make use of circuitry optimized for particular calculations, and in some cases the design must be revisited. Whether development costs can be recovered through sufficient processing volume, and whether the chip can be used over a long period, are part of the economics. From the reported volume of about 200,000 chips alone, it cannot be concluded that the scale is small or that mass production will be profitable.

In its April official announcement, Anthropic described a policy of using AWS Trainium, Google TPUs and NVIDIA GPUs according to the workload, and named Amazon as its primary cloud and training partner. Even if a custom chip is added, evaluating it as one candidate among computing platforms chosen by use would be consistent with this published policy. This report cannot be taken to foreshadow a complete departure from NVIDIA or a price cut for Claude.

If the tape-out, the point at which the final circuit data is handed to manufacturing, and the timing of mass production and deployment are announced, it will be possible to judge how the plan is progressing. Only once real-world processing performance with the model, quality and latency aligned is also shown can we assess how much a 2nm custom chip would increase Claude's supply and how it would change processing costs.