US semiconductor startup Volantis announced on October 1, 2026 that it has raised $88 million in a Series A round to develop an AI inference system that connects compute chips and memory with light. For its first system, the A-1, the company says in the announcement that it is targeting generation of up to 10,000 tokens per second per user for models exceeding 20 trillion parameters.
The idea is to replace the wiring that feeds large volumes of data to the compute units with optical links, raising memory capacity and data transfer speed at the same time. The company is still in development, aiming for its first delivery in 2027, so the precision and system configuration used to run such giant models will be key to evaluating these performance targets.
The round was co-led by Lachy Groom and Abstract Ventures, with participation from John Doerr, Susa Ventures and others. According to the company blog, total funding now stands at $97 million, and earlier backers include Sam Altman and Jeff Dean. The new capital will go toward developing and commercializing the A-1 and hiring engineers.
The "memory wall" that leaves compute chips waiting
When a large language model generates text, the compute units read the model's weights and earlier processing results from memory. Even if compute power increases, the units are left waiting if the necessary data does not arrive fast enough. The "memory wall" that Volantis aims to break down refers to this limit on data supply.
The hardware characteristics required also depend on which stage of inference you want to speed up. NVIDIA's technical explainer says the stage that processes the input text in bulk is easy to parallelize, whereas in the stage that generates tokens one at a time, memory data transfer speed tends to determine latency.
When a single user is waiting on long reasoning or code generation, the speed of that latter stage largely determines response time.
SRAM built into the chip is fast, but its capacity is limited. HBM, used in GPUs, offers relatively large capacity together with high memory bandwidth, but as models grow larger and per-user generation speed is pushed higher, its data supply can fall short.
In its company blog, Volantis says it aims to change these capacity and bandwidth constraints with optical interconnects.
Batching requests from multiple users lets the cost of reading the same model weights be shared across several tasks, making it easier to use GPU compute efficiently.
However, "total throughput," meaning how many tokens per second the whole system can process, is a different metric from "per-user generation speed," meaning how quickly tokens are returned to a single user.
The "10,000 tokens per second" the A-1 advertises is a per-user target. Without details such as the number of concurrent users, it cannot be simply compared with the total processing performance of existing systems.
Optical routing beyond 200mm, using mass-produced VCSELs
According to Volantis's technology page, its optical waveguides can carry signals over distances exceeding 200mm.
An optical waveguide is a microscopic channel that carries optical signals. Volantis builds these into the substrate that connects chips, allowing memory to be placed farther from the compute chip.
The company puts the distance for high-bandwidth electrical connections around HBM at roughly 2–5mm. By extending connection distance with optical routing, it envisions linking more than 220 memory chiplets as a single memory pool.
When HBM is placed immediately around the compute chip, the amount of memory that can be installed is limited by the connection area available around the chip and by wiring distance.
If light can connect over longer distances, many memory devices can be placed farther from the compute chip. The bandwidth of many memory chips can then be used in aggregate.
Note that Volantis's 2–5mm figure describes packaging around HBM; it does not mean electrical communication itself only reaches a few millimeters.
The light source is an array of tiny VCSELs built into the substrate.
VCSEL stands for vertical-cavity surface-emitting laser, a semiconductor laser that emits light perpendicular to the chip surface. The A-1 is designed without external lasers or optical fibers, forming short-range optical links from many tiny VCSELs and optical waveguides.
The company says it aims to combine many links to achieve bandwidth exceeding 200TB/s.
VCSELs themselves are not new components; a large manufacturing base has developed around them through uses such as 3D sensing in smartphones.
In December 2017, Apple announced a $390 million award to Finisar to support expanded production of VCSELs used in Face ID and other features.
Volantis says it plans to use the existing manufacturing and supply base for gallium arsenide (GaAs) VCSELs, a supply chain distinct from optical components that depend on indium phosphide materials.
Being able to use established mass-production technology is an advantage, but whether production can scale to include connecting to fine optical waveguides and mounting many optical components on large substrates will need to be confirmed separately.
For the A-1's compute engine, the company says it will license third-party design IP that has already been proven in working silicon.
Based on the product description, Volantis's own development centers not on the compute units themselves but on the optical interconnect linking them to large-capacity memory. This differs from optical computing, which uses light itself to perform matrix operations.
Can a 20-trillion-parameter model run in 10TB?
The A-1 product page lists a memory capacity of 10TB and memory bandwidth of 240TB/s.
The unit measures 15U with a 20kW power envelope, so it is not a unit that can be directly compared with a single GPU chip.
| Item | A-1 published value | Conditions for reading the figure |
|---|---|---|
| Memory capacity | 10TB | Runtime working space is needed in addition to model weights |
| Memory bandwidth | 240TB/s | Sustained performance of the finished system cannot be confirmed from public information |
| Off-wafer I/O bandwidth | 10TB/s | Refers to a different target than the 240TB/s memory-side figure |
| Power envelope | 20kW | Differs from measured power consumption during inference |
| System size | 15U | Specification of the integrated system housed in a rack |
The source is the Volantis product page. These are the published specifications as of October 2, 2026, not results from real-system benchmarks run under identical conditions.
The 10TB/s figure refers to off-wafer I/O bandwidth and covers a different target than the 240TB/s on the memory side. Public information does not detail what it connects to, so it cannot be assumed to be bandwidth for communication between systems.
Let's apply the 10TB memory capacity to a 20-trillion-parameter model.
Storing only the weights of a 20-trillion-parameter model requires 40TB at 16-bit, 20TB at 8-bit, and 10TB even at 4-bit.
| Storage precision per weight | Weight capacity for 20 trillion parameters |
|---|---|
| 16bit | 40TB |
| 8bit | 20TB |
| 4bit | 10TB |
The byte count is calculated as parameters × bits ÷ 8, with 1TB taken as 10^12 bytes.
For 4-bit, for example:
20 trillion × 4 ÷ 8 = 10 trillion bytes
The weights alone reach 10TB.
Here we use 20 trillion as the lower bound of the phrase "more than 20 trillion parameters" in the October 1 announcement, changing only the storage precision for the same parameter count.
Moreover, this is the theoretical capacity for storing model weights only.
It does not include the auxiliary information needed to handle quantized weights, the KV cache that holds past inputs and outputs, or temporary runtime data.
If a model exceeds 20 trillion parameters, its weights alone exceed 10TB even when quantized to 4-bit.
Therefore, the "10TB of memory" specification alone does not support concluding that a single A-1 can run a model of more than 20 trillion parameters unconditionally.
An explanation of the actual configuration is needed, such as what precision the weights are stored in and whether multiple A-1 units are combined.
It is also necessary to distinguish a model's total parameter count from the amount of weights actually used when generating one token.
The Volantis product page lists Mixture of Experts (MoE) models among its targets.
In MoE, even if the model as a whole has a very large number of parameters, not all weights are used in each inference pass.
Therefore, the capacity needed to store all of a model's weights does not match the amount of weights actually read when generating one token.
By the same token, the 240TB/s memory bandwidth alone cannot be used to calculate that a model exceeding 20 trillion parameters can achieve 10,000 tokens per second per user.
Optical link results, and the real-system verification awaited in 2027
On its technology page, Volantis publishes data including eye patterns showing optical signal quality and a bit error rate below 10^-12 at wafer scale.
However, these are results for the optical link the company developed, not evidence that the full A-1 has run a giant model and achieved 10,000 tokens per second.
From public information, details of the measurement conditions and long-duration operating data cannot be confirmed.
Energy consumption also requires distinguishing what is being evaluated.
The announcement describes the energy needed from sending to receiving an optical signal as under 1 picojoule per bit.
However, this is a metric for the data transmission portion only.
When a model actually generates a token, reading data from memory and the computation itself also consume power. The overall power efficiency of the system therefore cannot be assessed from this optical link figure alone.
The product page also makes the claim of "15x more tokens per dollar" compared with NVIDIA Rubin.
It also states "6x more tokens per watt" for MoE models of over 1 trillion parameters running at low latency.
However, public information does not reveal the specific model names, weight precision, system configurations or pricing conditions used in the comparison.
At this stage, these should therefore be treated as Volantis's performance and cost targets or comparative claims, not as performance differences confirmed by real-system benchmarks under identical conditions.
In 2027, when the first customer systems are scheduled for delivery, real-system tests with matched conditions will be an important basis for judgment. They would need to specify model names and weight precision, plus the number of systems used, input lengths and concurrent user counts.
If per-user generation speed, total system power consumption and cost become clear from such tests, it will be possible to evaluate concretely how far Volantis's claims can cut the wait times of AI agents that use giant models.
