Timothy Proctor and colleagues at Sandia National Laboratories in the US have proposed QUOPS, a new metric that measures the computational capability of quantum computers against a common standard. In a paper published on September 10, 2026, they tested four machines from Google, IBM, and Quantinuum, measuring how large a computation each can run and how fast it can process it.

The results show that the quantum computer able to run the largest circuits is not necessarily the one that processes them fastest. Compared with the computational scale required for representative practical problems, today's machines are still about five orders of magnitude short.

The proposal is to measure progress in quantum computing not by how many qubits a machine has, but by how large a computation it can carry through to the end correctly.

AD

Measuring the circuits a machine can run, not its qubit count

QUOPS evaluates the overall computational capability of a quantum computer as a system, not just its number of qubits or the performance of individual gates.

Even with many qubits, errors that accumulate mid-computation prevent long calculations from finishing. And when qubits that are far apart must interact, additional operations are needed, which increases the number of operations and gives errors more opportunities to creep in.

For that reason, qubit count or individual gate fidelity alone makes it hard to judge how large a program a machine can actually run.

The existing "quantum volume" metric also measures whole-system performance, but the research team points to a problem: the classical computing load needed to verify results grows exponentially with the number of qubits.

If evaluating a future large quantum computer required simulating its entire operation on a classical computer, the evaluator's own computing power would hit its limit first.

The QUOPS paper presents an evaluation method that avoids this problem while retaining a connection to practical computation.

The tests use gates that apply an arbitrary-angle rotation to a single qubit, and CNOT gates that act on two qubits. These are placed randomly, and performance is measured while varying the circuit's "width" and "size."

Width is the number of qubits used in the computation, and size is the amount of operations counted by a common standard. A single-qubit rotation counts as 1 and a CNOT as 2, so the instruction counts that differ from vendor to vendor are not compared directly. The overhead of converting a common circuit into instructions each machine can execute is also reflected in the results.

Under the standard evaluation, a circuit counts as a success if it can be confirmed, at a 95% confidence level, that the "average process polarization," which indicates how much of the ideal circuit's behavior is preserved, is at least about 61%.

However, this 61% does not mean that "the answer to a practical problem is correct with 61% probability."

The evaluation uses mirror circuits: after the computation, the inverse operations are appended and further randomized, so the ideal output is known in advance. Because the quality of the original circuit can be estimated from how far the output has degraded, there is no need to reproduce the entire large circuit on a classical computer.

To condense QUOPS into a single score, with circuit width w and size s, the largest s that runs successfully within the range

w² ≤ s ≤ w³

is adopted.

This excludes extremely long, thin circuits and, conversely, overly shallow ones, narrowing the scope to forms close to early practical quantum computation.

For example, at width 6, the candidate circuit sizes for the score run from 36 to 216.

But the maximum value alone does not reveal how far a machine can compute at other widths. QUOPS therefore requires checking not only the maximum score but also the breadth of the underlying "executable region."

Helios excels at large circuits, Willow at speed

Under standard conditions, Quantinuum's Helios-1 reached the largest circuit size. Google's Willow, on the other hand, was far ahead in effective throughput per second.

Machine Type QUOPS score Effective throughput (QUOPS/s) Circuit width at score
Google Willow Superconducting 216 20,000,000 6
IBM ibm_boston Superconducting 204 310,000 6
Quantinuum H2-1 Ion trap 1,320 353 12
Quantinuum Helios-1 Ion trap 1,504 303 16

The figures come from Table I of the paper published in September 2026. The comparison uses physical qubits directly, without discarding trials in which errors were detected.

Throughput is measured at the point where each machine achieved its QUOPS score, so it is not a direct comparison of speed when running exactly the same circuit.

The circuit width in the table is also not the total number of qubits a machine has. Willow, for example, has 105 qubits and Helios-1 has 98, but the scores in the table were obtained with circuits using only some of them.

The strength of superconducting systems is fast gate operations.

However, there are limits on which qubit pairs can be directly connected, and operations between distant qubits require additional steps.

In Quantinuum's ion-trap machines, by contrast, ions can be moved so that operations can be performed between a wide range of qubit pairs. The connectivity penalty is therefore small, but each operation takes longer.

In QUOPS, these differences between architectures show up in two aspects: the circuit size a machine can run and its processing speed.

The meaning of "per second" also deserves care.

QUOPS/s is not simple gate speed; it is an effective throughput that also reflects degraded circuit quality and the overhead of discarded trials. It does not indicate the time between submitting a job to a cloud service and receiving the result.

The paper says that waiting time and data transfer time, which can be excluded from the evaluation, are not counted.

The supplementary material adds that for Willow's speed measurement, 500,000 executions were run to reduce the influence of circuit-switching time.

At IBM, circuits of different shapes ran within the same batch, making it difficult to isolate exact execution times per circuit, so time was approximated per batch. The time actually used by the quantum processor was adopted instead of total elapsed time.

In other words, even with a common metric, the conditions for measuring processing time are not perfectly identical. When comparing quantum computers, these measurement conditions need to be checked alongside scores and speeds.

AD

Error mitigation enables larger circuits, but at lower speed

"Error mitigation," which runs a noisy quantum computation many times and processes the results statistically to suppress the effect of errors, may make it possible to extract useful information even from circuits that would not count as successes under standard conditions.

However, the weaker the remaining signal, the more trials are required.

In QUOPS, with the polarization threshold denoted α, this overhead is reflected in throughput as roughly 1/α². Lowering the threshold to 1% could require up to about 10,000 times as many trials.

Helios-1 evaluation condition Circuit size (QUOPS) Effective throughput (QUOPS/s) Nature of value
Standard evaluation, no trials discarded 1,504 303 Measured
Error mitigation assumed down to 1% polarization 20,442 0.29 Estimate combining measurement and extrapolation

For Helios-1, the standard conditions gave 1,504 QUOPS and 303 QUOPS/s, whereas assuming error mitigation the estimate is 20,442 QUOPS and 0.29 QUOPS/s.

The circuit size that can be handled expands by about 13.6 times, while effective throughput falls to about 1/1,045.

These ratios were calculated from Table I and the error-mitigation estimates in v1 of Proctor et al.'s paper as 20,442 ÷ 1,504 and 303 ÷ 0.29.

They compare the largest circuit size reachable on the same quantum computer under different success criteria. They therefore do not mean that the same circuit becomes about 1,000 times slower, nor do they directly indicate the processing speed of a practical application.

The values using error mitigation also include extrapolation beyond the measured range. The supplementary material likewise cautions against treating them with the same confidence as results from the region actually measured.

"Post-selection," which discards trials in which errors were detected, is a separate method.

On Helios-1, detecting that a qubit had left its intended computational state and excluding that trial yielded measured values of 1,824 QUOPS and 247 QUOPS/s.

Because part of the results is discarded, extra overhead arises, but the quality of the data obtained improves, and in some cases effective throughput improves as a result.

When a large QUOPS score is reported, it is therefore necessary to distinguish whether the quantum computer itself improved, the success criterion was relaxed, or trials were selected.

Practical computations need a scale on the order of 250 million QUOPS

To see how far current quantum computers are from practical computation, the research team chose two examples: factoring RSA-2048 and energy calculation for a molecular system called FeMoco.

FeMoco is the active center of an enzyme involved in nitrogen fixation, and the target here is a calculation of energy for selected electron orbitals. It does not mean the entire reaction mechanism of the enzyme would be elucidated on a quantum computer.

Computational task Logical data qubits Target scale in QUOPS Target throughput assuming completion within 5 days
Factoring RSA-2048 1,399 250 million About 5,700 QUOPS/s
FeMoco energy estimation 1,459 340 million About 800 QUOPS/s

The source is Supplementary Section V and Table II of v1 of the QUOPS paper.

Neither figure comes from an actual calculation on a quantum computer; they are resource estimates for existing algorithms converted into QUOPS.

The qubit counts in the table are also the number of logical qubits that hold data, not the total number of physical qubits needed for error correction and auxiliary processing.

For RSA, Craig Gidney's 2025 study on which the figure is based gives a design estimate that, assuming a physical gate error rate of 0.1% and an error-correction cycle of 1 microsecond, among other conditions, the computation could finish within a week using fewer than one million physical qubits.

The QUOPS paper calculated the required processing speed from the condition of running a circuit of about 12 hours an average of 9.2 times.

The study by Guang Hao Low and colleagues underlying the FeMoco figure does not specify a computation time. The "within 5 days" condition was therefore set by the QUOPS team for comparison.

Comparing 250 million and 340 million with Helios-1's standard score of 1,504 gives ratios of about 170,000 and 230,000, respectively.

This is what the paper means by a gap of "about five orders of magnitude."

That said, it is not simply a matter of increasing the number of qubits by a factor of 170,000 or 230,000.

On small circuits, Willow records QUOPS/s exceeding the practical computation targets shown in the table. However, it cannot run a huge circuit of the required width and size correctly to the end.

Calculating quickly and sustaining a large computation without errors over a long time are different capabilities.

The conversion to QUOPS also rests on several assumptions.

The RSA and FeMoco computations make heavy use of Toffoli gates, whereas the QUOPS evaluation circuits mostly use arbitrary-angle rotation gates.

The team compared the two by decomposing both into a common operation, the T gate, which is costly to implement in fault-tolerant quantum computing.

In doing so, they ignored errors arising in Clifford gates and errors that occur during idle time, and simplified the structural differences between real algorithms and random circuits.

The figure of 250 million QUOPS is therefore not a "practicality threshold" common to all quantum computers. The targets shown here are estimates based on the paper version published on September 10, 2026.

The definition of QUOPS reveals another important point.

A score below the target means that, in this conversion model, the machine has not reached the required computational capability. But merely exceeding the target does not guarantee that an RSA circuit using 1,399 logical qubits could actually be run.

The width of the circuit that produced the maximum score may differ from the width needed for practical computation.

To judge whether practical levels have been reached, one must check not just the maximum score but whether circuits of the target width and size fall within the executable region.

AD

Logical qubits can be evaluated by the same standard

In an experiment using logical qubits with error correction built in on Helios-1, the machine recorded 40 QUOPS and 4.9 QUOPS/s.

Using the Steane code, which distributes information across multiple physical qubits to protect it, the team built up to eight logical qubits, and the score came from a width-4 circuit.

This is far lower than the 1,504 QUOPS obtained using physical qubits directly, but that cannot be taken as a limit of error correction itself.

The configuration used was relatively simple and small, and not an implementation that extracts the maximum fault-tolerant performance possible on Helios-1.

What is noteworthy is that improvement on the same hardware could be tracked through QUOPS.

The score, initially 24 QUOPS, rose to 40 QUOPS by improving the quality of the auxiliary states needed for operations.

Furthermore, by revisiting where error correction is inserted and how operations are parallelized, while maintaining conditions under which errors do not spread in an uncorrectable form, effective throughput also improved from 1.3 QUOPS/s to 4.9 QUOPS/s.

This makes it possible to evaluate not merely whether logical qubits could be implemented, but how large a computation can be run and at what speed, including the extra processing error correction requires.

According to Sandia's explanation, QUOPS can be used from tools such as CUDA-Q and pyGSTi.

Quantinuum has also published example implementations using Guppy and pytket. However, the notice in the public repository states that the notebooks are intended for explanation and do not strictly reproduce all the experiments in the paper.

To calculate an official QUOPS score, one must follow the statistical evaluation procedure defined in the paper, not merely check whether individual circuits ran.

The QUOPS paper is currently pre-peer-review, and its authors include researchers from Quantinuum and NVIDIA.

Quantinuum itself says that benchmarks tailored to real applications and third-party verification will still be needed.

In an interview with IFLScience, Proctor also acknowledged that as quantum technology advances, the current QUOPS evaluation conditions may cease to be appropriate. For that reason, he said, the components used in evaluation are designed so they can be changed in the future.

In judging the performance of future quantum computers, what will matter is how far the executable region expands when logical-qubit configurations are improved, and whether sufficient processing speed can be maintained at that scale.

If third parties can reproduce the results under the same conditions, and scientific computations that fall within that executable region can actually be run, QUOPS could become one criterion for research institutions and companies adopting quantum computers to choose "a machine that can run the computations we need."