Reflection, a U.S. company backed by NVIDIA, announced on October 5 an AI model called "Beam," built for writing code and operating tools. The model has 501 billion parameters in total and uses 23 billion of them to process each token. Reflection says it ran four weeks of reinforcement learning on 10,500 GB300 GPUs.

The company says Beam can approach the reasoning ability of its competitor GLM 5.2 with less compute, and plans to release the model weights within October.

Can a large training investment be turned into a model that is easy to use in practice? To judge that, you need to look at the training infrastructure and also separate compute, memory use, and usage limits.

AD

Four weeks of large-scale reinforcement learning trials

Beam's pretraining used 23.8 trillion tokens. According to Reflection, it completed pretraining, which learns from web data, licensed data, and other sources, in under four weeks using 6,144 GB300 NVL72 GPUs.

It then went through mid-training, which handles long documents and complex tasks, before moving on to large-scale reinforcement learning. The "10,500 GPUs for four weeks" figure in the official announcement refers to this reinforcement learning stage. It does not mean the entire development of the model took four weeks.

If pretraining is the stage in which a model reads code and technical documents and acquires broad knowledge, reinforcement learning is the stage in which it actually works on tasks and sharpens how it solves them based on the results.

For example, when fixing a software bug, the model examines files, rewrites code, looks at test results, and revises further. What it learns from is not whether the answer text sounds plausible, but whether the task was actually completed.

The scale Reflection disclosed is easier to understand when you separate what each number counts.

Disclosed figure What it counts
About 1 million task environments Tasks in programming, science and technology, and other areas, along with their execution conditions
Over 100 million attempts Sequences of reasoning and actions in which the model worked on a task
About 1.3 billion sandboxes Isolated execution environments used for training and scoring
Up to 170,000 concurrent sandboxes Execution environments that could run in parallel at one time

Source: Reflection's Beam announcement. Task environments, attempts, and sandboxes are each counted in different units.

Dividing about 1.3 billion by four weeks, or 28 days, gives roughly 46.4 million per day.

However, this is a simple period average calculated from the disclosed figures, not a measured daily value. It also does not mean that 1.3 billion distinct problems were created, or that 1.3 billion execution environments ran at the same time.

The detailed breakdown of how sandboxes were divided between attempts and scoring has not been disclosed.

More important than the number of tasks is what the model is made to learn.

The roughly 1 million task environments were built mainly from synthetic data, combined with data from outside vendors and public materials.

Tasks the model could almost always solve, and tasks it could not solve no matter how it approached them, were excluded. Tasks with ambiguous instructions, and tasks that could appear to be solved correctly by exploiting loopholes in scoring, were also removed.

The company explains that lowering task quality stalled capability gains and also caused problems in training.

Even with a large number of GPUs, if tasks of appropriate difficulty and correct scoring are not in place, the compute invested will not translate into the intended capability gains.

Beam's training scale needs to be viewed together with the process of creating, testing, and re-filtering tasks.

Training moves on without waiting for long attempts to finish

During reinforcement learning, an average of 110,000 attempts were reportedly in progress at the same time.

In a system that waits for some long attempts to finish before moving on, GPUs and training processes spend more time idle.

Reflection therefore ran the process in which the model works on tasks and the process of updating weights from those results asynchronously. Tool execution and scoring were also separated from a structure in which every process waits on the others to finish.

Reducing waiting time, however, creates another problem.

Between the time a model starts an attempt and the time it returns a result, the training-side model may have been updated many times. In a long attempt, the version of the weights used for generation can even change midway.

If actions produced by an old model are used for learning in the same way as actions produced by the current model, training may become unstable.

In Beam's training infrastructure, each generated token is recorded with the version of the model used at that moment.

This is so the learning algorithm can handle the gap between the model at generation time and the model at training time.

The company reports that incorporating experience delayed by about one day, or 107 weight versions, was still numerically stable.

The details of the new algorithm that handles this gap are something to check in the technical report to be published later.

The mechanism for delivering updated weights to many inference models was also reworked.

It uses RoCE for transfers between racks and NVLink within a rack, first distributing hierarchically and then sharing among nearby models.

Compared with each model fetching weights individually, this method reportedly cut inter-rack traffic by 75% and sped up propagation to the whole system by 2.2 times.

The median time for new weights to reach the inference side was about 12 seconds.

There are also measures for keeping the large experimental environment running without stopping.

Ninety percent of sandboxes could be prepared in under 10 seconds, and even when 71 failures occurred in the inference system during training, the training job itself did not stop.

Inference processing capacity recovered in a median of 8 minutes, according to the company.

It also set up a mechanism in which an independent judge rechecks whether a correct answer to a task was obtained by exploiting a scoring loophole.

The MoE (mixture of experts) design itself also includes measures to stabilize training.

MoE selects, for each token, which parts to use from among multiple specialized components. If the selection becomes skewed toward some of them, training load and processing load become skewed as well.

Reflection built on the auxiliary-loss-free load balancing used in the DeepSeek-V3 technical report.

It also added a measure to weaken the mechanism that corrects expert assignment toward the end of training.

In addition, it adjusted the model so that internal values do not grow excessively large as layers get deeper, and reportedly used FP32 for calculations that apply small updates, to limit rounding error.

The result combines a model design that incorporates existing MoE techniques with an infrastructure that keeps feeding large numbers of attempts into training without stopping.

The number of GPUs alone does not determine capability. Large compute resources can be put to use in training only when the mechanisms for generating attempts, scoring them, and quickly distributing updated weights are working.

AD

What is the "3 to 4 times more efficient" figure comparing?

Reflection says that on benchmarks evaluating advanced reasoning, Beam achieved results close to GLM 5.2 while holding inference-time compute to about one-third to one-quarter.

This comparison, however, does not measure actual usage fees or processing time on GPUs.

The announcement estimates the compute needed to generate an answer with the following formula.

Estimated FLOPs ≈ 2 × active parameters × average generated tokens per attempt

FLOPs represent the number of floating-point operations needed for the computation.

Here, a multiply-accumulate operation is counted as two operations, and generated tokens include both the reasoning before the answer and the final answer.

In MoE, the multiplier is not the model's total parameter count but the number of parameters actually used to process that token.

This formula shows that, if tasks of similar difficulty can be solved, both reducing the parameters used per token and reducing unnecessary generation lead to lower compute.

Beam was trained using a penalty tied to generation length, giving a reward when a task succeeded while reducing extra tokens.

According to the company, early in training generation became shorter while scores rose, and later, using longer reasoning as needed raised capability further.

The design does not aim simply for short answers, but leaves room to think longer on hard problems.

On the other hand, this compute estimate does not include the computation for initially processing the input.

Attention computation that grows with context length, and the various processing needed to run the model as a service, are also excluded. The weight of the excluded parts differs between loading a long piece of code and generating a long answer to a short question.

For that reason, the "one-third to one-quarter compute" figure cannot be read directly as "one-third to one-quarter the price" or "one-third to one-quarter the processing time."

The capability comparison also varies by benchmark.

Pulling the coding-related metrics from the table Reflection published gives the following.

Benchmark Beam GLM 5.2 Qwen 3.8-Max
DeepSWE v1.1 44.4 44.0 51.0
SWE Bench Pro v1 65.5 62.1 67.7
Terminal Bench v2.1 80.1 81.0 86.6

Source: Reflection's announcement table. Competitor values use evaluations from Artificial Analysis and DataCurve. Higher is better. This table alone does not confirm whether hardware and execution conditions are uniform across models.

Beam is above GLM 5.2 on some benchmarks and below it on others, and it falls short of Qwen 3.8-Max on all three of these.

Beam's aim is less to compete for the highest scores than to obtain a given level of capability with less generation compute.

However, how far that advantage shows up in real work will depend heavily on verification after the model weights are released.

Using only 23 billion still means storing all 501 billion weights

The figure of 23 billion active parameters per token does not represent the capacity needed to store the whole model.

In MoE, the specialized components used change from token to token, so the whole model, including weights not in use at that moment, must be stored.

Hugging Face's explanation of MoE likewise notes that, compared with a same-size model that uses all parameters every time, compute can be reduced, but the memory needed to hold all the weights is larger.

Using the total parameter count Reflection published on October 5, the capacity needed for the weights alone can be calculated simply.

If all 501 billion parameters are stored, that comes to about 501 GB at 8 bits and about 250.5 GB even at 4 bits.

The calculation takes 501 billion × bits per parameter ÷ 8 to get bytes, then divides by 1 billion to convert to GB.

This assumes all weights are stored uniformly at 8 bits or 4 bits.

It does not include the additional information needed for quantization, the KV cache that holds the processed context, or working space during execution.

It is neither a measurement of the memory required to actually run Beam nor an indication of what format the planned weight release will take.

Therefore, looking only at the "23 billion active parameters" figure, you cannot conclude that it can run in memory comparable to a small model.

Where to place all 501 billion weights, and how to move data to the selected specialized components, are also important conditions for deployment.

Even with a design that reduces compute, the burden on memory and communication does not fall by the same proportion.

The significance of releasing the weights is that these conditions can be measured against your own equipment and data.

You can measure response times and the number of jobs that can be processed at once for your actual use, which a compute graph alone cannot show.

AD

After release, the question is how much it costs to finish your own work

The "Beam-501B-A23B" API, as presented to developers, has a 256K-token context window and a 128K-token maximum output window.

According to the model specifications, the context window means the upper limit on input and generated output combined.

This refers to a different stage from the statement in the announcement that "mid-training extended it to 1 million tokens."

Treating the length handled during training as the length the current API can accept as input could lead to mistakes in designing systems that load large codebases.

On generation as well, it is not only the final answer that consumes the output window.

According to the reasoning specifications, reasoning effort has five levels from low to max, with medium as the default.

Reasoning is always enabled, and tokens used during reasoning count toward the output limit.

If the limit is too low, it can be used up midway through reasoning, and the final answer may come back empty.

Making the model think longer does not necessarily mean it will solve harder tasks more reliably. Increasing reasoning effort also increases waiting time and generated tokens.

In real operation, alongside task success rates, you need to measure how long jobs take at different reasoning levels.

Comparing only the price of a single short answer makes it hard to see the cost of completing a job while calling tools many times.

Availability is through a beta opened in stages from a waitlist, and at the time of the announcement it is a limited preview.

Reflection plans to release the weights under the Apache 2.0 license in October, along with a technical report and development-related materials.

Final safety verification is also continuing, and the company says it will publish safety evaluation results in the technical report.

The company explains that it separated training aimed at improving capability from training to shape safety and behavior, and integrated them by distilling from multiple teacher models.

How that design behaves in actual tool operation and long-running tasks is also something to verify going forward.

Reflection has previously set out a policy of building its own pretraining and reinforcement learning infrastructure and developing open models.

Beam gives concrete form to that policy, and its structure shows large compute resources invested up front in training in order to keep inference-time compute down.

Once the weights are released, what will matter is whether you can prepare your own code and tools and measure the time and cost to complete tasks at the success rate you require.

If that can be confirmed, Beam could become an option for companies that want to run AI on their own equipment and tune the balance between compute and work quality themselves.