A research team including Jacky Kwok has released CLM-8B, a model that picks the most suitable action or answer from candidates given to an AI agent.
It processes the current situation and the options separately and compares them, and it can reuse the results already computed for the same options. Rather than having a large language model write out an answer every time, it narrows the job down to "choosing among the given candidates," aiming for fast decisions.
The team reports a speedup of up to about 9x compared with Jev, a decision model developed by TypeSafe. The code and trained model are released under the Apache 2.0 license.
However, the "9x" figure does not mean an AI agent's entire workload becomes 9 times faster. The actual effect varies greatly depending on which part of the work is handed to CLM, how much the candidates can be reused, and how much selection accuracy changes in exchange for speed.
Choosing from given candidates instead of generating text
CLM stands for "Contrastive Language Models," and it uses a method called contrastive learning.
The basic mechanism is to convert the current situation and the actions or answer candidates available in that situation into vectors, then compare them to see which candidate best fits the situation.
During training, the model is adjusted so that correct pairings of situation and action move closer together and incorrect pairings move apart. In use, it calculates how close each candidate is to the current situation and selects the one with the highest score.
Consider, for example, deciding whether to route a customer inquiry to the "billing team" or the "technical team."
The text of each inquiry changes every time, but the descriptions of the departments barely change:
- Billing team: handles fees, invoices, and refunds
- Technical team: handles malfunctions and outages
CLM can compute these candidate-side descriptions in advance and store them. There is no need to feed the same "billing team" and "technical team" descriptions back into the model each time a new inquiry arrives.
CLM's defining feature is that it can process the current situation and the options separately.
It is built on Qwen3-8B, which has 8 billion parameters. The weights of Qwen3-8B itself are frozen, and small projection heads that process the "state" and the "action" are added on top and trained.
The projection heads convert the text representations produced by Qwen3-8B into vectors that make it easy to compare situations and candidates.
The released CLM-8B was reportedly first trained on about 60 million question-answer pairs, then sharpened at telling similar candidates apart using about 30 million confusing wrong answers, and finally trained further on about 1 million AI agent action histories.
The research team calls a model that handles this kind of fast decision-making a "System One model."
Importantly, CLM itself does not create new text or draft answers.
What CLM returns is a ranking or probabilities for the candidates it is given. When the candidates themselves need to be created, another mechanism, such as a large language model, must be combined with it.
What the "up to 9x" figure actually measures
In the zero-shot evaluation published by the research team, meaning a comparison without additional training for each task, CLM-8B recorded shorter response times than Jev on every task.
On decision accuracy, however, it falls below Jev on some items.
| Evaluation task | CLM-8B latency | Jev latency | CLM-8B success rate / count | Jev success rate / count |
|---|---|---|---|---|
| T-Rex Game | 16.5ms | 149.8ms | 5/5 | 5/5 |
| Tool calling (BFCL v4) | 76.8ms | 125.5ms | 95.2% | 99.2% |
| WikiRacing | 79.8ms | 225ms | 26/30 | 30/30 |
| Super Mario | 33.5ms | 132.6ms | 5/5 | 5/5 |
Source: the research team's zero-shot evaluation. These are measurements by the authors themselves, not the results of independent third-party replication.
The figure corresponding to "up to 9x" is the T-Rex Game.
With 149.8ms against 16.5ms,
149.8 ÷ 16.5 ≒ 9.08
However, on BFCL v4 CLM-8B scored 95.2% versus Jev's 99.2%, and on WikiRacing CLM succeeded 26 times out of 30 while Jev succeeded all 30 times.
In other words, CLM being faster does not mean it produced better results than Jev in every use case.
In actual use, you need to look not only at the time a decision takes but also at how many more wrong selections occur as a result of the speedup.
The T-Rex "5/5" is not the model's score alone
The game results call for even more caution.
In the published T-Rex reproduction implementation, the task is to keep the dinosaur alive for 60 seconds in an environment imitating Chrome's dinosaur game.
However, the AI model alone is not looking at obstacles and judging when to jump.
A separate physics simulator calculates information in advance from factors such as the current speed, the distance to obstacles, and the model's response time, for example:
- Jump: safe
- Duck: dangerous
- Keep running: dangerous
The result is then passed to the model.
The model's job is to read those descriptions and choose from among the candidates.
A safety mechanism is also enabled in the default settings.
If the model selects an action the physics simulator judged dangerous, it is replaced with the safe action that has the highest probability. Emergency intervention may also occur just before a collision.
The public README itself states that the game's survival rate is a result for the whole system, including these assist mechanisms.
Therefore, the figure of 5 successes in 5 attempts on T-Rex cannot be interpreted as CLM itself fully understanding the game and beating it autonomously.
The measurement environments also differ.
CLM was run locally on an RTX 4090, while Jev was used through TypeSafe's hosted API, with response times measured on the client side.
This means it is not a benchmark comparing only the model's pure internal computation speed on the same hardware.
The "up to 9x" figure needs to be read as a comparison of end-to-end response times, including the published execution environments.
Not "coding ability" but "the ability to pick the right answer from candidates"
The DeepSWE and Terminal-Bench 2.1 results come from separate experiments from the zero-shot evaluation.
The 81.6% reported for DeepSWE is not a figure CLM achieved by writing code from scratch itself.
First, Opus 5 generates four candidate solutions for a single task. Then CLM, further trained for DeepSWE, evaluates the four candidates and picks the best one.
In an evaluation on 38 tasks not used in training, the answers CLM chose solved 31 tasks, giving
31 ÷ 38 ≒ 81.6%
The research team reports that when Jev was made to select candidates under the same conditions, the result was 71.1%.
Terminal-Bench 2.1 has the same structure.
Here, CLM or Jev picks one from five candidates generated by Fable 5. On 30 tasks not used in training, the authors report results of 87.6% for CLM and 83.1% for Jev.
Both are results on limited evaluation tasks, and do not mean that CLM alone achieved this performance on DeepSWE or Terminal-Bench as a whole.
For selection latency, CLM took 79ms versus Jev's 449ms on DeepSWE, and 32ms versus 131ms on Terminal-Bench 2.1.
The research team says it measured CLM on an H100.
What is shortened here is the time to choose one from several already-generated candidate solutions.
The time Opus 5 or Fable 5 takes to generate multiple candidates, along with the computation and API costs that requires, is not included in the 79ms or 32ms.
Furthermore, the model card for the released model states that reproducing the DeepSWE and Terminal-Bench results requires task-specific additional training.
It does not mean that running the base model released under Apache 2.0 as-is will produce 81.6% or 87.6%.
In addition, the research results published so far center on the authors' own technical materials and model card, and have not been confirmed by a peer-reviewed paper.
The more candidates can be reused, the easier the speedup
CLM's structure works especially well for processing that evaluates the same candidates over and over.
With an ordinary generative model, the current situation and the options are entered together, and an answer is generated each time.
Because CLM processes the current situation and the candidates separately, it can save the candidate-side results and use them repeatedly.
For inquiry routing, for example, candidates such as "billing team," "technical team," and "sales team" barely change. The candidate vectors, once created, can therefore be reused.
By contrast, when comparing long, entirely different candidate solutions each time, each one must be processed anew, so the benefit of caching is small.
In a separate measurement published by the research team, a fixed set of 50 candidates was evaluated on an RTX 4090.
When a different situation was entered each time, enabling an additional vector cache brought only a small reduction in median server-side latency, from 28.8ms to 28.1ms.
Under conditions where the same 20 situations were visited repeatedly, latency reportedly fell from 2.0ms to 0.7ms.
This is not a comparison with Jev but an experiment measuring the difference from adding a cache within CLM itself.
If the situation itself is new each time, its text must be processed by Qwen3-8B. Caching the candidate side therefore does not eliminate all of the computation.
Conversely, for processing where both situations and candidates can be reused repeatedly, the effect of caching may be larger.
The same "operation name" still requires recomputation if the description changes
Whether a candidate can be cached is not determined simply by whether the operation name is the same.
Looking at the public code, when a description is specified for an option, what CLM vectorizes is that description, not the operation name.
For example, even with the same operation name, "execute," if the content is rewritten each time, as in
- Execute this SQL
- Execute this Python code
- Delete this file
the candidate text itself changes, so it must be computed anew.
As a result, the size of the speedup differs between uses where the candidates are stable, such as department routing or a fixed list of tools, and uses that compare long, different candidate solutions each time.
CLM's "probability" is not absolute correctness
The meaning of the probabilities CLM returns also needs attention.
Even if candidate A gets 90%, that does not mean "A is correct with 90% probability in the real world."
CLM applies softmax to the scores of the given candidates and returns probabilities relative to that set of candidates.
Therefore, changing the content or number of candidates changes the probabilities.
The confidence returned by the public API is also a custom metric, calculated by subtracting the average probability of the other candidates from the probability of the top candidate.
Even if a high confidence is returned, it does not guarantee that the operation is actually safe and correct.
When executing processing autonomously, you need to combine it with mechanisms such as asking a human for confirmation when confidence falls below a certain level, and stopping dangerous operations with separate rules.
It doesn't run on just a "75MB model"
The required computing environment also needs attention when deploying.
The released CLM checkpoint is about 75MB, but this is not Qwen3-8B itself. It is data such as the projection heads added to the state side and the action side.
To actually run CLM-8B, the 8-billion-parameter Qwen3-8B must also be running underneath.
The official launch example sets up Qwen3-8B as an embedding model on a GPU with vLLM and places the CLM API server on top of it.
The standard launch example also sets an input limit of 2048 tokens.
Input beyond that is truncated, so if you want to use long conversation histories or large volumes of logs as material for decisions, the input limit and GPU memory settings need to be changed.
Even if candidate selection itself can be sped up, it is pointless if the information needed for the decision is dropped from the input.
Which kinds of processing benefit from CLM
To evaluate CLM's effect, you first need to find the "short decisions" your own AI agent repeats.
Examples include:
- Which department to send an inquiry to
- Which tool to call next
- Which of several search results to adopt
- Which of the answers created by multiple AIs to choose
- Deciding the next operation from routine action candidates
The more the candidates are fixed to some degree and the same decision is repeated in large volumes, the easier it is to take advantage of CLM's structure.
On the other hand, for work where the candidates themselves must be thought up from scratch each time, or where the correct answer can't be judged without long reasoning, it is not a replacement for a generative model.
When actually deploying it, you need to measure not only the decision time CLM can shorten but also the overall latency and cost including candidate generation, the rate of wrong selections, and the proportion of decisions deferred and handed off to a human or another model.
For uses that meet those conditions, CLM could be an option: instead of having a large language model generate a short decision each time, it carves out just the repeated "choosing work" and hands it to a fast, dedicated model.",
