On October 1, Amazon Web Services (AWS) released Strands Decider 2B, a small AI model that doesn't generate text. Instead, it evaluates a set of candidates and picks an answer.

The roughly 2-billion-parameter model can run on a local CPU or GPU. AWS has also published the code, the trained weights, and the procedures for retraining it.

TypeSafe's Jev and OpenAI's Decisions API offer similar judgment functions as cloud services. With Strands Decider 2B, the decision step itself runs in your own environment, and the model can be retrained if needed.

That flexibility comes with a cost: you have to verify for yourself whether the model correctly understands your questions and whether its judgments are stable enough to trust with automatic execution.

AD

A roughly 2-billion-parameter model that evaluates candidates instead of generating text

Strands Decider 2B is built on Alibaba's "Qwen3.5-2B-Base."

According to the design document, the team removed the language-generation output layer, which predicts the next word, from the base model of about 1.9 billion parameters. In its place, they added a "pointer head" of about 1 million parameters.

The model internally represents the input text and each option, compares their similarity, and calculates a score for each candidate.

A typical generative AI builds text by producing tokens one at a time. Strands Decider 2B generates no text; it computes a probability for each candidate in a single inference pass.

The base model has been further trained with LoRA. Rather than being able to produce free-form text or code, it is specialized for choosing an answer from a limited set of options, such as routing inquiries or selecting an agent's next action.

Three decision formats are available:

  • Choice: Select one option from several candidates
  • Score: Produce a rating on an ordered scale
  • Noul: Return the probability that the answer to a yes/no question is true

Noul is not a typo; it is the official name used in Strands Decider.

Candidate names and meanings can be specified per request, so you don't need to retrain the model just because you added a department to route inquiries to.

TypeSafe's Jev offers an API with a similar format. TypeSafe's documentation describes a mechanism that evaluates multiple questions about a single input at once.

However, similar input and output formats don't mean the underlying models or training methods are the same.

The confidence figures also call for caution.

In the example AWS gave, a payment-related inquiry is routed among three departments. The top candidate's probability is 0.845 and the confidence is about 0.768.

For Choice, confidence is calculated as follows, where N is the number of candidates and p is the highest probability:

(N × p − 1) ÷ (N − 1)

In this example:

(3 × 0.845 − 1) ÷ 2 = 0.7675

So a confidence of 0.768 does not mean "a 76.8% chance of being correct." It is a normalized value showing how concentrated the probability distribution is on a single answer: 0 if the probabilities are even, and 1 if they are completely concentrated on one candidate.

How it corresponds to actual accuracy has to be verified separately.

For Score, a different confidence value is calculated from the spread of the distribution. For Noul, the model simply returns the probability that the answer is true.

Compared with Jev and the Decisions API, the difference is running it yourself and retraining it

As of October 2, 2026, the publicly available information shows major differences in how Strands Decider 2B, TypeSafe's Jev, and OpenAI's Decisions API are offered.

Comparison Strands Decider 2B Jev OpenAI Decisions API
Delivery Obtain the weights and code and run them in your own environment TypeSafe-hosted API OpenAI-hosted API
Input Text, plus questions and candidates Text, JSON, etc. Text or images
Adapting to your business Questions and candidates can be changed; retraining is also possible using the published recipe Input, questions, and candidates specified per request Specify a question and a finite set of answer candidates
Availability For experimentation and local development Offered as a hosted service Limited preview as of September 29

For Strands Decider 2B, inference instructions and training instructions are published.

With Jev, users call the API and specify the input, questions, and candidates. With AWS's model, users can additionally change the training data and how the model is tuned.

This difference can be a reason to adopt it if you can't send internal documents to an external API, or if you want the model to learn judgments specific to your company.

On the other hand, if you want images to be used directly as decision input, the difference from OpenAI's Decisions API, which has been announced to support both text and images, matters. According to OpenAI, the Decisions API was in limited preview as of September 29, with general availability planned afterward.

The published training procedure is scaled so that it can run even on a consumer GPU.

According to the development team's records, training v19 from scratch took about 11 hours on a single RTX 3090, and about 1 hour 10 minutes on eight H100s with acceleration settings enabled.

However, the assumed environment for training the current Qwen3.5-family models is Linux or WSL2 with an NVIDIA GPU.

Inference works on Macs with Apple Silicon, but the same training procedure can't simply be run there. "Runs on a Mac" and "can be retrained on a Mac" need to be considered separately.

The same applies to the published data.

According to the data documentation, the repository includes some synthetic data and labels generated by models, while external public datasets are fetched and converted at the time of use.

Rerunning the existing training procedure doesn't require calling a paid model API. You do, however, need to obtain the base model, the teacher model, and the external datasets.

The code and distributed model-related files are released under Apache 2.0, but external datasets such as ContractNLI and MuSiQue have their own separate terms of use.

It would therefore be inaccurate to say that "both the code and the training data are all Apache 2.0."

AD

The "115 milliseconds" figure is a typical value for short inputs

According to the development team's measurements, v19 on an RTX 3090 processed one JevBench question in a median of 115 ms, with a 95th percentile of 299 ms.

A 95th percentile of 299 ms means that 95% of the measured runs finished within 299 ms.

Longer inputs, however, mean longer wait times.

Measurements of v19 on a Mac with an Apple M3 Pro and 36 GB of memory show a large gap between short and long inputs.

Input on M3 Pro Questions Median, first run Median, same input repeated immediately
Under 300 tokens 136 310 ms 153 ms
2,500–5,000 tokens 29 3,836 ms 2,514 ms
All 231 questions 231 622 ms 234 ms

Across all 231 questions, the 95th percentile was 3,862 ms on the first run, and 2,628 ms even when the same input was repeated immediately.

The test environment was macOS 26.6, MPS, and bf16.

One reason the first run is slower is that execution preparation for the Apple GPU occurs for each input size.

But with long documents, it still takes more than two seconds on subsequent runs. Even setting aside the one-time preparation, not every input runs as fast as a short text.

Classifying a short inquiry and making a judgment that requires reading several thousand tokens of internal documents are both a "single decision," but the amount of computation differs greatly.

Nor can these local measurements be compared directly with the response times of Jev or the Decisions API, which involve network communication.

Unless input length, number of questions, hardware, and network conditions are matched, the speed of the model itself can't be separated from differences in the execution environment.

A 72.3% accuracy rate alone doesn't show the model's true capability

During development, v19 correctly answered 167 of the 231 public JevBench questions at a 3,072-token input limit, an accuracy of 72.3%.

In the results by difficulty, accuracy was 100% on the 48 questions the development team classified as "easy," 87.5% on the 72 "standard" questions, and 50.5% on the 111 "hard" questions.

The model is highly accurate on easy classification, but its performance drops sharply on problems that require reading long texts and making multi-step judgments.

JevBench itself is an external public benchmark, but the 72.3% measurement was carried out by the Strands Decider development team. It should be distinguished from independent third-party replication.

The result also can't be taken to mean the model outperforms Jev or OpenAI's Decisions API, because the services were not compared directly under the same conditions.

The evaluation conditions for the distributed weights also need checking.

The v19 model published on Hugging Face uses 4,096 tokens as its standard input limit. The main pre-registered evaluation, however, was run at 3,072 tokens.

During development, v19 answered 167 questions correctly at 3,072 tokens, with a Brier score of 0.342 and an ECE of 0.052. Evaluated at 4,096 tokens, the published model answers 168 correctly.

Even for the same v19, numbers shouldn't be compared without matching input conditions.

Even more important is how the model responds to the wording of the question itself.

The evaluation materials describe tests in which only the question is changed while the same text and candidates are given, and in some cases the model did not adequately reflect the difference in the question.

For example, even when a question like "which one applies?" is inverted to "which one does not apply?", the model sometimes answers as it did for the original question.

This does not mean it "gets 94% of all questions wrong."

The problem is that when the condition in the question changes and the answer should change with it, the model doesn't adequately capture that change.

When you add business rules to a question, don't assume the model understood the condition just because you changed the wording. Test it with negations and opposite conditions as well.

The model card also indicates that performance may drop when the model is applied to types of binary questions or rating scales different from those used in training.

The same goes for confidence calibration.

For example, even if you decide on "automatic execution when confidence is 0.9 or higher" for one task, that threshold won't necessarily yield the same error rate for another task.

Indeed, the development team itself asks users to validate thresholds using real usage data.

AD

Separate the model's judgment from the conditions for actually executing an action

The official Strands integration example uses a scenario in which an AI agent asked about the weather tries to call a weather tool without confirming the city with the user.

The demo is set up so that the generative AI guesses the location and tries to run the tool.

Before execution, Strands Decider makes two yes/no judgments:

  • Is the value about to be passed to the tool grounded in something the user said?
  • Is it too early to run the tool at this point?

The Python code that receives those results then decides what to do next.

If there's a problem, it returns control to the generative AI so it can ask for the city; if not, it runs the tool. A mechanism for routing to human confirmation can also be built in if needed.

What matters is that Strands Decider itself does not make the final business decision.

The model returns a judgment on the question it was given.

Based on that result, the application's own rules determine which of the following to do:

  • Execute automatically
  • Ask the user again
  • Route to human confirmation
  • Reject the request

Keeping the model's judgment separate from the actual action makes it easier to change business policy and to trace the cause when a misjudgment occurs.

That said, the questions and thresholds used in the official weather demo are only examples, not recommended settings for production use.

The phrase "runs locally" also calls for care.

Strands Decider itself can run on a local PC, but the generative AI that handles the conversation in the official demo uses Amazon Bedrock.

The agent as a whole in the official example therefore doesn't run offline.

In addition, the bundled HTTP server is intended for local experimentation and has no authentication. What has been released is the decision-making model and its implementation, not a finished operational service with authentication and monitoring.

A "free model" doesn't mean zero operating cost

When comparing costs, it isn't enough to ask whether the model can be obtained for free.

Jev's published price is $0.042 per million input tokens, with no charge for output.

For example, if each input is 1,000 tokens and the API is called one million times, that totals one billion input tokens, and the input charge comes to $42 in a simple calculation.

This calculation uses only Jev's published unit price and excludes costs such as surrounding systems.

Running Strands Decider yourself, on the other hand, incurs costs for GPUs or CPUs, electricity, and server operations. Retraining the model also costs money for training GPUs, data preparation, and evaluation.

So the fact that the model itself can be obtained for free is not enough to conclude that total cost will be lower than a hosted API.

If you try Strands Decider 2B, a realistic starting point is a use case where right and wrong answers are easy to check, such as routing short inquiries to a handful of destinations.

Measure how often it wrongly permits automatic execution and how often it routes to human confirmation, including inputs in Japanese, with negations, and where no correct candidate exists.

Then check the wait times when handling long texts, and automate first the processes where the impact of a misjudgment is tolerable.

What sets Strands Decider 2B apart is not simply that it is a small, fast AI model. AWS has published the decision model's weights, code, and training procedures, giving users an environment in which they can repeatedly evaluate and retrain the model for their own purposes.