YFarmX logoYFarmX

AWS Strands Labs

Strands Decider 2B

a small open decision model that runs on a gaming GPU or a laptop

Released 1 October 20265 min readLarge Language Models

Editorial collage headed Strands Decider, with a cast-iron railway points lever beside a forking track whose chosen branch is picked out in blue, a graphics card on the ballast, the AWS logo and a tag reading 115 ms; the subtitle reads AWS, free 2B model.

Key facts

1 Oct 2026Strands Labs, AWS
Released
1.9BQwen3.5-2B base plus adapter
Size
FreeApache 2.0, run it yourself
Price
115 msNvidia RTX 3090, makers' test
Median latency
167 / 231the makers' own run
JevBench public tasks
91.5 MBwhole repository; base fetched on first use
Download

Strands Decider 2B is a small decision model from AWS's Strands Labs: you give it some text and a set of fixed questions, and it returns a probability for each allowed answer. It is free under the Apache 2.0 licence and runs on your own machine, answering in a median 115 milliseconds on an Nvidia RTX 3090 graphics card. AWS released it on 1 October 2026 with its training code and data recipe, as a tool for experimenting with decision models inside AI agents.

Strands Decider 2B is a decision model: an AI model that reads a piece of text and answers a set of fixed questions with a probability for each allowed answer. It is small enough to run on an ordinary graphics card or a laptop, and it is free. Strands Labs, AWS’s organisation for experimental open-source projects alongside its Strands Agents SDK, released it on 1 October 2026 with its weights, training code and data recipe.

Three AWS staff wrote the launch post: Marc Brooker, a vice-president and distinguished engineer at AWS, Mike Chambers and Fabio Nonato de Paula. They describe it as “optimized for fast experimentation, local development, and innovation”. The package on PyPI calls it “an experimental typed classifier”.

It answers yes, no, a choice or a score

Strands Decider 2B answers three kinds of question from one pass over the input: a yes-or-no question, a choice of one option from a list, or a score against a rubric. Each answer comes with a probability. In the makers’ own example, a customer writes “Help! My payouts have been failing for 3 days!”, and the model routes the message to billing with a probability of 0.845, ahead of retail at 0.091 and sales at 0.064.

The intended job is inside AI agents. The launch post shows the model checking an agent’s planned weather-tool call with two questions: are the arguments grounded in what the user said, and is it too early to call the tool. The project’s README lists model routing, tool selection, argument checks, triage, guardrails and evaluation as uses. The makers set out its limits in the post: answering everything in one parallel pass makes it “significantly worse at solving complex problems than reasoning models”, and it is unsuited to coding, chatbots or summaries.

Inside the model: a Qwen torso with a new head

The team took Alibaba’s Qwen3.5-2B, removed the part that generates text, and replaced it with a small “pointer head” that scores each allowed answer. The head holds just over a million parameters. The base model, which the team calls the torso, is fine-tuned with a rank-16 LoRA adapter, so the download on Hugging Face is about 91.5 MB, and the base weights arrive the first time it runs.

Block diagram titled Strands Decider 2B architecture. On the left, a standard decoder language model whose language-model head and generated text are marked discarded. On the right, a Qwen3.5-2B-Base torso with a rank-16 LoRA adapter passes hidden states for three options and the answer into a small blue pointer head, which produces one logit per option and a softmax of per-option probabilities. A legend marks trained, pretrained and discarded parts.
How Strands Decider 2B is built: a fine-tuned Qwen3.5-2B torso feeding a small pointer head that scores each answer. Select the diagram to enlarge. Diagram: Strands Agents, 1 October 2026.

The model counts 1.9 billion parameters in all. The version released is v19, published on Hugging Face as strands-decider-2B-hobson-v19. The training stage took 1,685 seconds on an AWS p5.48xlarge server with Nvidia H100 GPUs, and the makers say a retrain takes about 11 hours on one RTX 3090. The data is public datasets plus synthetic examples written by one Qwen model and checked by a larger one, and reproducing it “needs no paid model-API calls”, the data notes say.

It runs in about a tenth of a second

On an Nvidia RTX 3090 graphics card, Strands Decider 2B answers in a median 115 milliseconds, with the 95th percentile at 299 milliseconds, the makers report. On an Apple M3 Pro laptop chip, small tasks take a median 153 milliseconds.

Measure, makers’ tests Result
Median latency, RTX 3090 115 ms
95th percentile, RTX 3090 299 ms
Median latency, Apple M3 Pro, small tasks 153 ms
JevBench public tasks correct 167 of 231 (0.723)
Easy, standard and hard tiers 1.000, 0.875, 0.505
Brier score and expected calibration error 0.342 and 0.052
Scatter chart titled Strands Decider 2B v18 latency against prompt length, for 230 JevBench requests on an Nvidia RTX 3090: latency from 0 to 500 milliseconds against input tokens from 0 to 3,000. Most points sit near 100 ms below 700 tokens and rise to about 230 to 320 ms near 2,300 to 3,000 tokens; dashed lines mark the median at 106 ms and the 95th percentile at 296 ms.
Latency against prompt length for an earlier build, v18, over 230 requests on an RTX 3090. Select the chart to enlarge. Chart: Strands Agents, 1 October 2026.

Third of 33 in its size class

In the Strands team’s own run, the model answered 167 of the 231 public JevBench tasks correctly, which places it third of 33 systems in the 2-billion-parameter class on the board of 25 September 2026, and 50th of 89 overall. Three models in its class, the two ahead of it and one behind, are just over 2 billion parameters, so the team also calls it first of the 30 at or under that size. The top of the overall list is reasoning models, led by OpenAI’s GPT-6 Luna at 0.996, followed by Jev-style models on far larger bases, the team’s evaluation notes say.

It answered every easy task correctly and about half of the hard ones, and the makers say long, multi-step documents are its weak spot. Its probabilities hold up on short tasks: answers given at a confidence of 0.9 or more are right about 95% of the time, on classification tasks it had not seen. The team also warns that six retrains of an earlier recipe had a standard deviation of 3.2 tasks, so they treat a gap under about 10 tasks between two single runs as unresolved.

How to run it

Install it with pip install strands-decider, then ask it questions from the command line or serve it on your own machine. The local server answers in the same request format as TypeSafe’s Jev, so code written for Jev can point at it. It binds to your own machine only and has no authentication, which the model card says makes it a tool for local experiments. The model is free under Apache 2.0, the same licence as the Qwen base.

Small enough for one graphics card

At 1.9 billion parameters, Strands Decider 2B is a fraction of the size of the other decision models launched on 1 October 2026. Cloudflare’s Clef and Perplexity’s pplx-decider-v1-27b start from 27-billion-parameter Qwen bases and Clef-flash from a 9-billion one, and all three sell hosted access per token. The Strands team says it is working on libraries to plug decision models into agents, which is where this model is meant to sit.

Questions people ask

What is Strands Decider 2B?
Strands Decider 2B is an open decision model released on 1 October 2026 by Strands Labs, AWS's organisation for experimental open-source projects alongside its Strands Agents SDK. It reads a state, such as a user's message or an agent's planned tool call, and answers typed questions with probabilities: yes or no, one option from a list, or a score on a rubric. It does not generate text.
How fast is Strands Decider 2B?
Its makers report a median of about 115 milliseconds per decision on an Nvidia RTX 3090 and about 153 milliseconds for small tasks on an Apple M3 Pro laptop chip. Both figures are from the Strands team's own tests.
How accurate is Strands Decider 2B?
In the Strands team's own run of the public JevBench tasks, it answered 167 of 231 correctly, an accuracy of 0.723. On the JevBench board of 25 September 2026 that places it third of 33 systems in the 2-billion-parameter class and 50th of 89 overall. It answered every easy task correctly and about half of the hard ones.
Can I use Strands Decider 2B through an API?
Yes, on your own machine. Its command-line tool serves it locally in the same request format as TypeSafe's Jev, from a free Apache 2.0 download on Hugging Face and PyPI. The makers describe the local server as one for experiments, because it has no authentication.