AWS Strands Labs
Strands Decider 2B
a small open decision model that runs on a gaming GPU or a laptop

Key facts
- 1 Oct 2026Strands Labs, AWS
- Released
- 1.9BQwen3.5-2B base plus adapter
- Size
- FreeApache 2.0, run it yourself
- Price
- 115 msNvidia RTX 3090, makers' test
- Median latency
- 167 / 231the makers' own run
- JevBench public tasks
- 91.5 MBwhole repository; base fetched on first use
- Download
Strands Decider 2B is a small decision model from AWS's Strands Labs: you give it some text and a set of fixed questions, and it returns a probability for each allowed answer. It is free under the Apache 2.0 licence and runs on your own machine, answering in a median 115 milliseconds on an Nvidia RTX 3090 graphics card. AWS released it on 1 October 2026 with its training code and data recipe, as a tool for experimenting with decision models inside AI agents.
Strands Decider 2B is a decision model: an AI model that reads a piece of text and answers a set of fixed questions with a probability for each allowed answer. It is small enough to run on an ordinary graphics card or a laptop, and it is free. Strands Labs, AWS’s organisation for experimental open-source projects alongside its Strands Agents SDK, released it on 1 October 2026 with its weights, training code and data recipe.
Three AWS staff wrote the launch post: Marc Brooker, a vice-president and distinguished engineer at AWS, Mike Chambers and Fabio Nonato de Paula. They describe it as “optimized for fast experimentation, local development, and innovation”. The package on PyPI calls it “an experimental typed classifier”.
It answers yes, no, a choice or a score
Strands Decider 2B answers three kinds of question from one pass over the input: a yes-or-no question, a choice of one option from a list, or a score against a rubric. Each answer comes with a probability. In the makers’ own example, a customer writes “Help! My payouts have been failing for 3 days!”, and the model routes the message to billing with a probability of 0.845, ahead of retail at 0.091 and sales at 0.064.
The intended job is inside AI agents. The launch post shows the model checking an agent’s planned weather-tool call with two questions: are the arguments grounded in what the user said, and is it too early to call the tool. The project’s README lists model routing, tool selection, argument checks, triage, guardrails and evaluation as uses. The makers set out its limits in the post: answering everything in one parallel pass makes it “significantly worse at solving complex problems than reasoning models”, and it is unsuited to coding, chatbots or summaries.
Inside the model: a Qwen torso with a new head
The team took Alibaba’s Qwen3.5-2B, removed the part that generates text, and replaced it with a small “pointer head” that scores each allowed answer. The head holds just over a million parameters. The base model, which the team calls the torso, is fine-tuned with a rank-16 LoRA adapter, so the download on Hugging Face is about 91.5 MB, and the base weights arrive the first time it runs.

The model counts 1.9 billion parameters in all. The version released is v19, published on Hugging Face as strands-decider-2B-hobson-v19. The training stage took 1,685 seconds on an AWS p5.48xlarge server with Nvidia H100 GPUs, and the makers say a retrain takes about 11 hours on one RTX 3090. The data is public datasets plus synthetic examples written by one Qwen model and checked by a larger one, and reproducing it “needs no paid model-API calls”, the data notes say.
It runs in about a tenth of a second
On an Nvidia RTX 3090 graphics card, Strands Decider 2B answers in a median 115 milliseconds, with the 95th percentile at 299 milliseconds, the makers report. On an Apple M3 Pro laptop chip, small tasks take a median 153 milliseconds.
| Measure, makers’ tests | Result |
|---|---|
| Median latency, RTX 3090 | 115 ms |
| 95th percentile, RTX 3090 | 299 ms |
| Median latency, Apple M3 Pro, small tasks | 153 ms |
| JevBench public tasks correct | 167 of 231 (0.723) |
| Easy, standard and hard tiers | 1.000, 0.875, 0.505 |
| Brier score and expected calibration error | 0.342 and 0.052 |

Third of 33 in its size class
In the Strands team’s own run, the model answered 167 of the 231 public JevBench tasks correctly, which places it third of 33 systems in the 2-billion-parameter class on the board of 25 September 2026, and 50th of 89 overall. Three models in its class, the two ahead of it and one behind, are just over 2 billion parameters, so the team also calls it first of the 30 at or under that size. The top of the overall list is reasoning models, led by OpenAI’s GPT-6 Luna at 0.996, followed by Jev-style models on far larger bases, the team’s evaluation notes say.
It answered every easy task correctly and about half of the hard ones, and the makers say long, multi-step documents are its weak spot. Its probabilities hold up on short tasks: answers given at a confidence of 0.9 or more are right about 95% of the time, on classification tasks it had not seen. The team also warns that six retrains of an earlier recipe had a standard deviation of 3.2 tasks, so they treat a gap under about 10 tasks between two single runs as unresolved.
How to run it
Install it with pip install strands-decider, then ask it questions from the command line or serve it on your own machine. The local server answers in the same request format as TypeSafe’s Jev, so code written for Jev can point at it. It binds to your own machine only and has no authentication, which the model card says makes it a tool for local experiments. The model is free under Apache 2.0, the same licence as the Qwen base.
Small enough for one graphics card
At 1.9 billion parameters, Strands Decider 2B is a fraction of the size of the other decision models launched on 1 October 2026. Cloudflare’s Clef and Perplexity’s pplx-decider-v1-27b start from 27-billion-parameter Qwen bases and Clef-flash from a 9-billion one, and all three sell hosted access per token. The Strands team says it is working on libraries to plug decision models into agents, which is where this model is meant to sit.
Questions people ask
- What is Strands Decider 2B?
- Strands Decider 2B is an open decision model released on 1 October 2026 by Strands Labs, AWS's organisation for experimental open-source projects alongside its Strands Agents SDK. It reads a state, such as a user's message or an agent's planned tool call, and answers typed questions with probabilities: yes or no, one option from a list, or a score on a rubric. It does not generate text.
- How fast is Strands Decider 2B?
- Its makers report a median of about 115 milliseconds per decision on an Nvidia RTX 3090 and about 153 milliseconds for small tasks on an Apple M3 Pro laptop chip. Both figures are from the Strands team's own tests.
- How accurate is Strands Decider 2B?
- In the Strands team's own run of the public JevBench tasks, it answered 167 of 231 correctly, an accuracy of 0.723. On the JevBench board of 25 September 2026 that places it third of 33 systems in the 2-billion-parameter class and 50th of 89 overall. It answered every easy task correctly and about half of the hard ones.
- Can I use Strands Decider 2B through an API?
- Yes, on your own machine. Its command-line tool serves it locally in the same request format as TypeSafe's Jev, from a free Apache 2.0 download on Hugging Face and PyPI. The makers describe the local server as one for experiments, because it has no authentication.
More in Large Language Models
All LLMs →- TypeSafeJeva decision model that returns a typed answer and a confidence
- CloudflareClef and Clef-flashCloudflare's first trained models: open decision models that answer typed questions with probabilities
- PerplexityPerplexity Decider v1 27Bthe decision model behind Perplexity's Decisions API, with open weights
- AWSAmazon Novathe Bedrock native
- Google DeepMindGemini 4 Argonlonger reasoning, coding and cyber defence
- OpenAIGPT-5.6the flagship from 9 July to 3 September 2026