ARTIFICIAL INTELLIGENCE
Strands Labs Releases Strands Decider 2B, an Open Source Decision Model Built for Local Use
Strands Labs has released Strands Decider 2B, a 2-billion-parameter open source model that picks among options and scores its confidence rather than generating text. The team says it runs locally with median decisions in around 115ms.
Image: Strands · Uploaded by IntraGoals — usage rights confirmed
Strands Labs has released Strands Decider 2B, a small, open source decision model designed for fast experimentation and local development. The release was announced in a post by Marc Brooker, Mike Chambers and Fabio Nonato de Paula, and was reported by Hacker News. The authors said the model joins strands-labs, which they introduced earlier this year as a place to work hands-on with approaches to agentic AI.
Strands Decider belongs to a newer class of so-called decision models, also described as system one models. According to the authors, the category has drawn attention since TypeSafe AI launched Jev earlier this month. Unlike large language models, which can produce arbitrary output, decision models choose among a fixed set of options or assign a simple numerical score. Examples include deciding whether a string is about a coffee machine, identifying the language of a phrase from a list of candidates, or rating the sentiment of a sentence between 0 and 1.
The authors describe a trade-off. In exchange for less flexibility, decision models are faster and more capable at a given size, always return an answer from the offered options, and can run with very low latency. They are also significantly worse than reasoning models at complex problems, and because they cannot generate text, they are unsuited to coding, chatbots, document summarization and other common LLM tasks. The authors add that decision models attach a reliability score to each decision, which they say is not available through frontier LLM inference APIs, and that they make it efficient to ask several questions about the same prompt. They say these properties suit the agentic workflows many developers build with the Strands Harness SDK.
Strands Decider 2B has two billion parameters and is meant to run on a local CPU or GPU. The team says it can answer meaningful questions in tens of milliseconds and that its accuracy and calibration are competitive with other models in its class. The code is on GitHub and the weights are on Hugging Face, along with the training data and scripts used to build the model.
On architecture, the team started from a pre-trained LLM torso, Qwen3.5-2B, and removed its language-model head, which eliminates its ability to generate text. In its place is a pointer head of just over a million parameters. It scores the hidden state at each option position against the hidden state at the answer position. The torso is fine-tuned with a rank-16 LoRA adapter. The authors say this is the second major iteration of the design, after an earlier version that used a slot head and performed significantly worse. The released model is version 19, and the repository documents the changes made in each version.
The team evaluates models on three targets: accuracy, calibration and latency. Accuracy was measured on JevBench's public set, and calibration was measured using the Brier score on the same set. The authors report that strands-decider-2b ranked third of 33 models in the 2B class, and first of 30 when models just over 2B are excluded. They also say it answers all of JevBench's easy tasks correctly.
For latency, the authors report a median of about 115ms for local decisions on widely available hardware, rising roughly linearly with task size. Their graph was measured on a local Nvidia RTX 3090 against version 18 of the model. On an M3 MacBook, they say median latency for small tasks is around 153ms. They say they have ideas for further improvements, particularly in reducing the latency floor.
The authors say they chose 2 billion parameters as a sweet spot: small enough to use and even train on existing hardware, yet large enough to do meaningful work. They report early success using this class of model for model routing, tool selection, evaluations, guardrails, memory, context management and policy classification. They also point to hybrid agents, in which an LLM handles the hardest decisions and a decision model handles routine ones to reduce cost and latency, as well as experiments combining decision models with fixed workflow languages.
Getting started requires installing the strands-decider package with pip. The command-line tool can then take a state and a multiple-choice question. In the authors' example, a message saying payouts have been failing for three days is routed among billing, sales and retail teams, and the model chooses billing with a confidence of 0.768.
The repository also includes an example in which a Strands agent runs locally alongside the model and uses the default LLM from Amazon Bedrock. The agent has a demo weather tool and is deliberately prompted to guess a city when the user does not name one. Before the tool call runs, the decision model answers two yes/no questions: whether the arguments are grounded in what the user said, and whether it is premature to call the tool. Based on those predictions, the agent asks the user which city they meant. The check is implemented through the Strands intervention system, whose handler can return Proceed, Deny, Confirm or Guide actions. The authors stress that the example is an illustration rather than a recommendation, and that the questions, threshold and policy were chosen by hand.
The Strands team says it is working on libraries for decision model integration. The model, code and data are available now on GitHub and Hugging Face. Details in this report come from the authors' announcement, and the performance figures are the authors' own.