As a regular user of a bunch of specialized micromodels, I'll tell you this: you won't be happy with such a model (and its JEV counterparts) running permanently in the background on your PC's CPU. You need to offload their processing to the NPU. There are many pitfalls along the way, but the result is worth it.
NPU performance will be twice as high, while power consumption will be four times lower. No additional fan noise (if you know what I mean).
I'll wait another month until the first phase of the =battle royale= among models of this kind wraps up, put together a solution for the NPU/iGPU, and post it on HF.
What models are you running on CPU? Any repos or gists you can share? Curious about the models and your use case. Are you doing multi-language or single language?
I think under Lemonade you can do that, especially with AMD systems, using the models that allow for hybrid operation. Prefill happens on the NPU and token generation happens on the GPU. Not sure if this is faster or better than just doing it all on the GPU side.
It's slower but more power efficient. I think only some form of onyx models are available and I couldnt get anything to run on the first gen NPU's in the 7940HS cpu.
The new OpenAI Decisions API calls it "predicate". Also calling the API "decisions" rather than "system one". Usually I don't like inventing new standards but I hope the OpenAI schema takes over. We don't need this hype terminology.
My guess is if you pronouce this "nool" it sounds similar to "bool", and it's a slice of the name "Bernoulli" because the models generate Bernoulli distributions
oh god, you're completely right! nice and obvious on a reread. (I also just learnt about "garden path" sentences.)
I'd googled "bernoulli map" after reading that, and came across this[0], which I wasn't aware of and thought was somehow related so I entrenched my misunderstanding (I didn't dig in.)
Side note: I just wrote the sentence "you're completely right!", and almost changed it because it sounds like slop now. I wonder if we're gonna get an increase in these sentences in human written prose over time as people mimic these sentences, or a decrease as people shy away from them to not sound like AI. :)
On humans-for-humans, blog posts with my name on aren't written by AI (https://brooker.co.za/blog/2026/06/18/my-blog-and-ai.html) either for my personal blog or at work. I'm a heavy AI user, but this is something I think is best left to humans.
There are a ton a ways to use models like this. Model routing is a popular emerging one. The one that really interests me is hybrid agentic workflows - filling the gap between deterministic workflows (e.g. AWS StepFunctions) and fully agentic workflows that use frontier intelligence.
We're releasing more of that functionality in strands, and I expect an explosion of innovation in that area.
A use case I have in mind is creating a relatively simple home assistant app that would let me control smart home things like lightbulbs and playing music on some kitchen speakers. I can also imagine having it differentiate between commands like "set lights to 30% brightness" and more general questions like "what's the weather today" and piping the latter over to a cheap LLM to handle. It could probably all be handled by an LLM, but I like how fast these decision models end up being.
The ~2-week-old Intern-Decision family of models (0.8B, 2B and 4B) have the same Qwen3.5 base model family (albeit the instruction-tuned variants) and pointer head architecture.
I wish these would stop using JevBench. It focuses way too much on text classification tasks, and some of the models perform very poorly on tasks that need actual intelligence.
I appreciate what the JevBench folks are doing, but it's true the JevBench (like all benchmarks) isn't representative of every real-world workloads. I particularly don't love how JevBench handles confidence, and rewards being highly confident in certain cases. Benchmarking is hard, and JevBench is probably more representation of its target workloads than TPC-C is :)
As for actual intelligence, there's only so much you can expect from 2B without reasoning. One of the design challenges in training this model was to avoid forgetting too much, and the KL-to-frozen-base step partially exists for that purpose. Even then, Qwen3.5 2B base has fairly limited single-pass reasoning ability (which shows up for us in the performance on the JevBench hard set).
clef from cloudflare runs on llama.cpp - being locked-in to strands cli would be a bummer and will slow down adoption.
Since it's a LoRa on Qwen, I assume this is runnable via llama.cpp. Pity that the PEFT/LoRa->GGUF translation is left to the user. Anyone got past:
$ uv run --with transformers==5.19.0 convert_lora_to_gguf.py ~/Downloads/lora --dry-run --verbose
[...]
File "/Users/user/repos/llama.cpp/conversion/base.py", line 630, in map_tensor_name
raise ValueError(f"Can not map tensor {name!r}")
ValueError: Can not map tensor 'layers.0.linear_attn.in_proj_a.weight'
It's very doable, but we haven't done it yet (although some great community folks did an ONNX version of v21).
I haven't looked in depth, but it should be doable without modifications to llama.cpp. You can't just convert the LoRA, through - there's a whole pointer head and some custom layers that need to be correctly handled (in fact, the LoRA mostly exists to get the base model to behave the way the pointer head needs).
Does anyone know if it is worth fine-tuning one of these decision models on the shape of the questions you want it to work on, vs the more general versions? I'm using Jev pretty successfully at work at the moment, but am curious about what is doable
Yes, at this size (2B) you can get better performance and calibration fine tuning on your questions. You can also improve calibration on your problem type just be recalibrating without fine-tuning.
As to what's doable, that depends on what you expect. It's a single pass through a 2B model, no reasoning, so it's never going to be particularly 'intelligent'. On domain problems, though, you can get great in-distribution performance (and possibly better in-distribution performance than you'd get using the same volume of data to train a specialized classifier).
Everything you need to fine-tune, or even retrain from scratch, is in the github repo.
I'm interested in this as well and maybe to broaden the scope of the question a little:
If I have a sizeable amount of labeled data and need decisions calibrated to that data should I
1. Ignore the hype and train a traditional classifier
2. Finetune an LLM based decision model
3. Shoehorn (probably a small subset of) the data into the context of the LLM classifier somehow
If the answer is 3. where does the data belong? In the input content? Request wide state? In the question instructions? In the criteria? How much of my data can and should I use?
(1) would take the most data (probably, depends on your domain) and probably generalize worst, but would be closest to compute-optimal in the end.
(2) doesn't take much data, you can tweak calibration to your needs, and might only take a few minutes on a beefy GPU.
(3) is the way to go if you need a lot of general knowledge, any amount of reasoning, and don't need good calibration. Likely the least inference efficient of the three options for a given accuracy and calibration.
The hard part of (3) is keeping the answers to the questions independent. If you dump them all together with the state into the prompt, the answers to the second question will depend on the first (and the answer to the first if you it step-by-step). You can do it by tweaking the inference process with the right masking or use of batching, or by fiddling with cache control using a provider's API.
I think jev points to an interesting way to fine-tune open models
The goal isn’t to replace current models, it’s to train a model for a specific domain so it can handle multiple-choice and yes/no questions quickly, helping the overall system run faster and get better results
Are any of these multimodal yet? I'd love to try asking a model with calibrated probabilities to answer question like, "do these shapes match?". Sure, you can ask a LLM....
Can someone explain this architecture a bit more in depth? They say the pointer head scores the hidden state at each option against the hidden state of the answer. But the LLM produces hidden states per token, so an option can span multiple tokens, no?
On multiple tokens, we put each option (however many tokens it is) onto a line, and then read the hidden state at the end of that line to use at the option state. This is done option-by-option.
The hidden state at the end of the question is the query `q` (again, not sensitive to how many tokens the question is), each options end-of-line state is the key `k`, and the per-option logit is calculated as `q.k / sqrt(256)`.
as far as gpt-5.6 could help me understand the code in their repository (and as far as their architecture is accurate [0]), the "hidden state at each option" refers to the hidden state at the last token in each option, which is then scored by the (learned) pointer head against the hidden state at the <answer>-position.
To vaguely confirm that this is plausible, i tested the sample query from the post with differently arranged options (This should produce slightly different outputs since each option influences all subsequent hidden states, including those of the other options):
So each of these arrangements favor billing, but with pretty different confidence scores, and it looks like there is a pretty big bias toward the first option listed - which makes me wonder how viable it would be to have the torso process these options independently instead
Their docs say:
> An option's score depends on what it says, not where it sits. There is no per-option parameter anywhere in the model, so nothing can learn that "the first option is usually the right one"
This isn't quite correct. There might not be an explicit "per-option" parameter, but the hidden states themselves "contain" all prior options. My gut feeling tells me they messed up data augmentation by incorrectly permuting options during training.
Edit: the above numbers were obtained with the v19 version/checkpoint, v21 is much less sensitive to ordering, but still shows a first position bias.
Swapping out the text-generation head for a dedicated pointer head on a small footprint model is a pragmatic approach for low-latency local decision pipelines.
As a regular user of a bunch of specialized micromodels, I'll tell you this: you won't be happy with such a model (and its JEV counterparts) running permanently in the background on your PC's CPU. You need to offload their processing to the NPU. There are many pitfalls along the way, but the result is worth it.
NPU performance will be twice as high, while power consumption will be four times lower. No additional fan noise (if you know what I mean).
I'll wait another month until the first phase of the =battle royale= among models of this kind wraps up, put together a solution for the NPU/iGPU, and post it on HF.
Most LLMs cannot run efficiently on current NPUs (except for prefill stage), the hardware was built for a different kind of ML workload.
Decision models are prefill only
What models are you running on CPU? Any repos or gists you can share? Curious about the models and your use case. Are you doing multi-language or single language?
Btw, would love your opinion on this: pre trained classifiers that run and train on CPU https://github.com/nicobrenner/jeffy
I haven't seen any frameworks for running the NPU. my 395+ needs a buddy.
I think under Lemonade you can do that, especially with AMD systems, using the models that allow for hybrid operation. Prefill happens on the NPU and token generation happens on the GPU. Not sure if this is faster or better than just doing it all on the GPU side.
It's slower but more power efficient. I think only some form of onyx models are available and I couldnt get anything to run on the first gen NPU's in the 7940HS cpu.
At this point I can't wait for a comedian to release a decision model backed by humans.
Meet Jerry- it's literally a guy named Jerry answering your questions.
Oh! You probably want ChatTJB [1]
[1] - https://chattjb.org/about
Don't worry, Jerry will rig everything up for you.
Your comment made me chuckle.
Jerry, "The Decider" https://www.youtube.com/watch?v=r8VbzrZ9yHQ
Same, but my mind went straight to Rick and Morty… def don’t want Jerry deciding XD
Ah, easy to work around, you just iterate eliminating his choices one at a time, until you get the one he didn't pick (that's your correct pick).
RACE HIM JERRY!
Why is everyone calling binary choices `noul`? Does this have some meaning or is it just copying Jev’s API?
The new OpenAI Decisions API calls it "predicate". Also calling the API "decisions" rather than "system one". Usually I don't like inventing new standards but I hope the OpenAI schema takes over. We don't need this hype terminology.
My guess is if you pronouce this "nool" it sounds similar to "bool", and it's a slice of the name "Bernoulli" because the models generate Bernoulli distributions
The same reason why they’re calling this System One thinking. Everyone wants to be seen doing different things and smart ones.
It's from Bernoulli.
Huh. And here I was thinking the obvious thing is that it's a way to have "null" without it accidentally being parsed as null.
Interesting. Is that a unit he invented or is it just a reference to his last name that stuck?
From Bernoulli maps apparently. I still don't understand why[0].
[0] https://news.ycombinator.com/item?id=49723267 (see parent for reference)
A bit of a garden path path sentence... I believe "maps" in "maps to" was being used as a verb.
So "noul"--short for Bernoulli as in the Bernoulli distribution--"maps to if-statements."
oh god, you're completely right! nice and obvious on a reread. (I also just learnt about "garden path" sentences.)
I'd googled "bernoulli map" after reading that, and came across this[0], which I wasn't aware of and thought was somehow related so I entrenched my misunderstanding (I didn't dig in.)
Side note: I just wrote the sentence "you're completely right!", and almost changed it because it sounds like slop now. I wonder if we're gonna get an increase in these sentences in human written prose over time as people mimic these sentences, or a decrease as people shy away from them to not sound like AI. :)
[0] https://en.wikipedia.org/wiki/Dyadic_transformation
Claude generated it and it stuck.
Fantastically well written. It's rare for me to be able to understand what the AI gurus are talking about, and this was written by humans for humans.
It can technically be used for a lot of use cases, I'd like people to chime in on ideas on this?
One of the authors here. Thanks!
On humans-for-humans, blog posts with my name on aren't written by AI (https://brooker.co.za/blog/2026/06/18/my-blog-and-ai.html) either for my personal blog or at work. I'm a heavy AI user, but this is something I think is best left to humans.
There are a ton a ways to use models like this. Model routing is a popular emerging one. The one that really interests me is hybrid agentic workflows - filling the gap between deterministic workflows (e.g. AWS StepFunctions) and fully agentic workflows that use frontier intelligence.
We're releasing more of that functionality in strands, and I expect an explosion of innovation in that area.
A use case I have in mind is creating a relatively simple home assistant app that would let me control smart home things like lightbulbs and playing music on some kitchen speakers. I can also imagine having it differentiate between commands like "set lights to 30% brightness" and more general questions like "what's the weather today" and piping the latter over to a cheap LLM to handle. It could probably all be handled by an LLM, but I like how fast these decision models end up being.
Picking lunch menu..?
The ~2-week-old Intern-Decision family of models (0.8B, 2B and 4B) have the same Qwen3.5 base model family (albeit the instruction-tuned variants) and pointer head architecture.
https://huggingface.co/collections/internlm/intern-decision
I wish these would stop using JevBench. It focuses way too much on text classification tasks, and some of the models perform very poorly on tasks that need actual intelligence.
One of the authors here.
I appreciate what the JevBench folks are doing, but it's true the JevBench (like all benchmarks) isn't representative of every real-world workloads. I particularly don't love how JevBench handles confidence, and rewards being highly confident in certain cases. Benchmarking is hard, and JevBench is probably more representation of its target workloads than TPC-C is :)
As for actual intelligence, there's only so much you can expect from 2B without reasoning. One of the design challenges in training this model was to avoid forgetting too much, and the KL-to-frozen-base step partially exists for that purpose. Even then, Qwen3.5 2B base has fairly limited single-pass reasoning ability (which shows up for us in the performance on the JevBench hard set).
clef from cloudflare runs on llama.cpp - being locked-in to strands cli would be a bummer and will slow down adoption.
Since it's a LoRa on Qwen, I assume this is runnable via llama.cpp. Pity that the PEFT/LoRa->GGUF translation is left to the user. Anyone got past:
It's very doable, but we haven't done it yet (although some great community folks did an ONNX version of v21).
I haven't looked in depth, but it should be doable without modifications to llama.cpp. You can't just convert the LoRA, through - there's a whole pointer head and some custom layers that need to be correctly handled (in fact, the LoRA mostly exists to get the base model to behave the way the pointer head needs).
This is a great model. I've been running it on device in chrome extension to filter things like email.
It is just the right mix of size, capability and speed to make it generally useful for adhoc bulk classification tasks.
For those wanting to run it in browser: https://huggingface.co/alxnahas/strands-decider-2B-webgpu
Live demo: https://alxnahas.github.io/strands-decider-web/?backend=engi...
Does anyone know if it is worth fine-tuning one of these decision models on the shape of the questions you want it to work on, vs the more general versions? I'm using Jev pretty successfully at work at the moment, but am curious about what is doable
One of the authors here.
Yes, at this size (2B) you can get better performance and calibration fine tuning on your questions. You can also improve calibration on your problem type just be recalibrating without fine-tuning.
As to what's doable, that depends on what you expect. It's a single pass through a 2B model, no reasoning, so it's never going to be particularly 'intelligent'. On domain problems, though, you can get great in-distribution performance (and possibly better in-distribution performance than you'd get using the same volume of data to train a specialized classifier).
Everything you need to fine-tune, or even retrain from scratch, is in the github repo.
I'm interested in this as well and maybe to broaden the scope of the question a little:
If I have a sizeable amount of labeled data and need decisions calibrated to that data should I
1. Ignore the hype and train a traditional classifier
2. Finetune an LLM based decision model
3. Shoehorn (probably a small subset of) the data into the context of the LLM classifier somehow
If the answer is 3. where does the data belong? In the input content? Request wide state? In the question instructions? In the criteria? How much of my data can and should I use?
All are viable options.
(1) would take the most data (probably, depends on your domain) and probably generalize worst, but would be closest to compute-optimal in the end.
(2) doesn't take much data, you can tweak calibration to your needs, and might only take a few minutes on a beefy GPU.
(3) is the way to go if you need a lot of general knowledge, any amount of reasoning, and don't need good calibration. Likely the least inference efficient of the three options for a given accuracy and calibration.
The hard part of (3) is keeping the answers to the questions independent. If you dump them all together with the state into the prompt, the answers to the second question will depend on the first (and the answer to the first if you it step-by-step). You can do it by tweaking the inference process with the right masking or use of batching, or by fiddling with cache control using a provider's API.
I think jev points to an interesting way to fine-tune open models
The goal isn’t to replace current models, it’s to train a model for a specific domain so it can handle multiple-choice and yes/no questions quickly, helping the overall system run faster and get better results
Are any of these multimodal yet? I'd love to try asking a model with calibrated probabilities to answer question like, "do these shapes match?". Sure, you can ask a LLM....
One of the authors here.
v21 supports vision: https://huggingface.co/StrandsAgents/strands-decider-2B-hobs...
Cloudflare’s clef is multimodal (https://blog.cloudflare.com/clef-decision-models/)
(Disclaimer, I work at Cloudflare, but not on models)
Strands Decider can take vision in. How does it go with that question?
This is not advertised on their page, did you make this up?
I believe image classification/analysis by deciders (not just OCR, not everything is about text) is still lacking.
Cloudflare's Clef had fair results on my test, but it's larger and slower. Wondering about Strands.
v21 supports vision: https://huggingface.co/StrandsAgents/strands-decider-2B-hobs...
Can someone explain this architecture a bit more in depth? They say the pointer head scores the hidden state at each option against the hidden state of the answer. But the LLM produces hidden states per token, so an option can span multiple tokens, no?
I wrote a slightly deeper blog post here: https://brooker.co.za/blog/2026/09/28/engineering-system-one...
On multiple tokens, we put each option (however many tokens it is) onto a line, and then read the hidden state at the end of that line to use at the option state. This is done option-by-option.
The hidden state at the end of the question is the query `q` (again, not sensitive to how many tokens the question is), each options end-of-line state is the key `k`, and the per-option logit is calculated as `q.k / sqrt(256)`.
as far as gpt-5.6 could help me understand the code in their repository (and as far as their architecture is accurate [0]), the "hidden state at each option" refers to the hidden state at the last token in each option, which is then scored by the (learned) pointer head against the hidden state at the <answer>-position.
To vaguely confirm that this is plausible, i tested the sample query from the post with differently arranged options (This should produce slightly different outputs since each option influences all subsequent hidden states, including those of the other options):
"billing,sales,retail" produces billing -> 0.843, retail -> 0.092, sales -> 0.065
"retail,billing,sales" produces billing -> 0.470, retail -> 0.468, sales -> 0.062
"retail,sales,billing" produces billing -> 0.517, retail -> 0.415, sales -> 0.068
"billing,retail,sales" produces billing -> 0.803, retail -> 0.146, sales -> 0.051
"sales,retail,billing" produces billing -> 0.647, retail -> 0.127, sales -> 0.225
So each of these arrangements favor billing, but with pretty different confidence scores, and it looks like there is a pretty big bias toward the first option listed - which makes me wonder how viable it would be to have the torso process these options independently instead
Their docs say:
> An option's score depends on what it says, not where it sits. There is no per-option parameter anywhere in the model, so nothing can learn that "the first option is usually the right one"
This isn't quite correct. There might not be an explicit "per-option" parameter, but the hidden states themselves "contain" all prior options. My gut feeling tells me they messed up data augmentation by incorrectly permuting options during training.
Edit: the above numbers were obtained with the v19 version/checkpoint, v21 is much less sensitive to ordering, but still shows a first position bias.
[0] https://github.com/strands-labs/strands-decider/blob/main/do...
I wait for this open-source model.
I should let the development agent select a model and run it to improve token efficiency.
Is it being used this much these days?
Any idea how well this would run on a CPU?
On my M3 mac, it works okay inside a docker container with just CPU.
{ "model": "strands-decider-2B-hobson-v19", "answers": { "is_urgent": { "type": "noul", "noul": 0.8287 } }, "usage": { "input_tokens": 86, "output_tokens": 1 }, "latency_ms": 1732.17 }
This is how I got it running - https://gist.github.com/2891eb0db9ea92c1a4e860d44f556292
There's a lot more to be done if we optimize for MLX & let it run on a Mac mini instead of the docker wrapper.
About half a second per decision on six cores
Isnt that a bit slow for these?
Benchmark calibration does not establish reliability on unfamiliar production inputs.
Swapping out the text-generation head for a dedicated pointer head on a small footprint model is a pragmatic approach for low-latency local decision pipelines.
Looks really nice, I think I could use it on my Mac mini for some smaller automations
is jev commoditized now ?
a strand type game
2B being called small is such a sign of the times