All posts

Noël Vranckx • • 12 min read

Decision models in supply chain: Jev, Clef, d1 and the rest

In three weeks about 30 AI models appeared that cannot write a sentence. They pick an answer from your list, with a probability. How they differ, and where they fit.

Illustration of a 1950s customer service desk: three opened letters lie side by side on a leather blotter, two in blue and the third in red, with a letter opener, a candlestick telephone, a desk bell and a stack of unopened envelopes, morning light from a window on the left

Monday, 8:40. There are 28 new emails in a customer service inbox. Three of them:

"Hi, could you tell me when order 4471 will ship? Thanks, Marc."

"Hello, order 5203 was due last week. Can you give me a new date, please?"

"Great, the third delay this month. Thanks a lot."

All three ask the same question, and all three are polite, word by word. Only one comes from a customer who is about to try another supplier. Anyone who has worked in customer service sees which one in a second, once they get to it.

That second is a decision: one input, a short list of possible answers (neutral, frustrated, angry), made hundreds of times a day. People make it well and fast, but only once they open the email. If the third one is number 26 in the pile, it waits until the afternoon.

Until recently, handing this call to an AI meant a chatbot: slow, paid per word, and now and then an answer that is not on your list. On 15 September TypeSafe released Jev, a model that cannot write a sentence. It reads a record, picks an answer from a list you define, and says how sure it is. By 2 October an independent registry counted 30 of these decision models from 30 makers, and most of them accept Jev's request format.

I have written about Jev twice. This post steps back: what decision models are, how they differ from the chat models we know, which ones are on the market, and ten places they fit in supply chain work. At the end we go back to that Monday inbox.

What decision models are

A decision model takes two things: a block of information (an email, a record, a JSON object, sometimes a photo) and a set of questions whose answers you fix in advance. It answers every question in one pass and gives a probability for each allowed answer. It writes nothing else.

There are three kinds of question, and nearly every model in this post supports all three:

  • Choice: pick one option from a list, such as a reason code or a queue.
  • Score: place the input on an ordered scale, such as low, medium, high.
  • Yes/no: the probability that a statement is true.

The difference with a chat model is the shape of the answer, and everything follows from that.

Chat modelDecision model
OutputText you then have to parseOne of your answers, with a probability
SpeedSecondsTens to hundreds of milliseconds
CostYou pay for every word it writesInput only; output is free
Typical failureAn answer that is not on your listA confident answer that is wrong
Best atExplaining, drafting, reasoningSorting, routing, flagging at volume

That last line matters. A decision model cannot invent a category or misspell a status code. It can still be sure and wrong, which is why the probability is the most important part of the answer.

In practice, that gives them five jobs they do well:

  • Classify: put an email, a complaint or an item record in the right category.
  • Score: rate urgency, severity or risk on a scale you define.
  • Flag: say how likely it is that a statement is true ("this invoice is a duplicate").
  • Route: choose the queue, team or tool that handles the next step.
  • Look: some models also read photos, such as a damaged pallet or a scanned delivery note.

They do all of this in a single pass, so you can ask five questions about one record for the price of reading it once.

The models so far

This is the field as of 4 October 2026, when the System One Models registry listed 30 decision models from 30 makers. The table shows the ones a supply chain team is most likely to meet. Dates, prices and context sizes come from the makers or from the System One Models registry, and all benchmark figures are self-reported.

ModelMakerReleasedHow you get itReads up toPrice per million input tokens
Jev 1.13TypeSafe AI15 SepHosted API32k tokens, text$0.042
Decider 1meraGPT22 SepHosted API4k tokens$0.03
Solar DecideUpstage22 Sep, betaHosted API512k tokensBeta pricing
Tev1 (4B and 0.8B)Together AI23 SepHosted, and locally in OllamaNot published$0.042 hosted
GLiNER2.5-DecideFastino Labs24 SepOpen weights, 340M parametersNot publishedFree to run
Nimble (9B)Bespoke LabsLate SepOpen weights, locally in Ollama8k tokensFree to run
d1Liquid AI29 SepHosted API, no weightsNot publishedFree tier, paid price not published
Jeff 1.1 (0.8B to 2B)Independent29 SepOpen weightsSmall modelsFree to run
Decisions APIOpenAI29 Sep, previewLimited preview, on GPT-6 LunaNot publishedNot published
Mercury DecideInception30 SepHosted, via OpenRouter32k tokensFree in early access
Clef / Clef-flash (27B / 9B)Cloudflare1 OctOpen weights, and hosted on Workers AI64k tokens, text and images$0.24 / $0.09
Strands Decider 2BAWS Strands Labs1 OctOpen weights, Apache 2.0, runs on a laptopNot publishedFree to run
Kev 1.0 (0.8B to 27B)Independent1 OctOpen weights8k tokens (27B: 64k)Free to run

Two tools sit around the models. Ollama 0.35 (28 September) runs decision models on your own laptop. OpenRouter serves several hosted ones behind a single account. Both use the same request format as Jev, so in principle swapping models is a configuration change, not a rebuild. That shared format is what turns a single product into a category.

How they differ from each other

They all answer the same three kinds of question. They differ in five ways that matter for a supply chain team.

  • Where it runs. Jev, d1, Decider 1 and Solar Decide only run on the maker's servers. Clef, Kev, Nimble and Jeff can run on your own machine, so supplier prices or customer data never leave the building.
  • How much it reads. Decider 1 reads about 4,000 tokens, a long email. Solar Decide reads 512,000, a full contract with its annexes. Pick the model for your longest realistic document, not your average one.
  • Whether it sees pictures. Clef accepts up to four images per request on Workers AI. Cloudflare notes that Jev handles text only today. A damaged pallet or a scanned delivery note needs the first kind.
  • Whether you can teach it. Nimble was trained on 2,676 labelled examples. Jeff's maker reports one fine-tune on about 11,000 examples took half an hour on one GPU, lifting accuracy from 31.7% to 95.8%. The hosted models offer no fine-tuning today.
  • Speed and price. In Cloudflare's own test across 43 benchmarks, Clef-flash took a median 39 milliseconds, Clef 209 and Jev 524. AWS reports 115 milliseconds for Strands Decider on a gaming graphics card. Prices run from free to $0.24 per million input tokens. At those prices the model is never the expensive part of the project.
  • Sorting versus judging. An independent comparison by BERI found Clef far ahead on sorting work, with 94.2 against Jev's 79.7 on recognising banking intents. Jev stayed ahead on judgement calls, such as whether an agent should act at all (81.0 against 72.4). Strands Decider was the weakest on hard reasoning. A small model can sort emails well and still be the wrong choice for "should we release this order?".

What it means for supply chain

Here is the frame I use to decide where a decision model belongs. Ask two questions about the decision: how often do we make it, and what does a wrong answer cost?

Wrong answer is cheap to fixWrong answer is expensive
Hundreds a dayThe model decides, a person checks a sampleThe model sorts, a person signs
A few a weekNot worth buildingKeep it human

The top row is where these models earn their place.

Ten use cases in supply chain

Each of these is a decision made hundreds of times a day, with a fixed list of answers.

  1. Customer service mailbox. Intent (order status, change, complaint, quote), urgency, and "is an order number present?", answered for every message before a person opens it.
  2. Supplier order confirmations. Does the confirmation change the date, the quantity or the price compared with the PO? Only the changed ones reach the buyer.
  3. Invoice exceptions. Why did the three-way match fail: price, quantity, missing goods receipt or likely duplicate? Each reason goes to the person who can fix it.
  4. Carrier status messages. Turn free-text updates from carriers ("truck broke down near Lyon, new ETA tomorrow") into your own event codes, with an urgency score.
  5. Supplier risk news. Does this news item concern one of our supplier sites, and how severe is it on a three-point scale?
  6. Master data checks. Does this item description fit its commodity group and unit of measure? Thousands of records checked overnight.
  7. Purchase requisitions. Catalogue, existing contract or new sourcing event, and which category buyer takes it.
  8. Quality complaints with photos. Defect type and severity from the customer's picture, with an image-capable model like Clef.
  9. Licence and sanctions pre-screening. Does this product description or end use need a compliance check? Anything above a low probability goes to the compliance desk.
  10. Inside an agent. Choosing which tool, queue or team handles the next step, so the expensive chat model only runs when it must.

The worked example below goes back to Monday's inbox and builds the first of these, with one extra question that people usually answer by feel: how is this customer feeling?

Worked example: reading the tone of customer emails at Upshift

Upshift is a fictional bicycle maker. Its customer service team receives about 400 emails a day from dealers and consumers, in English, Dutch and French (all numbers here are illustrative). Most are routine. A few are from customers who are one bad reply away from cancelling an order or posting a review. Today those surface when someone happens to open them, which can be the next afternoon.

The team asks four questions of every email, in one call:

QuestionTypeAllowed answers
What does the customer want?ChoiceOrder status, change or cancel, complaint, quote, invoice, other
What is the tone?ScoreNeutral, frustrated, angry
Is there an escalation signal: a threat to cancel, a competitor, a lawyer, social media?Yes/noProbability
Is an order number present?Yes/noProbability

Notice what is not on the list: "How late is the order?" That is a lookup, not a decision. The system reads the order number, fetches the promised and current delivery dates from the ERP, and adds "order 6118, 6 days late, second delay" to the input as a fact. A polite email about an order that has slipped twice deserves different handling from a polite email about an order on time, and the model can only weigh that if it is told.

Tone is where a decision model needs testing most. The sarcastic email from Monday is angry, though every word in it is polite. Before going live, the team labels 300 past emails by hand: 100 per language, deliberately including the ones that later turned into a complaint or a cancellation. They check two things. Are answers given at 80% confidence right about eight times in ten? And does that hold in each language, not just on average? If the model is reliable in English and weak in French, French emails get a lower bar for review.

The test sets the routing:

RouteRuleEmails a day
Call back todayEscalation signal above 0.30, or tone angry above 0.6020
Priority queueTone frustrated above 0.60, or the order is late for the second time60
Standard queueEverything else, sorted by intent300
Automatic replyOrder status, neutral tone above 0.90, order number present and order on time20

Look at the escalation threshold. It sits at 0.30, not 0.50, because a missed angry customer costs far more than a team lead reading an email that turns out to be fine. Look also at what is automated: only the calmest, simplest case. An angry customer never gets an automatic reply, however sure the model is about the intent.

At about 800 tokens per email, 400 emails is 0.32 million tokens a day. That is under 8 cents on Clef or under 2 cents on Jev. The difference is not the cost. It is that the 20 customers most likely to walk away get a phone call the same day instead of a reply the next afternoon. And because the request format is shared, the team can run the same 300 labelled emails through a second model next month and switch if it reads tone better.

What it can't do, and the traps

  • It doesn't calculate or explain. No dates, no totals, no emails. Keep the maths in code and pair it with a chat model for anything that needs words.
  • Published benchmarks don't compare. The registry itself warns that makers report accuracy on different tests. Jeff's 2B model scores 82.0 against Jev's 83.0 on its maker's panel, yet one practitioner reported 70% against Jev's 94% on their own task.
  • Good averages hide blind spots. A September study of four decision models found Nimble ranked first at spotting prompt injection overall, while missing 92.2% of attacks without direct instructions. Test on your own awkward cases, by category.
  • A second model rarely catches the first one's confident mistakes. In the same study, a second judge corrected only 8 to 32 of 166 unsafe misses. Two models are not a control; a sample checked by a person is.
  • Thirty models in three weeks means many will disappear. Most are early access, and prices are launch prices. Build on the shared format, keep your labelled test set, and be ready to swap.

Pick the question before the model

This week, choose one decision your team makes more than a hundred times a day. Write down the exact answers people choose between, and label 200 past cases with the right one. That list and that test set are worth more than any model in the table above, because they let you test all of them.

The sarcastic email on Monday morning did not need a cleverer reader. It needed to be read first. That is the job a decision model does well: a fixed list of answers, a hit rate you have measured on your own cases, and a person for the doubtful ones. It will not replace the people who make these calls, but it can take the clear-cut ones and put the urgent ones on top, so people spend their attention where it is needed.

Which decision in your operation would you hand to one of these models first?

Sources

Illustration of a scale model of a small company on a workbench: two factories and a warehouse linked by roads with miniature lorries, beside a pair of callipers, a pencil and an open notebook

Read it, then try it

The posts on this blog are meant to be used, not only read. Each one ends in something to do: a prompt to run, a calculation to rebuild or a first step for this week. You learn what AI can do in supply chain by putting it to work on a real problem.

Meet Upshift

Most examples use Upshift, a fictional bicycle maker. It builds road and gravel bikes, each as a regular and an e-bike version, in two plants, and serves dealers and retail chains through five distribution hubs. Upshift isn’t a real company, and none of its data comes from one. Meet the company.

The dataset and the exercises

Behind Upshift sits a complete synthetic company dataset: products, suppliers, dealers, orders, stock, production and the transport network, all consistent with each other. Its dates follow a date you choose, so the data always looks current. The exercises on this site use the same company, each with a task, the files you need and a way to check your result.

A post keeps its example small enough to paste into a chat. The dataset and the exercises let you do the same work at the scale of a real company.

The dataset and the exercises are free with an account. The app is in early access for now: request an account from the sign-in screen and I’ll send you an invitation as soon as I can.

Pass it on

Know a colleague who should read this? Post it where they will see it, or send it to them directly.

More posts