Skip to main content

Overview

Decision models are a class of models built to decide rather than write. You give them a state — a support ticket, a transaction, a conversation — and typed questions about it, and they answer each question with probabilities instead of text. No prompt engineering to get a clean yes, no parsing of a sentence, no hallucinated label: the answer is one of the outcomes you listed, with a number you can put a threshold on.

They answer in a fraction of a second and cost a fraction of an LLM call, which makes them the right tool wherever an application needs a judgment in its hot path: routing a ticket, gating a tool call, scoring a lead, guarding a prompt.

The Otoroshi LLM extension makes decision models a first-class model type. It speaks the System One API introduced by TypeSafe with its Jev model and adopted since by Cloudflare, Liquid AI, Telnyx, Prem AI, OpenRouter and Vercel — so the same request reaches any of them, behind one endpoint, one set of API keys and one budget.

Features​

  • The System One API, as it is — your requests and the provider answers go through untouched, so the official TypeSafe SDKs work by changing their base URL
  • Many providers, one contract — TypeSafe, OpenRouter, Cloudflare Workers AI, Liquid AI, Telnyx, Prem AI, Vercel AI Gateway, and any self-hosted System One server (Laya, vLLM, LiteLLM...)
  • Swap the model without touching the clients — the model served is a gateway setting, not something hardcoded in every application
  • Fallback — another decision model takes over when one cannot answer
  • Decisions from your LLMs — let any text provider you already use answer System One questions
  • Cost tracking and budgets — every decision is priced per input token and counts against your budgets
  • Model constraints — restrict which models consumers can use, per API key or per user
  • Guardrail integration — use a decision model as a guardrail on your LLM providers
  • Workflows — decide inside a workflow, or let a decision model route it to one of its paths

API endpoints​

EndpointMethodDescription
/v1/systemonePOSTAsk typed questions about a state. The path the TypeSafe SDKs call
/v1/decisionsPOSTThe same endpoint, under the name of what it does

Request​

curl --request POST \
--url http://myroute.oto.tools:8080/v1/systemone \
--header 'content-type: application/json' \
--data '{
"model": "jev-latest",
"state": {
"ticket": "The checkout has been failing for every customer for the last hour.",
"customer_tier": "enterprise"
},
"questions": {
"urgent": {
"type": "noul",
"instructions": "Does this ticket need an urgent intervention?"
},
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Invoices and payments",
"technical": "Outages, bugs and integrations",
"sales": "Contracts and pricing"
}
},
"severity": {
"type": "score",
"instructions": "How severe is the incident?",
"criteria": ["Minor", "Degraded, with a workaround", "Blocking"]
}
}
}'

Request parameters​

ParameterTypeDescription
statestring, object or arrayWhat the questions are about
questionsobjectThe questions, by name. The answers come back under the same names
modelstringModel name. Optional: the model of the decision model entity is used when omitted. Can include a provider prefix for model routing

Question types​

Every question has a type and instructions. A choice and a score also have criteria.

TypeAsksCriteriaAnswer
noulA yes/no questionOptional: what true and false meanThe probability that the answer is yes
choiceOne option among severalAn object: each option and what it meansThe chosen option, the probability of every option, a confidence
scoreA position on an ordered scaleAn array: the levels, from the lowest to the highestThe score, the probability of every level, a confidence

All the questions of a request are answered in one call, against the same state.

Response​

{
"model": "jev-1.13.0",
"answers": {
"urgent": {
"type": "noul",
"noul": 0.95
},
"team": {
"type": "choice",
"choice": "technical",
"confidence": 0.78,
"probabilities": {
"billing": 0.15,
"technical": 0.85,
"sales": 0.0
}
},
"severity": {
"type": "score",
"score": 1.43,
"confidence": 0.35,
"legend": {
"0": "Minor",
"1": "Degraded, with a workaround",
"2": "Blocking"
},
"probabilities": {
"0": 0.0,
"1": 0.57,
"2": 0.43
}
}
},
"usage": {
"input_tokens": 350,
"output_tokens": 58
}
}

Response fields​

FieldTypeDescription
modelstringThe model that answered
answersobjectOne answer per question, under the name of the question
answers.*.noulnumberYes/no question: the probability of yes, from 0 to 1
answers.*.choicestringChoice: the most likely option
answers.*.scorenumberScore: the position on the scale, weighted by the probabilities
answers.*.probabilitiesobjectThe probability of every option or level
answers.*.confidencenumberHow far the answer stands out from the others: 0 is an even split, 1 a certainty
usageobjectThe tokens read and produced

Add ?embed_costs=true to a request to get what it cost in a costs field.

Use cases​

  • Routing — send a ticket, a call or a document to the right team, queue or model
  • Gating — decide whether an agent may call a tool, a refund may be issued, a message may be sent
  • Scoring — rate urgency, severity, sentiment or quality on your own scale
  • Guardrails — ask a yes/no question about a prompt or an answer as a guardrail, without the latency of an LLM judge
  • Workflows — branch a workflow on a probability rather than on a parsed sentence
  • Routing — send a workflow down the path a decision model picks, and only when it is sure of it
  • Model routing — let a decision model choose the LLM that answers each request, with the smart-router and the intent-router
tip

A probability is only comparable with the ones of the same model. When you move a decision to another model, tune its thresholds again on your own data rather than carrying them over.