Overview
Decision models are a class of models built to decide rather than write. You give them a state — a support ticket, a transaction, a conversation — and typed questions about it, and they answer each question with probabilities instead of text. No prompt engineering to get a clean yes, no parsing of a sentence, no hallucinated label: the answer is one of the outcomes you listed, with a number you can put a threshold on.
They answer in a fraction of a second and cost a fraction of an LLM call, which makes them the right tool wherever an application needs a judgment in its hot path: routing a ticket, gating a tool call, scoring a lead, guarding a prompt.
The Otoroshi LLM extension makes decision models a first-class model type. It speaks the System One API introduced by TypeSafe with its Jev model and adopted since by Cloudflare, Liquid AI, Telnyx, Prem AI, OpenRouter and Vercel — so the same request reaches any of them, behind one endpoint, one set of API keys and one budget.
Features
- The System One API, as it is — your requests and the provider answers go through untouched, so the official TypeSafe SDKs work by changing their base URL
- Many providers, one contract — TypeSafe, OpenRouter, Cloudflare Workers AI, Liquid AI, Telnyx, Prem AI, Vercel AI Gateway, and any self-hosted System One server (Laya, vLLM, LiteLLM...)
- Swap the model without touching the clients — the model served is a gateway setting, not something hardcoded in every application
- Fallback — another decision model takes over when one cannot answer
- Decisions from your LLMs — let any text provider you already use answer System One questions
- Cost tracking and budgets — every decision is priced per input token and counts against your budgets
- Model constraints — restrict which models consumers can use, per API key or per user
- Guardrail integration — use a decision model as a guardrail on your LLM providers
- Workflows — decide inside a workflow, or let a decision model route it to one of its paths
API endpoints
| Endpoint | Method | Description |
|---|---|---|
/v1/systemone | POST | Ask typed questions about a state. The path the TypeSafe SDKs call |
/v1/decisions | POST | The same endpoint, under the name of what it does |
Request
curl --request POST \
--url http://myroute.oto.tools:8080/v1/systemone \
--header 'content-type: application/json' \
--data '{
"model": "jev-latest",
"state": {
"ticket": "The checkout has been failing for every customer for the last hour.",
"customer_tier": "enterprise"
},
"questions": {
"urgent": {
"type": "noul",
"instructions": "Does this ticket need an urgent intervention?"
},
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Invoices and payments",
"technical": "Outages, bugs and integrations",
"sales": "Contracts and pricing"
}
},
"severity": {
"type": "score",
"instructions": "How severe is the incident?",
"criteria": ["Minor", "Degraded, with a workaround", "Blocking"]
}
}
}'
Request parameters
| Parameter | Type | Description |
|---|---|---|
state | string, object or array | What the questions are about |
questions | object | The questions, by name. The answers come back under the same names |
model | string | Model name. Optional: the model of the decision model entity is used when omitted. Can include a provider prefix for model routing |
Question types
Every question has a type and instructions. A choice and a score also have criteria.
| Type | Asks | Criteria | Answer |
|---|---|---|---|
noul | A yes/no question | Optional: what true and false mean | The probability that the answer is yes |
choice | One option among several | An object: each option and what it means | The chosen option, the probability of every option, a confidence |
score | A position on an ordered scale | An array: the levels, from the lowest to the highest | The score, the probability of every level, a confidence |
All the questions of a request are answered in one call, against the same state.
Response
{
"model": "jev-1.13.0",
"answers": {
"urgent": {
"type": "noul",
"noul": 0.95
},
"team": {
"type": "choice",
"choice": "technical",
"confidence": 0.78,
"probabilities": {
"billing": 0.15,
"technical": 0.85,
"sales": 0.0
}
},
"severity": {
"type": "score",
"score": 1.43,
"confidence": 0.35,
"legend": {
"0": "Minor",
"1": "Degraded, with a workaround",
"2": "Blocking"
},
"probabilities": {
"0": 0.0,
"1": 0.57,
"2": 0.43
}
}
},
"usage": {
"input_tokens": 350,
"output_tokens": 58
}
}
Response fields
| Field | Type | Description |
|---|---|---|
model | string | The model that answered |
answers | object | One answer per question, under the name of the question |
answers.*.noul | number | Yes/no question: the probability of yes, from 0 to 1 |
answers.*.choice | string | Choice: the most likely option |
answers.*.score | number | Score: the position on the scale, weighted by the probabilities |
answers.*.probabilities | object | The probability of every option or level |
answers.*.confidence | number | How far the answer stands out from the others: 0 is an even split, 1 a certainty |
usage | object | The tokens read and produced |
Add ?embed_costs=true to a request to get what it cost in a costs field.
Use cases
- Routing — send a ticket, a call or a document to the right team, queue or model
- Gating — decide whether an agent may call a tool, a refund may be issued, a message may be sent
- Scoring — rate urgency, severity, sentiment or quality on your own scale
- Guardrails — ask a yes/no question about a prompt or an answer as a guardrail, without the latency of an LLM judge
- Workflows — branch a workflow on a probability rather than on a parsed sentence
- Routing — send a workflow down the path a decision model picks, and only when it is sure of it
- Model routing — let a decision model choose the LLM that answers each request, with the smart-router and the intent-router
A probability is only comparable with the ones of the same model. When you move a decision to another model, tune its thresholds again on your own data rather than carrying them over.