How it works¶
A decision is a classification¶
A Jev call is a classification problem, not a generation problem. For every question the gateway:
- Renders the evidence (the
state) plus lettered options (A,B,C, …). - Asks the backend for exactly one next token with logprobs
(
max_tokens: 1,logprobs: true,top_logprobs: N). - Reads each candidate letter's logprob out of
choices[0].logprobs.content[0].top_logprobs. - Softmaxes the candidate letters into a probability distribution.
No text is generated, so a decision costs one forward pass over the prompt. That is what makes a small local model (0.6B–4B) viable as a router, a gate and a triage classifier.
The full distribution is kept, not just the argmax. For noul it is the answer; for
choice it is returned as probabilities; for score it is collapsed to a
probability-weighted mean so the answer interpolates between levels.
The three question types¶
Each question type maps a set of candidates onto the same letter machinery.
| type | candidates | answer |
|---|---|---|
noul |
criteria.true / criteria.false |
probability of true (defaults to Yes / No) |
choice |
the keys of criteria, in insertion order |
the argmax key plus the full distribution |
score |
the items of criteria, as levels 0..n-1 |
the probability-weighted mean of the level indices, plus legend |
Letters are assigned in order and run past Z: A…P for up to 16 candidates. Ties in
choice resolve to the first candidate, so the answer is deterministic.
Concurrency and failure semantics¶
All questions in a request run concurrently, bounded by
backend.max_concurrency. The response preserves the request's question order regardless
of completion order.
- Any failing question fails the whole request — there is no partial answer.
- The first error in question order is the one reported.
request.total_timeout_secondsbounds the whole request, not each question.backend.timeout_secondsbounds a single question.
Prompt layouts¶
The layout is chosen by request.prompt_layout.
fused¶
One user message containing {evidence, criterion, options} as compact JSON. The
messages are system, user — two in total.
fusedis byte-compatible with the reference gateway, so calibration measured against the hosted API transfers.splitputs the evidence (and any images) in the first user message and the criterion and options in a second. The first message is identical for every question in the request, so llama.cpp reuses the vision encoder work and the state prefill. Use it when latency matters more than byte-compatibility.
Reading the logprobs¶
The gateway maps every returned token onto a candidate letter. "A" and " A" both count
as A; the most likely variant wins. A candidate that falls outside the top-N window is
floored at the least likely logprob actually returned, and the question is flagged
truncated in the diagnostics — raise backend.top_logprobs (up to 4096) if you see that
flag.
The probabilities are renormalised over the candidates only. The rest of the vocabulary's mass is discarded, so a candidate that is unlikely in absolute terms can still hold a high share of the relative distribution.
Why not just call a tool¶
Asking the model to emit a tool call and reading the arguments is the obvious alternative, and it works. It also has to generate the call: ~18 decode steps per decision, ~55 output tokens. The gateway decodes exactly one token.
On the same evidence and rubric the two paths produced identical accuracy, but the gateway was ~11× faster with a warm prompt cache. Choose the direct path only when the model must reason in free text before deciding. See Limits & FAQ.
The system prompt¶
One fixed system message tells the backend to answer with a single letter. It is a constant in the code and is identical for all question types and layouts — the criterion and the options travel in the user message, so the system prompt stays cacheable.