Quick start¶
Install¶
Python 3.11 or newer is required. The runtime dependencies are fastapi, httpx,
pydantic, pyyaml and uvicorn[standard].
Configure¶
cp config.yaml config.local.yaml # optional, keeps the shipped file pristine
$EDITOR config.local.yaml
JEV_GATEWAY_CONFIG is the only environment variable the gateway reads; it defaults to
./config.yaml. The server then listens on server.host:server.port (default
0.0.0.0:8000).
The one line that matters is backend.base_url — the address of your OpenAI-compatible
server:
backend:
type: openai
base_url: http://127.0.0.1:8080
model: null # null = auto-detect from the server's /v1/models
Run¶
The console script jev-gateway is equivalent and needs no config path on the command
line if the environment variable is set.
Check that it is alive:
Send a decision¶
curl -X POST http://127.0.0.1:8000/v1/systemone \
-H 'content-type: application/json' \
-d '{
"state": {"error_rate": 0.42, "p99_latency_ms": 3100, "recent_deploy": true},
"questions": {
"is_healthy": {"type": "noul", "instructions": "Is the service healthy?"},
"severity": {
"type": "choice",
"instructions": "Pick the incident severity.",
"criteria": {"sev1": "a whole region is down", "sev2": "error_rate is at least 0.05", "sev3": "below 0.05 and users are unaffected"}
},
"urgency": {
"type": "score",
"instructions": "How urgent is the response?",
"criteria": ["can wait", "today", "right now"]
}
}
}'
{
"model": "Qwen3-4B-Instruct",
"answers": {
"is_healthy": {"type": "noul", "noul": 0.0314},
"severity": {
"type": "choice",
"choice": "sev2",
"probabilities": {"sev1": 0.024, "sev2": 0.921, "sev3": 0.055}
},
"urgency": {
"type": "score",
"score": 1.86,
"legend": {"0": "can wait", "1": "today", "2": "right now"},
"probabilities": {"0": 0.041, "1": 0.058, "2": 0.901}
}
}
}
One request, three decisions, three output tokens.
Docker¶
docker compose up -d # gateway alone
docker compose -f docker-compose.llamacpp.yml up -d # gateway + llama-server
Each template mounts the same ./config.yaml; edit the command block to change the
model, context length or port. Put a .gguf in ./models before starting a backend
stack. Vision models also need the projector: add
--mmproj /models/mmproj.gguf to the llama-server command and check
GET /props → modalities.vision is true.
Verify end to end¶
python scripts/smoke_test.py --url http://127.0.0.1:8000
python scripts/smoke_test.py --url http://127.0.0.1:8000 --image screenshot.png
The smoke test hits /healthz, /readyz and /v1/systemone, validates that every answer
has the right shape, and prints the usage and diagnostics blocks.
vLLM and SGLang
Only the llama.cpp compose stack ships with this repository. To reach a vLLM or
SGLang server, point backend.base_url at it and set backend.model to the name
the server was started with. See Backends.
Where to go next¶
- How it works — the classification trick and the prompt layouts.
- API reference — every field, endpoint and error code.
- Backends — the
base_urltable for llama.cpp, vLLM and SGLang.