The decision
A choice over your models, for example small, medium, frontier, each described by what it handles well. Or a score from 1 to 5 for how hard the prompt is.
How to build it
- Put a proxy in front of your LLM calls, either an OpenAI-compatible gateway or a LiteLLM router.
- For each request, ask the decision model which tier should answer.
- Send the request there. If the chosen tier's answer fails a check, escalate to the next tier.
- Measure cost per request and the escalation rate before and after.
Make it better
vLLM Semantic Router's Decision models were built for routing inside vLLM. Try them, and label a few hundred prompts with the cheapest tier that answered them well, to fine-tune your own router.