A project to build with System One models

Route each prompt to the cheapest LLM that can answer it

A decision model reads the prompt and picks the smallest model likely to handle it, cutting LLM spend without hurting quality.

LLM routing and guardrailsIntermediate4 stepsSuggested by @biplov

The decision

A choice over your models, for example small, medium, frontier, each described by what it handles well. Or a score from 1 to 5 for how hard the prompt is.

How to build it

  1. Put a proxy in front of your LLM calls, either an OpenAI-compatible gateway or a LiteLLM router.
  2. For each request, ask the decision model which tier should answer.
  3. Send the request there. If the chosen tier's answer fails a check, escalate to the next tier.
  4. Measure cost per request and the escalation rate before and after.

Make it better

vLLM Semantic Router's Decision models were built for routing inside vLLM. Try them, and label a few hundred prompts with the cheapest tier that answered them well, to fine-tune your own router.

More projects to build

All projects