A project to build with System One models

A cascade that answers with a small model first and escalates only when unsure

Answer every request with a decision model first, send only the uncertain ones to an LLM, and measure how much time and money it saves.

LLM routing and guardrailsAdvanced4 stepsSuggested by @biplov

The decision

The same typed question your LLM answers today, whether it is a classification, a yes/no or a rating, asked first of a decision model. When its top probability clears the threshold, that is the answer; otherwise the LLM answers.

How to build it

  1. Pick one LLM task in your product that has a fixed set of answers.
  2. Run the decision model and the LLM side by side on a week of traffic, and log both answers.
  3. Find the threshold where the decision model agrees with the LLM, for example 99% of the time.
  4. Switch: the decision model answers above the threshold, and the LLM handles the rest.

Make it better

Publish the numbers: the share of traffic that never reached the LLM, the latency, and the cost per 1,000 requests. That is the case for decision models in one chart.

More projects to build

All projects