A project to build with System One models

A RAG router that decides whether to retrieve, and from which index

Before a retrieval runs, a decision model reads the question and picks the index that holds the answer, or decides none is needed, so the LLM reads less noise and you run fewer searches.

LLM routing and guardrailsIntermediate5 stepsSuggested by @biplov

The decision

A choice over your sources, each described in one line: product docs, the support knowledge base, the changelog, and "no retrieval needed" for greetings and questions the LLM can answer alone. Add a noul question, "the answer depends on something from the last few days", to send those to a live search instead.

How to build it

  1. Describe each index in one line, the way you would tell a new colleague where things are.
  2. In your RAG pipeline, ask the decision model before the retriever runs. The state is the user's question and the conversation's last turn.
  3. Search only the chosen index. Below 0.7, search the top two and merge the results.
  4. Log each choice next to whether the answer was rated helpful.
  5. Run a week of questions through both pipelines, and compare answer quality, latency and the number of searches.

Make it better

Label a few hundred questions with the index that really held the answer, and fine-tune on them. vLLM Semantic Router's Decision model reads up to 16,384 tokens, so it can take the whole conversation when the history matters.

More projects to build

All projects