The decision
A choice over your sources, each described in one line: product docs, the support knowledge base, the changelog, and "no retrieval needed" for greetings and questions the LLM can answer alone. Add a noul question, "the answer depends on something from the last few days", to send those to a live search instead.
How to build it
- Describe each index in one line, the way you would tell a new colleague where things are.
- In your RAG pipeline, ask the decision model before the retriever runs. The state is the user's question and the conversation's last turn.
- Search only the chosen index. Below 0.7, search the top two and merge the results.
- Log each choice next to whether the answer was rated helpful.
- Run a week of questions through both pipelines, and compare answer quality, latency and the number of searches.
Make it better
Label a few hundred questions with the index that really held the answer, and fine-tune on them. vLLM Semantic Router's Decision model reads up to 16,384 tokens, so it can take the whole conversation when the history matters.