The decision
A score question on your rubric, from 1, "wrong or unsafe", to 5, "correct, complete and supported by the sources", with each level described in a sentence. Then a noul question for each check that matters in your product: "every claim is supported by the retrieved passages", "the answer follows the requested format", "the answer declines when it should".
How to build it
- Collect your eval set: the prompts, any retrieved context, and the answers your LLM gave.
- Build the state from the prompt, the context and the answer, and ask all the questions in one call.
- Report the mean score and the share of answers below 3 for every release, next to the previous one.
- Read the 20 answers the judge was least sure about after each run. That is where your rubric is ambiguous.
Make it better
Have people score 100 answers, and measure how often the judge agrees with them, level by level, before you trust it. Bespoke Nimble was trained to score rubric levels directly. AnyJev turns an open LLM into a judge without training, and calibrates from a few hundred labels.