A project to build with System One models

A judge that scores your LLM's answers on every eval run

Score every answer your LLM gives against a rubric, from 1 to 5 with a probability for each level, and check grounding and format with yes/no questions, fast enough for every test case and every release.

LLM routing and guardrailsIntermediate4 stepsSuggested by @biplov

The decision

A score question on your rubric, from 1, "wrong or unsafe", to 5, "correct, complete and supported by the sources", with each level described in a sentence. Then a noul question for each check that matters in your product: "every claim is supported by the retrieved passages", "the answer follows the requested format", "the answer declines when it should".

How to build it

  1. Collect your eval set: the prompts, any retrieved context, and the answers your LLM gave.
  2. Build the state from the prompt, the context and the answer, and ask all the questions in one call.
  3. Report the mean score and the share of answers below 3 for every release, next to the previous one.
  4. Read the 20 answers the judge was least sure about after each run. That is where your rubric is ambiguous.

Make it better

Have people score 100 answers, and measure how often the judge agrees with them, level by level, before you trust it. Bespoke Nimble was trained to score rubric levels directly. AnyJev turns an open LLM into a judge without training, and calibrates from a few hundred labels.

More projects to build

All projects