A project to build with System One models

A guardrail that checks every tool call before an agent runs it

Ask a decision model "is this command safe to run?" before a coding or ops agent executes it, and block or ask for approval below a threshold.

LLM routing and guardrailsIntermediate4 stepsSuggested by @biplov

The decision

A noul question for each tool call: "running this is safe and matches what the user asked for". Optionally a choice over allow, ask, deny. The state is the user's request, the proposed command or API call, and the working directory or target.

How to build it

  1. Wrap your agent's tool executor, whether Claude Code hooks, an MCP server or your own loop, so every call passes through one function first.
  2. Send the request and the call to the model; it answers in one forward pass, fast enough to check every call.
  3. Allow above 0.9, ask the user between 0.5 and 0.9, and block below.
  4. Keep a list of what was blocked and why, and review it weekly.

Make it better

Collect the calls a person approved or rejected, and fine-tune on them. A model tuned on your own infrastructure's commands will beat a general one.

More projects to build

All projects