The decision
A noul question for each tool call: "running this is safe and matches what the user asked for". Optionally a choice over allow, ask, deny. The state is the user's request, the proposed command or API call, and the working directory or target.
How to build it
- Wrap your agent's tool executor, whether Claude Code hooks, an MCP server or your own loop, so every call passes through one function first.
- Send the request and the call to the model; it answers in one forward pass, fast enough to check every call.
- Allow above 0.9, ask the user between 0.5 and 0.9, and block below.
- Keep a list of what was blocked and why, and review it weekly.
Make it better
Collect the calls a person approved or rejected, and fine-tune on them. A model tuned on your own infrastructure's commands will beat a general one.