The decision
One noul question per rule, such as "contains harassment" or "is spam or self-promotion", each returning the probability that the post breaks that rule. Well-calibrated probabilities matter here, because a threshold decides what gets hidden.
How to build it
- Turn each rule in your policy into one plain-language condition.
- Score every new post against all of them in one call: the questions travel together.
- Hide above 0.95, queue between 0.6 and 0.95, and publish below.
- Show moderators the rule and the score next to each queued post, so they can decide in seconds.
Make it better
Track how often moderators overturn the model per rule, and tune each rule's threshold separately. open-jev reports the lowest calibration error on the registry, which is a good start for thresholds you can trust.