The decision
Your labelling task as a typed question: a choice over the label set, a score for ordered labels such as severity, or a noul for yes/no labels. The probability decides who labels each row: the model when it is sure, a person when it is not.
How to build it
- Write each label with a one-line definition: the same guide you would give a person doing the labelling.
- Run the model over every row, and store its label and probability.
- Have a person label 200 random rows without seeing the model's answers, and find the probability above which the model agrees with them 98% of the time.
- Accept the model's labels above that line. Load the rest into your labelling tool, such as Label Studio or Argilla, with the model's answer shown as a suggestion.
- Spot check 50 accepted rows in every batch.
Make it better
Fine-tune on the labels people gave in System One Studio (systemone run studio), and run the remaining rows again: each round, fewer of them need a person. Publish the agreement rate with the dataset.