A project to build with System One models

Document intake that classifies each file and pulls out its key fields

Decide what kind of document arrived (invoice, contract, ID, receipt) and extract the fields that matter, with a confidence for each.

Documents and dataIntermediate4 stepsSuggested by @biplov

The decision

A choice over document types, then extraction of the fields for that type: amount, due date and vendor for an invoice; parties and term for a contract. GLiNER2.5-Decide both decides and extracts, which suits this pipeline.

How to build it

  1. OCR or parse each upload to text.
  2. Classify the type. Below a threshold, send it to a person.
  3. Extract the type's fields, and validate them with simple rules (dates parse, amounts add up).
  4. Write the result to your database, keeping the source page and the confidences.

Make it better

Show reviewers only the fields with low confidence instead of the whole document. That is where the time savings are.

More projects to build

All projects