FAQS
IDP is the pipeline that turns an unstructured document into structured data: capture, OCR or ICR, classification, field extraction and validation. The governed version adds a confidence score per field and an audit record of what was extracted and by whom.
Extraction runs automatically on ingest: OCR reads the page, classification identifies the document type, and type-specific rules populate the fields you defined - invoice number, effective date, counterparty - as metadata on the document record.
Accuracy varies with scan quality, which is why every extracted field carries a confidence score. Anything below your threshold routes to a review queue rather than entering the repository unverified - so a bad scan becomes a short review task, not a silent data error.
Usually a much smaller one, focused on the low-confidence exceptions rather than every document - the insurance example on this page processes 500,000 handwritten claim forms with a fraction of the original team.
Extracted fields become searchable metadata on the document record and feed the retention schedule, so a document's own contents determine how it is found later and how long it is kept.