Most businesses still have people retyping information that already exists on a page: invoice totals into the accounting system, application forms into a CRM, renewal dates from contracts into a spreadsheet. Document processing automation removes most of that retyping. It does not remove all of it, and the projects that disappoint are usually the ones that were sold as if it would.
This guide explains how the technology works step by step, where it is reliable and where it is not, how to choose between a cloud document AI service and a custom pipeline, what to settle on data protection, and how to run a pilot that tells you the truth before you commit a budget.
Quick answer
Document processing automation (often called intelligent document processing, or IDP) is software that takes in documents such as invoices, forms and contracts, reads them with optical character recognition (OCR), identifies what type of document each one is, extracts the specific fields you need, checks those values against business rules, sends doubtful cases to a person, and exports the approved data to an accounting, ERP or CRM system.
- It works best on high-volume documents with a predictable set of fields, such as supplier invoices, receipts, ID documents and standard forms.
- No extraction method is correct every time, so a human review queue for low-confidence or rule-failing documents is a permanent part of the design, not a temporary crutch.
- Cloud services from AWS, Google and Microsoft cover the reading and extraction. The validation rules, review screen and integration with your systems are usually where the custom work lies.
- The only dependable way to know what accuracy you will get is to test on a sample of your own real documents.
How intelligent document processing works
A working system is a pipeline of seven stages. Each stage can fail in its own way, which is why it helps to understand them separately.
1. Capture
Documents arrive by email attachment, supplier portal upload, mobile photo, scanner or shared folder. The capture stage collects them in one place, records where each came from, and rejects files that cannot be processed, such as password-protected PDFs. A PDF generated by accounting software already contains real text; a scan or phone photo is only an image. Good capture also splits a single scan containing several documents into separate items.
2. OCR and layout analysis
OCR converts the image of a page into text. Modern services also return layout: which words form a table, which label belongs to which value, where a signature sits. Amazon Textract, for example, is documented as detecting typed and handwritten text and returning text, forms, tables, query responses and signatures as separate categories of output. Image quality matters more here than anywhere else: skewed, faint or low-resolution scans cause errors that every later stage inherits.
3. Classification
Before extracting anything, the system decides what it is looking at: an invoice, a credit note, a purchase order, a W-9, a tenancy agreement. Classification routes each document to the right extraction logic. A credit note treated as an invoice is a classic and expensive mistake, because the amount has the wrong sign.
4. Field extraction: templates, machine learning or LLMs
This is the stage with the most design choice. There are three broad approaches, and mature systems often combine them.
| Approach | How it works | Strengths | Weaknesses | Good fit |
|---|---|---|---|---|
| Templates and rules | You define where each field sits on a known layout, or patterns that find it | Predictable, cheap to run, easy to explain | Breaks when a layout changes; one template per layout | Your own forms, a small number of fixed layouts |
| Trained machine learning models | A model learns from labeled examples to find fields across varying layouts | Handles many layouts without a template each; returns confidence scores | Needs labeled samples for custom document types; fixed field list | Invoices and receipts from many suppliers, ID documents, tax forms |
| Large language models (LLMs) | A general model reads the document and returns fields you describe in plain language | Flexible; copes with long, unstructured text such as contract clauses | Can return plausible but wrong values; output can vary between runs; cost and speed depend on document length | Contracts, correspondence, rare document types with too few samples to train on |
The cloud vendors' own documentation reflects this split. Microsoft describes separate custom template models for static layouts and custom neural models for mixed-type documents. LLM providers document direct document reading as well; Anthropic, for instance, lists extracting key information from legal documents and converting document information into structured formats among the uses of its PDF support.
5. Validation rules
Extraction tells you what the document says. Validation tells you whether to believe it. Typical rules:
- Arithmetic: line items add up to the subtotal; subtotal plus tax equals the total.
- Format: dates are real dates, tax IDs and bank details match the expected pattern.
- Cross-reference: the supplier exists in your vendor list, the purchase order number is open, the invoice number has not been seen before.
- Plausibility: the amount is within a normal range for that supplier, the date is not in the future.
These rules are ordinary deterministic code, and they catch a large share of extraction errors regardless of which extraction method produced them. They are also the part most specific to your business.
6. Human review queue
Documents that fail a rule, or where the extraction engine reports low confidence, go to a person. A good review screen shows the original page next to the extracted fields, highlights the doubtful value, and lets the reviewer correct it in a few seconds. Corrections should be stored, because they are your best data for improving the system later.
7. Export to accounting or ERP
Approved data is posted to the destination system through its API, or through a file import where no API exists. This stage needs care with duplicates, partial failures and retries so that an invoice is never posted twice, and it needs an audit trail linking each record back to its source document and to whoever approved it.
Accuracy, and why the review step stays
Vendors and integrators sometimes quote a single accuracy percentage. Treat any such figure with caution, for three reasons.
Field accuracy is not document accuracy. If a document has fifteen fields and each is right most of the time, the chance that all fifteen are right on the same document is noticeably lower. What matters operationally is the share of documents that pass through with no human touch and no error.
Accuracy depends on your documents. Clean digital PDFs from large suppliers behave very differently from faxed, stamped, handwritten or photographed pages. A number measured on someone else's sample tells you little.
Errors are not equally costly. A misspelled address line is an annoyance. A wrong bank account number or a misplaced decimal point in a payment amount is a loss.
The practical answer is to use confidence scores and validation rules to decide which documents a person must look at. This is a standard pattern rather than a workaround. AWS documentation for its human review integration with Textract describes triggering review when a field's confidence falls in a specified range, when specific fields are missing, or for a random sample of documents. (The same page notes that this particular AWS service, Amazon A2I, is no longer open to new customers, which is one reason many teams build the review queue themselves.)
The random sample matters. Confidence scores are a model's estimate of its own reliability, and a model can be confidently wrong. Regularly checking a sample of documents that passed automatically is the only way to detect that. LLM-based extraction needs this most, because a language model can produce a well-formed, plausible value that is not on the page.
Over time the review share should fall as rules and models improve. It should not be designed to reach zero.
Which documents suit automation
| Document type | Suitability | Why |
|---|---|---|
| Supplier invoices, receipts, utility bills | High | High volume, well-understood fields, pretrained models available, totals can be checked arithmetically |
| Your own printed or digital forms | High | You control the layout, so templates work well |
| ID documents, standard tax and payroll forms | High, with stricter data protection | Fixed formats and pretrained models, but sensitive personal data |
| Purchase orders, delivery notes, bank statements | Medium to high | Structured but layouts vary; long tables need testing |
| Handwritten forms | Medium | Handwriting recognition exists but error rates rise; plan for more review |
| Contracts and leases | Medium | Useful for finding parties, dates, renewal and termination terms; not a substitute for legal reading |
| Rare, one-off or heavily damaged documents | Low | Too few to justify setup; manual entry is cheaper |
A useful rule of thumb: automation pays off when volume is steady, the fields you need are the same each time, and there is something to validate against. Contract extraction is best treated as a triage and search aid. It can populate a register of key dates and flag unusual clauses for a lawyer, but the interpretation of a clause remains a professional judgment.
Cloud document AI services vs a custom pipeline
The three large cloud providers each offer a document service. Capabilities below are taken from their official documentation as checked on 2026-10-04; they change often, so confirm before you design around any of them.
| Service | What the vendor documents |
|---|---|
| Amazon Textract | Text, form and table extraction; a Queries feature for asking for specific information; an AnalyzeExpense API for invoices and receipts; an AnalyzeID API for US-issued ID documents; Custom Queries using adapters trained on your own documents |
| Google Cloud Document AI | Processors grouped as digitize (Enterprise Document OCR), extract (Form Parser, Layout Parser, Custom Extractor including a generative AI option, and pretrained processors such as Invoice Parser) and classify (Custom Classifier, Custom Splitter) |
| Azure Document Intelligence | Read and Layout models; prebuilt models including invoice, receipt, ID document, contract, bank statement and US tax forms; custom template, custom neural and composed models; custom classifiers. Microsoft positions it alongside Azure Content Understanding, which offers LLM-powered analyzers for unstructured content |
"Cloud service or custom build" is slightly the wrong question. Almost every real system is custom software wrapped around one or more of these engines. The real decision is how much of the pipeline you build.
| Option | Choose it when | Watch out for |
|---|---|---|
| Off-the-shelf product (accounts payable or expense tool with capture built in) | Your need is standard invoice or receipt processing and the product already connects to your accounting system | Limited control over rules, review flow and where data is processed |
| Cloud document AI plus custom workflow | You have your own validation rules, several document types, or a destination system that packaged tools do not support | Per-page usage charges, vendor dependence, changes to service features |
| Fully custom pipeline with self-hosted OCR and models | Documents cannot leave your environment, or volume is high enough that usage charges outweigh engineering cost | You own accuracy, infrastructure and model maintenance |
Cost is driven by page volume, the number of document types, how many systems you integrate with, how complex the validation rules are, and how polished the review screen needs to be. The extraction engine is rarely the largest part of the effort. If you plan to use an outside team for the build, our guide on how to outsource software development from the US or UK covers contracts, IP and data handling terms.
Data protection and retention
Invoices, forms and contracts contain personal data, bank details and commercially sensitive terms. The points below are a high-level summary, not legal advice; confirm your obligations with a qualified adviser.
- Know where documents are processed. Check which region a cloud service processes and stores data in, whether the provider retains inputs, and whether inputs can be used to improve its models. Get the answers from the provider's current terms, not from a sales summary.
- UK: a written contract with processors. The UK Information Commissioner's Office states that whenever a controller uses a processor there must be a written contract in place, and the same applies to sub-processors. A cloud document service and a development partner with access to live documents are both likely to fall into this category.
- UK: storage limitation. Under UK GDPR you must not keep personal data for longer than you need it and should be able to justify your retention periods.
- US: no single federal privacy law covers all businesses, and sector and state rules vary. As a baseline, the Federal Trade Commission's business guidance advises companies to keep only what they need, protect what they keep and properly dispose of what they no longer need.
- Tax record retention pulls the other way. In the UK, limited companies must keep records for 6 years from the end of the last company financial year they relate to, longer in some cases. In the US, the IRS describes retention periods that depend on the situation, commonly 3 years and longer in specific circumstances, with employment tax records kept for at least 4 years.
The design consequence is that retention should be set per data store, not once for the whole system. The source document and posted record may need to be kept for years in your system of record. Temporary copies in processing queues, OCR output, logs and review screens usually do not, and should be deleted on a short schedule. Restrict who can open the review queue, and log who viewed and changed what.
How to run a pilot on real documents
- Pick one document type and one destination. Supplier invoices into the accounting system is a common first choice.
- Collect a representative sample. Pull real documents from a normal period, including the ugly ones: scans, photos, multi-page invoices, credit notes, foreign currency. A sample of only clean PDFs will flatter every tool.
- Create the answer key. Have someone record the correct value of every field for every sample document. This is tedious and it is the most valuable work in the pilot, because without it you cannot measure anything.
- Agree metrics in advance. Field-level accuracy, the share of documents that pass with no human touch, the error rate among those auto-passed documents, and handling time per reviewed document.
- Test more than one extraction approach on the same sample, for example a pretrained invoice model and an LLM-based extractor.
- Add validation rules and set thresholds. Measure again. Rules usually change the picture more than switching engines does.
- Run in shadow mode. Process live documents in parallel with the existing manual process for a few weeks and compare results before anything is posted automatically.
- Decide with the numbers. Compare the measured time saved against build and running costs, and expand to the next document type only if the first one holds up.
Before sharing real documents with any vendor or development partner for a pilot, put confidentiality and data processing terms in place, or redact the sample.
For technical readers
- Process asynchronously. Use a queue between capture, extraction and export. Amazon Textract, for example, handles multi-page documents through asynchronous operations, and in any design you need retries and a dead-letter path.
- Make export idempotent. Derive a key from a document hash plus supplier and invoice number so a retry cannot create a second posting.
- Keep provenance. Store, per field, the raw extracted value, the confidence, the page location, the model or prompt version, and any human correction. This supports audit and regression testing.
- Constrain LLM output. Require a fixed schema, validate it in code, ask for the source text span for each value and verify that span exists in the OCR text. Never let model output reach the ERP without passing the same rules as every other path.
- Treat document content as untrusted input. Text inside a document can contain instructions aimed at a language model. The extraction step should have no ability to take actions; it returns data only.
- Abstract the engine. Put the OCR and extraction provider behind your own interface so you can switch or combine services, and keep the labeled pilot set as a regression suite to rerun whenever a model or prompt changes.
If the next step after extraction is having software act on the data, for example chasing a supplier about a mismatch, read our comparison of AI agents, chatbots and workflow automation before choosing an architecture.
Frequently asked questions
What is the difference between OCR and intelligent document processing?
OCR converts an image of a page into text. Intelligent document processing is the whole pipeline around it: classifying the document, extracting named fields, validating them, routing exceptions to people and exporting the result. OCR alone gives you words; IDP gives you a record you can post to a system.
Can document automation be 100 percent accurate?
No. Scan quality, unusual layouts and handwriting all cause errors, and LLMs can produce plausible wrong values. The realistic goal is that most documents pass automatically, the doubtful ones are reviewed quickly by a person, and a sampled check confirms that the auto-passed ones are correct.
Should we use an LLM instead of a traditional document AI service?
Not instead of, in most cases. Pretrained models for invoices, receipts and IDs are fast, consistent and return confidence scores. LLMs are most useful for long unstructured documents such as contracts, for rare document types, and as a second opinion on difficult fields. Test both on your own sample.
How many sample documents do we need for a pilot?
There is no fixed number. You need enough to cover the real variety: your main suppliers or form versions, plus poor scans and edge cases such as credit notes. A sample that represents a normal month is more useful than a larger one made up only of clean files.
Is it safe to send invoices and contracts to a cloud AI service?
It can be, if you have checked the provider's terms on processing location, retention and use of your data, have the required contract in place, and limit what you send to what is needed. Some organizations with strict confidentiality requirements choose self-hosted processing instead.
Can extracted data go directly into QuickBooks, Xero, SAP or another system?
Generally yes, where the destination offers an API or import format and your subscription allows access to it. The integration work is in mapping fields, matching suppliers and purchase orders, and preventing duplicates. Check the current API documentation for your specific system and edition before planning.
Conclusion
Document processing automation is a pipeline, not a single AI feature: capture, OCR, classification, extraction, validation, human review and export. The extraction engine gets the attention, but validation rules, the review queue and the integration with your accounting or ERP system decide whether the result can be trusted. Start with one document type, build an answer key from real documents, measure the no-touch rate and the error rate, and set retention rules for every place a document is stored.
Entrant Technologies builds web applications and custom software, including workflow and integration work of this kind. If you would like a second opinion on a pilot plan or on fitting document extraction into an existing system, you can get in touch here.