Document capture

Document capture is the automated extraction of structured data from invoices, receipts and bank statements – supplier, date, amount, VAT rates, invoice number, line items. Classical optical character recognition is increasingly complemented by learning models that also handle unfamiliar layouts without prior configuration.

How document capture works

The sequence has four steps: image preparation (deskewing, cropping, contrast), text recognition, mapping the recognised text fragments to fields, and a plausibility check against master data and arithmetic. The third step is the hard one. Whether “invoice number” sits top left or bottom right differs from supplier to supplier, and this is where learning models have brought the biggest progress: they recognise a field by its meaning in context rather than by its position.

Realistic hit rates

For header data – supplier, invoice number, date, gross amount – high hit rates are achievable today once the first documents from a supplier have been processed. Line items are considerably harder: article numbers, quantities, unit prices and discount lines drop off noticeably on multi-page invoices and unfamiliar layouts.

The figure that matters more in practice is not the hit rate but the reliability of the confidence score: does the system recognise when it is unsure? A system with a slightly lower hit rate that flags its doubtful cases correctly is better in operation than one that outputs everything with high confidence – including the wrong answers.

Where it reliably breaks

Handwritten documents and till receipts. Photographs with shadows, creases or skew. Foreign-language invoices with unfamiliar structures. Consolidated invoices with many line items. Credit notes recognised as invoices. And invoices without a clear reference, where the match to a transaction depends not on the document but only on context.

Why the e-invoice makes the step redundant

An e-invoice is a structured data record, not an image. All fields are already machine-readable – there is nothing left to extract. That removes precisely the step where most errors have originated.

Since 1 January 2025 all domestic German companies must be able to receive e-invoices in B2B transactions; the obligation to issue them applies from 2027 for companies with prior-year turnover above 800,000 euros and from 2028 for everyone. Document capture therefore becomes a transitional instrument for the shrinking volume of unstructured documents – foreign invoices, small-amount receipts, till slips – and loses importance in the core B2B process.

What matters at implementation

Not the advertised recognition rate, but three operational questions: how are doubtful cases presented and how quickly can they be corrected? Does the system learn from corrections? And how are the image and the extracted data archived together in a GoBD-compliant way?

Synonyme:
OCR, invoice capture, document recognition, data extraction
Englischer Begriff:
Document capture / OCR
Last updated:
September 2, 2026