Reading paper and PDF documents with OCR
OCR turns a scan or a PDF into invoice data you can work with, instead of a file someone retypes. It applies to purchase invoices, sales invoices and expense receipts, and each has its own screen under E-Invoice > OCR and under Receipt Expenses.
The flow
- Upload the file, or let it arrive as an e-mail attachment in the Email Inbox.
- The document waits its turn — the status is OCR Waiting.
- It is read, and the extracted data is attached to a document you can open.
- You check and correct the result before using it.
The queue matters: OCR is not instant and capacity is shared, so a batch of fifty documents does not finish in the time one does. A document sitting in OCR Waiting is normal; one that stays there for hours is worth reporting.
The statuses you will see
| Status | Meaning |
|---|---|
| OCR Waiting | Uploaded, queued, not read yet |
| OCR Processed / Scanned by OCR | Read successfully; the data is ready for review |
| OCR Failed | The document could not be read |
An OCR Failed result is usually the file rather than the platform: a photograph at an angle, a scan at low resolution, a PDF that is one flat image of poor quality, or several invoices in one file. Rescanning at 300 dpi, straight, one invoice per file, fixes most of them.
Always review the result
Treat the output as a draft. OCR reads what it sees, and what it sees on a bad scan can be wrong in ways that look plausible — a transposed amount, a date read from the wrong line, a VAT rate taken from a footer.
Check the amount, the VAT breakdown, the invoice number and the supplier before you accept a document into your records. A wrong figure accepted here becomes a wrong figure in your reports and your VAT position.
Two engines
The company setting Invoice OCR Source selects which engine reads invoices — the Docnova parser or Google's. They behave differently on the same document, so if a supplier's layout consistently fails, that setting is worth changing before you give up on the format.
Expense OCR needs credit
Reading expense receipts consumes Docnova AI tokens. When they run out the platform says so explicitly rather than failing quietly, and the fix is the AI allowance on your plan, not the document. Invoice OCR and expense OCR are counted separately, so expenses can stop while invoices carry on.
OCR does not make a document compliant. A read PDF is still a PDF, and under German law a sonstige Rechnung. OCR gets the data into your books; it does not replace the supplier's duty to issue a structured invoice once their phase of the mandate has started.