Back to News & insightsEngineering

OCR and document AI: recover the page before trusting the text

Preserve reading order, tables, and evidence when extracting documents.

Editorial guide · Updated September 28, 2026 · 7 min read
A translucent sheet passes through a silver scanning arch and separates into ordered strips.

A scanned report can contain perfectly recognizable letters and still produce unusable extracted text. The first column may be interleaved with the second. A table heading may become detached from its values. A footnote may appear in the middle of a sentence. These are not merely cosmetic problems: they change the relationships that give a document its meaning.

Optical character recognition, or OCR, converts visual text into machine-readable symbols. Document AI extends the problem to structure and interpretation. A reliable workflow needs both. Before asking a language model to summarize a document, establish what was recovered from the page, where it came from, and which parts remain uncertain.

Inspect the input before choosing a model

A PDF may contain selectable text, scanned page images, or a mixture of both. It may also contain an old OCR layer that looks plausible but has errors. Running every file through the same expensive visual model wastes resources and can replace accurate embedded text with a less reliable reconstruction.

Begin with a small inspection step. Identify page dimensions, available text, image coverage, rotation, and obvious extraction problems. Route clean digital text through a simpler path and reserve OCR for pages that need it. Preserve the original file so that later corrections can be traced back to the evidence rather than to another transformed copy.

Research treats layout as part of understanding

LayoutLMv3 combines textual and visual learning objectives for document tasks. Nougat studies converting academic document images into a structured markup representation. These approaches illustrate that reading a page requires more than recognizing characters: spatial relationships and output structure can be central to the task.

Read the original research paper on arXiv

Read the original research paper on arXiv

Neither paper establishes a universal best tool for every document collection. Receipts, historical scans, forms, and scientific papers have different failure patterns. The pipeline below is an original design example for a fictional archive of technical reports, with choices that should be evaluated on the actual documents being processed.

Preserve geometry alongside text

Store text blocks with page numbers and bounding regions where the extraction system provides them. Geometry lets a reviewer locate the original passage and helps downstream systems distinguish a title, a sidebar, and a table cell. Plain text alone discards this evidence and makes later repair unnecessarily difficult.

Use a consistent coordinate convention and retain the page dimensions. If a page is rotated or resized before recognition, record the transformation. Otherwise, a highlight may point to the wrong region even when the extracted words are correct. A trustworthy document viewer depends on this mundane alignment work as much as on recognition quality.

Reading order is a model of the page

For a two-column report, reading order usually follows one column before the next. A full-width heading, a floating figure, or a sidebar can interrupt that pattern. Sorting every text box by its vertical position may alternate between columns and create a paragraph that never existed in the source.

Represent the page as regions with relationships rather than as one flat list of words. Then test the proposed reading order on representative layouts. Include pages with headings across columns and captions below figures. When the layout is ambiguous, retain separate blocks instead of confidently merging them into a misleading continuous narrative.

Tables need relationships, not just numbers

Suppose a fictional report lists three test conditions across columns and several measurements down rows. If extraction returns only the numbers, the result is incomplete. Each value needs its row label, column label, units, and any note that changes its interpretation. A correct numeral in the wrong cell is still an incorrect extraction.

Evaluate tables at the relationship level. Ask whether a query for a particular condition returns the correct value and unit, not merely whether the page contains the expected digits. Preserve merged headings and multi-page continuation context when they matter. A language model should not be asked to reconstruct missing relationships solely from what seems plausible.

Distinguish recognition from interpretation

An OCR system may read a string that could be a product code or an ordinary word. A later interpretation step can use context to propose a normalized value, but it should not overwrite the original observation silently. Keep the raw extraction, the normalized value, and the reason for the change separately.

This distinction is especially useful for names, identifiers, and unusual technical notation. A spelling correction that improves ordinary prose can corrupt a valid part number. Let task-specific validation decide whether normalization is appropriate. When a value cannot be resolved, preserve uncertainty instead of replacing it with the most familiar-looking alternative.

Mathematical notation deserves its own evaluation

Superscripts, subscripts, fractions, and matrix layouts encode relationships that ordinary reading order cannot fully express. A missing minus sign can reverse a conclusion. A superscript rendered as a normal digit can turn an exponent into an unrelated number. High overall character accuracy does not guarantee reliable mathematical extraction.

For technical reports, create a separate sample of equations and symbol-heavy passages. Compare the recovered representation with the source visually and, where appropriate, structurally. Do not use a fluent natural-language summary as the only check. A summary may conceal the exact transcription error that a later calculation would depend on.

Confidence should direct review, not erase difficult pages

Recognition confidence can help prioritize human review, but its scale may differ across systems and document types. A highly confident error remains possible. Build a reviewed sample that includes both low-confidence regions and ordinary-looking regions so that the team can estimate what the score actually predicts.

Avoid dropping difficult pages without reporting the omission. An archive search that silently skips handwritten notes or faint scans presents an incomplete collection as if it were complete. Record processing status per page and expose meaningful gaps to downstream applications. An honest partial extraction is more useful than an apparently comprehensive but selective one.

Chunk after recovering structure

Search systems often divide documents into chunks for retrieval. If chunking happens before reading order and table relationships are repaired, the index preserves the extraction mistakes. A heading can be separated from the paragraph it describes, or a table note can be lost in another chunk.

Use document structure to guide chunk boundaries. Keep a short section with its heading and retain references to the original page regions. When a table is too large for one chunk, include enough repeated header context to interpret each part. Track the extraction revision so that an improved parser can trigger a controlled reindexing process.

Build a test set around downstream questions

Character error rate is useful, but the application may care more about finding the right specification or extracting a field accurately. Combine transcription checks with task-level questions. In the fictional archive, ask for a measurement under a named condition, the definition of a symbol, and the limitation attached to a result.

Include negative questions whose answers are absent from the document. This checks whether the system distinguishes missing evidence from extraction failure and from a genuinely stated answer. Keep these cases separate during review. They point to different improvements: better recognition, better retrieval, or more careful answer generation.

Plan for correction without rebuilding everything

Assign stable document and page identifiers. A reviewer should be able to correct one table or reading-order decision without replacing unrelated records. Keep an audit trail of the correction and regenerate only the derived artifacts that depend on it. This makes quality improvement practical as the collection grows.

Use a small release sample before processing the entire archive with a new model. Compare old and new extraction on difficult pages as well as average ones. A new system may improve prose while regressing on tables or technical symbols. Preserve the previous output until the new revision has passed the relevant checks.

A useful handoff includes the original page

When an assistant answers from an extracted document, let the reader inspect the source region. Show the page reference and avoid implying that normalized text is a verbatim transcription when it is not. For ambiguous passages, a direct view of the page can resolve questions that another generated explanation would only complicate.

Document AI succeeds when it carries meaning across representations without hiding what was lost. Accurate characters, coherent layout, preserved relationships, and traceable corrections form one connected system. Recover the page first; then ask the model to help people understand what the page actually says.

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.