Multimodal AI for documents: test the whole page
Build a document evaluation that checks visual detail, layout, extraction accuracy, and what happens when the evidence is unreadable.

A model that accepts images can be useful for documents containing tables, diagrams, and scanned text. Accepting an image does not guarantee that it will read every detail correctly. The quality of the input and the way the application prepares it are part of the system being evaluated.
Google's image-understanding documentation describes image inputs and image-based tasks for its API. Use the documentation for the exact model and endpoint you select. Input limits and supported options belong to that specific serving configuration, not to the word multimodal in general.
Start with the fields that matter
Imagine an assistant reading delivery slips. It must extract an order identifier, a date, and received quantities. A fluent description of the page is not the target. Correct field values and a clear indication of unreadable entries are the target.
Create reference answers for each field. Preserve leading zeros in identifiers and distinguish a blank quantity from zero items received. Decide how the output should represent uncertainty before testing. If the system must guess to satisfy a required field, the contract needs adjustment.
Evaluate input preparation
Test the same document through the pipeline users will use: phone capture, upload, resizing, compression, and page selection. A model evaluated on a clean original PDF may behave differently after the application converts it to a small image.
Inspect orientation, glare, and cropping. Confirm that important details remain legible after any size reduction. If processing a page in tiles, preserve their order and relationships. An isolated table cell may lose the row label needed to interpret it correctly.
Separate reading from interpretation
For a delivery slip, first inspect whether the model read “12” correctly. Then inspect whether it associated that value with the right product. Finally, check whether the application interpreted it as ordered or received quantity. These are distinct failure points.
Track exact matches for identifiers and appropriate tolerance rules for measurements. Do not use one similarity score for every field. A nearly correct order identifier may refer to another shipment, while a differently formatted date can still represent the same day.
Make evidence review convenient
Keep the page reference and, when the system supports a trustworthy location mapping, a view of the relevant region. Let a reviewer inspect the source alongside the extracted values. Test any coordinate mapping after resizing rather than assuming a box still points to the correct place.
When the input is unreadable, ask for a clearer capture or route the document to manual review. Avoid presenting a model's confident tone as proof that the text was visible. A useful exception path can matter more than squeezing another answer out of a poor scan.
Test document variety
Include several layouts, languages used by your audience, handwritten amendments, empty pages, and repeated page headers. Keep documents from the same template family together when designing holdout groups if near-duplicates would make the evaluation misleading.
Review failures by source and preparation step. If most errors come from excessive compression, changing models may be an expensive distraction. Improve the failing stage, rerun the same cases, and record whether the final field-level accuracy and review effort actually improved.