Documents · 9 min

Extract PDFs and scans with AI: check text, tables and numbers

A PDF that looks clear on screen may supply incomplete or disordered text to an AI tool. Before summarising an archive or filling a table, inspect what was actually extracted and preserve links to original pages.

Laptop and document folders on a desk for preparing document extraction.
AI-generated illustration.
Key points

Separate three operations

OCR recognises characters in an image. Layout analysis reconstructs blocks and tables. Structured extraction assigns values to fields. Success at one stage does not validate the next.

  • Original preserved
  • Values checked against the page
  • Uncertainty stays visible

From document to usable data

  1. 1. Identify the document

    Select and copy a passage, then compare it with the screen. Distinguish text PDFs, scans and mixed documents. Locate rotated pages, columns, annotations and tables. Check rights and sensitivity before upload.

  2. 2. Prepare pages

    Work on a copy. Straighten pages and inspect contrast, noise and clipped areas. Tesseract documentation describes the effect of image quality and segmentation on recognition. Do not alter the image to match presumed text.

  3. 3. Define fields

    Write a simple schema: identifier, value, unit, page and supporting excerpt. Allow missing values and uncertainty notes. Distinguish zero, empty, unreadable and not applicable.

  4. 4. Test difficult cases

    Include a clear page, a multipage table, footnote and unclear value. Compare with human transcription. A polished summary must not hide incorrect table reading.

  5. 5. Reconcile values

    Check signs, decimal separators, units, dates and headers. Recalculate totals where available. Keep raw values before conversion. Do not infer missing data merely to make totals agree.

  6. 6. Validate before automation

    Define critical fields requiring systematic review. Measure correction time per document and retain a control sample. Changes to engine, prompt or source format require another trial.

Put the method to work

Practical case

Prepare five authorised pages with two tables, a rotated page and an unreadable value. Define ten fields without filling missing values.

Evidence to keep

Keep originals, raw text, normalised values, page numbers and human corrections.

Make the decision

Accept the batch only after reviewing every critical field and leaving unknown values explicitly unknown.

Measure extraction

Fidelity

Exact values compared with the page before normalisation.

Traceability

Every field has a verifiable page and excerpt.

Structure

Cells remain linked to the correct headers.

Useful cost

Preparation and correction time per accepted document.

6 starting points

Explore document tools

These tools cover documents, local models and workflows. Check actual supported formats and quality on your own sample separately.

How is this selection produced?

Active services are distributed across guide-related categories, then ordered by editorial highlighting and internal score. This does not assess security, compliance or performance on your use case. Methodology.

Explore the full category

Explore tools for this task

  • Gemini Notebook — Explore a set of reports, prepare a synthesis or find useful passages in a defined corpus. Select relevant documents and remove obsolete versions first.
  • Ollama — Test a model on your computer or provide a backend for a local application. Check hardware compatibility and model licensing first.
  • n8n — Connect applications, transform data and orchestrate repeatable processes with AI steps. Identify inputs, outputs and the owner of each approval first.
  • LM Studio — Test a local model using authorized text and assess its quality on your hardware. Separate response speed, memory consumption and correctness.
  • Open WebUI — Provide a common interface for authorized models. Define users, available connections and documents each group may access.
  • DeepL — Produce a draft translation for review against defined terminology. Prepare proper names, preserved terms and number and date conventions.

All profiles organized by family →

Comparison frameworks and cost per accepted result →

Related tool families

Frequently asked questions

Does every PDF need OCR?

No. Some PDFs already contain text. First test extraction and reading order; OCR adds a stage that may introduce errors.

How should an unreadable cell be handled?

Leave it unknown, identify the page and request review. A plausible value is not a transcription.

Can important amounts be automated?

Define controls suited to their consequences, including critical-field review. A high overall score can hide a rare, costly error.

The references below expand on the concepts and checks discussed. Scenarios and trial frameworks remain editorial proposals; provider documentation describes its own product rather than an independent benchmark.

Official sources

Continue with another guide