Visual Document Understanding: Next-Gen OCR with LayoutLM and Donut

📅 January 1, 2026⏱️ 35 min read🏷️ Multimodal AI

For decades, "extracting data from a PDF" meant one thing: Tesseract OCR. You would get a wall of raw text, losing all spatial information. If "Total: \$500" was at the bottom right, and "Invoice #123" was at the top left, Tesseract just gave you a flat string. You had to write hundreds of fragile RegEx rules to parse it.

That era is over. Visual Document Understanding (VDU) models don't just "read" text; they "see" the document. They understand that bold text is a header. They understand that data in a grid is a table. They understand that a signature at the bottom validates the form.

In this guide, we will compare LayoutLM (which combines Text + Layout + Image) and Donut (OCR-free visual document understanding) and build a pipeline to automate invoice processing.

Why Traditional OCR Failed Enterprise

Traditional OCR (Optical Character Recognition) engines like Tesseract or AWS Textract (in raw mode) are purely unimodal. They look at pixel clusters and predict characters. They discard the most valuable signal in a document: Spatial Layout.

The "Key-Value" Problem

Consider an invoice. The word "Total" is a Key. The number "\$500.00" is the Value. They might be separated by 3 inches of whitespace.

To a human, the relationship is obvious because they are aligned horizontally. To Tesseract, they are just two disconnected strings. VDU models treat the 2D position (x, y coordinates) of text bounding boxes as a first-class input feature, allowing them to "link" Keys and Values spatially.

Theory: LayoutLM (Text + Layout + Image)

LayoutLM (by Microsoft) revolutionized this field by pre-training a BERT model on millions of scanned documents. It accepts three inputs:

LayoutLMv3 achieves state-of-the-art results on tasks like Invoice Information Extraction (finding the Date, Total, Vendor).

Evolution of LayoutLM

  • LayoutLMv1: Added Layout Embeddings (x,y coordinates) to BERT. Still relied on external OCR.
  • LayoutLMv2: Added a visual backbone (ResNet) to process image patches directly alongside text.
  • LayoutLMv3 (The King): Removed the dependency on CNNs. It uses a Vision Transformer (ViT) approach where image patches are linear embeddings just like text tokens. This "Unified Multimodal Transformer" approach is cleaner and faster.

Theory: Donut (Document Understanding Transformer)

LayoutLM still relies on an external OCR engine (like Tesseract) to get the text first. If the OCR fails (e.g., blurry text), LayoutLM fails.

Donut (Document Understanding Transformer) is OCR-Free. It is an Image-to-Text model.

It uses a Vision Encoder (Swin Transformer) to process the raw image pixels and a Text Decoder (BART) to generate the structured JSON output directly. It doesn't output a list of words; it outputs: {"invoice_no": "123", "total": "500"}.

Python Implementation: Fine-tuning Donut on Invoices

Let's see how to fine-tune Donut to extract data from a custom receipt dataset.

from transformers import DonutProcessor, VisionEncoderDecoderModel
from datasets import load_dataset

# 1. Load Pre-trained Donut
processor = DonutProcessor.from_pretrained("naver-clova-ix/donut-base-finetuned-cord-v2")
model = VisionEncoderDecoderModel.from_pretrained("naver-clova-ix/donut-base-finetuned-cord-v2")

# 2. Prepare Image
image = load_dataset("naver-clova-ix/cord-v2", split="validation")[0]["image"]

# 3. Inference
# Convert image to pixel values
pixel_values = processor(image, return_tensors="pt").pixel_values

# Generate JSON output
# We force the model to start with a special <s_cord-v2> token
task_prompt = "<s_cord-v2>"
decoder_input_ids = processor.tokenizer(task_prompt, add_special_tokens=False, return_tensors="pt").input_ids

outputs = model.generate(
    pixel_values,
    decoder_input_ids=decoder_input_ids,
    max_length=512,
    early_stopping=True,
    pad_token_id=processor.tokenizer.pad_token_id,
    eos_token_id=processor.tokenizer.eos_token_id,
    use_cache=True,
    num_beams=1,
    bad_words_ids=[[processor.tokenizer.unk_token_id]],
    return_dict_in_generate=True,
)

# 4. Parse Sequence to JSON
sequence = processor.batch_decode(outputs.sequences)[0]
sequence = sequence.replace(processor.tokenizer.eos_token, "").replace(processor.tokenizer.pad_token, "")
sequence = sequence.split("<s_cord-v2>", 1)[-1].strip()

import json
print(json.dumps(processor.token2json(sequence), indent=2))

The output isn't a mess of bounding boxes. It's clean JSON: {"menu": [{"nm": "Pizza", "price": "15.00"}]}.

Java Implementation: The Document Processor Service

In a real enterprise, you have a queue of PDFs coming in from emails. We need a Spring Boot service to consume these files, convert pages to images, and send them to the model.

package com.devmetrix.vdu;

import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.rendering.PDFRenderer;
import org.springframework.stereotype.Service;
import java.awt.image.BufferedImage;
import java.io.File;

@Service
public class DocumentProcessor {

    private final InferenceClient inferenceClient;

    public DocumentProcessor(InferenceClient inferenceClient) {
        this.inferenceClient = inferenceClient;
    }

    public void processInvoice(File pdfFile) throws Exception {
        try (PDDocument document = PDDocument.load(pdfFile)) {
            PDFRenderer pdfRenderer = new PDFRenderer(document);

            // Convert first page to image (Donut takes images)
            BufferedImage bim = pdfRenderer.renderImageWithDPI(0, 300);

            // Send to Python Inference Service
            InvoiceData data = inferenceClient.extractData(bim);

            // Save to DB
            saveToDatabase(data);
        }
    }

    private void saveToDatabase(InvoiceData data) {
        System.out.println("Vendor: " + data.getVendor());
        System.out.println("Total: " + data.getTotalAmount());
    }
}

Security: Redacting PII in Images

Uploading raw documents (drivers licenses, bank statements) to the cloud is risky.

Visual Redaction

Before sending a document to a 3rd party model API, you must redact PII visually. We can use a lightweight local model (like YOLOv8 trained on PII detection) to find bounding boxes of faces, credit card numbers, and SSNs.

Then, use OpenCV to draw a black rectangle over those coordinates. Only then do you send the image to the powerful (but untrusted) cloud VDU model. This ensures that even if the cloud provider is hacked, your users' PII is safe.

🌌
Purple Dream
Active Theme