Free Lesson

How to choose an OCR Model

Part of AI Product Engineering

45 min
Jul 10, 2026 1:00 PM
Virtual (Zoom)

In this video

What you'll learn

When an OCR API is enough

Understand when AWS Textract, Datalab, or a similar API is the right choice because you want OCR without infrastructure

What self-hosted OCR actually requires

See the basic pieces of an OCR stack: model serving, document ingestion, batching, storage, observability, and output

What open models make possible

Learn where models like LightOnOCR, Chandra, and DotsOCR fit and improve outputs

How output format changes model selection

Compare document structure extraction, plain text, markdown, and layout-aware outputs

What you gain by owning the stack

See what you get from hosting having your own scale to 0 infrastructure (like with modal)

Why this topic matters

OCR can start as a simple API call, and that is often the right choice. But teams hit limits when they need better structure, custom outputs, lower cost, or more control. Joe Barrow will show when to use a managed OCR API, what a manageable OCR stack looks like, and what open models make possible.

You'll learn from

Joe Barrow

Joe Barrow

Senior Research Scientist at Adobe Research

Hamel Husain

Hamel Husain

ML Engineer with 20+ years of experience

See all products from Hamel Husain & Shreya Shankar