Free Lesson
How to choose an OCR Model
Part of AI Product Engineering
45 min
Jul 10, 2026 1:00 PM
Virtual (Zoom)
In this video
What you'll learn
When an OCR API is enough
Understand when AWS Textract, Datalab, or a similar API is the right choice because you want OCR without infrastructure
What self-hosted OCR actually requires
See the basic pieces of an OCR stack: model serving, document ingestion, batching, storage, observability, and output
What open models make possible
Learn where models like LightOnOCR, Chandra, and DotsOCR fit and improve outputs
How output format changes model selection
Compare document structure extraction, plain text, markdown, and layout-aware outputs
What you gain by owning the stack
See what you get from hosting having your own scale to 0 infrastructure (like with modal)
Why this topic matters
OCR can start as a simple API call, and that is often the right choice. But teams hit limits when they need better structure, custom outputs, lower cost, or more control. Joe Barrow will show when to use a managed OCR API, what a manageable OCR stack looks like, and what open models make possible.
You'll learn from

Joe Barrow
Senior Research Scientist at Adobe Research

Hamel Husain
ML Engineer with 20+ years of experience