Overview
- Mistral launched OCR 4 on June 23 as its fourth-generation document intelligence model that converts images and PDFs into structured outputs formatted for retrieval-augmented generation workflows.
- OCR 4 adds paragraph- and block-level bounding boxes, typed block labels for elements like titles and tables, and per-word and per-page confidence scores to help systems and humans validate extracted content.
- The model supports 170 languages and claims improved handling of rare and low-resource tongues while running at high throughput—Mistral reports speeds up to about 2,000 pages per minute on a single GPU.
- Mistral published benchmark and evaluation figures that include an OlmOCRBench leaderboard score of 85.20 and a reported 72% average human-preference win rate in blind tests, though those metrics vary in independence and scope.
- OCR 4 is offered with tiered pricing and enterprise deployment options, including single-container on-premises installs and availability in Microsoft Foundry, which makes it easier for firms to keep sensitive documents in-house and feed structured data into downstream AI systems.