Mistral AI has launched OCR 4, its latest document intelligence model that goes beyond traditional text extraction to provide fully structured document representations. This model includes bounding boxes, block-type classification, and confidence scores at the word and page levels, supporting 170 languages across multiple formats. Designed for enterprise use, OCR 4 can be deployed on-premises, ensuring data sovereignty especially relevant for regulated industries avoiding U.S.-jurisdiction cloud APIs. The model’s structural approach allows detailed traceability and semantic organization of document content, reducing integration complexity and enabling human-in-the-loop verification workflows. Early tests show a 72% preference over competitors in real-world scenarios, though benchmark results vary. This release is timely amid geopolitical shifts affecting AI accessibility, reinforcing Mistral’s focus on European AI sovereignty. Positioned as a foundation for a broader AI stack, OCR 4 targets enterprises requiring compliance, cost-efficiency, and performance at scale, complementing Mistral’s expanding AI capabilities and growth ambitions.
Back