PaddleOCR-VL Now Integrated into llama.cpp for Multilingual OCR
2 min readThe integration of PaddleOCR-VL into llama.cpp extends the inference engine's capabilities from pure text generation into multimodal document understanding. At 900M parameters, this model is lightweight enough to run on modest hardware while delivering strong performance for multilingual optical character recognition—addressing a frequent bottleneck in local document processing pipelines.
For practitioners building local AI systems that need to process scanned documents, PDFs, or images, this is transformative. Previously, OCR often required either expensive cloud APIs or separate, specialized tools. Now, the entire pipeline—image-to-text via PaddleOCR-VL, followed by reasoning/summarization via larger LLMs—can run entirely on-device. The integration into llama.cpp means it works across Windows, macOS, Linux, and mobile platforms with the same optimized inference backend.
Community feedback suggests this is the strongest open-source multilingual OCR available, making it a critical building block for local knowledge workers, researchers, and enterprises handling sensitive documents. The addition to llama.cpp's latest release signals the ecosystem's maturation toward practical, multi-capability local AI.
Two things to know before you run it
Since writing this we have put the model through a 412-page job on a 16GB M1 Pro, and two properties of it are not obvious from the release notes.
It is an element-level model, so a full letter-size page is downscaled past legibility and it responds by inventing fluent, well-formed, entirely wrong text rather than returning noise or an error. And calling it through llama-mtmd-cli per crop instead of a resident llama-server cost 75x on identical inputs: 7,351 seconds against 98.
→ PaddleOCR-VL on Apple Silicon: Crop to Blocks, Keep the Model Resident has the segmenter, the measured timings, the --jinja flag that fails 0.6% of requests with the correct answer inside the error, and what was not tested.
Source: r/LocalLLaMA · Relevance: 8/10