kreuzberg
Kreuzberg is a high-performance, polyglot library designed to extract text and metadata from over 57 file formats, including comprehensive OCR capabilities. Built with a Rust core, it offers native speed processing, memory efficiency, and the ability to generate embeddings without requiring a GPU, making it highly versatile for various data extraction and processing tasks.
kreuzberg is currently grouped under Vision / Multimodal, which makes it easier to evaluate through workflow fit instead of isolated features alone. Based on the available data, it leans most heavily toward Extensible architecture with a plugin system for custom backends and processors. and Automated extraction of text, metadata, and structured data from diverse document types.. The listed license is MIT, which is useful when adoption constraints matter. It also shows measurable community traction with 8.7k GitHub stars.