kreuzberg: Kreuzberg is a high-performance, polyglot library designed to extract text and metadata from over 57 file formats, including comprehensive OCR capabilities. Built with a Rust core, it offers native speed processing, memory efficiency, and the ability to generate embeddings without requiring a GPU, making it highly versatile for various data extraction and processing tasks.; all2md: all2md is a universal Python library for converting over 40 document formats, including PDFs and Office files, into clean Markdown and back again. It features an AST-based architecture, a powerful CLI, and native integration for AI assistants and LLM workflows.
Automated extraction of text, metadata, and structured data from diverse document types.
Embed document conversion in Python applications, documentation systems, or CMS.