AgentIndex icon
AgentIndex
ToolsCategoriesTrendingNewCompare
Submit Tool
ToolsCategoriesTrendingNewCompare
Home/
Data Processing/
pullmd
pullmd logo

pullmd

Active·★ 197·AGPL-3.0·Updated 2026-07-10
★ Trending★ Hidden Gem

PullMD is a self-hosted service that converts various content types, including web pages, documents, images, audio, and YouTube videos, into clean, readable Markdown format. It supports integration with AI agents, offering a token-efficient output and multiple API endpoints.

pullmd is currently grouped under Data Processing, which makes it easier to evaluate through workflow fit instead of isolated features alone. Based on the available data, it leans most heavily toward Converts URLs, documents, images, audio, and YouTube videos to clean Markdown. and Providing clean, distilled content from web pages, documents, and media to AI agents for better context understanding and token efficiency.. The listed license is AGPL-3.0, which is useful when adoption constraints matter. It also shows measurable community traction with 197 GitHub stars.

© 2026 AgentIndex.app|Built by a 10-year iOS Developer.
QYSGitHubBuy me a coffee ☕

Browse by Category

Code AssistantWorkflow AutomationRAG / Knowledge BaseMulti-AgentBrowser AutomationLLM InfraDev ToolingObservability

Not affiliated with Anthropic, OpenAI or Microsoft.

#URL-to-Markdown#Self-hosted#AI Agent Tooling#Content Extraction#Document Conversion#Multimodal Processing#Web Scraping
$ Install
$ mkdir pullmd && cd pullmd && curl -O https://raw.githubusercontent.com/AeternaLabsHQ/pullmd/main/docker-compose.yml && docker compose up -d
↗ Visit site★ GitHub
01

Features

01Converts URLs, documents, images, audio, and YouTube videos to clean Markdown.
02Self-hosted with a PWA frontend, REST API, and MCP server for AI agent integration.
03Auto-detects and processes Reddit and Hacker News threads with full comment trees.
04Supports high-quality PDF OCR and media captioning/transcription via OpenAI-compatible models.
05Emits token-efficient Markdown body by default, with rich metadata in YAML frontmatter.
02

Why choose it

+Converts URLs, documents, images, audio, and YouTube videos to clean Markdown.
+Providing clean, distilled content from web pages, documents, and media to AI agents for better context understanding and token efficiency.
+Covers 5 supported environments or platforms, which is helpful for broader deployment needs.
+Ships with a public repository and a AGPL-3.0 license, which makes adoption and review easier.
03

Trade-offs

!There are at least 8 related tools in the same category, so the best choice is easier to make after side-by-side comparison.
04

Compatibility

Docker
Deployment
Verified via docs
AI Agents
Integration
Verified via docs
Web Browsers
Frontend
Verified via docs
Python
Sidecars
Verified via docs
Node.js
Backend Runtime
Verified via docs
05

Quick start

1
$ mkdir pullmd
2
$ cd pullmd
3
$ curl -O https://raw.githubusercontent.com/AeternaLabsHQ/pullmd/main/docker-compose.yml
4
$ docker compose up -d
06

Use cases

↳Providing clean, distilled content from web pages, documents, and media to AI agents for better context understanding and token efficiency.
↳Creating dynamic content feeds (e.g., subreddit feeds) using stable share IDs that automatically refresh.
↳Building a personal or team-wide knowledge base by converting various online and local content into a unified Markdown format.
↳Integrating with large language models and chat agents (like Claude, ChatGPT) as a specialized tool for robust web and document fetching.
07

How it compares

≈pullmd sits in the Data Processing category, so it makes more sense to evaluate it alongside tools like mirascope instead of in isolation.
≈If your main need is closer to "Providing clean, distilled content from web pages, documents, and media to AI agents for better context understanding and token efficiency.", that use case is a better lens for comparison than broad feature checklists alone.
≈pullmd uses a AGPL-3.0 license, and community traction are both easier to judge in category context.
08

Alternatives

mirascope logo
mirascope★ 1.5k
LLM abstractions that aren't obstructions
vs →
FinanceToolkit logo
FinanceToolkit★ 5.1k
Transparent and Efficient Financial Analysis
vs →
OpenClaw logo
OpenClaw★ 383.5k
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
vs →

Related searches

pullmd AlternativesBest Data Processing Tools 2026Open Source Data Processingpullmd Tutorialpullmd Vs CompetitorsURL-to-MarkdownSelf-hostedAI Agent Tooling

Comments

Log in to leave a comment

No comments yet. Be the first!

On this page
01Features02Why choose it03Trade-offs04Compatibility05Quick start06Use cases
GPT Researcher logo
GPT Researcher★ 28.4k
An LLM agent that conducts deep research (local and web) on any given topic and generates a long report with citations.
vs →
Scrapling logo
Scrapling★ 70.2k
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
vs →
firecrawl-mcp-server logo
firecrawl-mcp-server★ 7.0k
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
vs →
mcp-playwright logo
mcp-playwright★ 5.6k
Playwright Model Context Protocol Server - Tool to automate Browsers and APIs in Claude Desktop, Cline, Cursor IDE and More 🔌
vs →
browser-agent-py logo
browser-agent-py★ 1.4k
AI Browser Agent is an advanced Browser AI tool developed by Oxylabs AI Studio that automates real user browsing tasks using natural language instructions.
vs →
See all alternatives →
07
How it compares
08Alternatives
Stats
GitHub Stars★ 197
Last commit1w ago
StatusActive
LicenseAGPL-3.0
CategoryData Processing
Trend (30d)
+7.8↑ 2.4%
Links
Documentation↗Discussion↗Issues↗Releases↗

Deploy on DigitalOcean — Get $200 Free Credit

Ad