llmtrim: llmtrim is a local proxy that intelligently compresses LLM API requests by removing wasted tokens, significantly reducing billing costs without impacting the quality of the LLM's responses. It offers versatile deployment options, including a proxy, CLI, MCP server, or library for various programming languages.; headroom: Headroom compresses everything your AI agent reads — tool outputs, logs, RAG chunks, files, and conversation history — before it reaches the LLM. It achieves the same answers with a fraction of the tokens.
Reducing LLM API costs for AI agents (e.g., Claude Code, Cursor, Aider).
Reduce LLM token usage and API costs for AI agents.