by NameetP · MCP Server · ★ 76
pdfmux Self-healing PDF extraction with per-page confidence scoring. Open-source LlamaParse alternative for RAG pipelines, MCP server for Claude Desktop, LangChain + LlamaIndex loaders. Ranked #2 on opendataloader-bench (0.900). The only PDF extractor that audits its own output. Catches blank pages, scrambled columns, broken tables — re-extracts them with a stronger backend. So your LLM gets clean data, not silent garbage. Routes each page to the best of 5 rule-based backends + BYOK LLM fallback (Gemini / Claude / GPT-4o / Ollama). One CLI. One API. Zero config.
| Stars | 76 |
| Forks | 12 |
| Language | Python |
| Category | MCP Server |
| License | MIT |
| Quality Score | 68.870137129299/100 |
| Open Issues | 6 |
| Last Updated | 2026-07-16 |
| Created | 2026-03-03 |
| Platforms | cli, mcp, python |
| Est. Tokens | ~19k |
Explore other popular mcp server tools:
pdfmux is Self-healing PDF extraction that flags what it can't read instead of dropping it — and now certifies any extractor's output, catching silently-dropped pages. #2 of all tools, #1 free on opendataloader. It is categorized as a MCP Server with 76 GitHub stars.
pdfmux is primarily written in Python. It covers topics such as ai-agent, docling, document-parsing.
You can find installation instructions and usage details in the pdfmux GitHub repository at github.com/NameetP/pdfmux. The project has 76 stars and 12 forks, indicating an active community.
pdfmux is released under the MIT license, making it free to use and modify according to the license terms.