# MinerU

> Convert PDFs and Office docs to LLM-ready Markdown

Parses complex documents with layout, tables, formulas and images into Markdown or JSON for AI pipelines.

- Source: https://github.com/opendatalab/MinerU
- Homepage: https://opendatalab.github.io/MinerU/
- License: NOASSERTION
- Language: Python
- Stars: 81056
- Forks: 6757
- Contributors: 86
- Last commit: 2026-09-29
- Latest release: mineru-4.0.10-released (2026-09-29)
- Purpose: AI & LLM tooling, RAG & vector DBs
- Runs on: CLI, Library / SDK
- For: Personal, Enterprise

## Worth score: 78/100

How much DigGitHub recommends it, from activity, adoption, docs, license and security signals.

- Popularity: 25/25
- Momentum: 10/20
- Maintenance: 25/25
- Community: 12/15
- Readiness: 6/15

## Alternatives

- [Docling](https://diggithub.com/docling-project/docling.md): Get your documents ready for generative AI
- [PaddleOCR](https://diggithub.com/PaddlePaddle/PaddleOCR.md): OCR toolkit that turns documents into structured data
- [OpenDataLoader PDF](https://diggithub.com/opendataloader-project/opendataloader-pdf.md): PDF parser for AI-ready data
- [Unstructured](https://diggithub.com/Unstructured-IO/unstructured.md): Convert documents to structured data
- [MarkItDown](https://diggithub.com/microsoft/markitdown.md): Convert Office files and PDFs to Markdown for LLMs
- [Marker](https://diggithub.com/datalab-to/marker.md): Convert PDFs to Markdown and JSON with high accuracy
- [anydoc](https://diggithub.com/firecrawl/anydoc.md): Convert office files and PDFs to Markdown
- [Crawl4AI](https://diggithub.com/unclecode/crawl4ai.md): Web crawler that turns sites into LLM-ready Markdown

---

Source page: https://diggithub.com/opendatalab/MinerU
Updated: 2026-10-04
