# OpenDataLoader PDF

> PDF parser for AI-ready data

Extracts text, tables and layout with bounding boxes and helps with PDF accessibility.

- Source: https://github.com/opendataloader-project/opendataloader-pdf
- Homepage: https://opendataloader.org
- License: Apache-2.0
- Language: Java
- Stars: 29457
- Forks: 2806
- Contributors: 32
- Last commit: 2026-10-01
- Latest release: v2.5.12 (2026-10-01)
- Purpose: AI & LLM tooling, RAG & vector DBs
- Runs on: Library / SDK
- For: Personal, Enterprise

## Worth score: 79/100

How much DigGitHub recommends it, from activity, adoption, docs, license and security signals.

- Popularity: 22/25
- Momentum: 10/20
- Maintenance: 25/25
- Community: 11/15
- Readiness: 11/15

## Alternatives

- [Docling](https://diggithub.com/docling-project/docling.md): Get your documents ready for generative AI
- [MinerU](https://diggithub.com/opendatalab/MinerU.md): Convert PDFs and Office docs to LLM-ready Markdown
- [PaddleOCR](https://diggithub.com/PaddlePaddle/PaddleOCR.md): OCR toolkit that turns documents into structured data
- [Unstructured](https://diggithub.com/Unstructured-IO/unstructured.md): Convert documents to structured data
- [MarkItDown](https://diggithub.com/microsoft/markitdown.md): Convert Office files and PDFs to Markdown for LLMs
- [Crawl4AI](https://diggithub.com/unclecode/crawl4ai.md): Web crawler that turns sites into LLM-ready Markdown
- [Firecrawl](https://diggithub.com/firecrawl/firecrawl.md): Web data API that turns sites into LLM-ready data
- [LightRAG](https://diggithub.com/HKUDS/LightRAG.md): Simple and fast graph-based RAG

---

Source page: https://diggithub.com/opendataloader-project/opendataloader-pdf
Updated: 2026-10-04
