# Tesseract OCR

> Open-source OCR engine

A long-standing OCR engine that recognizes text in more than 100 languages.

- Source: https://github.com/tesseract-ocr/tesseract
- Homepage: https://tesseract-ocr.github.io/
- License: Apache-2.0
- Language: C++
- Stars: 76819
- Forks: 10818
- Contributors: 201
- Last commit: 2026-09-28
- Latest release: 5.5.3 (2026-07-24)
- Purpose: AI & LLM tooling
- Runs on: CLI, Library / SDK
- For: Personal, Enterprise

## Worth score: 84/100

How much DigGitHub recommends it, from activity, adoption, docs, license and security signals.

- Popularity: 24/25
- Momentum: 10/20
- Maintenance: 25/25
- Community: 14/15
- Readiness: 11/15

## Alternatives

- [OCRmyPDF](https://diggithub.com/ocrmypdf/OCRmyPDF.md): Make scanned PDFs searchable
- [MinerU](https://diggithub.com/opendatalab/MinerU.md): Convert PDFs and Office docs to LLM-ready Markdown

---

Source page: https://diggithub.com/tesseract-ocr/tesseract
Updated: 2026-10-04
