Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
- 言語
- Python
- ライセンス
- Apache-2.0
- スター
- 91k
- 最新リリース
- v3.7.0
- 最終プッシュ
- 2026-09-16
Recognise text in scanned documents and images.
5件のツール
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Tesseract Open Source OCR Engine (main repository)
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
OCR, layout analysis, reading order, table recognition in 90+ languages
docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.