Auto-generated — not yet edited · only the numbers are verified
- TESIGN / RADAR
- REPOSITORY CARD
- paddlepaddle/paddleocr
paddlepaddle/paddleocr
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
RANKS All-time #49
AT A GLANCE
- LANGUAGE
- Python
- LICENSE
- Apache-2.0
- USAGE
- Commercial use, redistribution OK. Keep notices; mark changes.
- ACTIVITY
- last commit 56 days ago ()
- TOPICS
- ai4science
- chineseocr
- document-parsing
- document-translation
- kie
- ocr
- paddleocr-vl
- pdf-extractor-rag
- pdf-parser
- pdf2markdown
- pp-ocr
- pp-structure
- HOMEPAGE
- https://www.paddleocr.com
- REPOSITORY
- GitHub ↗
- CATEGORY
- AI
A summary, not legal advice.
README EXCERPT
Global Leading OCR Toolkit & Document AI Engine English 简体中文 繁體中文 日本語 한국어 Français Русский Español العربية PaddleOCR converts PDF documents and images into structured, LLM-ready data (JSON/Markdown) with industry-leading accuracy. With 70k+ Stars and trusted by top-tier projects like Dify, RAGFlow, and Cherry Studio, PaddleOCR is the bedrock for building intelligent RAG and Agentic applications. 🚀 Key Features 📄 Intelligent Document Parsing (LLM-Ready) Transforming messy visuals into structured data for the LLM era. SOTA Document VLM : Featuring PaddleOCR-VL-1.6 (0.9B) , the industry's leading lightweight vision-language model for document parsing. It achieves 96.3% accuracy on OmniDocBench v1.6, leads in text, formula, and table recognition, and shows significantly enhanced capabilities in ancient documents, rare characters, seals, and charts, with structured outputs in Markdown and JSON formats. Structure-Aware Conversion : Powered by PP-StructureV3 , seamlessly convert complex PDFs and images into Markdown or JSON . Unlike the PaddleOCR-VL series models, it provides more fine-grained coordinate information, including table cell coordinates, text coordinates, and more. Production…
The opening of the GitHub README as stored, at most 1,200 characters. Markdown is not rendered.
TIMELINE
- Repository created
- Last push
- FIRST SEEN BY TESIGN ◌ BACK CATALOG
AI chooses the lists under the owner’s delegation. No human review is running in September 2026. Total stars, increases, cross-source signals, last updates and licences are shown as evidence. How ranks work →
If this repository gets editorial text (why, build, who, start, caveat) it becomes an edited entry. Until then the page shows only stored GitHub metadata and numbers.