TESIGN
SAVED

EN · KO

Auto-generated — not yet edited · only the numbers are verified

  1. TESIGN / RADAR
  2. REPOSITORY CARD
  3. paddlepaddle/paddleocr

paddlepaddle/paddleocr

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

★ 89.4k▲ +37 / 7dDaily star increases over 14 days. Hollow bars mean no data. The last day may be partial.2026-09-03 no data2026-09-04 +62026-09-05 +132026-09-06 +92026-09-07 +22026-09-08 no data2026-09-09 +22026-09-10 +22026-09-11 +102026-09-12 +82026-09-13 +132026-09-14 no data2026-09-15 no data2026-09-16 +4largest one-day increase in the 14 days+13last 14 days · stars per dayhollow = no data

RANKS All-time #49

AT A GLANCE

LANGUAGE
Python
LICENSE
Apache-2.0
USAGE
Commercial use, redistribution OK. Keep notices; mark changes.
ACTIVITY
last commit 56 days ago ()
TOPICS
  • ai4science
  • chineseocr
  • document-parsing
  • document-translation
  • kie
  • ocr
  • paddleocr-vl
  • pdf-extractor-rag
  • pdf-parser
  • pdf2markdown
  • pp-ocr
  • pp-structure
HOMEPAGE
https://www.paddleocr.com
REPOSITORY
GitHub ↗
CATEGORY
AI

A summary, not legal advice.

README EXCERPT

Global Leading OCR Toolkit & Document AI Engine English 简体中文 繁體中文 日本語 한국어 Français Русский Español العربية PaddleOCR converts PDF documents and images into structured, LLM-ready data (JSON/Markdown) with industry-leading accuracy. With 70k+ Stars and trusted by top-tier projects like Dify, RAGFlow, and Cherry Studio, PaddleOCR is the bedrock for building intelligent RAG and Agentic applications. 🚀 Key Features 📄 Intelligent Document Parsing (LLM-Ready) Transforming messy visuals into structured data for the LLM era. SOTA Document VLM : Featuring PaddleOCR-VL-1.6 (0.9B) , the industry's leading lightweight vision-language model for document parsing. It achieves 96.3% accuracy on OmniDocBench v1.6, leads in text, formula, and table recognition, and shows significantly enhanced capabilities in ancient documents, rare characters, seals, and charts, with structured outputs in Markdown and JSON formats. Structure-Aware Conversion : Powered by PP-StructureV3 , seamlessly convert complex PDFs and images into Markdown or JSON . Unlike the PaddleOCR-VL series models, it provides more fine-grained coordinate information, including table cell coordinates, text coordinates, and more. Production…

The opening of the GitHub README as stored, at most 1,200 characters. Markdown is not rendered.

TIMELINE

  1. Repository created
  2. Last push
  3. FIRST SEEN BY TESIGN ◌ BACK CATALOG

AI chooses the lists under the owner’s delegation. No human review is running in September 2026. Total stars, increases, cross-source signals, last updates and licences are shown as evidence. How ranks work →

If this repository gets editorial text (why, build, who, start, caveat) it becomes an edited entry. Until then the page shows only stored GitHub metadata and numbers.