Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Market Signal
Why It Has Market Pull
An institutionally-backed document-to-markdown/JSON engine (PDFs, Office docs) for LLM and agentic workflows, from OpenDataLab (Shanghai AI Lab). It is the most established and credible tool in this group, with deep engineering history and heavy production use.
- ~69.3k GitHub stars, supported by years of history
- Backed by OpenDataLab; reported production use by Google, Huawei, Alibaba and 100+ enterprises
- Mature release cadence: 175 releases, latest v3.4.0 (~June 18, 2026), with thousands of issues and discussions
- Benchmarked in independent 2026 PDF-to-markdown roundups against Marker, Docling and others
- License recently migrated from AGPLv3 to a custom Apache-2.0-based open-source license
feedbacks
What People Are Saying
"parsing quality approaches commercial tools; best for CJK documents and complex academic papers"Comparison roundup
"MinerU takes 3+ minutes before it actually starts processing a 120-page OCR PDF"GitHub issue
"video memory is still occupied and not released after the task completes"GitHub issue
"CUDA out of memory; tried to allocate 110 MiB"GitHub issue
"vLLM speed far below the official docs"GitHub issue
"Arabic, Hindi and Urdu were not reliable for text extraction in testing"Tutorial review






















