Toolkit for linearizing PDFs for LLM datasets/training

Market Signal
Why It Has Market Pull
AllenAI's production-grade toolkit for linearizing PDFs into clean plain text for LLM datasets, built on a fine-tuned 7B vision-language model with a novel document-anchoring prompt. It's backed by a top-tier research institute, ships a real accuracy benchmark, and is highly relevant to any content pipeline that ingests documents.
- Roughly 17.9k GitHub stars; maintained by the Allen Institute for AI with active releases
- Handles multi-column layouts, figures, and insets in natural reading order at under $200 per million pages
- Ships olmOCR-Bench: 7,000+ test cases across 1,400 documents
- Backed by a published paper; discussed on Hacker News and reviewed by LlamaIndex
- Directly usable for turning PDFs and papers into LLM-ready text in a content pipeline
feedbacks
What People Are Saying
"Open-source tool to extract plain text from PDFs"HN comment
"Achieves between 0.4 unoptimized and 4 pages per second"HN comment
"A great effort towards a general OCR benchmark, reproducible and low cost"GitHub issue
"Edit-distance metrics can let small errors get overshadowed by wrong reading order"GitHub issue
"Handles math from arXiv and historical scans, not just clean digital PDFs"Reddit r/MachineLearning
"Needing a GPU for the 7B model is the main barrier to just trying it"HN comment
"Coming from AllenAI it's clearly built for real dataset work, not a demo"Reddit r/LocalLLaMA



















