Command Palette
Search for a command to run...
olmOCR-bench Document Parsing Benchmark Dataset
olmOCR-bench is an OCR document parsing benchmark dataset released by AllenAI in 2025. This dataset aims to measure the accuracy of OCR systems in converting PDF documents to Markdown format, as well as their ability to ensure the complete preservation of key text and structural information. Related research papers include... olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models . This dataset contains 1,403 PDF files and 7,010 test case samples. The test categories cover a variety of document formats, including arXiv mathematical papers, old scans, tables, headers and footers, long texts, and multi-column layouts. The evaluation criteria include five core metrics: text presence, text exclusion verification, natural reading order, table accuracy, and mathematical formula formatting precision.
Dataset composition:
- arXiv_math: Contains 522 PDF source files and 2,927 test samples.
- old_scans_math: Contains 36 PDF source files and 458 test samples.
- tables_tests: Contains 188 PDF source files and 1,020 test samples.
- old_scans: Contains 98 PDF source files and 526 test samples.
- headers_footers: Contains 266 PDF source files and 753 test samples.
- multi_column: Contains 231 PDF source files and 884 test samples.
- long_tiny_text: Contains 62 PDF source files and 442 test samples.
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.