HyperAIHyperAI

Command Palette

Search for a command to run...

olmOCR-bench Document Parsing Benchmark Dataset

Date

an hour ago

Organization

Allen Institute for Artificial Intelligence

Paper URL

2502.18443

License

Other

olmOCR-bench is an OCR document parsing benchmark dataset released by AllenAI in 2025. This dataset aims to measure the accuracy of OCR systems in converting PDF documents to Markdown format, as well as their ability to ensure the complete preservation of key text and structural information. Related research papers include... olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models . This dataset contains 1,403 PDF files and 7,010 test case samples. The test categories cover a variety of document formats, including arXiv mathematical papers, old scans, tables, headers and footers, long texts, and multi-column layouts. The evaluation criteria include five core metrics: text presence, text exclusion verification, natural reading order, table accuracy, and mathematical formula formatting precision.

Dataset composition:

  • arXiv_math: Contains 522 PDF source files and 2,927 test samples.
  • old_scans_math: Contains 36 PDF source files and 458 test samples.
  • tables_tests: Contains 188 PDF source files and 1,020 test samples.
  • old_scans: Contains 98 PDF source files and 526 test samples.
  • headers_footers: Contains 266 PDF source files and 753 test samples.
  • multi_column: Contains 231 PDF source files and 884 test samples.
  • long_tiny_text: Contains 62 PDF source files and 442 test samples.

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp