HyperAIHyperAI

Command Palette

Search for a command to run...

ClimbMix Clustered Hybrid Pre-training Dataset

Date

Organization

NVIDIA

Paper URL

2504.13161

License

CC BY NC 4.0

The ClimbMix dataset is a high-quality dataset released by NVIDIA in 2025 specifically designed for efficient pre-training of large language models. Its associated paper can be found at 「CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training, aiming to deliver superior pre-training performance compared to existing public datasets under equivalent token budgets through innovative clustering and iterative mixing algorithms.

This dataset comprises 400 billion tokens, formatted as plain text stored in Parquet files, and utilizes the GPT-2 tokenizer for preprocessing and tokenization. During construction, raw data was initially categorized into 1,000 groups based on thematic information. Subsequently, two classifiers—advertising detection and educational value assessment—were applied to score each group, filtering out low-quality samples. Finally, the remaining high-quality data groups were mixed according to specific weights. This dataset is intended solely for research purposes, helping researchers explore more efficient model scaling trends.

Dataset Composition

The dataset primarily includes the following key details:

  • Data Scale: 400 Billion Tokens
  • Format: Text sequences stored in Parquet format
  • Tokenization Method: Utilizes the GPT-2 tokenizer
  • Source Origin: Automated collection and processing pipeline
  • Labeling Approach: Automatic labeling and filtering
数据集示例
数据集示例

Citation

@article{diao2025climb,
  author    = {Shizhe Diao and Yu Yang and Yonggan Fu and Xin Dong and Dan Su and Markus Kliegl and Zijia Chen and Peter Belcak and Yoshi Suhara and Hongxu Yin and Mostofa Patwary and Celine Lin and Jan Kautz and Pavlo Molchanov},
  title={CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training}, 
  journal   = {arXiv preprint},
  year      = {2025},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  url={https://arxiv.org/abs/2504.13161}, 
}

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp