Command Palette
Search for a command to run...
MMLU-CF Contamination-Free Multi-Task Language Understanding Benchmark Dataset
Date
Paper URL
License
Other
MMLU-CF is a pollution-free and more challenging multiple-choice benchmark dataset released by Microsoft in 2024, with the related research paper titled "MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark", aimed at addressing the issue of unreliable evaluation caused by data leakage from large language model training data in open-source benchmarks.
The dataset contains 20,000 multiple-choice questions, covering multiple subject areas such as biology, mathematics, chemistry, physics, law, and engineering. Through strict decontamination rules and a closed-source test set design, MMLU-CF effectively avoids model memorization effects, providing a cleaner benchmark for evaluating the true reasoning capabilities of large language models.
Dataset Composition
-
Validation set (val): Contains 10,000 multiple-choice questions, divided into subsets by subject, including Biology, Math, Chemistry, Physics, Law, Engineering, Other, Economics, Health, Psychology, Business, Philosophy, Computer Science, History, and more.
-
Development set (dev): Contains corresponding development data, divided into multiple subsets according to the same subject categories, used for model evaluation and experiments.
Citation
@misc{zhao2024mmlucfcontaminationfreemultitasklanguage,
title={MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark},
author={Qihao Zhao and Yangyu Huang and Tengchao Lv and Lei Cui and Qinzheng Sun and Shaoguang Mao and Xin Zhang and Ying Xin and Qiufeng Yin and Scarlett Li and Furu Wei},
year={2024},
eprint={2412.15194},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.15194},
}
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.