Command Palette
Search for a command to run...
AceMath RewardBench
AceMath RewardBench is a dataset released by NVIDIA in 2024 for evaluating mathematical reward models, associated with the paper 「AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling」. It aims to comprehensively measure the performance of math reasoning reward models using best-of-N (N=8) settings.
The dataset comprises samples from seven datasets: GSM8K (1,319 questions), Math500 (500 questions), Minerva Math (272 questions), Gaokao 2023 en (385 questions), OlympiadBench (675 questions), College Math (2,818 questions), and MMLU STEM (3,018 questions).
Each sample includes one mathematics problem, up to 64 answer attempts generated by eight different large language models including Qwen2/2.5-Math, Llama3.1, Mathstral, and DeepSeek-Math at varying quality levels, ground-truth scores for each response, and metadata such as difficulty level and topic domain.
This benchmark emphasizes diversity and robustness, averaging results over random sampling via 100 seeds.
Dataset Composition
The dataset consists of the following fields:
question: Text description of the math problemcode: List of full answers/solutions generated by modelsgt: Ground truth answerpred: List of predicted outputs extracted from model responsesscore: Boolean list indicating whether each response matches the ground truth answerreport: Related reportsidx: Indexgt_cot: Ground truth chain of thought
Citation
@article{acemath2024,
title={AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling},
author={Liu, Zihan and Chen, Yang and Shoeybi, Mohammad and Catanzaro, Bryan and Ping, Wei},
journal={arXiv preprint},
year={2024}
}
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.