HyperAIHyperAI

Command Palette

Search for a command to run...

Dataset Compilation | From Competition Math to Tool Applications: 9 Open-Source Math Datasets From MIT, NVIDIA, Huazhong University of Science and Technology, etc., Covering CoT, Multimodal Reasoning, and Long-Chain Thinking Training

Featured Image

Mathematical reasoning has become a core indicator for measuring the intelligence level of Large Language Models (LLMs). From arithmetic calculations to Olympiad-level problems, and then to multi-step planning and tool usage, the model is moving from "providing answers" to "understanding problems and completing reasoning."

Compared to traditional question-answering datasets,Mathematical reasoning data places greater emphasis on the thought process and the chain of reasoning.High-quality data not only needs to provide accurate answers, but also needs to provide standardized reasoning trajectories, multi-path solutions, and verifiable results to help the model break down the problem, build logical loops, and improve stability under complex tasks.

In recent years, teams such as OpenAI, NVIDIA, and Alibaba Cloud have accelerated the iteration of Reasoning models, and mathematical datasets have evolved from simple sets of problems into comprehensive resources covering long chain reasoning (CoT), tool-enhanced reasoning, multimodal understanding, reinforcement learning training, and knowledge structure modeling.

This article compiles nine representative mathematical reasoning datasets.It covers areas such as competition-level problems, multilingual and multimodal reasoning, model generation of reasoning trajectories, tool-assisted solution, and knowledge dependency modeling.These trends indicate that mathematical reasoning data is becoming the infrastructure for the evolution of reasoning models—the competition among large models in the future will depend not only on the scale of parameters, but also on the ability to learn reliable, verifiable, and transferable reasoning capabilities through high-quality data.

The following datasets are available on HyperAI, helping researchers and developers quickly explore training, evaluation, and applications.

Click to see more high-quality tutorials:

https://hyper.ai/datasets

1.Math-Graph Mathematical Theorem Dependency Graph Dataset

Run online:https://go.hyper.ai/oSSEs

Math-Graph is a dataset of mathematical theorem dependency graphs released in 2026 by the Mathematical Artificial Intelligence Laboratory at the University of Washington. It constructs a mathematical theorem dependency graph at the statement level. The related paper is titled TheoremGraph: Bridging Formal and Informal Mathematics.

This dataset contains 47,952 pairs of informal and formal mathematical statements, integrating the two types of mathematical knowledge into a coherent dependency graph that covers both informal and formal mathematics. It includes paper metadata, mathematical statements (formal/informal), dependency edges, and LLM-generated natural language descriptions.

2.Nemotron-SFT-Math-v4 Mathematical Inference SFT Dataset

Run online:https://go.hyper.ai/MU9v3

Nemotron-SFT-Math-v4 is a mathematical reasoning dataset released by NVIDIA in May 2026. The related paper is titled Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision. It aims to solve the problems of inconsistent quality, non-standard reasoning trajectories, low accuracy, and limited scenarios in traditional mathematical datasets. It effectively improves the model's structured reasoning, multi-trajectory reasoning, and answer verification capabilities. It is widely used for fine-tuning large-scale mathematical reasoning models, reasoning trajectory analysis, answer verification algorithm development, long-context reasoning system construction, and model reasoning robustness evaluation. 

This dataset contains 545,431 training samples, including 285,516 COT reasoning samples and 259,915 TIR tool reasoning samples. It covers mathematical scenarios in competitions and university research in algebra, geometry, number theory, combinatorics, etc. The data is annotated using a hybrid manual and automated method and includes standardized fields such as unique number, question text, multi-turn dialogue, standard answer, source, and protocol.

3.MathNet Multimodal Mathematical Benchmark Inference Dataset

Run online:https://go.hyper.ai/XA6AN

MathNet is a large-scale, multilingual, multimodal mathematical reasoning dataset released in 2026 by a team from MIT in collaboration with King Abdullah University of Science and Technology and other institutions. The related paper is titled "MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval."It aims to evaluate and improve the capabilities of large models in Olympic-level mathematical reasoning and structured retrieval tasks, and is widely used in mathematical reasoning evaluation, RAG research, and multimodal AI training.

This dataset, version v0, contains 27,817 expert-level math problems and their standard solutions. It covers official math competition problems from 58 countries and regions in 17 languages, including 5,148 illustrated problems with a total of 7,541 geometric and graphical illustrations. The dataset covers algebra, geometry, number theory, combinatorics, calculus, probability and statistics, and other Olympiad math knowledge systems. It supports three benchmark tasks: solving math problems, mathematical semantic retrieval (identifying structurally equivalent and similar problems), and retrieval enhancement problem solving.

4Nemotron-Math-v2 Mathematical Inference Dataset

Run online:https://go.hyper.ai/rYdfW

Nemotron-Math-v2 is a mathematical reasoning dataset released by NVIDIA Corporation in 2025. Related research is titled "Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision." It is primarily used to train LLMs to perform structured mathematical reasoning, to study the differences between tool-enhanced reasoning and pure language reasoning, and to build long-context or multi-path reasoning systems.

This dataset contains approximately 347,000 high-quality mathematical problems and 7 million model-generated inference trajectories. Each problem is solved in six configurations: high/medium/low inference depth and with or without Python TIR, and the answers are validated via a pipeline using an LLM as the arbiter.

5. CHIMERA General Inference Synthetic Dataset

Run online:https://go.hyper.ai/JskxG

CHIMERA is a synthetic reasoning dataset designed specifically for reasoning training. The related paper is titled CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning.

This dataset covers a wide range of STEM subjects and provides Long Chain Thinking (CoT) trajectories, containing 9,225 questions across 8 subjects (mathematics, computer science, chemistry, physics, literature, history, biology, and phonetics). All examples are generated by a large language model (LLM) and are automatically validated without manual annotation.

6. Open-RL Inference Problem Dataset

Run online:https://go.hyper.ai/WYDck

Open-RL is a multi-domain reasoning problem dataset released by Turing in 2026, containing independent, verifiable, and explicit STEM reasoning problems in physics, mathematics, biology, and chemistry. Each problem requires multi-step reasoning, involves symbolic operations and/or numerical computation, and has an objectively verifiable final answer.

This dataset is suitable for fine-tuning reinforcement learning, reward modeling, outcome-supervised training, and verifiable inference benchmarking.Each problem requires multiple steps of reasoning and involves symbolic operations and numerical computation, with a verifiable final answer.

7. GPT-5.4 Stepwise Inference Dataset

Run online:https://go.hyper.ai/CeBEp

The GPT-5.4 step-by-step reasoning dataset is a high-density synthetic reasoning dataset designed for long-chain reasoning (CoT) modeling and complex problem-solving tasks. The dataset is built on a "Master-Architect" workflow, using gemini 3 flash to generate challenging hints, and combining this with the GPT-5.4 Reasoning Core to complete multi-step reasoning and result generation.

This dataset contains approximately 1,500 elite-level samples, covering highly complex fields such as mathematics, programming, and medicine. The task difficulty is uniformly set at the "Grandmaster" and "Beyond-PhD" levels. The data samples include complete multi-step reasoning processes and verified final solutions, making them suitable for scenarios involving long-chain logical derivation and assessment of extreme reasoning abilities.

8.Sutra 10B Pretraining Teaching and Training Dataset

Run online:https://go.hyper.ai/skT3q

Sutra 10B Pretraining is a high-quality teaching dataset for pretraining large language models. Generated by the Sutra framework, it creates structured educational content and optimizes the pretraining of language models. This is the largest dataset in the Sutra series, designed to demonstrate how dense, well-curated datasets can provide optimal pretraining performance for small language models.

This dataset contains 10,193,029 teaching records, totaling over 10 billion tokens, covering nine major areas: interdisciplinary, technology, science, social studies, mathematics, life skills, arts and creativity, language arts, and philosophy and ethics. The data follows a well-established teaching paradigm, with 10 levels of difficulty from basic to advanced, demonstrating good hierarchy and systematic organization.

9. VisCoR-55K Visual Inference Dataset

Run online:https://go.hyper.ai/kIlWk

VisCoR-55K is a high-quality visual reasoning dataset released in 2026 by Huazhong University of Science and Technology in collaboration with Alibaba Cloud. This dataset contains approximately 55,000 visual reasoning samples, each of which generates a corresponding reasoning process using contrasting samples.A high-quality visual reasoning dataset covering five categories: general, reasoning, mathematical, graphing, and OCR.The aim is to promote research on visual language models in reliable and robust visual reasoning.