Command Palette
Search for a command to run...
PerceptionBench Visual Perception Benchmark Dataset
PerceptionBench is a benchmark dataset for visual perception capabilities of large models, released by Moonshot AI in 2026. It is specifically designed to evaluate the basic visual perception capabilities of multimodal large language models (MLLMs). Related research papers include... PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models . This dataset contains 3,000 validated test questions covering 10 basic perceptual abilities, including visual relationship understanding, counting, attribute recognition, depth and 3D perception, spatial localization, comparative reasoning, fine-grained recognition, context integration, OCR text understanding, and perception-related hallucinations. Of these, 1,800 questions are atomic-level sub-problems derived from failed samples of existing benchmarks, and 1,200 questions are newly created based on supplementary images. Evaluation uses standardized prompt word templates and the highest available inference budget to conduct open-ended question-answering tests on 16 cutting-edge multimodal large language models, including GPT-5.6-Sol, Kimi K3, and Claude-Fable-5.
Data fields:
- index: Unique index number of the sample
- Answer: Reference Answer
- Problem: Problem Description
- hint: Prompt message
- Image: Image information in the question
- error_category: The atomic sensing capability category corresponding to the sample.
- source_bmk: Source identifier of the upstream benchmark
- source_idx: The original index of the sample in the upstream benchmark.
Citation
@article{perceptionbench2026,
title = {PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models},
author = {Moonshot AI},
year = {2026}
}
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.