Command Palette
Search for a command to run...
Image Understanding Benchmark Dataset
Image Understanding Benchmark is a synthetic dataset for image understanding evaluation released by Microsoft in 2024, designed to assess multimodal models' capabilities in image understanding, spatial reasoning, visual prompting, object recognition, and detection.
The dataset contains four subtasks: object detection, object recognition, spatial reasoning, and visual prompting, generating approximately 10,240 images in total. The data is procedurally generated, combining COCO objects and Places365 backgrounds, with random rotations, positions, and scaling variations, and can be used to evaluate the visual understanding capabilities of multimodal models.
Dataset Composition
The dataset includes 8 configurations, divided into single-object and pairs conditions, covering the following four subtasks:
- Object Detection
- Object Recognition
- Spatial Reasoning
- Visual Prompting
Each subtask generates 1,280 images under both conditions, totaling approximately 10,240 images. Each configuration mainly contains the following fields:
- id: Unique identifier for each data sample.
- image: Input image.
- prompt: Prompt text used for model evaluation.
- ground_truth: Standard answer for certain tasks.
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.