Command Palette
Search for a command to run...
Multi Modal Self Instruct Multimodal Benchmark Dataset
Date
Size
Publish URL
Paper URL
License
CC BY-SA 4.0
Tags

This dataset was jointly launched by Zhejiang University, the Institute of Software of the Chinese Academy of Sciences, ShanghaiTech University and other institutions in 2024. The relevant paper results are "Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model". The dataset contains a total of 11,193 abstract images with relevant questions, covering 8 major categories including dashboards, roadmaps, charts, tables, flowcharts, relationship diagrams, visual puzzles and 2D floor plans, in addition to an additional 62,476 data for fine-tuning the model.
Citation
@inproceedings{zhang-etal-2024-multimodal,
title = "Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model",
author = "Zhang, Wenqi and
Cheng, Zhenglin and
He, Yuanyu and
Wang, Mengna and
Shen, Yongliang and
Tan, Zeqi and
Hou, Guiyang and
He, Mingqian and
Ma, Yanna and
Lu, Weiming and
Zhuang, Yueting",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
year = "2024",
address = "Miami, Florida, USA",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.emnlp-main.1072/",
pages = "19228--19252"}
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.