Command Palette
Search for a command to run...
WIT image-text Dataset
Date
Size
Publish URL
Paper URL
License
Other

WIT, short for Wikipedia-based Image Text, is a large multimodal and multilingual dataset. The dataset consists of a curated collection of 37.6 million entity-rich image-text examples, including 11.5 million unique images in 108 Wikipedia languages. The scale of the dataset allows it to be used as a pre-training dataset for multimodal machine learning models. WIT has four unique advantages:
- WIT is the largest multimodal dataset in terms of the number of image-text examples.
- Over 100 languages are covered (with at least 12,000 examples per language), and cross-lingual text is provided for many images.
- Relative to previous datasets, WIT represents a more diverse set of concepts and real-world entities.
- WIT provides a very challenging real-world test set.
Citation
@inproceedings{10.1145/3404835.3463257, author = {Srinivasan, Krishna and Raman, Karthik and Chen, Jiecao and Bendersky, Michael and Najork, Marc}, title = {WIT: Wikipedia-Based Image Text Dataset for Multimodal Multilingual Machine Learning}, year = {2021}, isbn = {9781450380379}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3404835.3463257}, doi = {10.1145/3404835.3463257}, booktitle = {Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval}, pages = {2443–2449}, numpages = {7}, keywords = {dataset, multimodal, machine learning, wikipedia, multilingual, image-text retrieval, neural networks}, location = {Virtual Event, Canada}, series = {SIGIR '21} }
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.