Command Palette
Search for a command to run...
CC12M image-text Pairs Dataset

CC12M (Conceptual 12M) is an image-text pair dataset specifically designed for vision and language pre-training. The dataset contains 12 million image-text pairs. Compared with CC3M, this dataset performs better in long-tail visual recognition for multiple downstream tasks.
Citation
@inproceedings{changpinyo2021cc12m, title = {{Conceptual 12M}: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts}, author = {Changpinyo, Soravit and Sharma, Piyush and Ding, Nan and Soricut, Radu}, booktitle = {CVPR}, year = {2021}, }
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.