HyperAIHyperAI

Command Palette

Search for a command to run...

AudioSetCaps Audio Subtitle Dataset

Date

2 years ago

Size

120.7 MB

Organization

University of Surrey
Nanyang Technological University (南洋理工大学)

Publish URL

github.com

License

CC BY 4.0

The dataset was released in 2024 by researchers from Northwestern Polytechnical University, Xi'an Lianfeng Acoustic Technology Co., Ltd., Nanyang Technological University, University of Surrey, and the Institute of Acoustics, Chinese Academy of Sciences.AudioSetCaps: Enriched Audio Captioning Dataset Generation Using Large Audio Language Models", has been accepted by NeurIPS 24. AudioSetCaps is an audio-caption dataset containing 6,117,099 10-second audio files. Each audio file is accompanied by a descriptive title and 3 Q&A pairs as metadata for generating the final caption (a total of 18,414,789 pairs of Q&A data). It is created using an automated generation pipeline of large audio and language models using data from three audio datasets: AudioSet, YouTube-8M, and VGGSound.

Citation

@article{bai2024audiosetcaps, title={AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models}, author={Bai, Jisheng and Liu, Haohe and Wang, Mou and Shi, Dongyuan and Wang, Wenwu and Plumbley, Mark D and Gan, Woon-Seng and Chen, Jianfeng} journal={arXiv preprint arXiv:2411.18953}, year={2024} }

AudioSetCaps.torrent
Seeding 1Downloading 0Completed 161Total Downloads 293
  • AudioSetCaps/
    • README.md
      1.63 KB
    • README.txt
      3.27 KB
      • data/
        • AudioSetCaps.zip
          120.7 MB

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp