Command Palette
Search for a command to run...
AF-Chat Audio Conversation Text Dataset
AF-Chat is an audio conversation text dataset released by NVIDIA in 2025. The related paper results are "Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models", which aims to train and evaluate dialogue generation models. The dataset contains about 75,000 multi-turn, multi-audio dialogues (average 4.6 segments and 6.2 rounds; range 2-8 segments and 2-10 rounds), covering speech, environmental sounds, and music. The dataset is divided into different subsets (sound, music 4ALL, million song datasets) according to the source dataset of each audio, and only text question-answer annotations are provided, not the audio files themselves.
Citation
@misc{ghoshaudioflamingonext,
title={Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music},
author={Sreyan Ghosh and Arushi Goel and Kaousheik Jayakumar and Lasha Koroshinadze and Nishit Anand and Zhifeng Kong and Siddharth Gururani and Sang-gil Lee and Jaehyeon Kim and Aya Aljafari and Chao-Han Huck Yang and Sungwon Kim and Ramani Duraiswami and Dinesh Manocha and Mohammad Shoeybi, Bryan Catanzaro and Ming-Yu Liu and Wei Ping},
year={2026},
eprint={},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={},
}
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.