Command Palette
Search for a command to run...
The GLOBE dataset is a high-quality English speech corpus released in 2024. Its associated paper, "[GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech]"(https://hyper.ai/papers/2406.14875), aims to address the poor generalization ability of zero-shot speaker adaptive TTS systems toward accented speakers.
This dataset comprises 535 hours of audio data sampled at 24 kHz, sourced from 23,519 speakers covering 164 global accents, along with detailed speaker metadata. GLOBE_V2 builds upon this foundation by providing audio sampled at 44.1 kHz, while removing approximately 5% of samples that exhibit volume anomalies or misalignments between audio and text, thereby further enhancing data quality.
Citation
@misc{wang2024globe,
title={GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech},
author={Wenbin Wang and Yang Song and Sanjay Jha},
year={2024},
eprint={2406.14875},
archivePrefix={arXiv},
}
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.