Subscribe
Hugging FaceSep 27, 20261 source

ESPnet releases YODAS v3, a 1.1-million-hour open speech dataset

The ESPnet team published YODAS v3, an open dataset of about 1.1 million hours of real-world speech in more than 100 languages, with 48kHz audio and timestamped transcripts.

An expanse of deep blue waves under a pink and purple sky

The ESPnet team released YODAS v3 on Hugging Face on September 27. It contains roughly 1.1 million hours of real-world speech across more than 100 languages, 34 of which have over 1,000 hours.

Each example includes 48kHz audio, a transcript with word- and utterance-level timestamps, the detected language and, where available, a timestamped English translation. The team says over 70% of the data has at least two distinct channels and estimates each file's effective sampling rate so users can filter by quality.

The authors say the dataset is large enough to train a Whisper-scale speech recognition model from scratch, and suits speech synthesis and spatial audio research.

Sources (1)

  1. Hugging Face blog — YODAS v3
← Back to AI World News