Context

A free audio dataset of spoken digits. Think MNIST for audio. (3,000 recordings, 6 speakers )A simple audio/speech dataset consisting of recordings of spoken digits in wav files at 8kHz. The recordings are trimmed so that they have near minimal silence at the beginnings and ends.

FSDD is an open dataset, which means it will grow over time as data is contributed. In order to enable reproducibility and accurate citation the dataset is versioned using Zenodo DOI as well as git tags.


Current status6 speakers3,000 recordings (50 of each digit per speaker)English pronunciations


Created by:Zohar Jackson, César Souza, Jason Flaks, Yuxin Pan, Hereman Nicolas, & Adhish Thite.


Link:https://github.com/Jakobovski/free-spoken-digit-dataset


Content

What's inside is more than just rows and columns. Make it easy for others to get started by describing how you acquired the data and what time period it represents, too.


Acknowledgements

Zohar Jackson, César Souza, Jason Flaks, Yuxin Pan, Hereman Nicolas, & Adhish Thite. (2018, August 9). Jakobovski/free-spoken-digit-dataset: v1.0.8 (Version v1.0.8). Zenodo.http://doi.org/10.5281/zenodo.1342401


Inspiration

A free audio dataset of spoken digits. Think MNIST for audio.