LibriSpeech Quick Language Model Dataset

Description

Data from transcribed LibriSpeech sentences for quick training and testing of language models.

Caveat emptor

This dataset is for educational and showcasing purposes only. The contents are provided 'as is' without any implied warranty, and disclaiming liability for damages resulting from using it.

Origin

The data was obtained from the original LibriSpeech ASR corpus, available fromOpenSLR.

Contents

The dataset is split in train and test with the following contents. All the data is capitalised and normalised.


A file listing the 10,419 most common words in the training set is available invocabulary/librispeech-top10k-words.txt