← 返回资源分享
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
论文
论文
发布时间2019-04-18
发表arXiv:1904.08779
作者:Yu Zhang,Chung-Cheng Chiu,Barret Zoph,Quoc V. Le,William Chan,Daniel S. Park,Ekin D. Cubuk
详细介绍
We present SpecAugment, a simple data augmentation method for speech recognition. SpecAugment is applied directly to the feature inputs of a neural network (i.e., filter bank coefficients). The augmentation policy consists of warping the features, masking blocks of frequency channels, and masking blocks of time steps. We apply SpecAugment on Listen, Attend and Spell networks for end-to-end speech recognition tasks. We achieve state-of-the-art performance on the LibriSpeech 960h and Swichboard 300h tasks, outperforming all prior work. On LibriSpeech, we achieve 6.8% WER on test-other without the use of a language model, and 5.8% WER with shallow fusion with a language model. This compares to the previous state-of-the-art hybrid system of 7.5% WER. For Switchboard, we achieve 7.2%/14.6% on the Switchboard/CallHome portion of the Hub5'00 test set without the use of a language model, and 6.8%/14.1% with shallow fusion, which compares to the previous state-of-the-art hybrid system at 8.3%/17.3% WER.
代码仓库 (29)
PaddlePaddle/PaddleSpeech官方PaddlePaddle
AmirmohammadRostami/KeywordsSpotting-EfficientNet-A0官方PyTorch
DemisEom/SpecAugmentPyTorch
shelling203/SpecAugmentPyTorch
sh951011/Korean-Speech-Recognition/blob/master/package/feature.pyPyTorch
ZhengkunTian/OpenTransformerPyTorch
iver56/audiomentationsPyTorch
audio-westlakeu/rctPyTorch
KimJeongSun/SpecAugment_numpy_scipyTensorFlow
google-research/leaf-audioTensorFlow
