← 返回资源分享
Deep Speech: Scaling up end-to-end speech recognition
论文
论文
发布时间2014-12-17
发表arXiv:1412.5567
作者:Carl Case,Sanjeev Satheesh,Erich Elsen,Adam Coates,Ryan Prenger,Shubho Sengupta,Andrew Y. Ng,Bryan Catanzaro,Awni Hannun,Jared Casper,Greg Diamos
详细介绍
We present a state-of-the-art speech recognition system developed using
end-to-end deep learning. Our architecture is significantly simpler than
traditional speech systems, which rely on laboriously engineered processing
pipelines; these traditional systems also tend to perform poorly when used in
noisy environments. In contrast, our system does not need hand-designed
components to model background noise, reverberation, or speaker variation, but
instead directly learns a function that is robust to such effects. We do not
need a phoneme dictionary, nor even the concept of a "phoneme." Key to our
approach is a well-optimized RNN training system that uses multiple GPUs, as
well as a set of novel data synthesis techniques that allow us to efficiently
obtain a large amount of varied data for training. Our system, called Deep
Speech, outperforms previously published results on the widely studied
Switchboard Hub5'00, achieving 16.0% error on the full test set. Deep Speech
also handles challenging noisy environments better than widely used,
state-of-the-art commercial speech systems.
代码仓库 (24)
PaddlePaddle/PaddleSpeech官方PaddlePaddle
Digital-Umuganda/Deepspeech-KinyarwandaTensorFlow
myrtleSoftware/deepspeechPyTorch
WalterJohnson0/DeepSpeech-KerasRebuildTensorFlow
robmsmt/KerasDeepSpeechTensorFlow
mozilla/DeepSpeechTensorFlow
pannous/caffe-speech-recognitionCaffe2
anssssss/Vietnamese-Speech-RecognitionTensorFlow
IBM/MAX-Speech-to-Text-ConverterTensorFlow
Loghijiaha/DeepSpeech-IndoTensorFlow
