← 返回资源分享
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
论文
论文
发布时间2015-12-08
发表arXiv:1512.02595
作者:Mike Chrzanowski,Carl Case,Chong Wang,Eric Battenberg,Sanjeev Satheesh,Erich Elsen,Sharan Narang,Jonathan Raiman,Yi Wang,Adam Coates,Zhenyao Zhu,David Seetapun,Ryan Prenger,Andrew Ng,Shubho Sengupta,Bryan Catanzaro,Linxi Fan,Awni Hannun,Dario Amodei,Rishita Anubhai,Jared Casper,Jingdong Chen,Greg Diamos,Jesse Engel,Christopher Fougner,Tony Han,Billy Jun,Patrick LeGresley,Libby Lin,Sherjil Ozair,Zhiqian Wang,Bo Xiao,Dani Yogatama,Jun Zhan
详细介绍
We show that an end-to-end deep learning approach can be used to recognize
either English or Mandarin Chinese speech--two vastly different languages.
Because it replaces entire pipelines of hand-engineered components with neural
networks, end-to-end learning allows us to handle a diverse variety of speech
including noisy environments, accents and different languages. Key to our
approach is our application of HPC techniques, resulting in a 7x speedup over
our previous system. Because of this efficiency, experiments that previously
took weeks now run in days. This enables us to iterate more quickly to identify
superior architectures and algorithms. As a result, in several cases, our
system is competitive with the transcription of human workers when benchmarked
on standard datasets. Finally, using a technique called Batch Dispatch with
GPUs in the data center, we show that our system can be inexpensively deployed
in an online setting, delivering low latency when serving users at scale.
代码仓库 (37)
PaddlePaddle/PaddleSpeech官方PaddlePaddle
raotnameh/End-to-end-E2E-Named-Entity-Recognition-from-English-Speech官方PyTorch
mangelroman/audio2scorePyTorch
GavinGuan95/Punctuator.PytorchPyTorch
sburud/masterPyTorch
myrtleSoftware/deepspeechPyTorch
robmsmt/KerasDeepSpeechTensorFlow
TensorSpeech/TensorFlowASRTensorFlow
UnofficialJuliaMirror/DeepMark-deepmarkPyTorch
tensorflow/models/tree/master/research/deep_speechTensorFlow
