← 返回资源分享
Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM
论文
论文
发布时间2017-06-08
发表arXiv:1706.02737
作者:Shinji Watanabe,Yu Zhang,Takaaki Hori,William Chan
详细介绍
We present a state-of-the-art end-to-end Automatic Speech Recognition (ASR)
model. We learn to listen and write characters with a joint Connectionist
Temporal Classification (CTC) and attention-based encoder-decoder network. The
encoder is a deep Convolutional Neural Network (CNN) based on the VGG network.
The CTC network sits on top of the encoder and is jointly trained with the
attention-based decoder. During the beam search process, we combine the CTC
predictions, the attention-based decoder predictions and a separately trained
LSTM language model. We achieve a 5-10\% error reduction compared to prior
systems on spontaneous Japanese and Chinese speech, and our end-to-end model
beats out traditional hybrid ASR systems.
代码仓库 (7)
mnm-rnd/elsa-voice-asrPyTorch
park-cheol/ASR-Transformer_with_MeanTeachersPyTorch
park-cheol/ASR-TransformerPyTorch
s3prl/End-to-end-ASR-PytorchPyTorch
Alexander-H-Liu/End-to-end-ASR-PytorchPyTorch
neil-zeng/asrPyTorch
sooftware/OpenSpeechPyTorch
