← 返回资源分享
Improved training of end-to-end attention models for speech recognition
论文
论文
发布时间2018-05-08
发表arXiv:1805.03294
作者:Albert Zeyer,Hermann Ney,Kazuki Irie,Ralf Schlüter
详细介绍
Sequence-to-sequence attention-based models on subword units allow simple
open-vocabulary end-to-end speech recognition. In this work, we show that such
models can achieve competitive results on the Switchboard 300h and LibriSpeech
1000h tasks. In particular, we report the state-of-the-art word error rates
(WER) of 3.54% on the dev-clean and 3.82% on the test-clean evaluation subsets
of LibriSpeech. We introduce a new pretraining scheme by starting with a high
time reduction factor and lowering it during training, which is crucial both
for convergence and final performance. In some experiments, we also use an
auxiliary CTC loss function to help the convergence. In addition, we train long
short-term memory (LSTM) language models on subword units. By shallow fusion,
we report up to 27% relative improvements in WER over the attention baseline
without a language model.
代码仓库 (16)
rwth-i6/returnn-experiments官方
pengchengguo/espnet官方PyTorch
creatorscan/espnet官方PyTorch
rwth-i6/returnnTensorFlow
victor45664/espnetPyTorch
roholazandie/espnetPyTorch
ElinaBaral/espnet_speech2speechPyTorch
bobchennan/espnetPyTorch
danoneata/espnetPyTorch
dhanya-e/Google_IndicPyTorch
