← 返回资源分享
Attention-Based Models for Speech Recognition
论文
论文
发布时间2015-06-24
发表NeurIPS 2015 12 · arXiv:1506.07503
作者:Yoshua Bengio,Dmitriy Serdyuk,Jan Chorowski,Dzmitry Bahdanau,Kyunghyun Cho
详细介绍
Recurrent sequence generators conditioned on input data through an attention
mechanism have recently shown very good performance on a range of tasks in-
cluding machine translation, handwriting synthesis and image caption gen-
eration. We extend the attention-mechanism with features needed for speech
recognition. We show that while an adaptation of the model used for machine
translation in reaches a competitive 18.7% phoneme error rate (PER) on the
TIMIT phoneme recognition task, it can only be applied to utterances which are
roughly as long as the ones it was trained on. We offer a qualitative
explanation of this failure and propose a novel and generic method of adding
location-awareness to the attention mechanism to alleviate this issue. The new
method yields a model that is robust to long inputs and achieves 18% PER in
single utterances and 20% in 10-times longer (repeated) utterances. Finally, we
propose a change to the at- tention mechanism that prevents it from
concentrating too much on single frames, which further reduces PER to 17.6%
level.
代码仓库 (14)
CKRC24/Listen-and-TranslateTensorFlow
jackjhliu/Pytorch-End-to-End-ASR-on-TIMITPyTorch
mnm-rnd/elsa-voice-asrPyTorch
jackjhliu/End-to-End-Mandarin-ASRPyTorch
sooftware/nlp-attentionsPyTorch
s3prl/End-to-end-ASR-PytorchPyTorch
30stomercury/Automatic_Speech_RecognitionTensorFlow
Alexander-H-Liu/End-to-end-ASR-PytorchPyTorch
neil-zeng/asrPyTorch
sooftware/OpenSpeechPyTorch
