← 返回资源分享
Listen, Attend and Spell
论文
论文
发布时间2015-08-05
发表arXiv:1508.01211
作者:Oriol Vinyals,Navdeep Jaitly,Quoc V. Le,William Chan
详细介绍
We present Listen, Attend and Spell (LAS), a neural network that learns to
transcribe speech utterances to characters. Unlike traditional DNN-HMM models,
this model learns all the components of a speech recognizer jointly. Our system
has two components: a listener and a speller. The listener is a pyramidal
recurrent network encoder that accepts filter bank spectra as inputs. The
speller is an attention-based recurrent network decoder that emits characters
as outputs. The network produces character sequences without making any
independence assumptions between the characters. This is the key improvement of
LAS over previous end-to-end CTC models. On a subset of the Google voice search
task, LAS achieves a word error rate (WER) of 14.1% without a dictionary or a
language model, and 10.3% with language model rescoring over the top 32 beams.
By comparison, the state-of-the-art CLDNN-HMM model achieves a WER of 8.0%.
代码仓库 (41)
custodio78/Speech_RecognitionTensorFlow
JJoving/SMLATPyTorch
nithinksath96/Speech_Recognition
neelrast/seq2seq-speech-recognitionTensorFlow
jackjhliu/Pytorch-End-to-End-ASR-on-TIMITPyTorch
foamliu/Listen-Attend-and-SpellPyTorch
foamliu/Listen-Attend-Spell-v2PyTorch
gargimahale/Speech-Recognition-Using-TensorflowTensorFlow
switiz/las.pytorchPyTorch
Garfield35/Speach-Recognition-Using-TensorflowTensorFlow
