← 返回资源分享
wav2letter++: The Fastest Open-source Speech Recognition System
论文
论文
发布时间2018-12-18
发表arXiv:1812.07625
作者:Vitaliy Liptchinsky,Gabriel Synnaeve,Ronan Collobert,Awni Hannun,Qiantong Xu,Vineel Pratap,Jeff Cai,Jacob Kahn
详细介绍
This paper introduces wav2letter++, the fastest open-source deep learning
speech recognition framework. wav2letter++ is written entirely in C++, and uses
the ArrayFire tensor library for maximum efficiency. Here we explain the
architecture and design of the wav2letter++ system and compare it to other
major open-source speech recognition systems. In some cases wav2letter++ is
more than 2x faster than other optimized frameworks for training end-to-end
neural networks for speech recognition. We also show that wav2letter++'s
training times scale linearly to 64 GPUs, the highest we tested, for models
with 100 million parameters. High-performance frameworks enable fast iteration,
which is often a crucial factor in successful research and model tuning on new
datasets and tasks.
代码仓库 (8)
facebookresearch/wav2letter
mailong25/wav2letter
jakeju/wav2letter
krantirk/Wav2letterPlus
lilei-John/wav2letter
flashlight/wav2letter
gcambara/wav2letter
bsridatta/wav2letter-Swedish
