← 返回资源分享
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
论文
论文
发布时间2024-07-05
发表arXiv:2407.04675
作者:Linhao Dong,Wei Chen,Jitong Chen,Rui Xia,Yang Zhang,Zhuo Chen,Yuxuan Wang,Lu Huang,Jun Zhang,Hao Wang,Chuang Ding,Ming Tu,Ye Bai,Yujiao Du,Bo wang,Qianqian Dong,Yi Guo,Rui Liu,Tian Tan,Wenchao Hu,Ting Han,Chen Shen,Yizhou Lu,Mingkun Huang,Minglun Han,Meng Yang,Hongmin Xu,Lu Lu,Yuping Wang,Yuxiang Hu,Xiaoyang Li,Bihong Zhang,Tianyu Li,Lu Gao,Shouda Liu,Jingping Chen,Kepan Gao,Xinying Hu,Deyu Hua,Youjia Huang,Jishuo Jin,Fanliu Kong,Zongwei Lan,Zeyang Li,Zehua Lin,Jingting Ma,Shengtao Ma,Yulin Pei,Xiaogang Tian,Hanzhang Xia,Shuangyi Xie,Wanyi Zhang,Yawei Zhang,Yijie Zheng,Ming Zou
详细介绍
Modern automatic speech recognition (ASR) model is required to accurately transcribe diverse speech signals (from different domains, languages, accents, etc) given the specific contextual information in various application scenarios. Classic end-to-end models fused with extra language models perform well, but mainly in data matching scenarios and are gradually approaching a bottleneck. In this work, we introduce Seed-ASR, a large language model (LLM) based speech recognition model. Seed-ASR is developed based on the framework of audio conditioned LLM (AcLLM), leveraging the capabilities of LLMs by inputting continuous speech representations together with contextual information into the LLM. Through stage-wise large-scale training and the elicitation of context-aware capabilities in LLM, Seed-ASR demonstrates significant improvement over end-to-end models on comprehensive evaluation sets, including multiple domains, accents/dialects and languages. Additionally, Seed-ASR can be further deployed to support specific needs in various scenarios without requiring extra language models. Compared to recently released large ASR models, Seed-ASR achieves 10%-40% reduction in word (or character, for Chinese) error rates on Chinese and English public test sets, further demonstrating its powerful performance.
