← 返回资源分享
Open-Source Conversational AI with SpeechBrain 1.0
论文
论文
发布时间2024-06-29
发表arXiv:2407.00463
作者:Titouan Parcollet,Mirco Ravanelli,Renato de Mori,Peter Plantinga,Xuechen Liu,Aku Rouhe,Mickael Rouvier,Yannick Esteve,Cem Subakan,Andreas Nautsch,Shucong Zhang,Juan Zuluaga-Gomez,Rudolf Braun,Salah Zaiem,Salima Mdhaffar,Jarod Duret,Yingzhi Wang,Pierre Champion,Sung-Lin Yeh,Artem Ploujnikov,Georgios Karakasidis,Adel Moumen,Francesco Paissan,Zeyu Zhao,Florian Mai,Sangeet Sagar,Luca Della Libera,Pooneh Mousavi,Seyed Mahed Mousavi,Sylvain de Langen,Davide Borra,Gaelle Laperriere
详细介绍
SpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more. It promotes transparency and replicability by releasing both the pre-trained models and the complete "recipes" of code and algorithms required for training them. This paper presents SpeechBrain 1.0, a significant milestone in the evolution of the toolkit, which now has over 200 recipes for speech, audio, and language processing tasks, and more than 100 models available on Hugging Face. SpeechBrain 1.0 introduces new technologies to support diverse learning modalities, Large Language Model (LLM) integration, and advanced decoding strategies, along with novel models, tasks, and modalities. It also includes a new benchmark repository, offering researchers a unified platform for evaluating models across diverse tasks.
