本文经arXiv每日学术速递授权转载
【1】 Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation
标题: 用于高效多语言语音到语音翻译的扩散合成器
作者:Nameer Hirschkind,Xiao Yu,Mahesh Kumar Nandwana,Joseph Liu,Eloi DuBois,Dao Le,Nicolas Thiebaut,Colin Sinclair,Kyle Spence,Charles Shang,Zoe Abrams,Morgan McGuire
备注:Published in Interspeech 2024
链接:点击下载PDF文件
【2】 One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model
标题: 使用一体化神经模型的一次多Conformer和Foundation语音系统压缩和量化
作者:Zhaoqing Li,Haoning Xu,Tianzi Wang,Shoukang Hu,Zengrui Jin,Shujie Hu,Jiajun Deng,Mingyu Cui,Mengzhe Geng,Xunying Liu
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【3】 Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
标题: 用于视听多通道语音分离和识别的联合说话人特征学习
作者:Guinan Li,Jiajun Deng,Youjun Chen,Mengzhe Geng,Shujie Hu,Zhe Li,Zengrui Jin,Tianzi Wang,Xurong Xie,Helen Meng,Xunying Liu
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【4】 On the Evaluation of Speech Foundation Models for Spoken Language Understanding
标题: 关于口语理解的言语基础模型的评估
作者:Siddhant Arora,Ankita Pasad,Chung-Ming Chien,Jionghao Han,Roshan Sharma,Jee-weon Jung,Hira Dhamyal,William Chen,Suwon Shon,Hung-yi Lee,Karen Livescu,Shinji Watanabe
备注:Accepted at ACL Findings 2024
链接:点击下载PDF文件
【5】 UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
标题: UniAudio 1.5:大型语言模型驱动的音频编解码器是一款只需几次的音频任务学习器
作者:Dongchao Yang,Haohan Guo,Yuanyuan Wang,Rongjie Huang,Xiang Li,Xu Tan,Xixin Wu,Helen Meng
链接:点击下载PDF文件
【6】 Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
标题: Sim-Whisper:具有截断检测的注意力引导流媒体Whisper
作者:Haoyu Wang,Guoqiang Hu,Guodong Lin,Wei-Qiang Zhang,Jian Li
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
【7】 Towards Effective and Efficient Non-autoregressive Decoding Using Block-based Attention Mask
标题: 使用基于块的注意力屏蔽实现有效且高效的非自回归解码
作者:Tianzi Wang,Xurong Xie,Zhaoqing Li,Shoukang Hu,Zengrui Jing,Jiajun Deng,Mingyu Cui,Shujie Hu,Mengzhe Geng,Guinan Li,Helen Meng,Xunying Liu
备注:5 pages, 2 figures, 2 tables, Interspeech24 conference
链接:点击下载PDF文件
【8】 Impact of Speech Mode in Automatic Pathological Speech Detection
标题: 语音模式对自动病理语音检测的影响
作者:Shakeel A. Sheikh,Ina Kodrasi
备注:Accepted in EUSIPCO 2024
链接:点击下载PDF文件
【9】 An efficient text augmentation approach for contextualized Mandarin speech recognition
标题: 一种有效的上下文化普通话语音识别文本增强方法
作者:Naijun Zheng,Xucheng Wan,Kai Liu,Ziqing Du,Zhou Huan
备注:accepted to interspeech2024
链接:点击下载PDF文件
【10】 What Does it Take to Generalize SER Model Across Datasets? A Comprehensive Benchmark
标题: 如何在数据集中推广BER模型?全面的基准
作者:Adham Ibrahim,Shady Shehata,Ajinkya Kulkarni,Mukhtar Mohamed,Muhammad Abdul-Mageed
备注:ACCEPTED AT INTERSPEECH 2024, GREECE
链接:点击下载PDF文件
【11】 Personalized Speech Enhancement Without a Separate Speaker Embedding Model
标题: 无需单独说话者嵌入模型的个性化语音增强
作者:Tanel Pärnamaa,Ando Saabas
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【12】 MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
标题: MMM:来自自监督学习模型的多层多残留多流离散语音表示
作者:Jiatong Shi,Xutai Ma,Hirofumi Inaguma,Anna Sun,Shinji Watanabe
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
【13】 Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
标题: Vec-Tok-VC+:双模式训练策略中具有渐进约束的剩余增强稳健Zero-Shot语音转换
作者:Linhan Ma,Xinfa Zhu,Yuanjun Lv,Zhichao Wang,Ziqian Wang,Wendi He,Hongbin Zhou,Lei Xie
备注:Accepted by INTERSPEECH2024
链接:点击下载PDF文件
【14】 SHMamba: Structured Hyperbolic State Space Model for Audio-Visual Question Answering
标题: SHMamba:视听问题回答的结构化双曲状态空间模型
作者:Zhe Yang,Wenrui Li,Guanghui Cheng
链接:点击下载PDF文件
【15】 Frequency-mix Knowledge Distillation for Fake Speech Detection
标题: 用于假语音检测的混频知识提炼
作者:Cunhang Fan,Shunbo Dong,Jun Xue,Yujie Chen,Jiangyan Yi,Zhao Lv
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【16】 Multi-Modal Retrieval For Large Language Model Based Speech Recognition
标题: 基于大语言模型的语音识别的多模式检索
作者:Jari Kolehmainen,Aditya Gourav,Prashanth Gurunath Shivakumar,Yile Gu,Ankur Gandhe,Ariya Rastrow,Grant Strimel,Ivan Bulyko
链接:点击下载PDF文件
【17】 Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
标题: Speech ReaLLM --通过教学时间流,使用多模式LLM进行实时流语音识别
作者:Frank Seide,Morrie Doulaty,Yangyang Shi,Yashesh Gaur,Junteng Jia,Chunyang Wu
链接:点击下载PDF文件
【18】 Analyzing phonetic structure of Mandarin using Audacity
标题: 用Audacity分析普通话的语音结构
作者:Shizheng Xu
备注:audio source: this https URL
链接:点击下载PDF文件
【19】 Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
标题: Whisper-Flamingo:将视觉特征集成到Whisper中以实现视听语音识别和翻译
作者:Andrew Rouditchenko,Yuan Gong,Samuel Thomas,Leonid Karlinsky,Hilde Kuehne,Rogerio Feris,James Glass
备注:Interspeech 2024. Code this https URL
链接:点击下载PDF文件
【20】 Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
标题: 检测法国电视和广播内容中言语互动的语音转向边界的终点
作者:Rémi Uro,Marie Tahon,David Doukhan,Antoine Laurent,Albert Rilliard
备注:keywords : Spoken interaction, Media, TV, Radio, Transition-Relevance Places, Turn Taking, Interruption. Accepted to InterSpeech 2024, Kos Island, Greece
链接:点击下载PDF文件
【21】 Understanding Pedestrian Movement Using Urban Sensing Technologies: The Promise of Audio-based Sensors
标题: 使用城市传感技术了解行人运动:基于音频的传感器的前景
作者:Chaeyeon Han,Pavan Seshadri,Yiwei Ding,Noah Posner,Bon Woo Koo,Animesh Agrawal,Alexander Lerch,Subhrajit Guhathakurta
备注:submitted to Urban Informatics
链接:点击下载PDF文件
【22】 Period Singer: Integrating Periodic and Aperiodic Variational Autoencoders for Natural-Sounding End-to-End Singing Voice Synthesis
标题: Period Singer:集成周期性和非周期性变分自动编码器,实现自然声音端到端歌唱声音合成
作者:Taewoo Kim,Choongsang Cho,Young Han Lee
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【23】 Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
标题: 感知器提示:Whisper中灵活的说话人适应中文无序语音识别
作者:Yicong Jiang,Tianzi Wang,Xurong Xie,Juan Liu,Wei Sun,Nan Yan,Hui Chen,Lan Wang,Xunying Liu,Feng Tian
备注:Accepted by interspeech 2024
链接:点击下载PDF文件
【24】 GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
标题: GenDistiller:基于自回归生成模型提取预训练语言模型
作者:Yingying Gao,Shilei Zhang,Chao Deng,Junlan Feng
备注:arXiv admin note: text overlap with arXiv:2310.13418
链接:点击下载PDF文件
标题: AlignNet:学习数据集得分对齐功能,以实现更好的语音质量估计器训练
作者:Jaden Pieper,Stephen Voran
备注:To be published in proc. of Interspeech 2024. 5 pages, 2 figures, 3 tables
链接:点击下载PDF文件
【2】 Inclusive ASR for Disfluent Speech: Cascaded Large-Scale Self-Supervised Learning with Targeted Fine-Tuning and Data Augmentation
标题: 针对不流利言语的包容性ASB:具有有针对性的微调和数据增强的级联大规模自我监督学习
作者:Dena Mujtaba,Nihar R. Mahapatra,Megan Arney,J. Scott Yaruss,Caryn Herring,Jia Bin
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
【3】 Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
标题: Whisper-Flamingo:将视觉特征集成到Whisper中以实现视听语音识别和翻译
作者:Andrew Rouditchenko,Yuan Gong,Samuel Thomas,Leonid Karlinsky,Hilde Kuehne,Rogerio Feris,James Glass
备注:Interspeech 2024. Code this https URL
链接:点击下载PDF文件
【4】 Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
标题: 检测法国电视和广播内容中言语互动的语音转向边界的终点
作者:Rémi Uro,Marie Tahon,David Doukhan,Antoine Laurent,Albert Rilliard
备注:keywords : Spoken interaction, Media, TV, Radio, Transition-Relevance Places, Turn Taking, Interruption. Accepted to InterSpeech 2024, Kos Island, Greece
链接:点击下载PDF文件
【5】 ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2Vec2.0 Based ASR
标题: ROAR:增强基于Wav2Vec2.0的ASB的原始数据与增强数据比率动态
作者:Vishwanath Pratap Singh,Federico Malato,Ville Hautamaki,Md. Sahidullah,Tomi Kinnunen
Journal-ref:Interspeech 2024
链接:点击下载PDF文件
【6】 Understanding Pedestrian Movement Using Urban Sensing Technologies: The Promise of Audio-based Sensors
标题: 使用城市传感技术了解行人运动:基于音频的传感器的前景
作者:Chaeyeon Han,Pavan Seshadri,Yiwei Ding,Noah Posner,Bon Woo Koo,Animesh Agrawal,Alexander Lerch,Subhrajit Guhathakurta
备注:submitted to Urban Informatics
链接:点击下载PDF文件
【7】 Period Singer: Integrating Periodic and Aperiodic Variational Autoencoders for Natural-Sounding End-to-End Singing Voice Synthesis
标题: Period Singer:集成周期性和非周期性变分自动编码器,实现自然声音端到端歌唱声音合成
作者:Taewoo Kim,Choongsang Cho,Young Han Lee
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【8】 Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
标题: 感知器提示:Whisper中灵活的说话人适应中文无序语音识别
作者:Yicong Jiang,Tianzi Wang,Xurong Xie,Juan Liu,Wei Sun,Nan Yan,Hui Chen,Lan Wang,Xunying Liu,Feng Tian
备注:Accepted by interspeech 2024
链接:点击下载PDF文件
【9】 Low algorithmic delay implementation of convolutional beamformer for online joint source separation and dereverberation
标题: 用于在线联合源分离和去回响的卷积束形成器的低算法延迟实现
作者:Kaien Mo,Xianrui Wang,Yichen Yang,Shoji Makino,Jingdong Chen
备注:4 pages, 4 figures. Accepted by EUSIPCO 2024
链接:点击下载PDF文件
【10】 Enhanced Deep Speech Separation in Clustered Ad Hoc Distributed Microphone Environments
标题: 小型Ad Hoc分布式麦克风环境中的增强深度语音分离
作者:Jihyun Kim,Stijn Kindt,Nilesh Madhu,Hong-Goo Kang
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【11】 A Multimodal Framework for the Assessment of the Schizophrenia Spectrum
标题: 精神分裂症谱系评估的多模式框架
作者:Gowtham Premananth,Yashish M. Siriwardena,Philip Resnik,Sonia Bansal,Deanna L. Kelly,Carol Espy-Wilson
备注:Accepted to be presented at Interspeech 2024
链接:点击下载PDF文件
【12】 Optimizing Byte-level Representation for End-to-end ASR
标题: 优化端到端ASB的字节级表示
作者:Roger Hsiao,Liuhui Deng,Erik McDermott,Ruchir Travadi,Xiaodan Zhuang
备注:5 pages, 1 figure
链接:点击下载PDF文件
【13】 Efficient Personalization of Amplification in Hearing Aids via Multi-band Bayesian Machine Learning
标题: 通过多频段Bayesian机器学习实现助听器放大的有效个性化
作者:Aoxin Ni,Edward Lobarinas,Nasser Kehtarnavaz
链接:点击下载PDF文件
【14】 Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
标题: 使用目标说话者独奏段的多通道多说话者ASB
作者:Yiwen Shao,Shi-Xiong Zhang,Yong Xu,Meng Yu,Dong Yu,Daniel Povey,Sanjeev Khudanpur
备注:Accepted for presentation at Interspeech 2024
链接:点击下载PDF文件
【15】 The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments
标题: 第二个DISPLACE挑战:对话环境中Speaker和LAnguage的二元化
作者:Shareef Babu Kalluri,Prachi Singh,Pratik Roy Chowdhuri,Apoorva Kulkarni,Shikha Baghel,Pradyoth Hegde,Swapnil Sontakke,Deepak K T,S. R. Mahadeva Prasanna,Deepu Vijayasenan,Sriram Ganapathy
备注:5 pages, 3 figures, Interspeech 2024
链接:点击下载PDF文件
【16】 GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
标题: GenDistiller:基于自回归生成模型提取预训练语言模型
作者:Yingying Gao,Shilei Zhang,Chao Deng,Junlan Feng
备注:arXiv admin note: text overlap with arXiv:2310.13418
链接:点击下载PDF文件
【17】 Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness
标题: 个性化语音活动检测系统的比较分析:评估现实世界的有效性
作者:Satyam Kumar,Sai Srujana Buddi,Utkarsh Oggy Sarawgi,Vineet Garg,Shivesh Ranjan,Ognjen,Rudovic,Ahmed Hussen Abdelaziz,Saurabh Adya
链接:点击下载PDF文件
【18】 Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation
标题: 用于高效多语言语音到语音翻译的扩散合成器
作者:Nameer Hirschkind,Xiao Yu,Mahesh Kumar Nandwana,Joseph Liu,Eloi DuBois,Dao Le,Nicolas Thiebaut,Colin Sinclair,Kyle Spence,Charles Shang,Zoe Abrams,Morgan McGuire
备注:Published in Interspeech 2024
链接:点击下载PDF文件
【19】 One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model
标题: 使用一体化神经模型的一次多Conformer和Foundation语音系统压缩和量化
作者:Zhaoqing Li,Haoning Xu,Tianzi Wang,Shoukang Hu,Zengrui Jin,Shujie Hu,Jiajun Deng,Mingyu Cui,Mengzhe Geng,Xunying Liu
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【20】 Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
标题: 用于视听多通道语音分离和识别的联合说话人特征学习
作者:Guinan Li,Jiajun Deng,Youjun Chen,Mengzhe Geng,Shujie Hu,Zhe Li,Zengrui Jin,Tianzi Wang,Xurong Xie,Helen Meng,Xunying Liu
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【21】 On the Evaluation of Speech Foundation Models for Spoken Language Understanding
标题: 关于口语理解的言语基础模型的评估
作者:Siddhant Arora,Ankita Pasad,Chung-Ming Chien,Jionghao Han,Roshan Sharma,Jee-weon Jung,Hira Dhamyal,William Chen,Suwon Shon,Hung-yi Lee,Karen Livescu,Shinji Watanabe
备注:Accepted at ACL Findings 2024
链接:点击下载PDF文件
【22】 UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
标题: UniAudio 1.5:大型语言模型驱动的音频编解码器是一款只需几次的音频任务学习器
作者:Dongchao Yang,Haohan Guo,Yuanyuan Wang,Rongjie Huang,Xiang Li,Xu Tan,Xixin Wu,Helen Meng
链接:点击下载PDF文件
【23】 Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
标题: Sim-Whisper:具有截断检测的注意力引导流媒体Whisper
作者:Haoyu Wang,Guoqiang Hu,Guodong Lin,Wei-Qiang Zhang,Jian Li
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
【24】 Towards Effective and Efficient Non-autoregressive Decoding Using Block-based Attention Mask
标题: 使用基于块的注意力屏蔽实现有效且高效的非自回归解码
作者:Tianzi Wang,Xurong Xie,Zhaoqing Li,Shoukang Hu,Zengrui Jing,Jiajun Deng,Mingyu Cui,Shujie Hu,Mengzhe Geng,Guinan Li,Helen Meng,Xunying Liu
备注:5 pages, 2 figures, 2 tables, Interspeech24 conference
链接:点击下载PDF文件
【25】 Impact of Speech Mode in Automatic Pathological Speech Detection
标题: 语音模式对自动病理语音检测的影响
作者:Shakeel A. Sheikh,Ina Kodrasi
备注:Accepted in EUSIPCO 2024
链接:点击下载PDF文件
【26】 An efficient text augmentation approach for contextualized Mandarin speech recognition
标题: 一种有效的上下文化普通话语音识别文本增强方法
作者:Naijun Zheng,Xucheng Wan,Kai Liu,Ziqing Du,Zhou Huan
备注:accepted to interspeech2024
链接:点击下载PDF文件
【27】 Personalized Speech Enhancement Without a Separate Speaker Embedding Model
标题: 无需单独说话者嵌入模型的个性化语音增强
作者:Tanel Pärnamaa,Ando Saabas
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【28】 MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
标题: MMM:来自自监督学习模型的多层多残留多流离散语音表示
作者:Jiatong Shi,Xutai Ma,Hirofumi Inaguma,Anna Sun,Shinji Watanabe
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
【29】 Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
标题: Vec-Tok-VC+:双模式训练策略中具有渐进约束的剩余增强稳健Zero-Shot语音转换
作者:Linhan Ma,Xinfa Zhu,Yuanjun Lv,Zhichao Wang,Ziqian Wang,Wendi He,Hongbin Zhou,Lei Xie
备注:Accepted by INTERSPEECH2024
链接:点击下载PDF文件
【30】 SHMamba: Structured Hyperbolic State Space Model for Audio-Visual Question Answering
标题: SHMamba:视听问题回答的结构化双曲状态空间模型
作者:Zhe Yang,Wenrui Li,Guanghui Cheng
链接:点击下载PDF文件
【31】 Frequency-mix Knowledge Distillation for Fake Speech Detection
标题: 用于假语音检测的混频知识提炼
作者:Cunhang Fan,Shunbo Dong,Jun Xue,Yujie Chen,Jiangyan Yi,Zhao Lv
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【32】 Multi-Modal Retrieval For Large Language Model Based Speech Recognition
标题: 基于大语言模型的语音识别的多模式检索
作者:Jari Kolehmainen,Aditya Gourav,Prashanth Gurunath Shivakumar,Yile Gu,Ankur Gandhe,Ariya Rastrow,Grant Strimel,Ivan Bulyko
链接:点击下载PDF文件
【33】 Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
标题: 用于设备定向语音检测的具有融合低等级自适应的多模式大语言模型
作者:Shruti Palaskar,Oggi Rudovic,Sameer Dharur,Florian Pesce,Gautam Krishna,Aswin Sivaraman,Jack Berkowitz,Ahmed Hussen Abdelaziz,Saurabh Adya,Ahmed Tewfik
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
【34】 Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
标题: Speech ReaLLM --通过教学时间流,使用多模式LLM进行实时流语音识别
作者:Frank Seide,Morrie Doulaty,Yangyang Shi,Yashesh Gaur,Junteng Jia,Chunyang Wu
链接:点击下载PDF文件
【35】 Analyzing phonetic structure of Mandarin using Audacity
标题: 用Audacity分析普通话的语音结构
作者:Shizheng Xu
备注:audio source: this https URL
链接:点击下载PDF文件
标题: 用于高效多语言语音到语音翻译的扩散合成器
作者:Nameer Hirschkind,Xiao Yu,Mahesh Kumar Nandwana,Joseph Liu,Eloi DuBois,Dao Le,Nicolas Thiebaut,Colin Sinclair,Kyle Spence,Charles Shang,Zoe Abrams,Morgan McGuire
备注:Published in Interspeech 2024
链接:点击下载PDF文件
摘要:我们介绍了DiffuseST,一个低延迟,直接语音到语音翻译系统,能够保留输入扬声器的声音zero-shot,同时从多个源语言翻译成英语。我们的实验与合成器组件的架构,比较基于Tacotron的合成器的一种新的扩散为基础的合成器。我们发现基于扩散的合成器可以将MOS和PESQ音频质量指标分别提高23%,扬声器相似性提高5%,同时保持可比较的BLEU分数。尽管有两倍多的参数计数,扩散合成器具有较低的延迟,使整个模型的运行速度比实时快5倍。摘要:We introduce DiffuseST, a low-latency, direct speech-to-speech translation system capable of preserving the input speaker's voice zero-shot while translating from multiple source languages into English. We experiment with the synthesizer component of the architecture, comparing a Tacotron-based synthesizer to a novel diffusion-based synthesizer. We find the diffusion-based synthesizer to improve MOS and PESQ audio quality metrics by 23 % each and speaker similarity by 5 % while maintaining comparable BLEU scores. Despite having more than double the parameter count, the diffusion synthesizer has lower latency, allowing the entire model to run more than 5$ times$ faster than real-time.
【2】 One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model
标题: 使用一体化神经模型的一次多Conformer和Foundation语音系统压缩和量化
作者:Zhaoqing Li,Haoning Xu,Tianzi Wang,Shoukang Hu,Zengrui Jin,Shujie Hu,Jiajun Deng,Mingyu Cui,Mengzhe Geng,Xunying Liu
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:我们提出了一种新的一次通过多个ASR系统联合压缩和量化的方法,使用一个全在一个神经模型。单个压缩周期允许同时构造具有不同编码器深度、宽度和量化精度设置的多个嵌套系统,而无需单独训练和存储各个目标系统。实验一致表明,在单个一体化模型中压缩的多个ASR系统产生的单词错误率(WER)与同等复杂度的单独训练系统相当,或比其低1.01%绝对值(6.98%相对值)。实现了3.4倍的整体系统压缩和训练时间加速。在基线Switchboard-300 hr Conformer和LibriSpeech-100 hr微调wav2vec2.0模型上分别获得了12.8x和3.93x的最大模型大小压缩比,未引起统计学显著的WER增加。摘要:We propose a novel one-pass multiple ASR systems joint compression and quantization approach using an all-in-one neural model. A single compression cycle allows multiple nested systems with varying Encoder depths, widths, and quantization precision settings to be simultaneously constructed without the need to train and store individual target systems separately. Experiments consistently demonstrate the multiple ASR systems compressed in a single all-in-one model produced a word error rate (WER) comparable to, or lower by up to 1.01 % absolute (6.98 % relative) than individually trained systems of equal complexity. A 3.4x overall system compression and training time speed-up was achieved. Maximum model size compression ratios of 12.8x and 3.93x were obtained over the baseline Switchboard-300hr Conformer and LibriSpeech-100hr fine-tuned wav2vec2.0 models, respectively, incurring no statistically significant WER increase.
【3】 Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
标题: 用于视听多通道语音分离和识别的联合说话人特征学习
作者:Guinan Li,Jiajun Deng,Youjun Chen,Mengzhe Geng,Shujie Hu,Zhe Li,Zengrui Jin,Tianzi Wang,Xurong Xie,Helen Meng,Xunying Liu
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:本文提出了一种用于视听多通道语音分离与识别系统中zero-shot自适应的联合说话人特征学习方法。xVector和ECAPA-TDNN扬声器编码器使用专用融合模块连接,并与完整的系统训练紧密集成。LRS 3-TED数据模拟多通道重叠语音进行的实验表明,联合说话人特征学习一贯提高语音分离和识别性能的基线没有联合说话人特征估计。进一步的分析表明,性能的改善与使用余弦相似性测量的说话人间辨别力的增加密切相关。性能最佳的联合扬声器特征学习适应系统优于基线微调WavLM模型,在结合WavLM特征和视频模态后,在Dev和Test集上的WER绝对值降低了21.6%和25.3%(相对值降低了67.5%和83.5%)。摘要:This paper proposes joint speaker feature learning methods for zero-shot adaptation of audio-visual multichannel speech separation and recognition systems. xVector and ECAPA-TDNN speaker encoders are connected using purpose-built fusion blocks and tightly integrated with the complete system training. Experiments conducted on LRS3-TED data simulated multichannel overlapped speech suggest that joint speaker feature learning consistently improves speech separation and recognition performance over the baselines without joint speaker feature estimation. Further analyses reveal performance improvements are strongly correlated with increased inter-speaker discrimination measured using cosine similarity. The best-performing joint speaker feature learning adapted system outperformed the baseline fine-tuned WavLM model by statistically significant WER reductions of 21.6% and 25.3% absolute (67.5% and 83.5% relative) on Dev and Test sets after incorporating WavLM features and video modality.
【4】 On the Evaluation of Speech Foundation Models for Spoken Language Understanding
标题: 关于口语理解的言语基础模型的评估
作者:Siddhant Arora,Ankita Pasad,Chung-Ming Chien,Jionghao Han,Roshan Sharma,Jee-weon Jung,Hira Dhamyal,William Chen,Suwon Shon,Hung-yi Lee,Karen Livescu,Shinji Watanabe
备注:Accepted at ACL Findings 2024
链接:点击下载PDF文件
摘要:最近推出了口语理解评估(SLUE)基准测试任务套件,以满足对开放资源的需求,并对复杂的口语理解(SLU)任务进行基准测试,包括自然语音的分类和序列生成任务。该基准测试已经证明了使用预先训练的语音基础模型(SFM)进行这些SLU任务的初步成功。然而,社会对不同可持续森林管理措施的比较效用仍缺乏深入的了解。受此启发,我们问:哪些SFM为这些复杂的SLU任务提供了最大的好处,以及整合这些SFM的最有效方法是什么?为了回答这个问题,我们使用几种评估协议对多个监督和自监督的SFM进行了广泛的评估:(i)具有轻量级预测头的冻结SFM,(ii)具有复杂预测头的冻结SFM,以及(iii)具有轻量级预测头的微调SFM。虽然监督的SFM是在更多的语音识别数据(带有标签)上进行预训练的,但它们并不总是优于自监督的SFM;后者的表现往往至少与监督的SFM一样好,有时甚至更好,特别是在SLUE中的序列生成任务上。虽然没有通用的最佳方法来合并SFM,但复杂的预测头为大多数任务提供了最佳性能,尽管它增加了推理时间。我们还介绍了一个开源的工具包和性能排行榜,SLUE-PERB,这些任务和建模策略。摘要:The Spoken Language Understanding Evaluation (SLUE) suite of benchmark tasks was recently introduced to address the need for open resources and benchmarking of complex spoken language understanding (SLU) tasks, including both classification and sequence generation tasks, on natural speech. The benchmark has demonstrated preliminary success in using pre-trained speech foundation models (SFM) for these SLU tasks. However, the community still lacks a fine-grained understanding of the comparative utility of different SFMs. Inspired by this, we ask: which SFMs offer the most benefits for these complex SLU tasks, and what is the most effective approach for incorporating these SFMs? To answer this, we perform an extensive evaluation of multiple supervised and self-supervised SFMs using several evaluation protocols: (i) frozen SFMs with a lightweight prediction head, (ii) frozen SFMs with a complex prediction head, and (iii) fine-tuned SFMs with a lightweight prediction head. Although the supervised SFMs are pre-trained on much more speech recognition data (with labels), they do not always outperform self-supervised SFMs; the latter tend to perform at least as well as, and sometimes better than, supervised SFMs, especially on the sequence generation tasks in SLUE. While there is no universally optimal way of incorporating SFMs, the complex prediction head gives the best performance for most tasks, although it increases the inference time. We also introduce an open-source toolkit and performance leaderboard, SLUE-PERB, for these tasks and modeling strategies.
【5】 UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
标题: UniAudio 1.5:大型语言模型驱动的音频编解码器是一款只需几次的音频任务学习器
作者:Dongchao Yang,Haohan Guo,Yuanyuan Wang,Rongjie Huang,Xiang Li,Xu Tan,Xixin Wu,Helen Meng
链接:点击下载PDF文件
摘要:大型语言模型(LLM)在文本理解和生成方面表现出了卓越的能力,但如果不进行微调,就不能直接应用于跨模态任务。本文提出了一种跨模态上下文学习方法,使冻结的LLM能够以Few-Shot风格实现多个音频任务,而无需任何参数更新。具体来说,我们提出了一种新的LLMs驱动的音频编解码器模型,LLM-Codec,将音频模态转移到文本空间,用LLM的词汇表中的词或子词来表示音频令牌,同时保持高的音频重构质量。其关键思想是通过将音频模态压缩到经过良好训练的LLM令牌空间中来减少文本和音频之间的模态异构性。因此,音频表示可以被视为一种新的 textit{foreign language},LLM可以通过几次演示来学习新的 textit{foreign language}。在实验中,我们研究了所提出的方法在多个音频理解和生成任务中的性能,语音情感分类、音频分类、文本到语音生成、语音增强等。实验结果表明,在简单的场景下,采用本文提出的LLM-Codec(UniAudio 1.5)的LLM能够实现预期的功能。实验结果验证了跨模态情境学习方法的可行性和有效性。为了促进对Few-Shot音频任务学习和多模式LLM的研究,我们开源了LLM-Codec模型。摘要:The Large Language models (LLMs) have demonstrated supreme capabilities in text understanding and generation, but cannot be directly applied to cross-modal tasks without fine-tuning. This paper proposes a cross-modal in-context learning approach, empowering the frozen LLMs to achieve multiple audio tasks in a few-shot style without any parameter update. Specifically, we propose a novel and LLMs-driven audio codec model, LLM-Codec, to transfer the audio modality into the textual space, textit{i.e.} representing audio tokens with words or sub-words in the vocabulary of LLMs, while keeping high audio reconstruction quality. The key idea is to reduce the modality heterogeneity between text and audio by compressing the audio modality into a well-trained LLMs token space. Thus, the audio representation can be viewed as a new textit{foreign language}, and LLMs can learn the new textit{foreign language} with several demonstrations. In experiments, we investigate the performance of the proposed approach across multiple audio understanding and generation tasks, textit{e.g.} speech emotion classification, audio classification, text-to-speech generation, speech enhancement, etc. The experimental results demonstrate that the LLMs equipped with the proposed LLM-Codec, named as UniAudio 1.5, prompted by only a few examples, can achieve the expected functions in simple scenarios. It validates the feasibility and effectiveness of the proposed cross-modal in-context learning approach. To facilitate research on few-shot audio task learning and multi-modal LLMs, we have open-sourced the LLM-Codec model.
【6】 Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
标题: Sim-Whisper:具有截断检测的注意力引导流媒体Whisper
作者:Haoyu Wang,Guoqiang Hu,Guodong Lin,Wei-Qiang Zhang,Jian Li
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:作为一个强大的大规模多语言语音识别模型,Whisper已经在许多低资源和非分布场景中展示了令人印象深刻的结果。然而,它的编解码器结构阻碍了其应用于流语音识别。在本文中,我们介绍了Simul-Whisper,它使用嵌入在Whisper的交叉注意中的时间对齐来指导自回归解码,并实现基于块的流式ASR,而无需对预训练模型进行任何微调。此外,我们观察到的负面影响,截断词在块边界上的解码结果,并提出了一个完整的和火灾为基础的截断检测模型来解决这个问题。在多种语言和Whisper架构上的实验表明,Simul-Whisper在1秒的块大小下平均绝对字错误率仅下降1.46%,显着优于当前最先进的基线。摘要:As a robust and large-scale multilingual speech recognition model, Whisper has demonstrated impressive results in many low-resource and out-of-distribution scenarios. However, its encoder-decoder structure hinders its application to streaming speech recognition. In this paper, we introduce Simul-Whisper, which uses the time alignment embedded in Whisper's cross-attention to guide auto-regressive decoding and achieve chunk-based streaming ASR without any fine-tuning of the pre-trained model. Furthermore, we observe the negative effect of the truncated words at the chunk boundaries on the decoding results and propose an integrate-and-fire-based truncation detection model to address this issue. Experiments on multiple languages and Whisper architectures show that Simul-Whisper achieves an average absolute word error rate degradation of only 1.46% at a chunk size of 1 second, which significantly outperforms the current state-of-the-art baseline.
【7】 Towards Effective and Efficient Non-autoregressive Decoding Using Block-based Attention Mask
标题: 使用基于块的注意力屏蔽实现有效且高效的非自回归解码
作者:Tianzi Wang,Xurong Xie,Zhaoqing Li,Shoukang Hu,Zengrui Jing,Jiajun Deng,Mingyu Cui,Shujie Hu,Mengzhe Geng,Guinan Li,Helen Meng,Xunying Liu
备注:5 pages, 2 figures, 2 tables, Interspeech24 conference
链接:点击下载PDF文件
摘要:本文提出了一种新的非自回归(NAR)块为基础的注意掩码解码器(AMD),灵活地平衡性能效率的权衡一致性ASR系统。AMD在使用注意掩码隐藏的输出标签的连续块内执行并行NAR推理,同时在块之间进行从左到右的AR预测和历史上下文融合。波束搜索算法被设计为利用CTC、AR解码器和AMD概率的动态融合。LibriSpeech-100 hr语料库上的实验表明,结合AMD模块的三方解码器产生了比基线CTC+AR解码1.73倍的最大解码加速比,同时在测试集上没有产生统计上显著的字错误率(WER)增加。当使用相同的解码实时因子操作时,在CTC+AR基线上获得了高达0.7%和0.3%的绝对(5.3%和6.1%的相对)统计学显著WER降低。摘要:This paper proposes a novel non-autoregressive (NAR) block-based Attention Mask Decoder (AMD) that flexibly balances performance-efficiency trade-offs for Conformer ASR systems. AMD performs parallel NAR inference within contiguous blocks of output labels that are concealed using attention masks, while conducting left-to-right AR prediction and history context amalgamation between blocks. A beam search algorithm is designed to leverage a dynamic fusion of CTC, AR Decoder, and AMD probabilities. Experiments on the LibriSpeech-100hr corpus suggest the tripartite Decoder incorporating the AMD module produces a maximum decoding speed-up ratio of 1.73x over the baseline CTC+AR decoding, while incurring no statistically significant word error rate (WER) increase on the test sets. When operating with the same decoding real time factors, statistically significant WER reductions of up to 0.7% and 0.3% absolute (5.3% and 6.1% relative) were obtained over the CTC+AR baseline.
【8】 Impact of Speech Mode in Automatic Pathological Speech Detection
标题: 语音模式对自动病理语音检测的影响
作者:Shakeel A. Sheikh,Ina Kodrasi
备注:Accepted in EUSIPCO 2024
链接:点击下载PDF文件
摘要:自动病理语音检测方法在识别各种病理方面产生了很好的结果。这些方法通常是针对语音控制的语音场景而设计和评估的,在语音控制的语音场景中,提示说话者清晰地表达相同的语音内容。虽然收集受控的语音记录可能是费力的,但随着潜在患者浏览他们的日常生活,可以方便地获得自发语音。此外,自发言语在检测病理性言语的微妙和抽象线索方面可能是有价值的。尽管如此,自动病理语音检测自发语音的有效性仍有待探索。本文分析了语音模式对病态语音检测方法的影响,研究了两种不同的方法,即,经典机器学习和深度学习。结果表明,经典的方法可能难以捕捉自发语音的病理判别线索。相比之下,深度学习方法表现出卓越的性能,能够提取以前在非自发语音中无法获得的额外线索摘要:Automatic pathological speech detection approaches yield promising results in identifying various pathologies. These approaches are typically designed and evaluated for phonetically-controlled speech scenarios, where speakers are prompted to articulate identical phonetic content. While gathering controlled speech recordings can be laborious, spontaneous speech can be conveniently acquired as potential patients navigate their daily routines. Further, spontaneous speech can be valuable in detecting subtle and abstract cues of pathological speech. Nonetheless, the efficacy of automatic pathological speech detection for spontaneous speech remains unexplored. This paper analyzes the influence of speech mode on pathological speech detection approaches, examining two distinct categories of approaches, i.e., classical machine learning and deep learning. Results indicate that classical approaches may struggle to capture pathology-discriminant cues in spontaneous speech. In contrast, deep learning approaches demonstrate superior performance, managing to extract additional cues that were previously inaccessible in non-spontaneous speech
【9】 An efficient text augmentation approach for contextualized Mandarin speech recognition
标题: 一种有效的上下文化普通话语音识别文本增强方法
作者:Naijun Zheng,Xucheng Wan,Kai Liu,Ziqing Du,Zhou Huan
备注:accepted to interspeech2024
链接:点击下载PDF文件
摘要:虽然上下文自动语音识别(ASR)系统通常用于提高识别不常见的单词,其有效性受到语音文本数据可用性的固有限制。为了应对这一挑战,我们的研究建议利用广泛的纯文本数据集,并使用简单的文本增强(TA)技术将预训练的ASR模型置于上下文中,同时保持计算成本最小。特别是,为了将预先训练的基于CIF的ASR置于上下文中,我们使用有限的语音文本数据构建码本。通过利用一个简单的码本查找过程,我们将可用的纯文本数据转换为潜在的文本嵌入。然后,这些嵌入增强了情境化ASR的输入。我们在不同的普通话测试集上的实验表明,我们的TA方法显着提高识别性能。表现最好的系统在罕见单词上显示出高达30%的相对CER改进,在所有单词上显示出15%的相对CER改进。摘要:Although contextualized automatic speech recognition (ASR) systems are commonly used to improve the recognition of uncommon words, their effectiveness is hindered by the inherent limitations of speech-text data availability. To address this challenge, our study proposes to leverage extensive text-only datasets and contextualize pre-trained ASR models using a straightforward text-augmentation (TA) technique, all while keeping computational costs minimal. In particular, to contextualize a pre-trained CIF-based ASR, we construct a codebook using limited speech-text data. By utilizing a simple codebook lookup process, we convert available text-only data into latent text embeddings. These embeddings then enhance the inputs for the contextualized ASR. Our experiments on diverse Mandarin test sets demonstrate that our TA approach significantly boosts recognition performance. The top-performing system shows relative CER improvements of up to 30% on rare words and 15% across all words in general.
【10】 What Does it Take to Generalize SER Model Across Datasets? A Comprehensive Benchmark
标题: 如何在数据集中推广BER模型?全面的基准
作者:Adham Ibrahim,Shady Shehata,Ajinkya Kulkarni,Mukhtar Mohamed,Muhammad Abdul-Mageed
备注:ACCEPTED AT INTERSPEECH 2024, GREECE
链接:点击下载PDF文件
摘要:语音情感识别(SER)对于增强基于语音的应用中的人机交互至关重要。尽管在特定的情感数据集的改进,仍然有SER的能力,在现实世界的情况下概括的研究差距。在本文中,我们研究的方法来概括SER系统在不同的情感数据集。特别是,我们将11个情感语音数据集,并说明了一个全面的基准SER任务。我们还解决了使用过采样方法组合SER数据集进行训练时数据分布不平衡的挑战。此外,我们探讨了各种评估协议的熟练程度在SER的泛化。在此基础上,我们探讨了潜在的耳语SER,强调全面评估的重要性。我们的方法旨在通过集成扬声器独立的方法来推进SER技术。摘要:Speech emotion recognition (SER) is essential for enhancing human-computer interaction in speech-based applications. Despite improvements in specific emotional datasets, there is still a research gap in SER's capability to generalize across real-world situations. In this paper, we investigate approaches to generalize the SER system across different emotion datasets. In particular, we incorporate 11 emotional speech datasets and illustrate a comprehensive benchmark on the SER task. We also address the challenge of imbalanced data distribution using over-sampling methods when combining SER datasets for training. Furthermore, we explore various evaluation protocols for adeptness in the generalization of SER. Building on this, we explore the potential of Whisper for SER, emphasizing the importance of thorough evaluation. Our approach is designed to advance SER technology by integrating speaker-independent methods.
【11】 Personalized Speech Enhancement Without a Separate Speaker Embedding Model
标题: 无需单独说话者嵌入模型的个性化语音增强
作者:Tanel Pärnamaa,Ando Saabas
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:个性化语音增强(PSE)模型可以通过适应说话人的语音特征来改善电话会议系统的音频质量。然而,大多数现有方法需要单独的说话人嵌入模型来从注册音频中提取说话人的向量表示,这增加了训练和部署过程的复杂性。我们建议使用PSE模型本身的内部表示作为说话人嵌入,从而避免了对单独模型的需要。我们表明,我们的方法在噪声抑制和回声消除任务上与使用预训练的说话人嵌入模型的标准方法一样好或更好。此外,我们的方法在平均意见得分上超过ICASSP 2023深度噪声抑制挑战赛冠军0.15。摘要:Personalized speech enhancement (PSE) models can improve the audio quality of teleconferencing systems by adapting to the characteristics of a speaker's voice. However, most existing methods require a separate speaker embedding model to extract a vector representation of the speaker from enrollment audio, which adds complexity to the training and deployment process. We propose to use the internal representation of the PSE model itself as the speaker embedding, thereby avoiding the need for a separate model. We show that our approach performs equally well or better than the standard method of using a pre-trained speaker embedding model on noise suppression and echo cancellation tasks. Moreover, our approach surpasses the ICASSP 2023 Deep Noise Suppression Challenge winner by 0.15 in Mean Opinion Score.
【12】 MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
标题: MMM:来自自监督学习模型的多层多残留多流离散语音表示
作者:Jiatong Shi,Xutai Ma,Hirofumi Inaguma,Anna Sun,Shinji Watanabe
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
摘要:语音离散表示已被证明是有效的,在各种下游应用,由于其优越的压缩率的波形,在训练过程中的快速收敛,并与其他形式的兼容性。从自监督学习(SSL)模型中提取的离散单元已经成为获得语音离散表示的一种重要方法。然而,虽然离散单元已经显示出与光谱特征相比的有效性,但它们仍然落后于连续SSL表示。在这项工作中,我们提出了MMM,一个多层多残差多流离散单元提取方法从SSL。具体来说,我们介绍了迭代残差矢量量化与K-均值的SSL模型中的不同层提取多流语音离散表示。通过在语音识别,语音再合成,和文本到语音的广泛实验,我们证明了建议的MMM可以超过或等同于神经编解码器的性能在各种条件下。摘要:Speech discrete representation has proven effective in various downstream applications due to its superior compression rate of the waveform, fast convergence during training, and compatibility with other modalities. Discrete units extracted from self-supervised learning (SSL) models have emerged as a prominent approach for obtaining speech discrete representation. However, while discrete units have shown effectiveness compared to spectral features, they still lag behind continuous SSL representations. In this work, we propose MMM, a multi-layer multi-residual multi-stream discrete units extraction method from SSL. Specifically, we introduce iterative residual vector quantization with K-means for different layers in an SSL model to extract multi-stream speech discrete representation. Through extensive experiments in speech recognition, speech resynthesis, and text-to-speech, we demonstrate the proposed MMM can surpass or on-par with neural codec's performance under various conditions.
【13】 Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
标题: Vec-Tok-VC+:双模式训练策略中具有渐进约束的剩余增强稳健Zero-Shot语音转换
作者:Linhan Ma,Xinfa Zhu,Yuanjun Lv,Zhichao Wang,Ziqian Wang,Wendi He,Hongbin Zhou,Lei Xie
备注:Accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:Zero-shot语音转换(Zero Shot Voice Conversion,VC)的目的是在保持语言内容不变的情况下,将源语音转换为任意不可见的目标语音。最近的VC方法已经取得了显着的进步,但在解耦过程中的语义损失以及训练推理不匹配仍然阻碍转换性能。在本文中,我们提出了Vec-Tok-VC+,一种新的基于MPEG-2的zero-shot VC模型的Vec-Tok Codec的改进,实现语音转换,只有一个3s的目标说话人提示。我们设计了一个残差增强的K-Means聚类器,通过两层聚类过程来增强语义内容提取。此外,我们采用教师指导的精化来模拟转换过程,以消除训练-推理失配,形成双模式训练策略。此外,我们设计了一个多码本渐进损失函数来约束模型的逐层输出从粗到细,以提高说话人相似度和内容准确性。客观和主观评价表明,Vec-Tok-VC+在自然度、可懂度和说话人相似度方面优于强基线。摘要:Zero-shot voice conversion (VC) aims to transform source speech into arbitrary unseen target voice while keeping the linguistic content unchanged. Recent VC methods have made significant progress, but semantic losses in the decoupling process as well as training-inference mismatch still hinder conversion performance. In this paper, we propose Vec-Tok-VC+, a novel prompt-based zero-shot VC model improved from Vec-Tok Codec, achieving voice conversion given only a 3s target speaker prompt. We design a residual-enhanced K-Means decoupler to enhance the semantic content extraction with a two-layer clustering process. Besides, we employ teacher-guided refinement to simulate the conversion process to eliminate the training-inference mismatch, forming a dual-mode training strategy. Furthermore, we design a multi-codebook progressive loss function to constrain the layer-wise output of the model from coarse to fine to improve speaker similarity and content accuracy. Objective and subjective evaluations demonstrate that Vec-Tok-VC+ outperforms the strong baselines in naturalness, intelligibility, and speaker similarity.
【14】 SHMamba: Structured Hyperbolic State Space Model for Audio-Visual Question Answering
标题: SHMamba:视听问题回答的结构化双曲状态空间模型
作者:Zhe Yang,Wenrui Li,Guanghui Cheng
链接:点击下载PDF文件
摘要:视听问答(AVQA)任务具有很大的应用潜力。与传统的单峰方法相比,AVQA的多模态输入使得特征提取和融合过程更具挑战性。欧氏空间难以有效地表示数据的多维关系。特别是当提取和处理具有树结构或层次结构的数据时,欧几里得空间不适合作为嵌入空间。此外,Transformers中的自注意机制在捕捉序列中元素之间的动态关系方面是有效的。然而,自注意机制在窗口建模和二次计算复杂性方面的局限性降低了其在长序列建模中的有效性。为了解决这些问题,我们提出了SHMamba:结构化双曲状态空间模型,以整合双曲几何和状态空间模型的优点。具体来说,SHMamba利用双曲空间的内在属性来表示视听数据中的层次结构和复杂关系。同时,状态空间模型通过对整个序列进行全局建模来捕获随时间的动态变化。此外,我们引入了一个自适应曲率双曲线对齐模块和交叉融合块,以提高层次结构的理解和跨模态信息的动态交换,分别。大量的实验表明,SHMamba优于以前的方法,更少的参数和计算成本。我们的可学习参数减少了78.12%,而平均性能提高了2.53%.实验结果表明,该方法在当前主流方法中具有一定的优越性,更适合于实际应用场景。摘要:The Audio-Visual Question Answering (AVQA) task holds significant potential for applications. Compared to traditional unimodal approaches, the multi-modal input of AVQA makes feature extraction and fusion processes more challenging. Euclidean space is difficult to effectively represent multi-dimensional relationships of data. Especially when extracting and processing data with a tree structure or hierarchical structure, Euclidean space is not suitable as an embedding space. Additionally, the self-attention mechanism in Transformers is effective in capturing the dynamic relationships between elements in a sequence. However, the self-attention mechanism's limitations in window modeling and quadratic computational complexity reduce its effectiveness in modeling long sequences. To address these limitations, we propose SHMamba: Structured Hyperbolic State Space Model to integrate the advantages of hyperbolic geometry and state space models. Specifically, SHMamba leverages the intrinsic properties of hyperbolic space to represent hierarchical structures and complex relationships in audio-visual data. Meanwhile, the state space model captures dynamic changes over time by globally modeling the entire sequence. Furthermore, we introduce an adaptive curvature hyperbolic alignment module and a cross fusion block to enhance the understanding of hierarchical structures and the dynamic exchange of cross-modal information, respectively. Extensive experiments demonstrate that SHMamba outperforms previous methods with fewer parameters and computational costs. Our learnable parameters are reduced by 78.12 %, while the average performance improves by 2.53 %. Experiments show that our method demonstrates superiority among all current major methods and is more suitable for practical application scenarios.
【15】 Frequency-mix Knowledge Distillation for Fake Speech Detection
标题: 用于假语音检测的混频知识提炼
作者:Cunhang Fan,Shunbo Dong,Jun Xue,Yujie Chen,Jiangyan Yi,Zhao Lv
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在电话场景中,对抗语音欺骗攻击的虚假语音检测(FSD)任务具有挑战性。数据增强(DA)方法被认为是解决电话场景中FSD任务的有效手段,通常分为时域和频域阶段。虽然每一种都有其优点,但两者都可能导致信息丢失。为了解决这个问题,我们提出了一种新的DA方法,频率混合(Freqmix),并引入Freqmix知识蒸馏(FKD),以提高模型信息提取和泛化能力。具体来说,我们使用Freqmix增强的数据作为教师模型的输入,而学生模型的输入经历时域DA方法。我们使用多层次的特征提取方法来恢复信息,提高模型的泛化能力。我们的方法在ASVspoof 2021 LA数据集上实现了最先进的结果,比基线提高了31%,并且在ASVspoof 2021 DF数据集上具有竞争力。摘要:In the telephony scenarios, the fake speech detection (FSD) task to combat speech spoofing attacks is challenging. Data augmentation (DA) methods are considered effective means to address the FSD task in telephony scenarios, typically divided into time domain and frequency domain stages. While each has its advantages, both can result in information loss. To tackle this issue, we propose a novel DA method, Frequency-mix (Freqmix), and introduce the Freqmix knowledge distillation (FKD) to enhance model information extraction and generalization abilities. Specifically, we use Freqmix-enhanced data as input for the teacher model, while the student model's input undergoes time-domain DA method. We use a multi-level feature distillation approach to restore information and improve the model's generalization capabilities. Our approach achieves state-of-the-art results on ASVspoof 2021 LA dataset, showing a 31 % improvement over baseline and performs competitively on ASVspoof 2021 DF dataset.
【16】 Multi-Modal Retrieval For Large Language Model Based Speech Recognition
标题: 基于大语言模型的语音识别的多模式检索
作者:Jari Kolehmainen,Aditya Gourav,Prashanth Gurunath Shivakumar,Yile Gu,Ankur Gandhe,Ariya Rastrow,Grant Strimel,Ivan Bulyko
链接:点击下载PDF文件
摘要:检索是利用外部信息改进语言模型的一种广泛采用的方法。随着该领域转向多模态大型语言模型,重要的是扩展基于纯文本的方法,以将其他模态纳入检索以及广泛的机器学习任务和数据类型的应用程序。在这项工作中,我们提出了多模态检索两种方法:kNN-LM和交叉注意技术。我们证明了我们的检索方法的有效性经验,将它们应用到自动语音识别任务与外部信息的访问。在这种设置下,我们表明,基于语音的多模态检索优于基于文本的检索,并产生高达50%的改善,在多模态语言模型基线的单词错误率。此外,我们在Spoken-Squad问答数据集上实现了最先进的识别结果。摘要:Retrieval is a widely adopted approach for improving language models leveraging external information. As the field moves towards multi-modal large language models, it is important to extend the pure text based methods to incorporate other modalities in retrieval as well for applications across the wide spectrum of machine learning tasks and data types. In this work, we propose multi-modal retrieval with two approaches: kNN-LM and cross-attention techniques. We demonstrate the effectiveness of our retrieval approaches empirically by applying them to automatic speech recognition tasks with access to external information. Under this setting, we show that speech-based multi-modal retrieval outperforms text based retrieval, and yields up to 50 % improvement in word error rate over the multi-modal language model baseline. Furthermore, we achieve state-of-the-art recognition results on the Spoken-Squad question answering dataset.
【17】 Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
标题: Speech ReaLLM --通过教学时间流,使用多模式LLM进行实时流语音识别
作者:Frank Seide,Morrie Doulaty,Yangyang Shi,Yashesh Gaur,Junteng Jia,Chunyang Wu
链接:点击下载PDF文件
摘要:我们引入了Speech ReaLLM,这是一种新的ASR架构,它将“仅解码器”ASR与RNN-T结合在一起,使多模式LLM架构能够实现实时流传输。这是第一个“仅解码器”ASR架构,旨在处理连续音频,而无需显式端点。Speech ReaLLM是更通用的ReaLLM(“实时LLM”)方法的特殊情况,也是第一次在这里介绍。这个想法受到RNN-T的启发:不是只在用户提示结束时生成响应,而是在实时接收到每个输入令牌后生成(通常是空的)。在Libripeech“测试”中,80 M Speech ReaLLM实时实现3.0%和7.4%的WER(没有外部LM或辅助损失)。这仅略高于3倍大的Attention-Encoder-Decoder基线。我们还表明,通过这种方式,LLM架构可以学习表示和再现时间流;并且可以对预训练的7 B LLM进行微调,以便在此任务上做得相当好。摘要:We introduce Speech ReaLLM, a new ASR architecture that marries "decoder-only" ASR with the RNN-T to make multimodal LLM architectures capable of real-time streaming. This is the first "decoder-only" ASR architecture designed to handle continuous audio without explicit end-pointing. Speech ReaLLM is a special case of the more general ReaLLM ("real-time LLM") approach, also introduced here for the first time. The idea is inspired by RNN-T: Instead of generating a response only at the end of a user prompt, generate after every input token received in real time (it is often empty). On Librispeech "test", an 80M Speech ReaLLM achieves WERs of 3.0% and 7.4% in real time (without an external LM or auxiliary loss). This is only slightly above a 3x larger Attention-Encoder-Decoder baseline. We also show that this way, an LLM architecture can learn to represent and reproduce the flow of time; and that a pre-trained 7B LLM can be fine-tuned to do reasonably well on this task.
【18】 Analyzing phonetic structure of Mandarin using Audacity
标题: 用Audacity分析普通话的语音结构
作者:Shizheng Xu
备注:audio source: this https URL
链接:点击下载PDF文件
摘要:普通话是中国、台湾和新加坡的官方语言。它也是主要的非官方语言,主要在多伦多和温哥华的家中使用。本文运用音频软件Audacity,运用理论知识对汉语普通话进行了全面的分析。本研究首先概述了普通话语音的基本原则,旨在提供对普通话语音结构的见解。摘要:Mandarin Chinese is the official language in China, Taiwan, and Singapore. It is also the main non-official language spoken predominantly at home in Toronto and Vancouver. This article employs the audio software Audacity and leverages theoretical knowledge to conduct a comprehensive analysis of Mandarin Chinese. The study initiates with an overview of the fundamental principles underlying Mandarin pronunciation, aiming to provide insights into its phonetic structure.
【19】 Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
标题: Whisper-Flamingo:将视觉特征集成到Whisper中以实现视听语音识别和翻译
作者:Andrew Rouditchenko,Yuan Gong,Samuel Thomas,Leonid Karlinsky,Hilde Kuehne,Rogerio Feris,James Glass
备注:Interspeech 2024. Code this https URL
链接:点击下载PDF文件
摘要:视听语音识别(AVSR)使用基于嘴唇的视频来提高噪声中的性能。由于视频比音频更难获得,AVSR模型的视频训练数据通常限于几千小时。相比之下,像Whisper这样的语音模型是用数十万小时的数据训练的,因此可以学习更好的语音到文本解码器。巨大的训练数据差异促使我们调整Whisper来处理视频输入。受Flamingo将视觉特征注入语言模型的启发,我们提出了Whisper-Flamingo,它将视觉特征集成到具有门控交叉注意的Whisper语音识别和翻译模型。我们的视听Whisper-Flamingo在嘈杂环境下的英语语音识别和6种语言的En-X翻译方面优于仅音频Whisper。此外,Whisper-Flamingo是一个多功能模型,使用一组参数执行所有这些任务,而之前的方法是在每种语言上单独训练的。摘要:Audio-Visual Speech Recognition (AVSR) uses lip-based video to improve performance in noise. Since videos are harder to obtain than audio, the video training data of AVSR models is usually limited to a few thousand hours. In contrast, speech models such as Whisper are trained with hundreds of thousands of hours of data, and thus learn a better speech-to-text decoder. The huge training data difference motivates us to adapt Whisper to handle video inputs. Inspired by Flamingo which injects visual features into language models, we propose Whisper-Flamingo which integrates visual features into the Whisper speech recognition and translation model with gated cross attention. Our audio-visual Whisper-Flamingo outperforms audio-only Whisper on English speech recognition and En-X translation for 6 languages in noisy conditions. Moreover, Whisper-Flamingo is a versatile model and conducts all of these tasks using one set of parameters, while prior methods are trained separately on each language.
【20】 Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
标题: 检测法国电视和广播内容中言语互动的语音转向边界的终点
作者:Rémi Uro,Marie Tahon,David Doukhan,Antoine Laurent,Albert Rilliard
备注:keywords : Spoken interaction, Media, TV, Radio, Transition-Relevance Places, Turn Taking, Interruption. Accepted to InterSpeech 2024, Kos Island, Greece
链接:点击下载PDF文件
摘要:过渡相关性地点被定义为话语的结尾,其中对话者可以在不打断当前说话者的情况下发言--即,转弯是终点的地方。分析话轮终结性有助于研究会话中话轮转换的动态性。本文提出了一种在多说话人环境中自动将语音分类为终结语和非终结语的方法。我们比较了音频,文本,和融合的两种方法在法国语料库的电视和广播提取物在每个扬声器的变化与回合终端信息注释。我们的模型基于预先训练的自监督表示。我们报告不同的融合策略和不同的上下文大小的结果。本研究还通过分析随机初始化的多次训练运行结果的差异来质疑性能变异性的问题。测量的准确性将允许使用这些模型进行大规模的话轮转换分析。摘要:Transition Relevance Places are defined as the end of an utterance where the interlocutor may take the floor without interrupting the current speaker --i.e., a place where the turn is terminal. Analyzing turn terminality is useful to study the dynamic of turn-taking in spontaneous conversations. This paper presents an automatic classification of spoken utterances as Terminal or Non-Terminal in multi-speaker settings. We compared audio, text, and fusions of both approaches on a French corpus of TV and Radio extracts annotated with turn-terminality information at each speaker change. Our models are based on pre-trained self-supervised representations. We report results for different fusion strategies and varying context sizes. This study also questions the problem of performance variability by analyzing the differences in results for multiple training runs with random initialization. The measured accuracy would allow the use of these models for large-scale analysis of turn-taking.
【21】 Understanding Pedestrian Movement Using Urban Sensing Technologies: The Promise of Audio-based Sensors
标题: 使用城市传感技术了解行人运动:基于音频的传感器的前景
作者:Chaeyeon Han,Pavan Seshadri,Yiwei Ding,Noah Posner,Bon Woo Koo,Animesh Agrawal,Alexander Lerch,Subhrajit Guhathakurta
备注:submitted to Urban Informatics
链接:点击下载PDF文件
摘要:虽然已经部署了各种传感器来监测车辆流量,但感知行人运动仍处于起步阶段。然而,在许多城市,特别是欧洲、非洲和亚洲的城市,步行是一种重要的旅行方式。了解行人数量和流量对于设计更安全和更具吸引力的行人基础设施以及控制周期性过度拥挤至关重要。这项研究讨论了一种新的方法,以扩大城市感知的人的帮助下,新的音频为基础的技术。它评估了与其他形式的行人感应相比,基于麦克风的传感器的优点和局限性。介绍了一个名为ASPED的大规模数据集,其中包括用于标记行人计数数据的高质量音频记录和视频记录。基线分析强调了使用音频传感器进行行人跟踪的前景,尽管算法和技术改进使传感器实际可用仍在继续。这项研究还展示了如何利用这些数据来预测行人轨迹。最后,它讨论了基于音频的行人感知可以支持更好的城市和交通规划的用例和场景。摘要:While various sensors have been deployed to monitor vehicular flows, sensing pedestrian movement is still nascent. Yet walking is a significant mode of travel in many cities, especially those in Europe, Africa, and Asia. Understanding pedestrian volumes and flows is essential for designing safer and more attractive pedestrian infrastructure and for controlling periodic overcrowding. This study discusses a new approach to scale up urban sensing of people with the help of novel audio-based technology. It assesses the benefits and limitations of microphone-based sensors as compared to other forms of pedestrian sensing. A large-scale dataset called ASPED is presented, which includes high-quality audio recordings along with video recordings used for labeling the pedestrian count data. The baseline analyses highlight the promise of using audio sensors for pedestrian tracking, although algorithmic and technological improvements to make the sensors practically usable continue. This study also demonstrates how the data can be leveraged to predict pedestrian trajectories. Finally, it discusses the use cases and scenarios where audio-based pedestrian sensing can support better urban and transportation planning.
【22】 Period Singer: Integrating Periodic and Aperiodic Variational Autoencoders for Natural-Sounding End-to-End Singing Voice Synthesis
标题: Period Singer:集成周期性和非周期性变分自动编码器,实现自然声音端到端歌唱声音合成
作者:Taewoo Kim,Choongsang Cho,Young Han Lee
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了周期歌手,一种新的端到端的歌唱声音合成(SVS)模型,利用变分推理的周期性和非周期性成分,旨在产生自然的声音波形。最近的端到端SVS模型已经证明了合成高保真歌唱声音的能力。然而,由于确定性的音高调节,它们不能完全解决一对多的问题。为了解决这个问题,我们提出了周期歌手架构,它集成了变分自动编码器的周期性和非周期性组件。此外,我们的方法消除了依赖于外部对准估计的音素对齐通过单调对齐搜索内注意到边界。我们的实证评估表明,Period Singer在普通话和韩语数据集上的性能优于现有的端到端SVS模型。消融研究进一步证实了所提出的方法的有效性。摘要:In this paper, we present Period Singer, a novel end-to-end singing voice synthesis (SVS) model that utilizes variational inference for periodic and aperiodic components, aimed at producing natural-sounding waveforms. Recent end-to-end SVS models have demonstrated the capability of synthesizing high-fidelity singing voices. However, owing to deterministic pitch conditioning, they do not fully address the one-to-many problem. To address this problem, we present the Period Singer architecture, which integrates variational autoencoders for the periodic and aperiodic components. Additionally, our methodology eliminates the dependency on an external aligner by estimating the phoneme alignment through a monotonic alignment search within note boundaries. Our empirical evaluations show that Period Singer outperforms existing end-to-end SVS models on Mandarin and Korean datasets. The efficacy of the proposed method was further corroborated by ablation studies.
【23】 Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
标题: 感知器提示:Whisper中灵活的说话人适应中文无序语音识别
作者:Yicong Jiang,Tianzi Wang,Xurong Xie,Juan Liu,Wei Sun,Nan Yan,Hui Chen,Lan Wang,Xunying Liu,Feng Tian
备注:Accepted by interspeech 2024
链接:点击下载PDF文件
摘要:语音识别障碍对改善构音障碍患者的生活质量具有深远的意义。构音障碍的语音识别遇到的挑战,包括有限的数据,构音障碍和非构音障碍的扬声器之间的实质性差异,以及显着的扬声器的变化源于障碍。本文介绍了Perceiver-Prompt,一种在Whisper大规模模型上利用P-Tuning的说话人自适应方法。我们首先使用LoRA微调Whisper,然后集成可训练的Perceiver,从可变长度的输入中生成固定长度的说话人提示,以提高汉语构音障碍语音的模型识别。从我们的中国构音障碍语音数据集的实验结果表明,识别性能与Perceiver提示一致的改善。在微调Whisper上获得了CER的相对降低高达13.04%。摘要:Disordered speech recognition profound implications for improving the quality of life for individuals afflicted with, for example, dysarthria. Dysarthric speech recognition encounters challenges including limited data, substantial dissimilarities between dysarthric and non-dysarthric speakers, and significant speaker variations stemming from the disorder. This paper introduces Perceiver-Prompt, a method for speaker adaptation that utilizes P-Tuning on the Whisper large-scale model. We first fine-tune Whisper using LoRA and then integrate a trainable Perceiver to generate fixed-length speaker prompts from variable-length inputs, to improve model recognition of Chinese dysarthric speech. Experimental results from our Chinese dysarthric speech dataset demonstrate consistent improvements in recognition performance with Perceiver-Prompt. Relative reduction up to 13.04% in CER is obtained over the fine-tuned Whisper.
【24】 GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
标题: GenDistiller:基于自回归生成模型提取预训练语言模型
作者:Yingying Gao,Shilei Zhang,Chao Deng,Junlan Feng
备注:arXiv admin note: text overlap with arXiv:2310.13418
链接:点击下载PDF文件
摘要:预训练的语音语言模型(如HuBERT和WavLM)利用未标记的语音数据进行自我监督学习,并为许多下游任务提供强大的表示。尽管这些模型取得了成功,但它们对存储器和计算资源的高要求阻碍了它们在资源受限设备上的应用。因此,本文介绍了GenDistiller,一种新的知识蒸馏框架,它通过一个小得多的学生网络直接生成预训练教师模型的隐藏表示。该方法以前一个隐含层为历史,实现教师模型的逐层自回归预测。在SUPERB上的实验表明,与没有自回归框架的基线提取方法相比,GenDistiller具有33%的参数减少,相似的时间消耗和更好的性能。最终,GenDistiller将WavLM的大小减少了82%。摘要:Pre-trained speech language models such as HuBERT and WavLM leverage unlabeled speech data for self-supervised learning and offer powerful representations for numerous downstream tasks. Despite the success of these models, their high requirements for memory and computing resource hinder their application on resource restricted devices. Therefore, this paper introduces GenDistiller, a novel knowledge distillation framework which generates the hidden representations of the pre-trained teacher model directly by a much smaller student network. The proposed method takes the previous hidden layer as history and implements a layer-by-layer prediction of the teacher model autoregressively. Experiments on SUPERB reveal the advantage of GenDistiller over the baseline distilling method without an autoregressive framework, with 33% fewer parameters, similar time consumption and better performance on most of the SUPERB tasks. Ultimately, the proposed GenDistiller reduces the size of WavLM by 82%.
eeess.AS音频处理
【1】 AlignNet: Learning dataset score alignment functions to enable better training of speech quality estimators
作者:Jaden Pieper,Stephen Voran
备注:To be published in proc. of Interspeech 2024. 5 pages, 2 figures, 3 tables
链接:点击下载PDF文件
摘要:我们开发了两个互补的进步训练无参考(NR)语音质量估计独立的数据集。多数据集微调(Multi-dataset finetuning,缩写为DFI)在单个数据集上预训练NR估计器,然后同时在多个数据集上对其进行微调,包括用于预训练的数据集。AlignNet使用AudioNet生成中间分数估计值,然后使用Aligner将中间估计值映射到适当的分数范围。AlignNet对AudioNet的选择是不可知的,因此任何成功的NR语音质量估计器都可以从其Aligner中受益。这些方法可以串联使用,我们使用两项研究来证明它们改进了当前的解决方案:一项研究使用了九个较小的数据集,另一项研究使用了四个较大的数据集。AlignNet与其他解决方案相比有所改进,因为它有效地消除了影响学习过程的不一致,从而能够使用更大量更多样化的数据进行成功的训练。摘要:We develop two complementary advances for training no-reference (NR) speech quality estimators with independent datasets. Multi-dataset finetuning (MDF) pretrains an NR estimator on a single dataset and then finetunes it on multiple datasets at once, including the dataset used for pretraining. AlignNet uses an AudioNet to generate intermediate score estimates before using the Aligner to map intermediate estimates to the appropriate score range. AlignNet is agnostic to the choice of AudioNet so any successful NR speech quality estimator can benefit from its Aligner. The methods can be used in tandem, and we use two studies to show that they improve on current solutions: one study uses nine smaller datasets and the other uses four larger datasets. AlignNet with MDF improves on other solutions because it efficiently and effectively removes misalignments that impair the learning process, and thus enables successful training with larger amounts of more diverse data.
【2】 Inclusive ASR for Disfluent Speech: Cascaded Large-Scale Self-Supervised Learning with Targeted Fine-Tuning and Data Augmentation
标题: 针对不流利言语的包容性ASB:具有有针对性的微调和数据增强的级联大规模自我监督学习
作者:Dena Mujtaba,Nihar R. Mahapatra,Megan Arney,J. Scott Yaruss,Caryn Herring,Jia Bin
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:自动语音识别(ASR)系统在处理与口吃相关的不流利(如无意识的块和单词重复)时经常会出现问题,从而产生不准确的转录。进展的一个关键障碍是缺乏大型的、带注释的不流利语音数据集。因此,我们提出了一种包容性的ASR设计方法,利用对标准语音的大规模自我监督学习,然后在较小的非流利语音数据集上进行有针对性的微调和数据增强。我们的数据增强技术丰富了具有各种不流利性的训练数据集,增强了这些语音模式的ASR处理。结果表明,即使是相对较小的标记数据集,以及数据增强,微调wav2vec 2.0也可以显着降低不流利语音的单词错误率。我们的方法不仅提高了口吃者的ASR包容性,而且为能够适应更广泛的语音变化的ASR铺平了道路。摘要:Automatic speech recognition (ASR) systems often falter while processing stuttering-related disfluencies -- such as involuntary blocks and word repetitions -- yielding inaccurate transcripts. A critical barrier to progress is the scarcity of large, annotated disfluent speech datasets. Therefore, we present an inclusive ASR design approach, leveraging large-scale self-supervised learning on standard speech followed by targeted fine-tuning and data augmentation on a smaller, curated dataset of disfluent speech. Our data augmentation technique enriches training datasets with various disfluencies, enhancing ASR processing of these speech patterns. Results show that fine-tuning wav2vec 2.0 with even a relatively small, labeled dataset, alongside data augmentation, can significantly reduce word error rates for disfluent speech. Our approach not only advances ASR inclusivity for people who stutter, but also paves the way for ASRs that can accommodate wider speech variations.
【3】 Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
标题: Whisper-Flamingo:将视觉特征集成到Whisper中以实现视听语音识别和翻译
作者:Andrew Rouditchenko,Yuan Gong,Samuel Thomas,Leonid Karlinsky,Hilde Kuehne,Rogerio Feris,James Glass
备注:Interspeech 2024. Code this https URL
链接:点击下载PDF文件
摘要:视听语音识别(AVSR)使用基于嘴唇的视频来提高噪声中的性能。由于视频比音频更难获得,AVSR模型的视频训练数据通常限于几千小时。相比之下,像Whisper这样的语音模型是用数十万小时的数据训练的,因此可以学习更好的语音到文本解码器。巨大的训练数据差异促使我们调整Whisper来处理视频输入。受Flamingo将视觉特征注入语言模型的启发,我们提出了Whisper-Flamingo,它将视觉特征集成到具有门控交叉注意的Whisper语音识别和翻译模型。我们的视听Whisper-Flamingo在嘈杂环境下的英语语音识别和6种语言的En-X翻译方面优于仅音频Whisper。此外,Whisper-Flamingo是一个多功能模型,使用一组参数执行所有这些任务,而之前的方法是在每种语言上单独训练的。摘要:Audio-Visual Speech Recognition (AVSR) uses lip-based video to improve performance in noise. Since videos are harder to obtain than audio, the video training data of AVSR models is usually limited to a few thousand hours. In contrast, speech models such as Whisper are trained with hundreds of thousands of hours of data, and thus learn a better speech-to-text decoder. The huge training data difference motivates us to adapt Whisper to handle video inputs. Inspired by Flamingo which injects visual features into language models, we propose Whisper-Flamingo which integrates visual features into the Whisper speech recognition and translation model with gated cross attention. Our audio-visual Whisper-Flamingo outperforms audio-only Whisper on English speech recognition and En-X translation for 6 languages in noisy conditions. Moreover, Whisper-Flamingo is a versatile model and conducts all of these tasks using one set of parameters, while prior methods are trained separately on each language.
【4】 Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
标题: 检测法国电视和广播内容中言语互动的语音转向边界的终点
作者:Rémi Uro,Marie Tahon,David Doukhan,Antoine Laurent,Albert Rilliard
备注:keywords : Spoken interaction, Media, TV, Radio, Transition-Relevance Places, Turn Taking, Interruption. Accepted to InterSpeech 2024, Kos Island, Greece
链接:点击下载PDF文件
摘要:过渡相关性地点被定义为话语的结尾,其中对话者可以在不打断当前说话者的情况下发言--即,转弯是终点的地方。分析话轮终结性有助于研究会话中话轮转换的动态性。本文提出了一种在多说话人环境中自动将语音分类为终结语和非终结语的方法。我们比较了音频,文本,和融合的两种方法在法国语料库的电视和广播提取物在每个扬声器的变化与回合终端信息注释。我们的模型基于预先训练的自监督表示。我们报告不同的融合策略和不同的上下文大小的结果。本研究还通过分析随机初始化的多次训练运行结果的差异来质疑性能变异性的问题。测量的准确性将允许使用这些模型进行大规模的话轮转换分析。摘要:Transition Relevance Places are defined as the end of an utterance where the interlocutor may take the floor without interrupting the current speaker --i.e., a place where the turn is terminal. Analyzing turn terminality is useful to study the dynamic of turn-taking in spontaneous conversations. This paper presents an automatic classification of spoken utterances as Terminal or Non-Terminal in multi-speaker settings. We compared audio, text, and fusions of both approaches on a French corpus of TV and Radio extracts annotated with turn-terminality information at each speaker change. Our models are based on pre-trained self-supervised representations. We report results for different fusion strategies and varying context sizes. This study also questions the problem of performance variability by analyzing the differences in results for multiple training runs with random initialization. The measured accuracy would allow the use of these models for large-scale analysis of turn-taking.
【5】 ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2Vec2.0 Based ASR
标题: ROAR:增强基于Wav2Vec2.0的ASB的原始数据与增强数据比率动态
作者:Vishwanath Pratap Singh,Federico Malato,Ville Hautamaki,Md. Sahidullah,Tomi Kinnunen
Journal-ref:Interspeech 2024
链接:点击下载PDF文件
摘要:虽然自动语音识别(ASR)极大地受益于数据增强,但增强配方本身往往是启发式的。在本文中,我们通过引入基于强化学习(RL)的原始增强数据比(OAR)的动态调整,解决了与ASR训练中平衡适量增强数据相关的启发式方法之一。与传统数据增强中的固定OAR方法不同,我们提出的方法采用深度Q网络(DQN)作为RL机制,在基于wav2vec2.0的ASR训练中学习OAR的最佳动态。我们使用LibriSpeech数据集进行实验,训练数据量不同,具体来说,10Min,1H,10H和100H分裂,以评估所提出的方法在不同数据条件下的有效性。我们提出的方法,平均而言,实现了4.96%的相对改善,比开源wav2vec2.0基础模型的标准LibriSpeech测试集。摘要:While automatic speech recognition (ASR) greatly benefits from data augmentation, the augmentation recipes themselves tend to be heuristic. In this paper, we address one of the heuristic approach associated with balancing the right amount of augmented data in ASR training by introducing a reinforcement learning (RL) based dynamic adjustment of original-to-augmented data ratio (OAR). Unlike the fixed OAR approach in conventional data augmentation, our proposed method employs a deep Q-network (DQN) as the RL mechanism to learn the optimal dynamics of OAR throughout the wav2vec2.0 based ASR training. We conduct experiments using the LibriSpeech dataset with varying amounts of training data, specifically, the 10Min, 1H, 10H, and 100H splits to evaluate the efficacy of the proposed method under different data conditions. Our proposed method, on average, achieves a relative improvement of 4.96% over the open-source wav2vec2.0 base model on standard LibriSpeech test sets.
【6】 Understanding Pedestrian Movement Using Urban Sensing Technologies: The Promise of Audio-based Sensors
标题: 使用城市传感技术了解行人运动:基于音频的传感器的前景
作者:Chaeyeon Han,Pavan Seshadri,Yiwei Ding,Noah Posner,Bon Woo Koo,Animesh Agrawal,Alexander Lerch,Subhrajit Guhathakurta
备注:submitted to Urban Informatics
链接:点击下载PDF文件
摘要:虽然已经部署了各种传感器来监测车辆流量,但感知行人运动仍处于起步阶段。然而,在许多城市,特别是欧洲、非洲和亚洲的城市,步行是一种重要的旅行方式。了解行人数量和流量对于设计更安全和更具吸引力的行人基础设施以及控制周期性过度拥挤至关重要。这项研究讨论了一种新的方法,以扩大城市感知的人的帮助下,新的音频为基础的技术。它评估了与其他形式的行人感应相比,基于麦克风的传感器的优点和局限性。介绍了一个名为ASPED的大规模数据集,其中包括用于标记行人计数数据的高质量音频记录和视频记录。基线分析强调了使用音频传感器进行行人跟踪的前景,尽管算法和技术改进使传感器实际可用仍在继续。这项研究还展示了如何利用这些数据来预测行人轨迹。最后,它讨论了基于音频的行人感知可以支持更好的城市和交通规划的用例和场景。摘要:While various sensors have been deployed to monitor vehicular flows, sensing pedestrian movement is still nascent. Yet walking is a significant mode of travel in many cities, especially those in Europe, Africa, and Asia. Understanding pedestrian volumes and flows is essential for designing safer and more attractive pedestrian infrastructure and for controlling periodic overcrowding. This study discusses a new approach to scale up urban sensing of people with the help of novel audio-based technology. It assesses the benefits and limitations of microphone-based sensors as compared to other forms of pedestrian sensing. A large-scale dataset called ASPED is presented, which includes high-quality audio recordings along with video recordings used for labeling the pedestrian count data. The baseline analyses highlight the promise of using audio sensors for pedestrian tracking, although algorithmic and technological improvements to make the sensors practically usable continue. This study also demonstrates how the data can be leveraged to predict pedestrian trajectories. Finally, it discusses the use cases and scenarios where audio-based pedestrian sensing can support better urban and transportation planning.
【7】 Period Singer: Integrating Periodic and Aperiodic Variational Autoencoders for Natural-Sounding End-to-End Singing Voice Synthesis
标题: Period Singer:集成周期性和非周期性变分自动编码器,实现自然声音端到端歌唱声音合成
作者:Taewoo Kim,Choongsang Cho,Young Han Lee
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了周期歌手,一种新的端到端的歌唱声音合成(SVS)模型,利用变分推理的周期性和非周期性成分,旨在产生自然的声音波形。最近的端到端SVS模型已经证明了合成高保真歌唱声音的能力。然而,由于确定性的音高调节,它们不能完全解决一对多的问题。为了解决这个问题,我们提出了周期歌手架构,它集成了变分自动编码器的周期性和非周期性组件。此外,我们的方法消除了依赖于外部对准估计的音素对齐通过单调对齐搜索内注意到边界。我们的实证评估表明,Period Singer在普通话和韩语数据集上的性能优于现有的端到端SVS模型。消融研究进一步证实了所提出的方法的有效性。摘要:In this paper, we present Period Singer, a novel end-to-end singing voice synthesis (SVS) model that utilizes variational inference for periodic and aperiodic components, aimed at producing natural-sounding waveforms. Recent end-to-end SVS models have demonstrated the capability of synthesizing high-fidelity singing voices. However, owing to deterministic pitch conditioning, they do not fully address the one-to-many problem. To address this problem, we present the Period Singer architecture, which integrates variational autoencoders for the periodic and aperiodic components. Additionally, our methodology eliminates the dependency on an external aligner by estimating the phoneme alignment through a monotonic alignment search within note boundaries. Our empirical evaluations show that Period Singer outperforms existing end-to-end SVS models on Mandarin and Korean datasets. The efficacy of the proposed method was further corroborated by ablation studies.
【8】 Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
标题: 感知器提示:Whisper中灵活的说话人适应中文无序语音识别
作者:Yicong Jiang,Tianzi Wang,Xurong Xie,Juan Liu,Wei Sun,Nan Yan,Hui Chen,Lan Wang,Xunying Liu,Feng Tian
备注:Accepted by interspeech 2024
链接:点击下载PDF文件
摘要:语音识别障碍对改善构音障碍患者的生活质量具有深远的意义。构音障碍的语音识别遇到的挑战,包括有限的数据,构音障碍和非构音障碍的扬声器之间的实质性差异,以及显着的扬声器的变化源于障碍。本文介绍了Perceiver-Prompt,一种在Whisper大规模模型上利用P-Tuning的说话人自适应方法。我们首先使用LoRA微调Whisper,然后集成可训练的Perceiver,从可变长度的输入中生成固定长度的说话人提示,以提高汉语构音障碍语音的模型识别。从我们的中国构音障碍语音数据集的实验结果表明,识别性能与Perceiver提示一致的改善。在微调Whisper上获得了CER的相对降低高达13.04%。摘要:Disordered speech recognition profound implications for improving the quality of life for individuals afflicted with, for example, dysarthria. Dysarthric speech recognition encounters challenges including limited data, substantial dissimilarities between dysarthric and non-dysarthric speakers, and significant speaker variations stemming from the disorder. This paper introduces Perceiver-Prompt, a method for speaker adaptation that utilizes P-Tuning on the Whisper large-scale model. We first fine-tune Whisper using LoRA and then integrate a trainable Perceiver to generate fixed-length speaker prompts from variable-length inputs, to improve model recognition of Chinese dysarthric speech. Experimental results from our Chinese dysarthric speech dataset demonstrate consistent improvements in recognition performance with Perceiver-Prompt. Relative reduction up to 13.04% in CER is obtained over the fine-tuned Whisper.
【9】 Low algorithmic delay implementation of convolutional beamformer for online joint source separation and dereverberation
标题: 用于在线联合源分离和去回响的卷积束形成器的低算法延迟实现
作者:Kaien Mo,Xianrui Wang,Yichen Yang,Shoji Makino,Jingdong Chen
备注:4 pages, 4 figures. Accepted by EUSIPCO 2024
链接:点击下载PDF文件
摘要:盲音频源分离(BASS)技术,特别是那些具有低延迟的技术,在广泛的实时系统中起着重要作用,例如,助听器、车内免提语音通信、实时人机交互等。大多数现有的BASS算法被推导为运行在批处理模式下,因此不可避免地存在较大的延迟。最近,开发了一些在线算法,其在短时傅里叶变换(STFT)域中逐帧地实现分离,并且与那些批处理方法相比,延迟显著降低。然而,这些算法的延迟对于许多实时系统来说可能仍然太长。为了进一步减少延迟,同时实现良好的分离性能,我们建议在这项工作中将加权预测误差(WPE)模块集成到基于非因果样本截断的独立向量分析(NST-IVA)中。仿真结果表明,在适当控制WPE时延的情况下,该算法可以保持NST-IVA算法时延,同时获得更好的性能。摘要:Blind-audio-source-separation (BASS) techniques, particularly those with low latency, play an important role in a wide range of real-time systems, e.g., hearing aids, in-car hand-free voice communication, real-time human-machine interaction, etc. Most existing BASS algorithms are deduced to run on batch mode, and therefore large latency is unavoidable. Recently, some online algorithms were developed, which achieve separation on a frame-by-frame basis in the short-time-Fourier-transform (STFT) domain and the latency is significantly reduced as compared to those batch methods. However, the latency with these algorithms may still be too long for many real-time systems to bear. To further reduce latency while achieving good separation performance, we propose in this work to integrate a weighted prediction error (WPE) module into a non-causal sample-truncating-based independent vector analysis (NST-IVA). The resulting algorithm can maintain the algorithmic delay as NST-IVA if the delay with WPE is appropriately controlled while achieving significantly better performance, which is validated by simulations.
【10】 Enhanced Deep Speech Separation in Clustered Ad Hoc Distributed Microphone Environments
标题: 小型Ad Hoc分布式麦克风环境中的增强深度语音分离
作者:Jihyun Kim,Stijn Kindt,Nilesh Madhu,Hong-Goo Kang
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:由于麦克风的位置和数量不可预测,因此Ad-hoc分布式麦克风环境对传统的深度学习模型提出了挑战,因为传统的深度学习模型通常需要固定的架构。为了定制深度学习模型以适应任意阵列配置,之前引入了变换平均级联(TAC)层。在这项工作中,我们将TAC层与双路Transformers的语音分离,从两个同时在现实的设置。然而,分布式特性使得很难有效地跨麦克风融合信息。因此,我们探讨了在增强之前,围绕感兴趣的源盲聚类麦克风的功效。实验结果表明,这种深度集群通知的方法显着提高了系统的能力,以应付在ad-hoc分布式麦克风环境中观察到的固有的可变性。摘要:Ad-hoc distributed microphone environments, where microphone locations and numbers are unpredictable, present a challenge to traditional deep learning models, which typically require fixed architectures. To tailor deep learning models to accommodate arbitrary array configurations, the Transform-Average-Concatenate (TAC) layer was previously introduced. In this work, we integrate TAC layers with dual-path transformers for speech separation from two simultaneous talkers in realistic settings. However, the distributed nature makes it hard to fuse information across microphones efficiently. Therefore, we explore the efficacy of blindly clustering microphones around sources of interest prior to enhancement. Experimental results show that this deep cluster-informed approach significantly improves the system's capacity to cope with the inherent variability observed in ad-hoc distributed microphone environments.
【11】 A Multimodal Framework for the Assessment of the Schizophrenia Spectrum
标题: 精神分裂症谱系评估的多模式框架
作者:Gowtham Premananth,Yashish M. Siriwardena,Philip Resnik,Sonia Bansal,Deanna L. Kelly,Carol Espy-Wilson
备注:Accepted to be presented at Interspeech 2024
链接:点击下载PDF文件
摘要:本文提出了一种新的多模态框架来区分不同的症状类别的受试者在精神分裂症谱和健康对照使用音频,视频和文本形式。我们实现了基于卷积神经网络和长短期记忆的单峰模型,并对各种多模态融合方法进行了实验,以提出所提出的框架。我们利用最小门控多模态单元(mGMU)来获得从输入模态提取的特征的双峰中间融合,然后最终融合双峰融合的输出以执行主题分类。在多模态框架中使用mGMU单位改善了加权f1评分和加权AUC-ROC评分的性能。摘要:This paper presents a novel multimodal framework to distinguish between different symptom classes of subjects in the schizophrenia spectrum and healthy controls using audio, video, and text modalities. We implemented Convolution Neural Network and Long Short Term Memory based unimodal models and experimented on various multimodal fusion approaches to come up with the proposed framework. We utilized a minimal Gated multimodal unit (mGMU) to obtain a bi-modal intermediate fusion of the features extracted from the input modalities before finally fusing the outputs of the bimodal fusions to perform subject-wise classifications. The use of mGMU units in the multimodal framework improved the performance in both weighted f1-score and weighted AUC-ROC scores.
【12】 Optimizing Byte-level Representation for End-to-end ASR
标题: 优化端到端ASB的字节级表示
作者:Roger Hsiao,Liuhui Deng,Erik McDermott,Ruchir Travadi,Xiaodan Zhuang
备注:5 pages, 1 figure
链接:点击下载PDF文件
摘要:我们提出了一种新的方法来优化字节级表示的端到端的自动语音识别(ASR)。当所支持的语言的字符集很大时,字节级表示通常被大规模的多语言ASR系统所使用。字节级表示的紧凑性和通用性允许ASR模型使用更小的输出词汇表,从而提供更大的灵活性。UTF-8是多语言ASR常用的字节级表示,但它并不是直接用于优化机器学习任务。通过使用自动编码器和矢量量化,我们表明,我们可以优化ASR的字节级表示,并实现更好的准确性。我们提出的框架可以将来自不同模态的信息,并提供了一个纠错机制。在一个英语 汉语听写任务中,我们表明,用这种方法建立的双语ASR模型可以超过UTF-8表示5%的相对错误率。摘要:We propose a novel approach to optimizing a byte-level representation for end-to-end automatic speech recognition (ASR). Byte-level representation is often used by large scale multilingual ASR systems when the character set of the supported languages is large. The compactness and universality of byte-level representation allow the ASR models to use smaller output vocabularies and therefore, provide more flexibility. UTF-8 is a commonly used byte-level representation for multilingual ASR, but it is not designed to optimize machine learning tasks directly. By using auto-encoder and vector quantization, we show that we can optimize a byte-level representation for ASR and achieve better accuracy. Our proposed framework can incorporate information from different modalities, and provides an error correction mechanism. In an English Mandarin dictation task, we show that a bilingual ASR model built with this approach can outperform UTF-8 representation by 5% relative in error rate.
【13】 Efficient Personalization of Amplification in Hearing Aids via Multi-band Bayesian Machine Learning
标题: 通过多频段Bayesian机器学习实现助听器放大的有效个性化
作者:Aoxin Ni,Edward Lobarinas,Nasser Kehtarnavaz
链接:点击下载PDF文件
摘要:在之前的研究中,助听器放大功能的个性化已被证明对助听器用户有益。文献中介绍了几种基于机器学习的个性化方法。本文提出了一种机器学习个性化方法,其优点是基于配对比较的训练效率高,使其具有实用性和现场可部署性。这种方法的训练效率是相互独立地处理频带并在所有频带的每个频带中同时执行贝叶斯机器学习的结果。仿真结果表明,这种方法导致估计的听力偏好函数接近真实的听力偏好函数在较少数量的配对比较相对于以前的机器学习方法。此外,对八名听力受损受试者进行的临床实验表明,这种训练有效的个性化方法提供了个性化的增益设置,其平均比标准规定的增益设置优选六倍。摘要:Personalization of the amplification function of hearing aids has been shown to be of benefit to hearing aid users in previous studies. Several machine learning-based personalization approaches have been introduced in the literature. This paper presents a machine learning personalization approach with the advantage of being efficient in its training based on paired comparisons which makes it practical and field deployable. The training efficiency of this approach is the result of treating frequency bands independent of one another and by simultaneously carrying out Bayesian machine learning in each band across all of the frequency bands. Simulation results indicate that this approach leads to an estimated hearing preference function close to the true hearing preference function in fewer number of paired comparisons relative to the previous machine learning approaches. In addition, a clinical experiment conducted on eight subjects with hearing impairment indicate that this training efficient personalization approach provides personalized gain settings which are on average six times more preferred over the standard prescriptive gain settings.
【14】 Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
标题: 使用目标说话者独奏段的多通道多说话者ASB
作者:Yiwen Shao,Shi-Xiong Zhang,Yong Xu,Meng Yu,Dong Yu,Daniel Povey,Sanjeev Khudanpur
备注:Accepted for presentation at Interspeech 2024
链接:点击下载PDF文件
摘要:在多通道、多说话人的自动语音识别(ASR)领域中,在背景噪声中识别并准确地转录目标说话人的语音的任务仍然是一个艰巨的挑战。传统的方法通常依赖于麦克风阵列配置和目标说话人的位置或声纹信息。本研究介绍了独奏空间功能(独奏SF),一种创新的方法,利用目标扬声器的孤立的语音段,以提高ASR的性能,从而规避了传统的输入,如麦克风阵列布局的需要。我们探索有效的策略,选择最佳的独奏部分,独奏SF的成功的一个重要方面。通过对AliMeeting数据集和AISHELL-1模拟进行的评估,Solo-SF表现出优于现有技术的性能,在各种测试条件下显着降低了字符错误率(CER)。我们的研究结果突出了Solo-SF作为解决多通道,多扬声器ASR任务的复杂性的有效解决方案的潜力。摘要:In the field of multi-channel, multi-speaker Automatic Speech Recognition (ASR), the task of discerning and accurately transcribing a target speaker's speech within background noise remains a formidable challenge. Traditional approaches often rely on microphone array configurations and the information of the target speaker's location or voiceprint. This study introduces the Solo Spatial Feature (Solo-SF), an innovative method that utilizes a target speaker's isolated speech segment to enhance ASR performance, thereby circumventing the need for conventional inputs like microphone array layouts. We explore effective strategies for selecting optimal solo segments, a crucial aspect for Solo-SF's success. Through evaluations conducted on the AliMeeting dataset and AISHELL-1 simulations, Solo-SF demonstrates superior performance over existing techniques, significantly lowering Character Error Rates (CER) in various test conditions. Our findings highlight Solo-SF's potential as an effective solution for addressing the complexities of multi-channel, multi-speaker ASR tasks.
【15】 The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments
标题: 第二个DISPLACE挑战:对话环境中Speaker和LAnguage的二元化
作者:Shareef Babu Kalluri,Prachi Singh,Pratik Roy Chowdhuri,Apoorva Kulkarni,Shikha Baghel,Pradyoth Hegde,Swapnil Sontakke,Deepak K T,S. R. Mahadeva Prasanna,Deepu Vijayasenan,Sriram Ganapathy
备注:5 pages, 3 figures, Interspeech 2024
链接:点击下载PDF文件
摘要:对话环境中的演讲者和语言的日记化(DIPLACE)2024挑战赛是DIPLACE系列挑战赛中的第二个挑战赛,该挑战赛涉及具有挑战性的多语言对话语音数据集上的演讲者日记化(SD)和语言日记化(LD)任务。在DISPLACE 2024挑战赛中,我们还介绍了在该数据集上进行自动语音识别(ASR)的任务。该数据集包含158小时的语音,包括有监督和无监督的单通道远场录音,已发布用于LD和SD音轨。此外,还为用5种印度语言进行的ASR跟踪提供了12小时的近场单通道录音。本文重点介绍了数据集、基线系统和排行榜结果的详细情况。我们还比较了我们的基线模型和团队在DISPLACE-2023评估数据上的表现,以强调第二版挑战中所取得的进步。摘要:The DIarization of SPeaker and LAnguage in Conversational Environments (DISPLACE) 2024 challenge is the second in the series of DISPLACE challenges, which involves tasks of speaker diarization (SD) and language diarization (LD) on a challenging multilingual conversational speech dataset. In the DISPLACE 2024 challenge, we also introduced the task of automatic speech recognition (ASR) on this dataset. The dataset containing 158 hours of speech, consisting of both supervised and unsupervised mono-channel far-field recordings, was released for LD and SD tracks. Further, 12 hours of close-field mono-channel recordings were provided for the ASR track conducted on 5 Indian languages. The details of the dataset, baseline systems and the leader board results are highlighted in this paper. We have also compared our baseline models and the team's performances on evaluation data of DISPLACE-2023 to emphasize the advancements made in this second version of the challenge.
【16】 GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
标题: GenDistiller:基于自回归生成模型提取预训练语言模型
作者:Yingying Gao,Shilei Zhang,Chao Deng,Junlan Feng
备注:arXiv admin note: text overlap with arXiv:2310.13418
链接:点击下载PDF文件
摘要:预训练的语音语言模型(如HuBERT和WavLM)利用未标记的语音数据进行自我监督学习,并为许多下游任务提供强大的表示。尽管这些模型取得了成功,但它们对存储器和计算资源的高要求阻碍了它们在资源受限设备上的应用。因此,本文介绍了GenDistiller,一种新的知识蒸馏框架,它通过一个小得多的学生网络直接生成预训练教师模型的隐藏表示。该方法以前一个隐含层为历史,实现教师模型的逐层自回归预测。在SUPERB上的实验表明,与没有自回归框架的基线提取方法相比,GenDistiller具有33%的参数减少,相似的时间消耗和更好的性能。最终,GenDistiller将WavLM的大小减少了82%。摘要:Pre-trained speech language models such as HuBERT and WavLM leverage unlabeled speech data for self-supervised learning and offer powerful representations for numerous downstream tasks. Despite the success of these models, their high requirements for memory and computing resource hinder their application on resource restricted devices. Therefore, this paper introduces GenDistiller, a novel knowledge distillation framework which generates the hidden representations of the pre-trained teacher model directly by a much smaller student network. The proposed method takes the previous hidden layer as history and implements a layer-by-layer prediction of the teacher model autoregressively. Experiments on SUPERB reveal the advantage of GenDistiller over the baseline distilling method without an autoregressive framework, with 33% fewer parameters, similar time consumption and better performance on most of the SUPERB tasks. Ultimately, the proposed GenDistiller reduces the size of WavLM by 82%.
【17】 Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness
标题: 个性化语音活动检测系统的比较分析:评估现实世界的有效性
作者:Satyam Kumar,Sai Srujana Buddi,Utkarsh Oggy Sarawgi,Vineet Garg,Shivesh Ranjan,Ognjen,Rudovic,Ahmed Hussen Abdelaziz,Saurabh Adya
链接:点击下载PDF文件
摘要:语音活动检测(VAD)是诸如语音识别、语音增强和免提通信系统等各种应用中的关键组件。随着对个性化和上下文感知技术的需求不断增长,对有效的个性化VAD系统的需求变得至关重要。在本文中,我们提出了一个个性化的语音活动检测(PVAD)系统的比较分析,以评估其在现实世界中的有效性。我们引入了一种全面的方法来评估PVAD系统,结合各种性能指标,如帧级和话语级错误率,检测延迟和准确性,以及用户级分析。通过广泛的实验和评估,我们对各种PVAD变体的优势和局限性有了全面的了解。本文通过使用一套全面的指标来深入了解PVAD技术在实际应用中的功效和可行性,从而促进了对PVAD技术的理解。摘要:Voice activity detection (VAD) is a critical component in various applications such as speech recognition, speech enhancement, and hands-free communication systems. With the increasing demand for personalized and context-aware technologies, the need for effective personalized VAD systems has become paramount. In this paper, we present a comparative analysis of Personalized Voice Activity Detection (PVAD) systems to assess their real-world effectiveness. We introduce a comprehensive approach to assess PVAD systems, incorporating various performance metrics such as frame-level and utterance-level error rates, detection latency and accuracy, alongside user-level analysis. Through extensive experimentation and evaluation, we provide a thorough understanding of the strengths and limitations of various PVAD variants. This paper advances the understanding of PVAD technology by offering insights into its efficacy and viability in practical applications using a comprehensive set of metrics.
【18】 Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation
标题: 用于高效多语言语音到语音翻译的扩散合成器
作者:Nameer Hirschkind,Xiao Yu,Mahesh Kumar Nandwana,Joseph Liu,Eloi DuBois,Dao Le,Nicolas Thiebaut,Colin Sinclair,Kyle Spence,Charles Shang,Zoe Abrams,Morgan McGuire
备注:Published in Interspeech 2024
链接:点击下载PDF文件
摘要:我们介绍了DiffuseST,一个低延迟,直接语音到语音翻译系统,能够保留输入扬声器的声音zero-shot,同时从多个源语言翻译成英语。我们的实验与合成器组件的架构,比较基于Tacotron的合成器的一种新的扩散为基础的合成器。我们发现基于扩散的合成器可以将MOS和PESQ音频质量指标分别提高23%,扬声器相似性提高5%,同时保持可比较的BLEU分数。尽管有两倍多的参数计数,扩散合成器具有较低的延迟,使整个模型的运行速度比实时快5倍。摘要:We introduce DiffuseST, a low-latency, direct speech-to-speech translation system capable of preserving the input speaker's voice zero-shot while translating from multiple source languages into English. We experiment with the synthesizer component of the architecture, comparing a Tacotron-based synthesizer to a novel diffusion-based synthesizer. We find the diffusion-based synthesizer to improve MOS and PESQ audio quality metrics by 23 % each and speaker similarity by 5 % while maintaining comparable BLEU scores. Despite having more than double the parameter count, the diffusion synthesizer has lower latency, allowing the entire model to run more than 5$ times$ faster than real-time.
【19】 One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model
标题: 使用一体化神经模型的一次多Conformer和Foundation语音系统压缩和量化
作者:Zhaoqing Li,Haoning Xu,Tianzi Wang,Shoukang Hu,Zengrui Jin,Shujie Hu,Jiajun Deng,Mingyu Cui,Mengzhe Geng,Xunying Liu
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:我们提出了一种新的一次通过多个ASR系统联合压缩和量化的方法,使用一个全在一个神经模型。单个压缩周期允许同时构造具有不同编码器深度、宽度和量化精度设置的多个嵌套系统,而无需单独训练和存储各个目标系统。实验一致表明,在单个一体化模型中压缩的多个ASR系统产生的单词错误率(WER)与同等复杂度的单独训练系统相当,或比其低1.01%绝对值(6.98%相对值)。实现了3.4倍的整体系统压缩和训练时间加速。在基线Switchboard-300 hr Conformer和LibriSpeech-100 hr微调wav2vec2.0模型上分别获得了12.8x和3.93x的最大模型大小压缩比,未引起统计学显著的WER增加。摘要:We propose a novel one-pass multiple ASR systems joint compression and quantization approach using an all-in-one neural model. A single compression cycle allows multiple nested systems with varying Encoder depths, widths, and quantization precision settings to be simultaneously constructed without the need to train and store individual target systems separately. Experiments consistently demonstrate the multiple ASR systems compressed in a single all-in-one model produced a word error rate (WER) comparable to, or lower by up to 1.01 % absolute (6.98 % relative) than individually trained systems of equal complexity. A 3.4x overall system compression and training time speed-up was achieved. Maximum model size compression ratios of 12.8x and 3.93x were obtained over the baseline Switchboard-300hr Conformer and LibriSpeech-100hr fine-tuned wav2vec2.0 models, respectively, incurring no statistically significant WER increase.
【20】 Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
标题: 用于视听多通道语音分离和识别的联合说话人特征学习
作者:Guinan Li,Jiajun Deng,Youjun Chen,Mengzhe Geng,Shujie Hu,Zhe Li,Zengrui Jin,Tianzi Wang,Xurong Xie,Helen Meng,Xunying Liu
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:本文提出了一种用于视听多通道语音分离与识别系统中zero-shot自适应的联合说话人特征学习方法。xVector和ECAPA-TDNN扬声器编码器使用专用融合模块连接,并与完整的系统训练紧密集成。LRS 3-TED数据模拟多通道重叠语音进行的实验表明,联合说话人特征学习一贯提高语音分离和识别性能的基线没有联合说话人特征估计。进一步的分析表明,性能的改善与使用余弦相似性测量的说话人间辨别力的增加密切相关。在结合WavLM功能和视频模式后,性能最佳的联合扬声器特征学习自适应系统在Dev和Test集上的WER绝对下降了21.6%和25.3%(相对下降了67.5%和83.5%),表现优于基线微调WavLM模型。摘要:This paper proposes joint speaker feature learning methods for zero-shot adaptation of audio-visual multichannel speech separation and recognition systems. xVector and ECAPA-TDNN speaker encoders are connected using purpose-built fusion blocks and tightly integrated with the complete system training. Experiments conducted on LRS3-TED data simulated multichannel overlapped speech suggest that joint speaker feature learning consistently improves speech separation and recognition performance over the baselines without joint speaker feature estimation. Further analyses reveal performance improvements are strongly correlated with increased inter-speaker discrimination measured using cosine similarity. The best-performing joint speaker feature learning adapted system outperformed the baseline fine-tuned WavLM model by statistically significant WER reductions of 21.6% and 25.3% absolute (67.5% and 83.5% relative) on Dev and Test sets after incorporating WavLM features and video modality.
【21】 On the Evaluation of Speech Foundation Models for Spoken Language Understanding
标题: 关于口语理解的言语基础模型的评估
作者:Siddhant Arora,Ankita Pasad,Chung-Ming Chien,Jionghao Han,Roshan Sharma,Jee-weon Jung,Hira Dhamyal,William Chen,Suwon Shon,Hung-yi Lee,Karen Livescu,Shinji Watanabe
备注:Accepted at ACL Findings 2024
链接:点击下载PDF文件
摘要:最近推出了口语理解评估(SLUE)基准测试任务套件,以满足对开放资源的需求,并对复杂的口语理解(SLU)任务进行基准测试,包括自然语音的分类和序列生成任务。该基准测试已经证明了使用预先训练的语音基础模型(SFM)进行这些SLU任务的初步成功。然而,社会对不同可持续森林管理措施的比较效用仍缺乏深入的了解。受此启发,我们问:哪些SFM为这些复杂的SLU任务提供了最大的好处,以及整合这些SFM的最有效方法是什么?为了回答这个问题,我们使用几种评估协议对多个监督和自监督的SFM进行了广泛的评估:(i)具有轻量级预测头的冻结SFM,(ii)具有复杂预测头的冻结SFM,以及(iii)具有轻量级预测头的微调SFM。虽然监督的SFM是在更多的语音识别数据(带有标签)上进行预训练的,但它们并不总是优于自监督的SFM;后者的表现往往至少与监督的SFM一样好,有时甚至更好,特别是在SLUE中的序列生成任务上。虽然没有通用的最佳方法来合并SFM,但复杂的预测头为大多数任务提供了最佳性能,尽管它增加了推理时间。我们还介绍了一个开源的工具包和性能排行榜,SLUE-PERB,这些任务和建模策略。摘要:The Spoken Language Understanding Evaluation (SLUE) suite of benchmark tasks was recently introduced to address the need for open resources and benchmarking of complex spoken language understanding (SLU) tasks, including both classification and sequence generation tasks, on natural speech. The benchmark has demonstrated preliminary success in using pre-trained speech foundation models (SFM) for these SLU tasks. However, the community still lacks a fine-grained understanding of the comparative utility of different SFMs. Inspired by this, we ask: which SFMs offer the most benefits for these complex SLU tasks, and what is the most effective approach for incorporating these SFMs? To answer this, we perform an extensive evaluation of multiple supervised and self-supervised SFMs using several evaluation protocols: (i) frozen SFMs with a lightweight prediction head, (ii) frozen SFMs with a complex prediction head, and (iii) fine-tuned SFMs with a lightweight prediction head. Although the supervised SFMs are pre-trained on much more speech recognition data (with labels), they do not always outperform self-supervised SFMs; the latter tend to perform at least as well as, and sometimes better than, supervised SFMs, especially on the sequence generation tasks in SLUE. While there is no universally optimal way of incorporating SFMs, the complex prediction head gives the best performance for most tasks, although it increases the inference time. We also introduce an open-source toolkit and performance leaderboard, SLUE-PERB, for these tasks and modeling strategies.
【22】 UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
标题: UniAudio 1.5:大型语言模型驱动的音频编解码器是一款只需几次的音频任务学习器
作者:Dongchao Yang,Haohan Guo,Yuanyuan Wang,Rongjie Huang,Xiang Li,Xu Tan,Xixin Wu,Helen Meng
链接:点击下载PDF文件
摘要:大型语言模型(LLM)在文本理解和生成方面表现出了卓越的能力,但如果不进行微调,就不能直接应用于跨模态任务。本文提出了一种跨模态上下文学习方法,使冻结的LLM能够以Few-Shot风格实现多个音频任务,而无需任何参数更新。具体来说,我们提出了一种新的LLMs驱动的音频编解码器模型,LLM-Codec,将音频模态转移到文本空间,用LLM的词汇表中的词或子词来表示音频令牌,同时保持高的音频重构质量。其关键思想是通过将音频模态压缩到经过良好训练的LLM令牌空间中来减少文本和音频之间的模态异构性。因此,音频表示可以被视为一种新的 textit{foreign language},LLM可以通过几次演示来学习新的 textit{foreign language}。在实验中,我们研究了所提出的方法在多个音频理解和生成任务中的性能,语音情感分类、音频分类、文本到语音生成、语音增强等。实验结果表明,在简单的场景下,采用本文提出的LLM-Codec(UniAudio 1.5)的LLM能够实现预期的功能。实验结果验证了跨模态情境学习方法的可行性和有效性。为了促进对Few-Shot音频任务学习和多模式LLM的研究,我们开源了LLM-Codec模型。摘要:The Large Language models (LLMs) have demonstrated supreme capabilities in text understanding and generation, but cannot be directly applied to cross-modal tasks without fine-tuning. This paper proposes a cross-modal in-context learning approach, empowering the frozen LLMs to achieve multiple audio tasks in a few-shot style without any parameter update. Specifically, we propose a novel and LLMs-driven audio codec model, LLM-Codec, to transfer the audio modality into the textual space, textit{i.e.} representing audio tokens with words or sub-words in the vocabulary of LLMs, while keeping high audio reconstruction quality. The key idea is to reduce the modality heterogeneity between text and audio by compressing the audio modality into a well-trained LLMs token space. Thus, the audio representation can be viewed as a new textit{foreign language}, and LLMs can learn the new textit{foreign language} with several demonstrations. In experiments, we investigate the performance of the proposed approach across multiple audio understanding and generation tasks, textit{e.g.} speech emotion classification, audio classification, text-to-speech generation, speech enhancement, etc. The experimental results demonstrate that the LLMs equipped with the proposed LLM-Codec, named as UniAudio 1.5, prompted by only a few examples, can achieve the expected functions in simple scenarios. It validates the feasibility and effectiveness of the proposed cross-modal in-context learning approach. To facilitate research on few-shot audio task learning and multi-modal LLMs, we have open-sourced the LLM-Codec model.
【23】 Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
标题: Sim-Whisper:具有截断检测的注意力引导流媒体Whisper
作者:Haoyu Wang,Guoqiang Hu,Guodong Lin,Wei-Qiang Zhang,Jian Li
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:作为一个强大的大规模多语言语音识别模型,Whisper已经在许多低资源和非分布场景中展示了令人印象深刻的结果。然而,它的编解码器结构阻碍了其应用于流语音识别。在本文中,我们介绍了Simul-Whisper,它使用嵌入在Whisper的交叉注意中的时间对齐来指导自回归解码,并实现基于块的流式ASR,而无需对预训练模型进行任何微调。此外,我们观察到的负面影响,截断词在块边界上的解码结果,并提出了一个完整的和火灾为基础的截断检测模型来解决这个问题。在多种语言和Whisper架构上的实验表明,Simul-Whisper在1秒的块大小下平均绝对字错误率仅下降1.46%,显着优于当前最先进的基线。摘要:As a robust and large-scale multilingual speech recognition model, Whisper has demonstrated impressive results in many low-resource and out-of-distribution scenarios. However, its encoder-decoder structure hinders its application to streaming speech recognition. In this paper, we introduce Simul-Whisper, which uses the time alignment embedded in Whisper's cross-attention to guide auto-regressive decoding and achieve chunk-based streaming ASR without any fine-tuning of the pre-trained model. Furthermore, we observe the negative effect of the truncated words at the chunk boundaries on the decoding results and propose an integrate-and-fire-based truncation detection model to address this issue. Experiments on multiple languages and Whisper architectures show that Simul-Whisper achieves an average absolute word error rate degradation of only 1.46% at a chunk size of 1 second, which significantly outperforms the current state-of-the-art baseline.
【24】 Towards Effective and Efficient Non-autoregressive Decoding Using Block-based Attention Mask
标题: 使用基于块的注意力屏蔽实现有效且高效的非自回归解码
作者:Tianzi Wang,Xurong Xie,Zhaoqing Li,Shoukang Hu,Zengrui Jing,Jiajun Deng,Mingyu Cui,Shujie Hu,Mengzhe Geng,Guinan Li,Helen Meng,Xunying Liu
备注:5 pages, 2 figures, 2 tables, Interspeech24 conference
链接:点击下载PDF文件
摘要:本文提出了一种新的非自回归(NAR)块为基础的注意掩码解码器(AMD),灵活地平衡性能效率的权衡一致性ASR系统。AMD在使用注意掩码隐藏的输出标签的连续块内执行并行NAR推理,同时在块之间进行从左到右的AR预测和历史上下文融合。波束搜索算法被设计为利用CTC、AR解码器和AMD概率的动态融合。LibriSpeech-100 hr语料库上的实验表明,结合AMD模块的三方解码器产生了比基线CTC+AR解码1.73倍的最大解码加速比,同时在测试集上没有产生统计上显著的字错误率(WER)增加。当使用相同的解码实时因子操作时,在CTC+AR基线上获得了高达0.7%和0.3%的绝对(5.3%和6.1%的相对)统计学显著WER降低。摘要:This paper proposes a novel non-autoregressive (NAR) block-based Attention Mask Decoder (AMD) that flexibly balances performance-efficiency trade-offs for Conformer ASR systems. AMD performs parallel NAR inference within contiguous blocks of output labels that are concealed using attention masks, while conducting left-to-right AR prediction and history context amalgamation between blocks. A beam search algorithm is designed to leverage a dynamic fusion of CTC, AR Decoder, and AMD probabilities. Experiments on the LibriSpeech-100hr corpus suggest the tripartite Decoder incorporating the AMD module produces a maximum decoding speed-up ratio of 1.73x over the baseline CTC+AR decoding, while incurring no statistically significant word error rate (WER) increase on the test sets. When operating with the same decoding real time factors, statistically significant WER reductions of up to 0.7% and 0.3% absolute (5.3% and 6.1% relative) were obtained over the CTC+AR baseline.
【25】 Impact of Speech Mode in Automatic Pathological Speech Detection
标题: 语音模式对自动病理语音检测的影响
作者:Shakeel A. Sheikh,Ina Kodrasi
备注:Accepted in EUSIPCO 2024
链接:点击下载PDF文件
摘要:自动病理语音检测方法在识别各种病理方面产生了很好的结果。这些方法通常是针对语音控制的语音场景而设计和评估的,在语音控制的语音场景中,提示说话者清晰地表达相同的语音内容。虽然收集受控的语音记录可能是费力的,但随着潜在患者浏览他们的日常生活,可以方便地获得自发语音。此外,自发言语在检测病理性言语的微妙和抽象线索方面可能是有价值的。尽管如此,自动病理语音检测自发语音的有效性仍有待探索。本文分析了语音模式对病态语音检测方法的影响,研究了两种不同的方法,即,经典机器学习和深度学习。结果表明,经典的方法可能难以捕捉自发语音的病理判别线索。相比之下,深度学习方法表现出卓越的性能,能够提取以前在非自发语音中无法获得的额外线索摘要:Automatic pathological speech detection approaches yield promising results in identifying various pathologies. These approaches are typically designed and evaluated for phonetically-controlled speech scenarios, where speakers are prompted to articulate identical phonetic content. While gathering controlled speech recordings can be laborious, spontaneous speech can be conveniently acquired as potential patients navigate their daily routines. Further, spontaneous speech can be valuable in detecting subtle and abstract cues of pathological speech. Nonetheless, the efficacy of automatic pathological speech detection for spontaneous speech remains unexplored. This paper analyzes the influence of speech mode on pathological speech detection approaches, examining two distinct categories of approaches, i.e., classical machine learning and deep learning. Results indicate that classical approaches may struggle to capture pathology-discriminant cues in spontaneous speech. In contrast, deep learning approaches demonstrate superior performance, managing to extract additional cues that were previously inaccessible in non-spontaneous speech
【26】 An efficient text augmentation approach for contextualized Mandarin speech recognition
标题: 一种有效的上下文化普通话语音识别文本增强方法
作者:Naijun Zheng,Xucheng Wan,Kai Liu,Ziqing Du,Zhou Huan
备注:accepted to interspeech2024
链接:点击下载PDF文件
摘要:虽然上下文自动语音识别(ASR)系统通常用于提高识别不常见的单词,其有效性受到语音文本数据可用性的固有限制。为了应对这一挑战,我们的研究建议利用广泛的纯文本数据集,并使用简单的文本增强(TA)技术将预训练的ASR模型置于上下文中,同时保持计算成本最小。特别是,为了将预先训练的基于CIF的ASR置于上下文中,我们使用有限的语音文本数据构建码本。通过利用一个简单的码本查找过程,我们将可用的纯文本数据转换为潜在的文本嵌入。然后,这些嵌入增强了情境化ASR的输入。我们在不同的普通话测试集上的实验表明,我们的TA方法显着提高识别性能。表现最好的系统在罕见单词上显示出高达30%的相对CER改进,在所有单词上显示出15%的相对CER改进。摘要:Although contextualized automatic speech recognition (ASR) systems are commonly used to improve the recognition of uncommon words, their effectiveness is hindered by the inherent limitations of speech-text data availability. To address this challenge, our study proposes to leverage extensive text-only datasets and contextualize pre-trained ASR models using a straightforward text-augmentation (TA) technique, all while keeping computational costs minimal. In particular, to contextualize a pre-trained CIF-based ASR, we construct a codebook using limited speech-text data. By utilizing a simple codebook lookup process, we convert available text-only data into latent text embeddings. These embeddings then enhance the inputs for the contextualized ASR. Our experiments on diverse Mandarin test sets demonstrate that our TA approach significantly boosts recognition performance. The top-performing system shows relative CER improvements of up to 30% on rare words and 15% across all words in general.
【27】 Personalized Speech Enhancement Without a Separate Speaker Embedding Model
标题: 无需单独说话者嵌入模型的个性化语音增强
作者:Tanel Pärnamaa,Ando Saabas
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:个性化语音增强(PSE)模型可以通过适应说话人的语音特征来改善电话会议系统的音频质量。然而,大多数现有方法需要单独的说话人嵌入模型来从注册音频中提取说话人的向量表示,这增加了训练和部署过程的复杂性。我们建议使用PSE模型本身的内部表示作为说话人嵌入,从而避免了对单独模型的需要。我们表明,我们的方法在噪声抑制和回声消除任务上与使用预训练的说话人嵌入模型的标准方法一样好或更好。此外,我们的方法在平均意见得分上超过ICASSP 2023深度噪声抑制挑战赛冠军0.15。摘要:Personalized speech enhancement (PSE) models can improve the audio quality of teleconferencing systems by adapting to the characteristics of a speaker's voice. However, most existing methods require a separate speaker embedding model to extract a vector representation of the speaker from enrollment audio, which adds complexity to the training and deployment process. We propose to use the internal representation of the PSE model itself as the speaker embedding, thereby avoiding the need for a separate model. We show that our approach performs equally well or better than the standard method of using a pre-trained speaker embedding model on noise suppression and echo cancellation tasks. Moreover, our approach surpasses the ICASSP 2023 Deep Noise Suppression Challenge winner by 0.15 in Mean Opinion Score.
【28】 MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
标题: MMM:来自自监督学习模型的多层多残留多流离散语音表示
作者:Jiatong Shi,Xutai Ma,Hirofumi Inaguma,Anna Sun,Shinji Watanabe
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
摘要:语音离散表示已被证明是有效的,在各种下游应用,由于其优越的压缩率的波形,在训练过程中的快速收敛,并与其他形式的兼容性。从自监督学习(SSL)模型中提取的离散单元已经成为获得语音离散表示的一种重要方法。然而,虽然离散单元已经显示出与光谱特征相比的有效性,但它们仍然落后于连续SSL表示。在这项工作中,我们提出了MMM,一个多层多残差多流离散单元提取方法从SSL。具体来说,我们介绍了迭代残差矢量量化与K-均值的SSL模型中的不同层提取多流语音离散表示。通过在语音识别,语音再合成,和文本到语音的广泛实验,我们证明了建议的MMM可以超过或等同于神经编解码器的性能在各种条件下。摘要:Speech discrete representation has proven effective in various downstream applications due to its superior compression rate of the waveform, fast convergence during training, and compatibility with other modalities. Discrete units extracted from self-supervised learning (SSL) models have emerged as a prominent approach for obtaining speech discrete representation. However, while discrete units have shown effectiveness compared to spectral features, they still lag behind continuous SSL representations. In this work, we propose MMM, a multi-layer multi-residual multi-stream discrete units extraction method from SSL. Specifically, we introduce iterative residual vector quantization with K-means for different layers in an SSL model to extract multi-stream speech discrete representation. Through extensive experiments in speech recognition, speech resynthesis, and text-to-speech, we demonstrate the proposed MMM can surpass or on-par with neural codec's performance under various conditions.
【29】 Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
标题: Vec-Tok-VC+:双模式训练策略中具有渐进约束的剩余增强稳健Zero-Shot语音转换
作者:Linhan Ma,Xinfa Zhu,Yuanjun Lv,Zhichao Wang,Ziqian Wang,Wendi He,Hongbin Zhou,Lei Xie
备注:Accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:Zero-shot语音转换(Zero Shot Voice Conversion,VC)的目的是在保持语言内容不变的情况下,将源语音转换为任意不可见的目标语音。最近的VC方法已经取得了显着的进步,但在解耦过程中的语义损失以及训练推理不匹配仍然阻碍转换性能。在本文中,我们提出了Vec-Tok-VC+,一种新的基于MPEG-2的zero-shot VC模型的Vec-Tok Codec的改进,实现语音转换,只有一个3s的目标说话人提示。我们设计了一个残差增强的K-Means聚类器,通过两层聚类过程来增强语义内容提取。此外,我们采用教师指导的精化来模拟转换过程,以消除训练-推理失配,形成双模式训练策略。此外,我们设计了一个多码本渐进损失函数来约束模型的逐层输出从粗到细,以提高说话人相似度和内容准确性。客观和主观评价表明,Vec-Tok-VC+在自然度、可懂度和说话人相似度方面优于强基线。摘要:Zero-shot voice conversion (VC) aims to transform source speech into arbitrary unseen target voice while keeping the linguistic content unchanged. Recent VC methods have made significant progress, but semantic losses in the decoupling process as well as training-inference mismatch still hinder conversion performance. In this paper, we propose Vec-Tok-VC+, a novel prompt-based zero-shot VC model improved from Vec-Tok Codec, achieving voice conversion given only a 3s target speaker prompt. We design a residual-enhanced K-Means decoupler to enhance the semantic content extraction with a two-layer clustering process. Besides, we employ teacher-guided refinement to simulate the conversion process to eliminate the training-inference mismatch, forming a dual-mode training strategy. Furthermore, we design a multi-codebook progressive loss function to constrain the layer-wise output of the model from coarse to fine to improve speaker similarity and content accuracy. Objective and subjective evaluations demonstrate that Vec-Tok-VC+ outperforms the strong baselines in naturalness, intelligibility, and speaker similarity.
【30】 SHMamba: Structured Hyperbolic State Space Model for Audio-Visual Question Answering
标题: SHMamba:视听问题回答的结构化双曲状态空间模型
作者:Zhe Yang,Wenrui Li,Guanghui Cheng
链接:点击下载PDF文件
摘要:视听问答(AVQA)任务具有很大的应用潜力。与传统的单峰方法相比,AVQA的多模态输入使得特征提取和融合过程更具挑战性。欧氏空间难以有效地表示数据的多维关系。特别是当提取和处理具有树结构或层次结构的数据时,欧几里得空间不适合作为嵌入空间。此外,Transformers中的自注意机制在捕捉序列中元素之间的动态关系方面是有效的。然而,自注意机制在窗口建模和二次计算复杂性方面的局限性降低了其在长序列建模中的有效性。为了解决这些问题,我们提出了SHMamba:结构化双曲状态空间模型,以整合双曲几何和状态空间模型的优点。具体来说,SHMamba利用双曲空间的内在属性来表示视听数据中的层次结构和复杂关系。同时,状态空间模型通过对整个序列进行全局建模来捕获随时间的动态变化。此外,我们引入了一个自适应曲率双曲线对齐模块和交叉融合块,以提高层次结构的理解和跨模态信息的动态交换,分别。大量的实验表明,SHMamba优于以前的方法,更少的参数和计算成本。我们的可学习参数减少了78.12%,而平均性能提高了2.53%.实验结果表明,该方法在当前主流方法中具有一定的优越性,更适合于实际应用场景。摘要:The Audio-Visual Question Answering (AVQA) task holds significant potential for applications. Compared to traditional unimodal approaches, the multi-modal input of AVQA makes feature extraction and fusion processes more challenging. Euclidean space is difficult to effectively represent multi-dimensional relationships of data. Especially when extracting and processing data with a tree structure or hierarchical structure, Euclidean space is not suitable as an embedding space. Additionally, the self-attention mechanism in Transformers is effective in capturing the dynamic relationships between elements in a sequence. However, the self-attention mechanism's limitations in window modeling and quadratic computational complexity reduce its effectiveness in modeling long sequences. To address these limitations, we propose SHMamba: Structured Hyperbolic State Space Model to integrate the advantages of hyperbolic geometry and state space models. Specifically, SHMamba leverages the intrinsic properties of hyperbolic space to represent hierarchical structures and complex relationships in audio-visual data. Meanwhile, the state space model captures dynamic changes over time by globally modeling the entire sequence. Furthermore, we introduce an adaptive curvature hyperbolic alignment module and a cross fusion block to enhance the understanding of hierarchical structures and the dynamic exchange of cross-modal information, respectively. Extensive experiments demonstrate that SHMamba outperforms previous methods with fewer parameters and computational costs. Our learnable parameters are reduced by 78.12 %, while the average performance improves by 2.53 %. Experiments show that our method demonstrates superiority among all current major methods and is more suitable for practical application scenarios.
【31】 Frequency-mix Knowledge Distillation for Fake Speech Detection
标题: 用于假语音检测的混频知识提炼
作者:Cunhang Fan,Shunbo Dong,Jun Xue,Yujie Chen,Jiangyan Yi,Zhao Lv
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在电话场景中,对抗语音欺骗攻击的虚假语音检测(FSD)任务具有挑战性。数据增强(DA)方法被认为是解决电话场景中FSD任务的有效手段,通常分为时域和频域阶段。虽然每一种都有其优点,但两者都可能导致信息丢失。为了解决这个问题,我们提出了一种新的DA方法,频率混合(Freqmix),并引入Freqmix知识蒸馏(FKD),以提高模型信息提取和泛化能力。具体来说,我们使用Freqmix增强的数据作为教师模型的输入,而学生模型的输入经历时域DA方法。我们使用多层次的特征提取方法来恢复信息,提高模型的泛化能力。我们的方法在ASVspoof 2021 LA数据集上实现了最先进的结果,比基线提高了31%,并且在ASVspoof 2021 DF数据集上具有竞争力。摘要:In the telephony scenarios, the fake speech detection (FSD) task to combat speech spoofing attacks is challenging. Data augmentation (DA) methods are considered effective means to address the FSD task in telephony scenarios, typically divided into time domain and frequency domain stages. While each has its advantages, both can result in information loss. To tackle this issue, we propose a novel DA method, Frequency-mix (Freqmix), and introduce the Freqmix knowledge distillation (FKD) to enhance model information extraction and generalization abilities. Specifically, we use Freqmix-enhanced data as input for the teacher model, while the student model's input undergoes time-domain DA method. We use a multi-level feature distillation approach to restore information and improve the model's generalization capabilities. Our approach achieves state-of-the-art results on ASVspoof 2021 LA dataset, showing a 31 % improvement over baseline and performs competitively on ASVspoof 2021 DF dataset.
【32】 Multi-Modal Retrieval For Large Language Model Based Speech Recognition
标题: 基于大语言模型的语音识别的多模式检索
作者:Jari Kolehmainen,Aditya Gourav,Prashanth Gurunath Shivakumar,Yile Gu,Ankur Gandhe,Ariya Rastrow,Grant Strimel,Ivan Bulyko
链接:点击下载PDF文件
摘要:检索是利用外部信息改进语言模型的一种广泛采用的方法。随着该领域转向多模态大型语言模型,重要的是扩展基于纯文本的方法,以将其他模态纳入检索以及广泛的机器学习任务和数据类型的应用程序。在这项工作中,我们提出了多模态检索两种方法:kNN-LM和交叉注意技术。我们证明了我们的检索方法的有效性经验,将它们应用到自动语音识别任务与外部信息的访问。在这种设置下,我们表明,基于语音的多模态检索优于基于文本的检索,并产生高达50%的改善,在多模态语言模型基线的单词错误率。此外,我们在Spoken-Squad问答数据集上实现了最先进的识别结果。摘要:Retrieval is a widely adopted approach for improving language models leveraging external information. As the field moves towards multi-modal large language models, it is important to extend the pure text based methods to incorporate other modalities in retrieval as well for applications across the wide spectrum of machine learning tasks and data types. In this work, we propose multi-modal retrieval with two approaches: kNN-LM and cross-attention techniques. We demonstrate the effectiveness of our retrieval approaches empirically by applying them to automatic speech recognition tasks with access to external information. Under this setting, we show that speech-based multi-modal retrieval outperforms text based retrieval, and yields up to 50 % improvement in word error rate over the multi-modal language model baseline. Furthermore, we achieve state-of-the-art recognition results on the Spoken-Squad question answering dataset.
【33】 Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
标题: 用于设备定向语音检测的具有融合低等级自适应的多模式大语言模型
作者:Shruti Palaskar,Oggi Rudovic,Sameer Dharur,Florian Pesce,Gautam Krishna,Aswin Sivaraman,Jack Berkowitz,Ahmed Hussen Abdelaziz,Saurabh Adya,Ahmed Tewfik
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:虽然大型语言模型(LLM)已经显示出类似人类对话的前景,但它们主要是在文本数据上进行预训练的。播放音频或视频可以提高性能,但收集大规模多模式数据和预训练多模式LLM是一项挑战。为此,我们提出了一种融合低秩自适应(FLoRA)技术,该技术可以有效地调整预训练的单峰LLM,以通过低秩自适应来消耗新的、以前看不见的模态。对于设备导向的语音检测,使用FLoRA,多模态LLM实现了22%的相对减少等错误率(EER)的纯文本的方法,并达到性能平价与其完全微调(FFT)对应,而只需要调整其参数的一小部分。此外,通过新引入的适配器丢弃,FLoRA对丢失数据具有鲁棒性,比FFT提高了20%的EER和56%的错误接受率。所提出的方法可以很好地扩展从16 M到3B参数的模型大小。摘要:Although Large Language Models (LLMs) have shown promise for human-like conversations, they are primarily pre-trained on text data. Incorporating audio or video improves performance, but collecting large-scale multimodal data and pre-training multimodal LLMs is challenging. To this end, we propose a Fusion Low Rank Adaptation (FLoRA) technique that efficiently adapts a pre-trained unimodal LLM to consume new, previously unseen modalities via low rank adaptation. For device-directed speech detection, using FLoRA, the multimodal LLM achieves 22% relative reduction in equal error rate (EER) over the text-only approach and attains performance parity with its full fine-tuning (FFT) counterpart while needing to tune only a fraction of its parameters. Furthermore, with the newly introduced adapter dropout, FLoRA is robust to missing data, improving over FFT by 20% lower EER and 56% lower false accept rate. The proposed approach scales well for model sizes from 16M to 3B parameters.
【34】 Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
标题: Speech ReaLLM --通过教学时间流,使用多模式LLM进行实时流语音识别
作者:Frank Seide,Morrie Doulaty,Yangyang Shi,Yashesh Gaur,Junteng Jia,Chunyang Wu
链接:点击下载PDF文件
摘要:我们引入了Speech ReaLLM,这是一种新的ASR架构,它将“仅解码器”ASR与RNN-T结合在一起,使多模式LLM架构能够实现实时流传输。这是第一个“仅解码器”ASR架构,旨在处理连续音频,而无需显式端点。Speech ReaLLM是更通用的ReaLLM(“实时LLM”)方法的特殊情况,也是第一次在这里介绍。这个想法受到RNN-T的启发:不是只在用户提示结束时生成响应,而是在实时接收到每个输入令牌后生成(通常是空的)。在Libripeech“测试”中,80 M Speech ReaLLM实时实现3.0%和7.4%的WER(没有外部LM或辅助损失)。这仅略高于3倍大的Attention-Encoder-Decoder基线。我们还表明,通过这种方式,LLM架构可以学习表示和再现时间流;并且可以对预训练的7 B LLM进行微调,以便在此任务上做得相当好。摘要:We introduce Speech ReaLLM, a new ASR architecture that marries "decoder-only" ASR with the RNN-T to make multimodal LLM architectures capable of real-time streaming. This is the first "decoder-only" ASR architecture designed to handle continuous audio without explicit end-pointing. Speech ReaLLM is a special case of the more general ReaLLM ("real-time LLM") approach, also introduced here for the first time. The idea is inspired by RNN-T: Instead of generating a response only at the end of a user prompt, generate after every input token received in real time (it is often empty). On Librispeech "test", an 80M Speech ReaLLM achieves WERs of 3.0% and 7.4% in real time (without an external LM or auxiliary loss). This is only slightly above a 3x larger Attention-Encoder-Decoder baseline. We also show that this way, an LLM architecture can learn to represent and reproduce the flow of time; and that a pre-trained 7B LLM can be fine-tuned to do reasonably well on this task.
【35】 Analyzing phonetic structure of Mandarin using Audacity
标题: 用Audacity分析普通话的语音结构
作者:Shizheng Xu
备注:audio source: this https URL
链接:点击下载PDF文件
摘要:普通话是中国、台湾和新加坡的官方语言。它也是主要的非官方语言,主要在多伦多和温哥华的家中使用。本文运用音频软件Audacity,运用理论知识对汉语普通话进行了全面的分析。本研究首先概述了普通话语音的基本原则,旨在提供对普通话语音结构的见解。摘要:Mandarin Chinese is the official language in China, Taiwan, and Singapore. It is also the main non-official language spoken predominantly at home in Toronto and Vancouver. This article employs the audio software Audacity and leverages theoretical knowledge to conduct a comprehensive analysis of Mandarin Chinese. The study initiates with an overview of the fundamental principles underlying Mandarin pronunciation, aiming to provide insights into its phonetic structure.
机器翻译,仅供参考
