本文经arXiv每日学术速递授权转载
【1】 Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech
标题: 面部表情增强的TTC:结合面部表示和情感强度实现自适应语音
作者:Yunji Chu,Yunseob Shim,Unsang Park
备注:13 pages, 3 figures, accepted to ECCV Workshop ABAW(Affective Behavior Analysis in-the-wild)7 (to be appear)
链接:点击下载PDF文件
【2】 Leveraging Mixture of Experts for Improved Speech Deepfake Detection
标题: 利用专家混合改进语音深度伪造检测
作者:Viola Negroni,Davide Salvi,Alessandro Ilic Mezza,Paolo Bestagini,Stefano Tubaro
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【3】 Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
标题: 弥合语音和文本:通过LLM中的拼音到字符预训练增强ASB
作者:Yang Yuhang,Peng Yizhou,Eng Siong Chng,Xionghu Zhong
备注:Accepted by ISCSLP2024-Special session-Speech Processing in LLM Era
链接:点击下载PDF文件
【4】 Disentangling Age and Identity with a Mutual Information Minimization Approach for Cross-Age Speaker Verification
标题: 用互信息最小化方法解开年龄和身份,用于跨年龄段说话人验证
作者:Fengrun Zhang,Wangjin Zhou,Yiming Liu,Wang Geng,Yahui Shan,Chen Zhang
备注:Interspeech 2024
链接:点击下载PDF文件
【5】 ASD-Diffusion: Anomalous Sound Detection with Diffusion Models
标题: ASD扩散:使用扩散模型检测异常声音
作者:Fengrun Zhang,Xiang Xie,Kai Guo
备注:This paper will appear at ICPR 2024
链接:点击下载PDF文件
【6】 A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation
标题: 基于模块的语音同步翻译中梯度冲突缓解策略
作者:Xiaoqian Liu,Yangfan Du,Jianjin Wang,Yuan Ge,Chen Xu,Tong Xiao,Guocheng Chen,Jingbo Zhu
链接:点击下载PDF文件
【7】 Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
标题: 通过混合专家提高代码转换ASB增强的言语条件LLM
作者:Fengrun Zhang,Wang Geng,Hukai Huang,Cheng Yi,He Qu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【8】 On the calibration of powerset speaker diarization models
标题: 关于Powerset扬声器Dialogue模型的校准
作者:Alexis Plaquet,Hervé Bredin
Journal-ref:Interspeech 2024, Sep 2024, Kos, Greece. pp.3764-3768, ⟨10.21437Interspeech.2024-1060⟩
链接:点击下载PDF文件
【9】 NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
标题: NanoVoice:适合多个扬声器的高效扬声器自适应文本到语音
作者:Nohil Park,Heeseung Kim,Che Hyun Lee,Jooyoung Choi,Jiheum Yeom,Sungroh Yoon
备注:Submitted to ICASSP 2025, Demo Page: this https URL
链接:点击下载PDF文件
【10】 VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance
标题: VoiceGuider:通过自动引导增强参数高效扬声器自适应文本到语音的域外性能
作者:Jiheum Yeom,Heeseung Kim,Jooyoung Choi,Che Hyun Lee,Nohil Park,Sungroh Yoon
备注:Submitted to ICASSP 2025, Demo Page: this https URL
链接:点击下载PDF文件
【11】 Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
标题: 假设聚集和合并:具有说话者令牌的新型多说话者语音识别
作者:Yosuke Kashiwagi,Hayato Futami,Emiru Tsunoo,Siddhant Arora,Shinji Watanabe
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【12】 Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
标题: 超越回合制接口:同步LLM作为全速对话代理
作者:Bandhav Veluri,Benjamin N Peloquin,Bokai Yu,Hongyu Gong,Shyamnath Gollakota
备注:EMNLP Main 2024
链接:点击下载PDF文件
【13】 Generalization in birdsong classification: impact of transfer learning methods and dataset characteristics
标题: 鸟鸣分类的推广:迁移学习方法和数据集特征的影响
作者:Burooj Ghani,Vincent J. Kalkman,Bob Planqué,Willem-Pier Vellinga,Lisa Gill,Dan Stowell
备注:25 pages
链接:点击下载PDF文件
【14】 Efficient learning-based sound propagation for virtual and real-world audio processing applications
标题: 针对虚拟和现实世界音频处理应用程序的高效基于学习的声音传播
作者:Anton Jeran Ratnarajah
备注:PhD thesis
链接:点击下载PDF文件
【15】 An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
标题: 用于相重建和语音增强的显式一致性保持损失函数
作者:Pin-Jui Ku,Chun-Wei Ho,Hao Yen,Sabato Marco Siniscalchi,Chin-Hui Lee
备注:5 pages, Submitted to ICASSP 2025
链接:点击下载PDF文件
【16】 Evaluation of state-of-the-art ASR Models in Child-Adult Interactions
标题: 儿童与成人互动中最先进的ASB模型的评估
作者:Aditya Ashvin,Rimita Lahiri,Aditya Kommineni,Somer Bishop,Catherine Lord,Sudarsana Reddy Kadiri,Shrikanth Narayanan
备注:5 pages, 3 figures, 4 tables
链接:点击下载PDF文件
【17】 Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
标题: 生成性语音基础模型预训练用于高质量语音提取和恢复
作者:Pin-Jui Ku,Alexander H. Liu,Roman Korostik,Sung-Feng Huang,Szu-Wei Fu,Ante Jukić
备注:5 pages, Submitted to ICASSP 2025. The implementation and configuration could be found in this https URL The audio demo page could be found in this https URL
链接:点击下载PDF文件
【18】 Scenario of Use Scheme: Threat Model Specification for Speaker Privacy Protection in the Medical Domain
标题: 使用场景方案:医疗领域演讲者隐私保护的威胁模型规范
作者:Mehtab Ur Rahman,Martha Larson,Louis ten Bosch,Cristian Tejedor-García
备注:Accepted and published at SPSC Symposium 2024 4th Symposium on Security and Privacy in Speech Communication. Interspeech 2024
链接:点击下载PDF文件
【19】 StyleSinger 2: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
标题: StyleSinger 2:具有风格转移和多层风格控制的Zero-Shot歌唱声音合成
作者:Yu Zhang,Ziyue Jiang,Ruiqi Li,Changhao Pan,Jinzheng He,Rongjie Huang,Chuxin Wang,Zhou Zhao
备注:Accepted by EMNLP 2024
链接:点击下载PDF文件
【20】 ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
标题: ESPnet-Codec:音频、音乐和语音神经编解码器的全面训练和评估
作者:Jiatong Shi,Jinchuan Tian,Yihan Wu,Jee-weon Jung,Jia Qi Yip,Yoshiki Masuyama,William Chen,Yuning Wu,Yuxun Tang,Massa Baali,Dareen Alharhi,Dong Zhang,Ruifan Deng,Tejes Srivastava,Haibin Wu,Alexander H. Liu,Bhiksha Raj,Qin Jin,Ruihua Song,Shinji Watanabe
备注:Accepted by SLT
链接:点击下载PDF文件
【21】 Interpolation filter design for sample rate independent audio effect RNNs
标题: 独立于采样率的音频效果RNN的内插过滤器设计
作者:Alistair Carson,Alec Wright,Stefan Bilbao
链接:点击下载PDF文件
【22】 Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
标题: 美杜莎耳朵里的低语:基于变形器的ASB的多头高效解码
作者:Yael Segal-Feldman,Aviv Shamsian,Aviv Navon,Gill Hetz,Joseph Keshet
备注:Under Review
链接:点击下载PDF文件
【23】 WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
标题: WeSep:一个可扩展且灵活的工具包,可推广的目标说话者提取
作者:Shuai Wang,Ke Zhang,Shaoxiong Lin,Junjie Li,Xuefei Wang,Meng Ge,Jianwei Yu,Yanmin Qian,Haizhou Li
备注:Interspeech 2024
链接:点击下载PDF文件
【24】 M-Vec: Matryoshka Speaker Embeddings with Flexible Dimensions
标题: M-Vec:具有灵活尺寸的Matryoshka扬声器嵌入式
作者:Shuai Wang,Pengcheng Zhu,Haizhou Li
备注:ICSR 2024, Shenzhen
链接:点击下载PDF文件
【25】 Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
标题: 利用随机选择策略最小化表示损失以实现高效的环境假音频检测
作者:Orchid Chetia Phukan,Girish,Mohd Mujtaba Akhtar,Swarup Ranjan Behera,Nitin Choudhury,Arun Balaji Buduru,Rajesh Sharma,S. R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【26】 Enhancing Open-Set Speaker Identification through Rapid Tuning with Speaker Reciprocal Points and Negative Sample
标题: 通过使用说话者互反点和负样本快速调整来增强开集说话者识别
作者:Zhiyong Chen,Zhiqi Ai,Xinnuo Li,Shugong Xu
Journal-ref:IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
【27】 StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
标题: StyleFusion TTC:用于Zero-Shot文本到语音合成的多模式风格控制和增强的特征融合
作者:Zhiyong Chen,Xinnuo Li,Zhiqi Ai,Shugong Xu
Journal-ref:The 7th Chinese Conference on Pattern Recognition and Computer Vision PRCV 2024
链接:点击下载PDF文件
【28】 Language-based Audio Moment Retrieval
标题: 基于百分比的音频时刻检索
作者:Hokuto Munakata,Taichi Nishimura,Shota Nakada,Tatsuya Komatsu
链接:点击下载PDF文件
【29】 Safe Guard: an LLM-agent for Real-time Voice-based Hate Speech Detection in Social Virtual Reality
标题: Safe Guard:用于社交虚拟现实中实时基于语音的仇恨言语检测的LLM代理
作者:Yiwen Xu,Qinyang Hou,Hongyu Wan,Mirjana Prpa
链接:点击下载PDF文件
【30】 Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
标题: 修改、推理和识别:通过特定描述和ASB错误纠正的基于LLM的情感识别
作者:Yuanchao Li,Yuan Gong,Chao-Han Huck Yang,Peter Bell,Catherine Lai
链接:点击下载PDF文件
【31】 Rethinking Emotion Bias in Music via Frechet Audio Distance
标题: 通过Frechet音频距离重新思考音乐中的情感偏见
作者:Yuanchao Li,Azalea Gui,Dimitra Emmanouilidou,Hannes Gamper
链接:点击下载PDF文件
【32】 Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
标题: Speech 2 rtMRI:语音引导扩散模型,用于语音期间人声实时MRI视频
作者:Hong Nguyen,Sean Foley,Kevin Huang,Xuan Shi,Tiantian Feng,Shrikanth Narayanan
备注:4 pages
链接:点击下载PDF文件
【33】 Blind Localization of Early Room Reflections with Arbitrary Microphone Array
标题: 用任意麦克风阵列实现早期房间反射的盲定位
作者:Yogev Hadadi,Vladimir Tourbabin,Zamir Ben-Hur,David Lou Alon,Boaz Rafaely
链接:点击下载PDF文件
【34】 The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
标题: 议会议事录中自动生成的语音和文本数据集的ParlaSpeech集合
作者:Nikola Ljubešić,Peter Rupnik,Danijel Koržinek
备注:Submitted to SPECOM 2024
链接:点击下载PDF文件
【35】 Toward Automated Clinical Transcriptions
标题: 迈向自动化临床Transit
作者:Mitchell A. Klusty,W. Vaiden Logan,Samuel E. Armstrong,Aaron D. Mullen,Caroline N. Leach,Jeff Talbert,V. K. Cody Bumgardner
备注:7 pages, 6 figures
链接:点击下载PDF文件
【36】 A Joint Spectro-Temporal Relational Thinking Based Acoustic Modeling Framework
标题: 基于频谱-时间关系思维的联合声学建模框架
作者:Zheng Nan,Ting Dang,Vidhyasaharan Sethu,Beena Ahmed
链接:点击下载PDF文件
【37】 TCG CREST System Description for the Second DISPLACE Challenge
标题: 第二次DISPLACE挑战赛的TCG CREST系统描述
作者:Nikhil Raghav,Subhajit Saha,Md Sahidullah,Swagatam Das
链接:点击下载PDF文件
【38】 Contextualization of ASR with LLM using phonetic retrieval-based augmentation
标题: 使用基于语音检索的增强将ASC与LLM进行上下文化
作者:Zhihong Lei,Xingyu Na,Mingbin Xu,Ernest Pusateri,Christophe Van Gysel,Yuanyuan Zhang,Shiyi Han,Zhen Huang
链接:点击下载PDF文件
【39】 Equivariance-based self-supervised learning for audio signal recovery from clipped measurements
标题: 基于等效性的自我监督学习,从剪辑测量中恢复音频信号
作者:Victor Sechaud,Laurent Jacques,Patrice Abry,Julián Tachella
Journal-ref:EUSIPCO, Aug 2024, Lyon, France
链接:点击下载PDF文件
标题: 面部表情增强的TTC:结合面部表示和情感强度实现自适应语音
作者:Yunji Chu,Yunseob Shim,Unsang Park
备注:13 pages, 3 figures, accepted to ECCV Workshop ABAW(Affective Behavior Analysis in-the-wild)7 (to be appear)
链接:点击下载PDF文件
摘要:我们提出FEIM-TTS,一个创新的zero-shot文本到语音(TTS)模型,合成情感表达的语音,与面部图像对齐,并通过情绪强度调制。利用深度学习,FEIM-TTS超越了传统的TTS系统,通过解释面部线索和调整情绪的细微差别,而不依赖于标记的数据集。为了解决稀疏的视听情感数据,该模型使用LRS 3,CREMA-D和MELD数据集进行训练,证明了其适应性。FEIM-TTS的独特能力,以产生高品质,说话者不可知的语音,使其适合创造适应性的声音,为虚拟人物。此外,FEIM-TTS大大提高了视力障碍者或视力有问题的人的可访问性。通过将情感的细微差别整合到TTS中,我们的模型为网络漫画提供了动态和引人入胜的听觉体验,使视障用户能够更充分地享受这些叙事。综合评价表明,它在调节情感和强度,提高情感语音合成和可达性方面表现出色。样品可在以下网址获得:https: feim-tts.github.io 。摘要:We propose FEIM-TTS, an innovative zero-shot text-to-speech (TTS) model that synthesizes emotionally expressive speech, aligned with facial images and modulated by emotion intensity. Leveraging deep learning, FEIM-TTS transcends traditional TTS systems by interpreting facial cues and adjusting to emotional nuances without dependence on labeled datasets. To address sparse audio-visual-emotional data, the model is trained using LRS3, CREMA-D, and MELD datasets, demonstrating its adaptability. FEIM-TTS's unique capability to produce high-quality, speaker-agnostic speech makes it suitable for creating adaptable voices for virtual characters. Moreover, FEIM-TTS significantly enhances accessibility for individuals with visual impairments or those who have trouble seeing. By integrating emotional nuances into TTS, our model enables dynamic and engaging auditory experiences for webcomics, allowing visually impaired users to enjoy these narratives more fully. Comprehensive evaluation evidences its proficiency in modulating emotion and intensity, advancing emotional speech synthesis and accessibility. Samples are available at: https: feim-tts.github.io .
【2】 Leveraging Mixture of Experts for Improved Speech Deepfake Detection
标题: 利用专家混合改进语音深度伪造检测
作者:Viola Negroni,Davide Salvi,Alessandro Ilic Mezza,Paolo Bestagini,Stefano Tubaro
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:语音深度伪造对个人安全和内容真实性构成重大威胁。文献中已经提出了几种检测器,这些系统必须面对的主要挑战之一是对看不见的数据进行泛化,以识别各种数据集中的假信号。在本文中,我们介绍了一种使用混合专家架构来增强语音深度伪造检测性能的新方法。专家混合框架非常适合语音深度伪造检测任务,因为它能够专门处理不同的输入类型并有效地处理数据变化。与传统的单一模型或集成方法相比,这种方法对未知数据具有更好的泛化能力和适应性。此外,它的模块化结构支持可扩展的更新,使其在管理深度伪造技术不断变化的复杂性方面更加灵活,同时保持高检测精度。我们提出了一个高效,轻量级的门控机制,动态分配专家的权重为每个输入,优化检测性能。在多个数据集上的实验结果证明了我们所提出的方法的有效性和潜力。摘要:Speech deepfakes pose a significant threat to personal security and content authenticity. Several detectors have been proposed in the literature, and one of the primary challenges these systems have to face is the generalization over unseen data to identify fake signals across a wide range of datasets. In this paper, we introduce a novel approach for enhancing speech deepfake detection performance using a Mixture of Experts architecture. The Mixture of Experts framework is well-suited for the speech deepfake detection task due to its ability to specialize in different input types and handle data variability efficiently. This approach offers superior generalization and adaptability to unseen data compared to traditional single models or ensemble methods. Additionally, its modular structure supports scalable updates, making it more flexible in managing the evolving complexity of deepfake techniques while maintaining high detection accuracy. We propose an efficient, lightweight gating mechanism to dynamically assign expert weights for each input, optimizing detection performance. Experimental results across multiple datasets demonstrate the effectiveness and potential of our proposed approach.
【3】 Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
标题: 弥合语音和文本:通过LLM中的拼音到字符预训练增强ASB
作者:Yang Yuhang,Peng Yizhou,Eng Siong Chng,Xionghu Zhong
备注:Accepted by ISCSLP2024-Special session-Speech Processing in LLM Era
链接:点击下载PDF文件
摘要:大型语言模型(LLM)与预训练语音模型的集成为自动语音识别(ASR)开辟了新的途径。虽然LLM擅长多模态理解任务,但有效利用他们的ASR能力仍然是一个重大挑战。本文提出了一种新颖的培训方法来提高LLM在ASR任务中的表现。我们提出了预训练LLM拼音嵌入序列,这代表发音特征,以产生相应的汉字。该步骤使得LLM能够在遇到真实语音数据之前适应于从发音特征生成文本。此外,我们微调的LoRA参数,以提高LLM的语音模态信息的理解。在AISHELL-1语料库中,与没有拼音到字符预训练的基线相比,我们的方法在ASR任务中产生了9.5%的相对改善。此外,将辅助文本数据用于拼音到字符的预训练进一步提高了性能,实现了19.0%的相对提高。摘要:The integration of large language models (LLMs) with pre-trained speech models has opened up new avenues in automatic speech recognition (ASR). While LLMs excel in multimodal understanding tasks, effectively leveraging their capabilities for ASR remains a significant challenge. This paper presents a novel training approach to enhance LLM performance in ASR tasks. We propose pre-training LLMs on Pinyin embedding sequences, which represent pronunciation features, to generate corresponding Chinese characters. This step enables the LLM to adapt to generating text from pronunciation features before encountering real speech data. Furthermore, we fine-tune the LoRA parameters to enhance the LLM's understanding of speech modality information. In AISHELL-1 corpus, our approach yields a 9.5% relative improvement in ASR tasks compared to the baseline without Pinyi-to-Character pre-training. Additionally, incorporating auxiliary text data for Pinyi-to-Character pre-training further boosts performance, achieving a 19.0% relative improvement.
【4】 Disentangling Age and Identity with a Mutual Information Minimization Approach for Cross-Age Speaker Verification
标题: 用互信息最小化方法解开年龄和身份,用于跨年龄段说话人验证
作者:Fengrun Zhang,Wangjin Zhou,Yiming Liu,Wang Geng,Yahui Shan,Chen Zhang
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:跨年龄说话人确认(CASV)是近年来的研究热点。然而,现有的说话人确认系统在CASV中表现不佳,这是由于年龄造成的语音个体差异很大。本文提出了一个基于互信息最小化的CASV表示学习框架。在我们的方法中,骨干模型被训练来从说话人信息中分离身份和年龄相关的嵌入,并且MI估计器被训练来通过MI最小化来最小化年龄和身份相关的嵌入之间的相关性,从而产生年龄不变的说话人嵌入。此外,通过使用积极和消极的样本之间的年龄差距,我们提出了一个老化意识的MI最小化损失函数,使骨干模型更专注于声音的变化与大的年龄差距。实验结果表明,该方法在Vox-CA的多个跨年龄测试集上的性能优于其他方法。摘要:There has been an increasing research interest in cross-age speaker verification~(CASV). However, existing speaker verification systems perform poorly in CASV due to the great individual differences in voice caused by aging. In this paper, we propose a disentangled representation learning framework for CASV based on mutual information~(MI) minimization. In our method, a backbone model is trained to disentangle the identity- and age-related embeddings from speaker information, and an MI estimator is trained to minimize the correlation between age- and identity-related embeddings via MI minimization, resulting in age-invariant speaker embeddings. Furthermore, by using the age gaps between positive and negative samples, we propose an aging-aware MI minimization loss function that allows the backbone model to focus more on the vocal changes with large age gaps. Experimental results show that the proposed method outperforms other methods on multiple Cross-Age test sets of Vox-CA.
【5】 ASD-Diffusion: Anomalous Sound Detection with Diffusion Models
标题: ASD扩散:使用扩散模型检测异常声音
作者:Fengrun Zhang,Xiang Xie,Kai Guo
备注:This paper will appear at ICPR 2024
链接:点击下载PDF文件
摘要:无监督异常声音检测(ASD)的目的是设计一种通用的方法,可以用来检测异常时,只有正常的声音。本文提出了一种基于扩散模型的异常声音检测方法(ASD-扩散),用于实际工厂中的ASD。在我们的流水线中,声学特征中的异常从其噪声损坏的特征重建成其近似正常的模式。其次,提出了一种后处理异常滤波算法,用于检测重构后与原始输入存在显著偏差的异常。此外,引入去噪扩散隐式模型,通过延长去噪过程的采样间隔,加快了推理速度。该方法是扩散模型应用的一种新方法。在DCASE 2023挑战任务2的开发集上的实验结果比基线高出7.75%,证明了该方法的有效性。摘要:Unsupervised Anomalous Sound Detection (ASD) aims to design a generalizable method that can be used to detect anomalies when only normal sounds are given. In this paper, Anomalous Sound Detection based on Diffusion Models (ASD-Diffusion) is proposed for ASD in real-world factories. In our pipeline, the anomalies in acoustic features are reconstructed from their noisy corrupted features into their approximate normal pattern. Secondly, a post-processing anomalies filter algorithm is proposed to detect anomalies that exhibit significant deviation from the original input after reconstruction. Furthermore, denoising diffusion implicit model is introduced to accelerate the inference speed by a longer sampling interval of the denoising process. The proposed method is innovative in the application of diffusion models as a new scheme. Experimental results on the development set of DCASE 2023 challenge task 2 outperform the baseline by 7.75%, demonstrating the effectiveness of the proposed method.
【6】 A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation
标题: 基于模块的语音同步翻译中梯度冲突缓解策略
作者:Xiaoqian Liu,Yangfan Du,Jianjin Wang,Yuan Ge,Chen Xu,Tong Xiao,Guocheng Chen,Jingbo Zhu
链接:点击下载PDF文件
摘要:同步语音翻译(SimulST)涉及在连续处理流式语音输入的同时生成目标语言文本,这带来了重大的实时挑战。多任务学习通常用于增强SimulST性能,但会在主任务和辅助任务之间引入优化冲突,从而可能影响整体效率。现有的模型级冲突解决方法并不适合此任务,这加剧了效率低下并导致高GPU内存消耗。为了解决这些挑战,我们提出了一个模块化梯度冲突缓解(MGCM)的策略,检测冲突在一个更细粒度的模块化水平,并解决它们利用梯度投影。实验结果表明,MGCM显着提高SimulST的性能,特别是在中等和高延迟条件下,实现了0.68 BLEU分数增益离线任务。此外,与其他冲突缓解方法相比,MGCM将GPU内存消耗减少了95%以上,使其成为SimulST任务的强大解决方案。摘要:Simultaneous Speech Translation (SimulST) involves generating target language text while continuously processing streaming speech input, presenting significant real-time challenges. Multi-task learning is often employed to enhance SimulST performance but introduces optimization conflicts between primary and auxiliary tasks, potentially compromising overall efficiency. The existing model-level conflict resolution methods are not well-suited for this task which exacerbates inefficiencies and leads to high GPU memory consumption. To address these challenges, we propose a Modular Gradient Conflict Mitigation (MGCM) strategy that detects conflicts at a finer-grained modular level and resolves them utilizing gradient projection. Experimental results demonstrate that MGCM significantly improves SimulST performance, particularly under medium and high latency conditions, achieving a 0.68 BLEU score gain in offline tasks. Additionally, MGCM reduces GPU memory consumption by over 95 % compared to other conflict mitigation methods, establishing it as a robust solution for SimulST tasks.
【7】 Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
标题: 通过混合专家提高代码转换ASB增强的言语条件LLM
作者:Fengrun Zhang,Wang Geng,Hukai Huang,Cheng Yi,He Qu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在本文中,我们介绍了一个语音条件的大语言模型(LLM)集成的混合专家(MoE)为基础的连接器,以解决自动语音识别(ASR)中的代码切换(CS)的挑战。具体来说,我们提出了一个插入和删除中断令牌(IDIT)机制,以更好地转移LLM的文本生成能力的语音识别任务。我们还提出了一个连接器与MoE架构,有效地管理多种语言。为了进一步增强多个专家的协作并利用LLM的理解能力,我们提出了一种两阶段渐进式训练策略:1)连接器被解冻并与语言专业专家一起训练,以将语音表示映射到文本空间。2)连接器和LLM LoRA适配器使用建议的IDIT机制进行培训,所有专家都被激活以学习一般表示。实验结果表明,我们的方法显着优于国家的最先进的模型,包括端到端和大规模的音频语言模型。摘要:In this paper, we introduce a speech-conditioned Large Language Model (LLM) integrated with a Mixture of Experts (MoE) based connector to address the challenge of Code-Switching (CS) in Automatic Speech Recognition (ASR). Specifically, we propose an Insertion and Deletion of Interruption Token (IDIT) mechanism for better transfer text generation ability of LLM to speech recognition task. We also present a connecter with MoE architecture that manages multiple languages efficiently. To further enhance the collaboration of multiple experts and leverage the understanding capabilities of LLM, we propose a two-stage progressive training strategy: 1) The connector is unfrozen and trained with language-specialized experts to map speech representations to the text space. 2) The connector and LLM LoRA adaptor are trained with the proposed IDIT mechanism and all experts are activated to learn general representations. Experimental results demonstrate that our method significantly outperforms state-of-the-art models, including end-to-end and large-scale audio-language models.
【8】 On the calibration of powerset speaker diarization models
标题: 关于Powerset扬声器Dialogue模型的校准
作者:Alexis Plaquet,Hervé Bredin
Journal-ref:Interspeech 2024, Sep 2024, Kos, Greece. pp.3764-3768, ⟨10.21437Interspeech.2024-1060⟩
链接:点击下载PDF文件
摘要:端到端神经日志模型通常依赖于说话人日志问题的多标签分类公式。最近,我们提出了一个powerset多类公式,它在多个数据集上击败了最先进的技术。在本文中,我们提出了研究的幂集扬声器日记模型的校准,并探讨其使用的一些。我们研究了域内和域外的校准,并探索低置信度区域中的数据。然后在实践中测试模型置信度的可靠性:我们使用预训练模型的置信度来选择性地从未注释的数据中创建训练和验证子集,并将其与随机选择进行比较。我们发现,顶标签置信度可以用来可靠地预测高误差区域。此外,在低置信度区域上进行训练可以提供更好的校准模型,并且在低置信度区域上进行验证可以比随机区域更具注释效率。摘要:End-to-end neural diarization models have usually relied on a multilabel-classification formulation of the speaker diarization problem. Recently, we proposed a powerset multiclass formulation that has beaten the state-of-the-art on multiple datasets. In this paper, we propose to study the calibration of a powerset speaker diarization model, and explore some of its uses. We study the calibration in-domain, as well as out-of-domain, and explore the data in low-confidence regions. The reliability of model confidence is then tested in practice: we use the confidence of the pretrained model to selectively create training and validation subsets out of unannotated data, and compare this to random selection. We find that top-label confidence can be used to reliably predict high-error regions. Moreover, training on low-confidence regions provides a better calibrated model, and validating on low-confidence regions can be more annotation-efficient than random regions.
【9】 NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
标题: NanoVoice:适合多个扬声器的高效扬声器自适应文本到语音
作者:Nohil Park,Heeseung Kim,Che Hyun Lee,Jooyoung Choi,Jiheum Yeom,Sungroh Yoon
备注:Submitted to ICASSP 2025, Demo Page: this https URL
链接:点击下载PDF文件
摘要:我们提出了NanoVoice,一个个性化的文本到语音的模型,有效地构建语音适配器,同时为多个扬声器。NanoVoice引入了一种批量扬声器自适应技术,能够并行微调多个参考,显著减少训练时间。除了为每个扬声器构建单独的适配器之外,我们还提出了一种参数共享技术,该技术减少了用于扬声器自适应的参数数量。通过整合一个新的可训练的尺度矩阵,NanoVoice减轻了参数共享过程中潜在的性能下降。NanoVoice实现了与基线相当的性能,同时训练速度提高了4倍,使用40个参考语音进行扬声器自适应的参数减少了45%。广泛的消融研究和分析进一步验证了我们的模型的效率。摘要:We present NanoVoice, a personalized text-to-speech model that efficiently constructs voice adapters for multiple speakers simultaneously. NanoVoice introduces a batch-wise speaker adaptation technique capable of fine-tuning multiple references in parallel, significantly reducing training time. Beyond building separate adapters for each speaker, we also propose a parameter sharing technique that reduces the number of parameters used for speaker adaptation. By incorporating a novel trainable scale matrix, NanoVoice mitigates potential performance degradation during parameter sharing. NanoVoice achieves performance comparable to the baselines, while training 4 times faster and using 45 percent fewer parameters for speaker adaptation with 40 reference voices. Extensive ablation studies and analysis further validate the efficiency of our model.
【10】 VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance
标题: VoiceGuider:通过自动引导增强参数高效扬声器自适应文本到语音的域外性能
作者:Jiheum Yeom,Heeseung Kim,Jooyoung Choi,Che Hyun Lee,Nohil Park,Sungroh Yoon
备注:Submitted to ICASSP 2025, Demo Page: this https URL
链接:点击下载PDF文件
摘要:当通过LoRA将参数高效微调应用于说话人自适应文本到语音模型时,与完全微调的对应物相比,自适应性能可能会下降,特别是对于域外说话人。在这里,我们提出了VoiceGuider,一个参数高效的扬声器自适应文本到语音系统,通过自动指导来增强扬声器自适应性能,减少与全微调模型的差距。我们仔细探索各种加强自动导航的方法,最终找到最佳策略。VoiceGuider作为结果显示了强大的适应性能,特别是在极端的域外语音数据。我们在演示页面中提供音频样本。摘要:When applying parameter-efficient finetuning via LoRA onto speaker adaptive text-to-speech models, adaptation performance may decline compared to full-finetuned counterparts, especially for out-of-domain speakers. Here, we propose VoiceGuider, a parameter-efficient speaker adaptive text-to-speech system reinforced with autoguidance to enhance the speaker adaptation performance, reducing the gap against full-finetuned models. We carefully explore various ways of strengthening autoguidance, ultimately finding the optimal strategy. VoiceGuider as a result shows robust adaptation performance especially on extreme out-of-domain speech data. We provide audible samples in our demo page.
【11】 Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
标题: 假设聚集和合并:具有说话者令牌的新型多说话者语音识别
作者:Yosuke Kashiwagi,Hayato Futami,Emiru Tsunoo,Siddhant Arora,Shinji Watanabe
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在许多现实世界的场景中,例如会议,多个发言者与未知数量的参与者一起出现,并且他们的话语经常重叠。我们解决了这些多扬声器的挑战,一种新的基于注意力的编码器-解码器的方法与扬声器聚类获得的特殊扬声器类令牌增强。在推理过程中,我们选择多个识别假设条件下的预测说话人集群令牌,这些假设合并凝聚层次聚类(AHC)的基础上归一化的编辑距离。聚类假设的结果在多说话人transmittance与AHC确定的适当数量的发言人。我们在LibriMix数据集上的实验表明,我们提出的方法在复杂的3混合环境中特别有效,与传统的序列化输出训练相比,在干净数据上实现了55%的相对误差减少,在嘈杂数据上实现了36%的相对误差减少。摘要:In many real-world scenarios, such as meetings, multiple speakers are present with an unknown number of participants, and their utterances often overlap. We address these multi-speaker challenges by a novel attention-based encoder-decoder method augmented with special speaker class tokens obtained by speaker clustering. During inference, we select multiple recognition hypotheses conditioned on predicted speaker cluster tokens, and these hypotheses are merged by agglomerative hierarchical clustering (AHC) based on the normalized edit distance. The clustered hypotheses result in the multi-speaker transcriptions with the appropriate number of speakers determined by AHC. Our experiments on the LibriMix dataset demonstrate that our proposed method was particularly effective in complex 3-mix environments, achieving a 55% relative error reduction on clean data and a 36% relative error reduction on noisy data compared with conventional serialized output training.
【12】 Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
标题: 超越回合制接口:同步LLM作为全速对话代理
作者:Bandhav Veluri,Benjamin N Peloquin,Bokai Yu,Hongyu Gong,Shyamnath Gollakota
备注:EMNLP Main 2024
链接:点击下载PDF文件
摘要:尽管对口语对话代理建模有广泛的兴趣,但大多数方法本质上是“半双工”的-仅限于基于回合的交互,其响应需要用户的显式提示或对中断或沉默事件的隐式跟踪。相比之下,人类的对话是“全双工”的,允许以快速和动态的话轮转换,重叠的演讲和反向引导的形式实现丰富的同步性。从技术上讲,实现与LLM的全双工对话的挑战在于对同步进行建模,因为预先训练的LLM没有“时间”感。为了弥合这一差距,我们提出了同步LLM全双工口语对话建模。我们设计了一种新的机制,将时间信息集成到Llama 3 -8b中,使它们与现实世界的时钟同步运行。我们还介绍了一个训练配方,它使用从文本对话数据生成的212 k小时的合成口语对话数据来创建一个模型,该模型可以生成有意义和自然的口语对话,而只有2k小时的真实口语对话数据。同步LLM在保持自然性的同时,在对话的意义上优于最先进的水平。最后,我们证明了模型的能力,通过模拟两个代理之间的互动,在不同的数据集上训练,同时考虑互联网规模的延迟高达240毫秒。https: syncllm.cs.washington.edu 摘要:Despite broad interest in modeling spoken dialogue agents, most approaches are inherently "half-duplex" -- restricted to turn-based interaction with responses requiring explicit prompting by the user or implicit tracking of interruption or silence events. Human dialogue, by contrast, is "full-duplex" allowing for rich synchronicity in the form of quick and dynamic turn-taking, overlapping speech, and backchanneling. Technically, the challenge of achieving full-duplex dialogue with LLMs lies in modeling synchrony as pre-trained LLMs do not have a sense of "time". To bridge this gap, we propose Synchronous LLMs for full-duplex spoken dialogue modeling. We design a novel mechanism to integrate time information into Llama3-8b so that they run synchronously with the real-world clock. We also introduce a training recipe that uses 212k hours of synthetic spoken dialogue data generated from text dialogue data to create a model that generates meaningful and natural spoken dialogue, with just 2k hours of real-world spoken dialogue data. Synchronous LLMs outperform state-of-the-art in dialogue meaningfulness while maintaining naturalness. Finally, we demonstrate the model's ability to participate in full-duplex dialogue by simulating interaction between two agents trained on different datasets, while considering Internet-scale latencies of up to 240 ms. Webpage: https: syncllm.cs.washington.edu .
【13】 Generalization in birdsong classification: impact of transfer learning methods and dataset characteristics
标题: 鸟鸣分类的推广:迁移学习方法和数据集特征的影响
作者:Burooj Ghani,Vincent J. Kalkman,Bob Planqué,Willem-Pier Vellinga,Lisa Gill,Dan Stowell
备注:25 pages
链接:点击下载PDF文件
摘要:机器学习可以自动识别动物的声音,这在生物多样性监测中发挥着重要作用。然而,尽管越来越令人印象深刻的能力,生物声学物种分类器仍然表现出跨物种和栖息地的不平衡性能,特别是在复杂的音景。在本研究中,我们探索了迁移学习在各种条件下(包括单标签和多标签场景)以及不同模型架构(例如CNN和Transformers)的大规模鸟类声音分类中的有效性。我们的实验表明,微调和知识蒸馏产生强大的性能,交叉蒸馏证明特别有效地提高在领域内的性能Xeno-canto数据。然而,当推广到音景时,与知识蒸馏相比,浅层微调表现出更优越的性能,突出了其鲁棒性和受约束的性质。我们的研究进一步探讨了如何使用多物种标签,在这些情况下,存在但不完整。我们提倡在动物声音社区内进行更全面的标记实践,包括注释背景物种和提供时间细节,以加强对强大的鸟类声音分类器的训练。这些发现为预训练模型的最佳重用提供了见解,以推进自动生物声学识别。摘要:Animal sounds can be recognised automatically by machine learning, and this has an important role to play in biodiversity monitoring. Yet despite increasingly impressive capabilities, bioacoustic species classifiers still exhibit imbalanced performance across species and habitats, especially in complex soundscapes. In this study, we explore the effectiveness of transfer learning in large-scale bird sound classification across various conditions, including single- and multi-label scenarios, and across different model architectures such as CNNs and Transformers. Our experiments demonstrate that both fine-tuning and knowledge distillation yield strong performance, with cross-distillation proving particularly effective in improving in-domain performance on Xeno-canto data. However, when generalizing to soundscapes, shallow fine-tuning exhibits superior performance compared to knowledge distillation, highlighting its robustness and constrained nature. Our study further investigates how to use multi-species labels, in cases where these are present but incomplete. We advocate for more comprehensive labeling practices within the animal sound community, including annotating background species and providing temporal details, to enhance the training of robust bird sound classifiers. These findings provide insights into the optimal reuse of pretrained models for advancing automatic bioacoustic recognition.
【14】 Efficient learning-based sound propagation for virtual and real-world audio processing applications
标题: 针对虚拟和现实世界音频处理应用程序的高效基于学习的声音传播
作者:Anton Jeran Ratnarajah
备注:PhD thesis
链接:点击下载PDF文件
摘要:声传播是声能通过介质(例如空气)以声波形式传播到周围环境的过程。房间脉冲响应(RIR)描述了这一过程,并受到声源和听众的位置、房间的几何形状及其材料的影响。几十年来,基于物理的声学模拟器一直用于计算特定声学环境的精确RIR。然而,我们遇到了现有的声学模拟器的局限性。为了解决这些问题,我们提出了三个新的解决方案。首先,我们介绍了一个基于学习的RIR生成器,它比交互式光线跟踪模拟器快两个数量级。我们的方法可以被训练为直接输入统计和传统参数,并且它可以为重建和合成的3D场景生成单声道和双耳RIR。我们生成的RIR在语音处理应用中的性能优于交互式光线跟踪模拟器,包括ASR,语音增强和语音分离。其次,我们提出了估计RIR混响语音信号和视觉线索,没有一个3D表示的环境。通过从混响语音中估计RIR,我们可以增加训练数据以匹配测试数据,从而改善ASR系统的单词错误率。在远场ASR任务中,我们估计的RIR比以前基于学习的RIR估计器提高了6.9%。我们证明,我们的视听RIR估计艾滋病的任务,如视觉声学匹配,新颖的视图声学合成,语音配音,通过感知评估验证。最后,我们引入IR-GAN来使用真实的RIR来增强准确的RIR。IR-GAN参数化控制从真实RIR中学习的声学参数,以生成模拟不同声学环境的新RIR,在远场ASR基准测试中,其性能优于光线跟踪模拟器8.95%。摘要:Sound propagation is the process by which sound energy travels through a medium, such as air, to the surrounding environment as sound waves. The room impulse response (RIR) describes this process and is influenced by the positions of the source and listener, the room's geometry, and its materials. Physics-based acoustic simulators have been used for decades to compute accurate RIRs for specific acoustic environments. However, we have encountered limitations with existing acoustic simulators. To address these limitations, we propose three novel solutions. First, we introduce a learning-based RIR generator that is two orders of magnitude faster than an interactive ray-tracing simulator. Our approach can be trained to input both statistical and traditional parameters directly, and it can generate both monaural and binaural RIRs for both reconstructed and synthetic 3D scenes. Our generated RIRs outperform interactive ray-tracing simulators in speech-processing applications, including ASR, Speech Enhancement, and Speech Separation. Secondly, we propose estimating RIRs from reverberant speech signals and visual cues without a 3D representation of the environment. By estimating RIRs from reverberant speech, we can augment training data to match test data, improving the word error rate of the ASR system. Our estimated RIRs achieve a 6.9% improvement over previous learning-based RIR estimators in far-field ASR tasks. We demonstrate that our audio-visual RIR estimator aids tasks like visual acoustic matching, novel-view acoustic synthesis, and voice dubbing, validated through perceptual evaluation. Finally, we introduce IR-GAN to augment accurate RIRs using real RIRs. IR-GAN parametrically controls acoustic parameters learned from real RIRs to generate new RIRs that imitate different acoustic environments, outperforming Ray-tracing simulators on the far-field ASR benchmark by 8.95%.
【15】 An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
标题: 用于相重建和语音增强的显式一致性保持损失函数
作者:Pin-Jui Ku,Chun-Wei Ho,Hao Yen,Sabato Marco Siniscalchi,Chin-Hui Lee
备注:5 pages, Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在这项工作中,我们提出了一种新的一致性保持损失函数恢复相位重建(PR)和语音增强(SE)的背景下的相位信息。与使用深度模型直接估计相位的传统技术不同,我们的想法是利用ad-hoc约束直接生成一致的幅度和相位对。具体地,所提出的损失迫使一组复数成为一致的短时傅立叶变换(STFT)表示,即,是真实信号的谱图。因此,我们的方法避免了估计原始相位的困难,这是高度非结构化和敏感的时移。我们提出的损失的影响,首先评估PR任务,实验证明,我们的方法是可行的。接下来,我们使用VB-DMD和WSJ 0-CHiME 3数据集展示了其在SE任务上的有效性。在VB-DMD上,我们的方法与传统解决方案相比具有竞争力。在具有挑战性的WSJ 0-CHiME 3集上,所提出的框架与显式估计相位的技术相比是有利的。摘要:In this work, we propose a novel consistency-preserving loss function for recovering the phase information in the context of phase reconstruction (PR) and speech enhancement (SE). Different from conventional techniques that directly estimate the phase using a deep model, our idea is to exploit ad-hoc constraints to directly generate a consistent pair of magnitude and phase. Specifically, the proposed loss forces a set of complex numbers to be a consistent short-time Fourier transform (STFT) representation, i.e., to be the spectrogram of a real signal. Our approach thus avoids the difficulty of estimating the original phase, which is highly unstructured and sensitive to time shift. The influence of our proposed loss is first assessed on a PR task, experimentally demonstrating that our approach is viable. Next, we show its effectiveness on an SE task, using both the VB-DMD and WSJ0-CHiME3 data sets. On VB-DMD, our approach is competitive with conventional solutions. On the challenging WSJ0-CHiME3 set, the proposed framework compares favourably over those techniques that explicitly estimate the phase.
【16】 Evaluation of state-of-the-art ASR Models in Child-Adult Interactions
标题: 儿童与成人互动中最先进的ASB模型的评估
作者:Aditya Ashvin,Rimita Lahiri,Aditya Kommineni,Somer Bishop,Catherine Lord,Sudarsana Reddy Kadiri,Shrikanth Narayanan
备注:5 pages, 3 figures, 4 tables
链接:点击下载PDF文件
摘要:在临床环境中可靠地转录儿童-成人对话的能力对于诊断和理解许多发育障碍如自闭症谱系障碍是有价值的。深度学习架构的最新进展和大规模转录数据的可用性已经导致语音基础模型的开发,这些模型在ASR性能方面表现出显着的改进。然而,这些模型的能力,以及转化为对话的儿童与成人的互动正在研究中。在这项工作中,我们使用Whisper,Wav 2 Vec 2,HuBERT和WavLM对包含自闭症诊断会话中儿童-成人互动的数据集进行了ASR性能的全面评估。我们发现,语音基础模型显示出显着的性能下降(15-20%的绝对WER)的儿童语音相比,成人语音在会话设置。然后,我们采用LoRA的最佳性能的zero shot模型(耳语大),以探讨在低资源设置微调的有效性,导致约8%的绝对WER改善儿童语音和约13%的绝对WER改善成人语音。摘要:The ability to reliably transcribe child-adult conversations in a clinical setting is valuable for diagnosis and understanding of numerous developmental disorders such as Autism Spectrum Disorder. Recent advances in deep learning architectures and availability of large scale transcribed data has led to development of speech foundation models that have shown dramatic improvements in ASR performance. However, the ability of these models to translate well to conversational child-adult interactions is under studied. In this work, we provide a comprehensive evaluation of ASR performance on a dataset containing child-adult interactions from autism diagnostic sessions, using Whisper, Wav2Vec2, HuBERT, and WavLM. We find that speech foundation models show a noticeable performance drop (15-20% absolute WER) for child speech compared to adult speech in the conversational setting. Then, we employ LoRA on the best performing zero shot model (whisper-large) to probe the effectiveness of fine-tuning in a low resource setting, resulting in ~8% absolute WER improvement for child speech and ~13% absolute WER improvement for adult speech.
【17】 Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
标题: 生成性语音基础模型预训练用于高质量语音提取和恢复
作者:Pin-Jui Ku,Alexander H. Liu,Roman Korostik,Sung-Feng Huang,Szu-Wei Fu,Ante Jukić
备注:5 pages, Submitted to ICASSP 2025. The implementation and configuration could be found in this https URL The audio demo page could be found in this https URL
链接:点击下载PDF文件
摘要:本文提出了一种高质量语音恢复任务的生成式预训练基础模型。通过直接对复值短时傅里叶变换系数进行运算,我们的模型不依赖于任何声码器进行时域信号重构。其结果是,我们的模型简化了合成过程,并删除了质量上限引入的任何梅尔频谱声码器相比,以前的工作SpeechFlow。该方法在多个语音恢复任务上进行了评估,包括语音去噪、带宽扩展、编解码器伪影去除和目标说话人提取。在所有情况下,微调我们的预训练模型都会导致优于强基线的性能。值得注意的是,在目标说话人提取任务中,我们的模型优于现有系统,包括那些利用SSL预训练编码器(如WavLM)的系统。代码和预训练的检查点在NVIDIA NeMo框架中公开提供。摘要:This paper proposes a generative pretraining foundation model for high-quality speech restoration tasks. By directly operating on complex-valued short-time Fourier transform coefficients, our model does not rely on any vocoders for time-domain signal reconstruction. As a result, our model simplifies the synthesis process and removes the quality upper-bound introduced by any mel-spectrogram vocoder compared to prior work SpeechFlow. The proposed method is evaluated on multiple speech restoration tasks, including speech denoising, bandwidth extension, codec artifact removal, and target speaker extraction. In all scenarios, finetuning our pretrained model results in superior performance over strong baselines. Notably, in the target speaker extraction task, our model outperforms existing systems, including those leveraging SSL-pretrained encoders like WavLM. The code and the pretrained checkpoints are publicly available in the NVIDIA NeMo framework.
【18】 Scenario of Use Scheme: Threat Model Specification for Speaker Privacy Protection in the Medical Domain
标题: 使用场景方案:医疗领域演讲者隐私保护的威胁模型规范
作者:Mehtab Ur Rahman,Martha Larson,Louis ten Bosch,Cristian Tejedor-García
备注:Accepted and published at SPSC Symposium 2024 4th Symposium on Security and Privacy in Speech Communication. Interspeech 2024
链接:点击下载PDF文件
摘要:语音记录越来越多地用于检测和监测疾病,导致隐私问题。除了密码学之外,还可以通过扰动、解纠缠和重新合成等方法来解决语音保护问题,这些方法可以消除说话者的敏感信息,留下医学分析所需的信息。为了制定这种隐私保护办法,有必要对有关医疗环境和医疗专业人员需求的假设作出明确和系统的说明。在本文中,我们提出了一个场景的使用计划,其中包括一个攻击者模型,其特点是对手对谁的发言人的隐私必须捍卫,和一个保护者模型,指定的防御。我们讨论了该计划与以前的工作语音隐私的连接。最后,我们提出了一个具体的例子,一个指定的使用场景和一组实验,保护说话人数据免受性别推断攻击,同时保持实用程序帕金森氏症的检测。摘要:Speech recordings are being more frequently used to detect and monitor disease, leading to privacy concerns. Beyond cryptography, protection of speech can be addressed by approaches, such as perturbation, disentanglement, and re-synthesis, that eliminate sensitive information of the speaker, leaving the information necessary for medical analysis purposes. In order for such privacy protective approaches to be developed, clear and systematic specifications of assumptions concerning medical settings and the needs of medical professionals are necessary. In this paper, we propose a Scenario of Use Scheme that incorporates an Attacker Model, which characterizes the adversary against whom the speaker's privacy must be defended, and a Protector Model, which specifies the defense. We discuss the connection of the scheme with previous work on speech privacy. Finally, we present a concrete example of a specified Scenario of Use and a set of experiments about protecting speaker data against gender inference attacks while maintaining utility for Parkinson's detection.
【19】 StyleSinger 2: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
标题: StyleSinger 2:具有风格转移和多层风格控制的Zero-Shot歌唱声音合成
作者:Yu Zhang,Ziyue Jiang,Ruiqi Li,Changhao Pan,Jinzheng He,Rongjie Huang,Chuxin Wang,Zhou Zhao
备注:Accepted by EMNLP 2024
链接:点击下载PDF文件
摘要:带风格转换和风格控制的Zero-shot歌唱声音合成(SVS)旨在从音频和文本提示中生成具有不可见音色和风格(包括演唱方法、情感、节奏、技术和发音)的高质量歌唱声音。然而,歌唱风格的多面性对有效的建模,转移和控制提出了重大挑战。此外,目前的SVS模型往往无法为看不见的歌手生成丰富的风格细微差别的歌声。为了应对这些挑战,我们引入了StyleSinger 2,这是第一个zero-shot SVS模型,用于跨语言语音和演唱风格的风格转移,以及多级风格控制。具体地说,StyleSinger 2提出了三个主要模块:1)聚类风格编码器采用聚类矢量量化模型,将风格信息稳定地压缩到一个紧凑的潜在空间中; 2)风格和时长语言模型(S &D-LM)同时预测风格信息和音素时长,这对两者都有好处; 3)风格自适应解码器使用新颖的Mel风格自适应归一化方法来生成具有增强的细节的歌声。实验结果表明,StyleSinger 2在合成质量、歌手相似度和风格可控性方面优于所有基线模型,包括zero-shot风格迁移、多级风格控制、跨语言风格迁移和语音到歌唱风格迁移。可以在https: stylesinger2.github.io 上访问唱歌的声音样本。摘要:Zero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunciation) from audio and text prompts. However, the multifaceted nature of singing styles poses a significant challenge for effective modeling, transfer, and control. Furthermore, current SVS models often fail to generate singing voices rich in stylistic nuances for unseen singers. To address these challenges, we introduce StyleSinger 2, the first zero-shot SVS model for style transfer across cross-lingual speech and singing styles, along with multi-level style control. Specifically, StyleSinger 2 proposes three primary modules: 1) the clustering style encoder employs a clustering vector quantization model to stably condense style information into a compact latent space; 2) the Style and Duration Language Model (S &D-LM) concurrently predicts style information and phoneme duration, which benefits both; 3) the style adaptive decoder uses a novel mel-style adaptive normalization method to generate singing voices with enhanced details. Experimental results show that StyleSinger 2 outperforms all baseline models in synthesis quality, singer similarity, and style controllability across various tasks, including zero-shot style transfer, multi-level style control, cross-lingual style transfer, and speech-to-singing style transfer. Singing voice samples can be accessed at https: stylesinger2.github.io .
【20】 ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
标题: ESPnet-Codec:音频、音乐和语音神经编解码器的全面训练和评估
作者:Jiatong Shi,Jinchuan Tian,Yihan Wu,Jee-weon Jung,Jia Qi Yip,Yoshiki Masuyama,William Chen,Yuning Wu,Yuxun Tang,Massa Baali,Dareen Alharhi,Dong Zhang,Ruifan Deng,Tejes Srivastava,Haibin Wu,Alexander H. Liu,Bhiksha Raj,Qin Jin,Ruihua Song,Shinji Watanabe
备注:Accepted by SLT
链接:点击下载PDF文件
摘要:神经编解码器已经成为最近的语音和音频生成研究的关键。除了信号压缩能力之外,还发现离散编解码器可以增强下游训练效率和与自回归语言模型的兼容性。然而,随着广泛的下游应用程序的调查,在确保不同应用程序之间的公平比较方面出现了挑战。为了解决这些问题,我们提出了一个新的开源平台ESPnet-Codec,该平台构建在ESPnet上,专注于神经编解码器的训练和评估。ESPnet-Codec提供音频、音乐和语音方面的各种配方,用于使用几种广泛采用的编解码器模型进行训练和评估。与ESPnet-Codec一起,我们提出了VERSA,一个独立的评估工具包,它提供了超过20个音频评估指标的编解码器性能的全面评估。值得注意的是,我们证明了ESPnet-Codec可以集成到六个ESPnet任务中,支持不同的应用程序。摘要:Neural codecs have become crucial to recent speech and audio generation research. In addition to signal compression capabilities, discrete codecs have also been found to enhance downstream training efficiency and compatibility with autoregressive language models. However, as extensive downstream applications are investigated, challenges have arisen in ensuring fair comparisons across diverse applications. To address these issues, we present a new open-source platform ESPnet-Codec, which is built on ESPnet and focuses on neural codec training and evaluation. ESPnet-Codec offers various recipes in audio, music, and speech for training and evaluation using several widely adopted codec models. Together with ESPnet-Codec, we present VERSA, a standalone evaluation toolkit, which provides a comprehensive evaluation of codec performance over 20 audio evaluation metrics. Notably, we demonstrate that ESPnet-Codec can be integrated into six ESPnet tasks, supporting diverse applications.
【21】 Interpolation filter design for sample rate independent audio effect RNNs
标题: 独立于采样率的音频效果RNN的内插过滤器设计
作者:Alistair Carson,Alec Wright,Stefan Bilbao
链接:点击下载PDF文件
摘要:递归神经网络(RNN)在模拟模拟吉他放大器的非线性、有状态行为和失真效果方面是有效的。与直接电路仿真的情况不同,RNN具有编码在其模型权重中的固定采样率,使得采样率在推理期间不可调整。最近的工作提出了通过增加样本中的反馈延迟长度来增加RNN在推理(过采样)时的采样率,使用分数延迟滤波器进行非整数转换。在这里,我们调查的任务,降低采样率的推断(欠采样),并建议使用外推滤波器来近似所需的分数信号提前。我们考虑了两种滤波器设计方法,并分析了滤波器阶数对音频质量的影响。我们的研究结果表明,滤波器的正确选择可以提供高质量的过采样和欠采样的结果,但是,在某些情况下,采样率的调整会导致输出信号中不必要的伪影。我们通过线性稳定性分析来分析这些失效情况,表明它们是由固定点周围的不稳定性引起的。这种方法能够在运行之前为给定的RNN模型提供合适的插值滤波器的知情预测。摘要:Recurrent neural networks (RNNs) are effective at emulating the non-linear, stateful behavior of analog guitar amplifiers and distortion effects. Unlike the case of direct circuit simulation, RNNs have a fixed sample rate encoded in their model weights, making the sample rate non-adjustable during inference. Recent work has proposed increasing the sample rate of RNNs at inference (oversampling) by increasing the feedback delay length in samples, using a fractional delay filter for non-integer conversions. Here, we investigate the task of lowering the sample rate at inference (undersampling), and propose using an extrapolation filter to approximate the required fractional signal advance. We consider two filter design methods and analyze the impact of filter order on audio quality. Our results show that the correct choice of filter can give high quality results for both oversampling and undersampling; however, in some cases the sample rate adjustment leads to unwanted artefacts in the output signal. We analyse these failure cases through linearised stability analysis, showing that they result from instability around a fixed point. This approach enables an informed prediction of suitable interpolation filters for a given RNN model before runtime.
【22】 Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
标题: 美杜莎耳朵里的低语:基于变形器的ASB的多头高效解码
作者:Yael Segal-Feldman,Aviv Shamsian,Aviv Navon,Gill Hetz,Joseph Keshet
备注:Under Review
链接:点击下载PDF文件
摘要:基于变换器的大型模型在语音转录和翻译方面具有巨大的潜力。它们的自我注意机制和并行处理使它们能够捕捉音频序列中的复杂模式和依赖关系。然而,这种潜力带来了挑战,因为这些大型和计算密集型模型导致推理速度缓慢。已经提出了各种优化策略来提高性能,包括有效的硬件利用率和算法增强。在本文中,我们介绍了耳语美杜莎,一种新的方法,旨在提高处理速度,最小的影响字错误率(WER)。该模型扩展了OpenAI的Whisper架构,每次迭代预测多个令牌,从而减少了50%的延迟。我们展示了Whisper-Medusa在不同学习设置和数据集上的有效性。摘要:Large transformer-based models have significant potential for speech transcription and translation. Their self-attention mechanisms and parallel processing enable them to capture complex patterns and dependencies in audio sequences. However, this potential comes with challenges, as these large and computationally intensive models lead to slow inference speeds. Various optimization strategies have been proposed to improve performance, including efficient hardware utilization and algorithmic enhancements. In this paper, we introduce Whisper-Medusa, a novel approach designed to enhance processing speed with minimal impact on Word Error Rate (WER). The proposed model extends the OpenAI's Whisper architecture by predicting multiple tokens per iteration, resulting in a 50% reduction in latency. We showcase the effectiveness of Whisper-Medusa across different learning setups and datasets.
【23】 WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
标题: WeSep:一个可扩展且灵活的工具包,可推广的目标说话者提取
作者:Shuai Wang,Ke Zhang,Shaoxiong Lin,Junjie Li,Xuefei Wang,Meng Ge,Jianwei Yu,Yanmin Qian,Haizhou Li
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:目标说话人提取(TSE)的核心是从多说话人重叠的语音中分离出特定目标说话人的语音,这是鸡尾酒会问题中的一个典型设置。近年来,TSE由于其在用户定制界面和助听器等各种应用中的潜力,或者作为语音识别和说话人识别等基本任务的关键前端处理技术,引起了越来越多的关注。然而,目前很少有开源工具包或可供现成使用的预训练模型。在这项工作中,我们介绍了WeSep,一个工具包,专为研究和实际应用在TSE。WeSep具有灵活的目标扬声器建模,可扩展的数据管理,有效的实时数据模拟,结构化配方和部署支持。该工具包在 url{https: github.com wenet-e2e WeSep.}上公开提供。摘要:Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increasing attention due to its potential for various applications such as user-customized interfaces and hearing aids, or as a crutial front-end processing technologies for subsequential tasks such as speech recognition and speaker recongtion. However, there are currently few open-source toolkits or available pre-trained models for off-the-shelf usage. In this work, we introduce WeSep, a toolkit designed for research and practical applications in TSE. WeSep is featured with flexible target speaker modeling, scalable data management, effective on-the-fly data simulation, structured recipes and deployment support. The toolkit is publicly avaliable at url{https: github.com wenet-e2e WeSep.}
【24】 M-Vec: Matryoshka Speaker Embeddings with Flexible Dimensions
标题: M-Vec:具有灵活尺寸的Matryoshka扬声器嵌入式
作者:Shuai Wang,Pengcheng Zhu,Haizhou Li
备注:ICSR 2024, Shenzhen
链接:点击下载PDF文件
摘要:固定维度的说话人嵌入已经成为说话人建模的主要方法,通常跨越数百到数千个维度。这些维度是超参数,它们不是专门挑选的,也不是按照重要性进行层次排序的。在大规模说话人表征数据库中,降低嵌入维数可以显著降低存储和计算成本。然而,直接训练低维表示通常会产生次优性能。在本文中,我们介绍了Matryoshka扬声器嵌入,一种方法,允许动态提取子维度的嵌入,同时保持性能。我们的方法在VoxCeleb数据集上进行了验证,证明它可以实现极低维的嵌入,如8维,同时保持较高的说话人验证性能。摘要:Fixed-dimensional speaker embeddings have become the dominant approach in speaker modeling, typically spanning hundreds to thousands of dimensions. These dimensions are hyperparameters that are not specifically picked, nor are they hierarchically ordered in terms of importance. In large-scale speaker representation databases, reducing the dimensionality of embeddings can significantly lower storage and computational costs. However, directly training low-dimensional representations often yields suboptimal performance. In this paper, we introduce the Matryoshka speaker embedding, a method that allows dynamic extraction of sub-dimensions from the embedding while maintaining performance. Our approach is validated on the VoxCeleb dataset, demonstrating that it can achieve extremely low-dimensional embeddings, such as 8 dimensions, while preserving high speaker verification performance.
【25】 Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
标题: 利用随机选择策略最小化表示损失以实现高效的环境假音频检测
作者:Orchid Chetia Phukan,Girish,Mohd Mujtaba Akhtar,Swarup Ranjan Behera,Nitin Choudhury,Arun Balaji Buduru,Rajesh Sharma,S. R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:基础模型的适应性极大地推进了环境音频深度伪造检测(EADD),这是一个快速增长的研究领域。这些模型通常被微调或在其冻结状态下用于下游任务。然而,其表示的维度可以实质上导致下游模型的高参数计数,从而导致更高的计算需求。因此,一般的方法是通过利用最先进的(SOTA)无监督降维技术(PCA,SVD,KPCA,GRP)来压缩这些表示,以实现有效的EADD。然而,随着这些技术的应用,我们观察到性能下降。因此,在本文中,我们证明了表示向量包含冗余信息,随机选择40-50%的表示值并在其上构建下游模型可以保留甚至有时可以提高性能。我们表明,这种随机选择保留更多的性能比SOTA降维技术,同时减少模型参数和推理时间几乎一半以上。摘要:The adaptation of foundation models has significantly advanced environmental audio deepfake detection (EADD), a rapidly growing area of research. These models are typically fine-tuned or utilized in their frozen states for downstream tasks. However, the dimensionality of their representations can substantially lead to a high parameter count of downstream models, leading to higher computational demands. So, a general way is to compress these representations by leveraging state-of-the-art (SOTA) unsupervised dimensionality reduction techniques (PCA, SVD, KPCA, GRP) for efficient EADD. However, with the application of such techniques, we observe a drop in performance. So in this paper, we show that representation vectors contain redundant information, and randomly selecting 40-50% of representation values and building downstream models on it preserves or sometimes even improves performance. We show that such random selection preserves more performance than the SOTA dimensionality reduction techniques while reducing model parameters and inference time by almost over half.
【26】 Enhancing Open-Set Speaker Identification through Rapid Tuning with Speaker Reciprocal Points and Negative Sample
标题: 通过使用说话者互反点和负样本快速调整来增强开集说话者识别
作者:Zhiyong Chen,Zhiqi Ai,Xinnuo Li,Shugong Xu
Journal-ref:IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
摘要:本文介绍了一种新的框架开放集说话人识别在家庭环境中,发挥了至关重要的作用,促进无缝的人机交互。针对当前说话人模型和分类方法的局限性,我们的工作将预训练的WavLM前端与用于注册的Few-Shot快速调谐神经网络(NN)后端集成在一起,采用任务优化的说话人倒数点学习(SRPL)来增强对多个目标说话人的区分。此外,我们提出了一个增强版的SRPL(SRPL+),它结合了负样本学习与语音合成和真正的负样本,以显着提高开集SID的准确性。我们的方法在各种多语言文本相关的说话人识别数据集进行了全面评估,证明了其在复杂的家庭多说话人识别场景中实现高可用性的有效性。与直接使用高效的WavLM基+模型相比,该系统将开集性能提高了27%.摘要:This paper introduces a novel framework for open-set speaker identification in household environments, playing a crucial role in facilitating seamless human-computer interactions. Addressing the limitations of current speaker models and classification approaches, our work integrates an pretrained WavLM frontend with a few-shot rapid tuning neural network (NN) backend for enrollment, employing task-optimized Speaker Reciprocal Points Learning (SRPL) to enhance discrimination across multiple target speakers. Furthermore, we propose an enhanced version of SRPL (SRPL+), which incorporates negative sample learning with both speech-synthesized and real negative samples to significantly improve open-set SID accuracy. Our approach is thoroughly evaluated across various multi-language text-dependent speaker recognition datasets, demonstrating its effectiveness in achieving high usability for complex household multi-speaker recognition scenarios. The proposed system enhanced open-set performance by up to 27 % over the directly use of efficient WavLM base+ model.
【27】 StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
标题: StyleFusion TTC:用于Zero-Shot文本到语音合成的多模式风格控制和增强的特征融合
作者:Zhiyong Chen,Xinnuo Li,Zhiqi Ai,Shugong Xu
Journal-ref:The 7th Chinese Conference on Pattern Recognition and Computer Vision PRCV 2024
链接:点击下载PDF文件
摘要:我们介绍StyleFusion-TTS,一个提示和 或音频参考,风格和说话人可控,zero-shot文本到语音(TTS)合成系统,旨在提高当前研究文献的可编辑性和自然性。我们提出了一个通用的前端编码器作为一个紧凑和有效的模块,利用多模态输入,包括文本提示,音频参考,扬声器音色参考在一个完全zero-shot的方式,并产生解开的风格和扬声器控制嵌入。我们的新方法还利用了一个分层的构象结构的风格和扬声器控制嵌入的融合,旨在实现最佳的功能融合在当前先进的TTS架构。StyleFusion-TTS通过主观和客观的多个指标进行评估。该系统在我们的评估中表现出良好的性能,这表明它有潜力为zero-shot文本到语音合成领域的进步做出贡献。摘要:We introduce StyleFusion-TTS, a prompt and or audio referenced, style and speaker-controllable, zero-shot text-to-speech (TTS) synthesis system designed to enhance the editability and naturalness of current research literature. We propose a general front-end encoder as a compact and effective module to utilize multimodal inputs including text prompts, audio references, and speaker timbre references in a fully zero-shot manner and produce disentangled style and speaker control embeddings. Our novel approach also leverages a hierarchical conformer structure for the fusion of style and speaker control embeddings, aiming to achieve optimal feature fusion within the current advanced TTS architecture. StyleFusion-TTS is evaluated through multiple metrics, both subjectively and objectively. The system shows promising performance across our evaluations, suggesting its potential to contribute to the advancement of the field of zero-shot text-to-speech synthesis.
【28】 Language-based Audio Moment Retrieval
标题: 基于百分比的音频时刻检索
作者:Hokuto Munakata,Taichi Nishimura,Shota Nakada,Tatsuya Komatsu
链接:点击下载PDF文件
摘要:在本文中,我们提出并设计了一个新的任务,称为音频时刻检索(AMR)。与从音频数据库中搜索短音频片段的传统基于语言的音频检索任务不同,AMR旨在基于文本查询预测未修剪长音频中的相关时刻。鉴于AMR之前缺乏相关工作,我们首先构建了一个专用数据集Clotho-Moment,由带有时刻注释的大规模模拟音频记录组成。然后,我们提出了一个基于DETR的模型,命名为音频时刻DETR(AM-DETR),AMR任务的基本框架。该模型捕获音频特征的时间依赖性,灵感来自类似的视频时刻检索任务,从而超越了传统的剪辑级音频检索方法。此外,我们提供手动注释的数据集,以正确衡量我们的方法在真实数据上的有效性和鲁棒性。实验结果表明,使用Clotho-Moment训练的AM-DETR的性能优于在所有指标上应用带有滑动窗口的剪辑级音频检索方法的基线模型,特别是将Recall1@0.7提高了9.00个点。我们的数据集和代码可在https: h-munakata.github.io Language-based-Audio-Moment-Retrieval上公开获取。摘要:In this paper, we propose and design a new task called audio moment retrieval (AMR). Unlike conventional language-based audio retrieval tasks that search for short audio clips from an audio database, AMR aims to predict relevant moments in untrimmed long audio based on a text query. Given the lack of prior work in AMR, we first build a dedicated dataset, Clotho-Moment, consisting of large-scale simulated audio recordings with moment annotations. We then propose a DETR-based model, named Audio Moment DETR (AM-DETR), as a fundamental framework for AMR tasks. This model captures temporal dependencies within audio features, inspired by similar video moment retrieval tasks, thus surpassing conventional clip-level audio retrieval methods. Additionally, we provide manually annotated datasets to properly measure the effectiveness and robustness of our methods on real data. Experimental results show that AM-DETR, trained with Clotho-Moment, outperforms a baseline model that applies a clip-level audio retrieval method with a sliding window on all metrics, particularly improving Recall1@0.7 by 9.00 points. Our datasets and code are publicly available in https: h-munakata.github.io Language-based-Audio-Moment-Retrieval.
【29】 Safe Guard: an LLM-agent for Real-time Voice-based Hate Speech Detection in Social Virtual Reality
标题: Safe Guard:用于社交虚拟现实中实时基于语音的仇恨言语检测的LLM代理
作者:Yiwen Xu,Qinyang Hou,Hongyu Wan,Mirjana Prpa
链接:点击下载PDF文件
摘要:在本文中,我们提出了安全卫士,一个LLM代理,用于检测社交VR(VRChat)中基于语音的交互中的仇恨言论。我们的系统利用Open AI GPT和音频特征提取进行实时语音交互。我们贡献了一个系统的设计和评估系统,证明了我们的方法在检测仇恨言论的能力,并减少误报相比,目前可用的方法。我们的研究结果表明,基于LLM的代理在创建更安全的虚拟环境中的潜力,并为LLM驱动的适度方法的进一步发展奠定了基础。摘要:In this paper, we present Safe Guard, an LLM-agent for the detection of hate speech in voice-based interactions in social VR (VRChat). Our system leverages Open AI GPT and audio feature extraction for real-time voice interactions. We contribute a system design and evaluation of the system that demonstrates the capability of our approach in detecting hate speech, and reducing false positives compared to currently available approaches. Our results indicate the potential of LLM-based agents in creating safer virtual environments and set the groundwork for further advancements in LLM-driven moderation approaches.
【30】 Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
标题: 修改、推理和识别:通过特定描述和ASB错误纠正的基于LLM的情感识别
作者:Yuanchao Li,Yuan Gong,Chao-Han Huck Yang,Peter Bell,Catherine Lai
链接:点击下载PDF文件
摘要:随着大语言模型(LLM)的发展,最近出现了使用提示工程来注释和识别语音情感的方法,但其有效性和可靠性仍然值得怀疑。在本文中,我们对这一主题进行了系统的研究,首先提出了新的提示,结合声学,语言学和心理学的情感特定的知识。随后,我们研究了基于LLM的提示对自动语音识别(ASR)转录的有效性,将其与地面实况转录进行对比。此外,我们提出了一个修改原因识别提示管道,用于从具有ASR错误的口语中进行鲁棒的基于LLM的情感识别。此外,上下文感知学习,在上下文学习和指令调整的实验进行检查LLM培训计划在这个方向上的有用性。最后,我们调查的灵敏度LLM轻微提示变化。实验结果表明,情感特定的提示,ASR纠错,和LLM训练计划的LLM为基础的情感识别的功效。我们的研究旨在改进LLM在情感识别和相关领域的使用。摘要:Annotating and recognizing speech emotion using prompt engineering has recently emerged with the advancement of Large Language Models (LLMs), yet its efficacy and reliability remain questionable. In this paper, we conduct a systematic study on this topic, beginning with the proposal of novel prompts that incorporate emotion-specific knowledge from acoustics, linguistics, and psychology. Subsequently, we examine the effectiveness of LLM-based prompting on Automatic Speech Recognition (ASR) transcription, contrasting it with ground-truth transcription. Furthermore, we propose a Revise-Reason-Recognize prompting pipeline for robust LLM-based emotion recognition from spoken language with ASR errors. Additionally, experiments on context-aware learning, in-context learning, and instruction tuning are performed to examine the usefulness of LLM training schemes in this direction. Finally, we investigate the sensitivity of LLMs to minor prompt variations. Experimental results demonstrate the efficacy of the emotion-specific prompts, ASR error correction, and LLM training schemes for LLM-based emotion recognition. Our study aims to refine the use of LLMs in emotion recognition and related domains.
【31】 Rethinking Emotion Bias in Music via Frechet Audio Distance
标题: 通过Frechet音频距离重新思考音乐中的情感偏见
作者:Yuanchao Li,Azalea Gui,Dimitra Emmanouilidou,Hannes Gamper
链接:点击下载PDF文件
摘要:音乐情感的主观性质在识别和生成中引入了固有的偏差,特别是当依赖于单个音频编码器、情感分类器或评估度量时。在这项工作中,我们对音乐情感识别(MER)和情感音乐生成(EMG)进行了研究,采用了不同的音频编码器以及Frechet音频距离(FAD),这是一种无参考的评估指标。我们的研究从MER的基准评估开始,强调了与使用单个音频编码器相关的限制以及在不同测量中观察到的差异。然后,我们建议使用来自多个编码器的FAD来评估MER性能,以提供更客观的音乐情感测量。此外,我们引入了一个增强的EMG方法,旨在提高所产生的音乐情感的变化和突出,从而提高现实主义。此外,我们研究了真实音乐和合成音乐中传达的情感之间的现实主义差异,将我们的EMG模型与两个基线模型进行比较。实验结果强调了MER和EMG中的情感偏差问题,并展示了使用FAD和各种音频编码器来客观评估音乐情感的潜力。摘要:The subjective nature of music emotion introduces inherent bias in both recognition and generation, especially when relying on a single audio encoder, emotion classifier, or evaluation metric. In this work, we conduct a study on Music Emotion Recognition (MER) and Emotional Music Generation (EMG), employing diverse audio encoders alongside the Frechet Audio Distance (FAD), a reference-free evaluation metric. Our study begins with a benchmark evaluation of MER, highlighting the limitations associated with using a single audio encoder and the disparities observed across different measurements. We then propose assessing MER performance using FAD from multiple encoders to provide a more objective measure of music emotion. Furthermore, we introduce an enhanced EMG approach designed to improve both the variation and prominence of generated music emotion, thus enhancing realism. Additionally, we investigate the realism disparities between the emotions conveyed in real and synthetic music, comparing our EMG model against two baseline models. Experimental results underscore the emotion bias problem in both MER and EMG and demonstrate the potential of using FAD and diverse audio encoders to evaluate music emotion objectively.
【32】 Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
标题: Speech 2 rtMRI:语音引导扩散模型,用于语音期间人声实时MRI视频
作者:Hong Nguyen,Sean Foley,Kevin Huang,Xuan Shi,Tiantian Feng,Shrikanth Narayanan
备注:4 pages
链接:点击下载PDF文件
摘要:从视觉上和运动学上理解语音产生可以为第二语言学习系统的设计提供信息,以及在视频游戏和动画中创建说话的角色。在这项工作中,我们介绍了一种数据驱动的方法来直观地表示发音关节运动的磁共振成像(MRI)视频中的人类声道在讲话的基础上,任意的音频或语音输入。我们利用嵌入先验知识的大型预训练语音模型,使用语音到视频扩散模型将视觉域推广到看不见的数据。我们的研究结果表明,视觉生成显着受益于预训练的语音表示。我们还观察到,孤立地评估音素是具有挑战性的,但在口语的背景下进行评估时变得更加简单。目前的研究结果的局限性包括舌头接触上颚时存在舌头运动不平滑和视频失真。摘要:Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven method to visually represent articulator motion in Magnetic Resonance Imaging (MRI) videos of the human vocal tract during speech based on arbitrary audio or speech input. We leverage large pre-trained speech models, which are embedded with prior knowledge, to generalize the visual domain to unseen data using a speech-to-video diffusion model. Our findings demonstrate that the visual generation significantly benefits from the pre-trained speech representations. We also observed that evaluating phonemes in isolation is challenging but becomes more straightforward when assessed within the context of spoken words. Limitations of the current results include the presence of unsmooth tongue motion and video distortion when the tongue contacts the palate.
【33】 Blind Localization of Early Room Reflections with Arbitrary Microphone Array
标题: 用任意麦克风阵列实现早期房间反射的盲定位
作者:Yogev Hadadi,Vladimir Tourbabin,Zamir Ben-Hur,David Lou Alon,Boaz Rafaely
链接:点击下载PDF文件
摘要:在没有房间脉冲响应或源信号的先验知识的情况下,盲目地估计早期房间反射的到达方向(DoA)在音频信号处理应用中是非常有价值的。FF-PHALCOR(频率聚焦相位对准相关)方法最近开发了用于此目的,扩展了原来的PHALCOR方法工作与任意阵列,而不仅仅是球形的。以前的研究仅提供了对其性能的初步了解。这项研究提供了一个全面的分析方法的性能和局限性,研究反射特性,如延迟,幅度和空间密度如何影响其有效性。该研究还提出了改进措施来克服这些限制,提高检测质量并减少误报。此外,该研究还研究了使用估计的反射信息生成房间脉冲响应如何影响空间感知。研究结果表明,所提出的方法在基线上具有感知优势,当使用具有32个麦克风的球形阵列时具有特别高的感知质量。然而,当使用仅具有6个麦克风的半圆形阵列时,质量有所降低。摘要:Blindly estimating the direction of arrival (DoA) of early room reflections without prior knowledge of the room impulse response or source signal is highly valuable in audio signal processing applications. The FF-PHALCOR (Frequency Focusing PHase ALigned CORrelation) method was recently developed for this purpose, extending the original PHALCOR method to work with arbitrary arrays rather than just spherical ones. Previous studies have provided only initial insights into its performance. This study offers a comprehensive analysis of the method's performance and limitations, examining how reflection characteristics such as delay, amplitude, and spatial density affect its effectiveness. The research also proposes improvements to overcome these limitations, enhancing detection quality and reducing false alarms. Additionally, the study examined how spatial perception is affected by generating room impulse responses using estimated reflection information. The findings suggest a perceptual advantage of the proposed approach over the baseline, with particularly high perceptual quality when using the spherical array with 32 microphones. However, the quality is somewhat reduced when using a semi-circular array with only 6 microphones.
【34】 The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
标题: 议会议事录中自动生成的语音和文本数据集的ParlaSpeech集合
作者:Nikola Ljubešić,Peter Rupnik,Danijel Koržinek
备注:Submitted to SPECOM 2024
链接:点击下载PDF文件
摘要:最近语音和语言技术的重大改进来自于对原始语言数据的自我监督方法以及各种类型的显式监督。为了确保对语音数据的高质量处理,最有用的显式监督类型仍然是语音信号与其对应的文本转录之间的对齐,这是一种不适用于许多语言的数据类型。在本文中,我们提出了一种基于议会会议记录及其录音的方法来构建资源较少的语言的大型开放式语音和文本对齐数据集。我们的出发点是ParlaMint可比语料库的26个国家的欧洲议会的议会议事记录。在扩大ParlaMint语料库的试点运行中,我们将重点放在三种斯拉夫语言上,即克罗地亚语、波兰语和塞尔维亚语。我们的方法的主要挑战是ParlaMint文本和可用记录之间缺乏任何全局对齐,以及每个模式中有时不同的数据顺序,这需要一种新的方法来在大搜索空间中对齐长序列的文本和音频。这次试运行的结果是三个高质量的数据集,涵盖了超过5,000小时的演讲和附带的文本转录。虽然这些数据集已经在三种语言的口语和文本数据的可用性方面产生了巨大的差异,但我们希望强调所提出的方法在为更多语言构建类似数据集方面的潜力。摘要:Recent significant improvements in speech and language technologies come both from self-supervised approaches over raw language data as well as various types of explicit supervision. To ensure high-quality processing of spoken data, the most useful type of explicit supervision is still the alignment between the speech signal and its corresponding text transcript, which is a data type that is not available for many languages. In this paper, we present our approach to building large and open speech-and-text-aligned datasets of less-resourced languages based on transcripts of parliamentary proceedings and their recordings. Our starting point are the ParlaMint comparable corpora of transcripts of parliamentary proceedings of 26 national European parliaments. In the pilot run on expanding the ParlaMint corpora with aligned publicly available recordings, we focus on three Slavic languages, namely Croatian, Polish, and Serbian. The main challenge of our approach is the lack of any global alignment between the ParlaMint texts and the available recordings, as well as the sometimes varying data order in each of the modalities, which requires a novel approach in aligning long sequences of text and audio in a large search space. The results of this pilot run are three high-quality datasets that span more than 5,000 hours of speech and accompanying text transcripts. Although these datasets already make a huge difference in the availability of spoken and textual data for the three languages, we want to emphasize the potential of the presented approach in building similar datasets for many more languages.
【35】 Toward Automated Clinical Transcriptions
标题: 迈向自动化临床Transit
作者:Mitchell A. Klusty,W. Vaiden Logan,Samuel E. Armstrong,Aaron D. Mullen,Caroline N. Leach,Jeff Talbert,V. K. Cody Bumgardner
备注:7 pages, 6 figures
链接:点击下载PDF文件
摘要:行政文件是医疗保健成本上升的主要驱动力,并与不良后果有关,包括医生倦怠和护理质量下降。本文介绍了一种安全的系统,适用于语音到文本转录和说话人标记(日记)的最新进展,病人提供者的对话。该系统经过优化,可生成准确的transmittance并突出显示潜在的错误,以促进快速的人工验证,进一步减少必要的人工工作。应用于超过40个小时的模拟对话,该系统提供了一个有前途的基础,自动化临床transmittance。摘要:Administrative documentation is a major driver of rising healthcare costs and is linked to adverse outcomes, including physician burnout and diminished quality of care. This paper introduces a secure system that applies recent advancements in speech-to-text transcription and speaker-labeling (diarization) to patient-provider conversations. This system is optimized to produce accurate transcriptions and highlight potential errors to promote rapid human verification, further reducing the necessary manual effort. Applied to over 40 hours of simulated conversations, this system offers a promising foundation for automating clinical transcriptions.
【36】 A Joint Spectro-Temporal Relational Thinking Based Acoustic Modeling Framework
标题: 基于频谱-时间关系思维的联合声学建模框架
作者:Zheng Nan,Ting Dang,Vidhyasaharan Sethu,Beena Ahmed
链接:点击下载PDF文件
摘要:关系思维是指人类对感觉信号和先验知识之间的关系形成心理印象,并随后将其纳入其世界模型的固有能力。尽管关系思维在人类理解语音方面发挥着至关重要的作用,但它尚未在任何人工语音识别系统中得到利用。最近,已经有一些尝试来纠正这种疏忽,但这些尝试仅限于仅在时域中操作的粗略话语级模型。为了缩小人工系统和人类能力之间的差距,本文提出了一种新的基于频谱-时间关系思维的声学建模框架。具体来说,它首先生成大量的概率图来模拟跨时域和频域的语音段之间的关系。然后,将这些图中每对节点中的关系信息聚合并嵌入到可以由下游任务使用的潜在表示中。建立在这个框架上的模型优于最先进的系统,在TIMIT数据集上的音素识别任务提高了7.82%。深入的分析进一步表明,我们提出的关系思维建模主要提高了模型的识别能力,元音,这是最容易混淆的音素识别器。摘要:Relational thinking refers to the inherent ability of humans to form mental impressions about relations between sensory signals and prior knowledge, and subsequently incorporate them into their model of their world. Despite the crucial role relational thinking plays in human understanding of speech, it has yet to be leveraged in any artificial speech recognition systems. Recently, there have been some attempts to correct this oversight, but these have been limited to coarse utterance-level models that operate exclusively in the time domain. In an attempt to narrow the gap between artificial systems and human abilities, this paper presents a novel spectro-temporal relational thinking based acoustic modeling framework. Specifically, it first generates numerous probabilistic graphs to model the relationships among speech segments across both time and frequency domains. The relational information rooted in every pair of nodes within these graphs is then aggregated and embedded into latent representations that can be utilized by downstream tasks. Models built upon this framework outperform state-of-the-art systems with a 7.82 % improvement in phoneme recognition tasks over the TIMIT dataset. In-depth analyses further reveal that our proposed relational thinking modeling mainly improves the model's ability to recognize vowels, which are the most likely to be confused by phoneme recognizers.
【37】 TCG CREST System Description for the Second DISPLACE Challenge
标题: 第二次DISPLACE挑战赛的TCG CREST系统描述
作者:Nikhil Raghav,Subhajit Saha,Md Sahidullah,Swagatam Das
链接:点击下载PDF文件
摘要:在本报告中,我们描述了我们的团队为2024年第二次DISPLACE挑战赛开发的说话者日记(SD)和语言日记(LD)系统。我们的贡献致力于多语言和多发言者场景中的可持续发展轨道1和LD轨道2。我们研究了不同的语音增强技术,语音活动检测(VAD)技术,无监督域分类,神经嵌入提取架构。我们还利用了各种嵌入提取模型的融合。我们使用开源SpeechBrain工具包实现了我们的系统。我们最终的意见书使用频谱聚类的扬声器和语言日记。我们在轨道1的挑战基线上实现了约7 %$的相对改进。我们没有获得超过轨道2中挑战基线的改进。摘要:In this report, we describe the speaker diarization (SD) and language diarization (LD) systems developed by our team for the Second DISPLACE Challenge, 2024. Our contributions were dedicated to Track 1 for SD and Track 2 for LD in multilingual and multi-speaker scenarios. We investigated different speech enhancement techniques, voice activity detection (VAD) techniques, unsupervised domain categorization, and neural embedding extraction architectures. We also exploited the fusion of various embedding extraction models. We implemented our system with the open-source SpeechBrain toolkit. Our final submissions use spectral clustering for both the speaker and language diarization. We achieve about $7 %$ relative improvement over the challenge baseline in Track 1. We did not obtain improvement over the challenge baseline in Track 2.
【38】 Contextualization of ASR with LLM using phonetic retrieval-based augmentation
标题: 使用基于语音检索的增强将ASC与LLM进行上下文化
作者:Zhihong Lei,Xingyu Na,Mingbin Xu,Ernest Pusateri,Christophe Van Gysel,Yuanyuan Zhang,Shiyi Han,Zhen Huang
链接:点击下载PDF文件
摘要:大型语言模型(LLM)已经显示出对包括音频和文本在内的多模态信号进行建模的卓越能力,允许模型在给定语音输入的情况下生成口语或文本响应。然而,它仍然是一个挑战的模型,以识别个人命名的实体,如在电话簿中的联系人,当输入模态是语音。在这项工作中,我们从语音识别任务开始,并提出了一个基于检索的解决方案,以上下文的LLM:我们首先让LLM检测命名实体的语音没有任何上下文,然后使用这个命名实体作为查询检索语音相似的命名实体从个人数据库和饲料的LLM,最后运行上下文感知的LLM解码。在语音助理任务中,与没有上下文的基线系统相比,我们的解决方案实现了高达30.2%的相对单词错误率降低和73.6%的相对命名实体错误率降低。值得注意的是,我们的解决方案通过设计避免了用完整的命名实体数据库提示LLM,使其高效并适用于大型命名实体数据库。摘要:Large language models (LLMs) have shown superb capability of modeling multimodal signals including audio and text, allowing the model to generate spoken or textual response given a speech input. However, it remains a challenge for the model to recognize personal named entities, such as contacts in a phone book, when the input modality is speech. In this work, we start with a speech recognition task and propose a retrieval-based solution to contextualize the LLM: we first let the LLM detect named entities in speech without any context, then use this named entity as a query to retrieve phonetically similar named entities from a personal database and feed them to the LLM, and finally run context-aware LLM decoding. In a voice assistant task, our solution achieved up to 30.2% relative word error rate reduction and 73.6% relative named entity error rate reduction compared to a baseline system without contextualization. Notably, our solution by design avoids prompting the LLM with the full named entity database, making it highly efficient and applicable to large named entity databases.
【39】 Equivariance-based self-supervised learning for audio signal recovery from clipped measurements
标题: 基于等效性的自我监督学习,从剪辑测量中恢复音频信号
作者:Victor Sechaud,Laurent Jacques,Patrice Abry,Julián Tachella
Journal-ref:EUSIPCO, Aug 2024, Lyon, France
链接:点击下载PDF文件
摘要:
标题: 用于相重建和语音增强的显式一致性保持损失函数
作者:Pin-Jui Ku,Chun-Wei Ho,Hao Yen,Sabato Marco Siniscalchi,Chin-Hui Lee
备注:5 pages, Submitted to ICASSP 2025
链接:点击下载PDF文件
【2】 Evaluation of state-of-the-art ASR Models in Child-Adult Interactions
标题: 儿童与成人互动中最先进的ASB模型的评估
作者:Aditya Ashvin,Rimita Lahiri,Aditya Kommineni,Somer Bishop,Catherine Lord,Sudarsana Reddy Kadiri,Shrikanth Narayanan
备注:5 pages, 3 figures, 4 tables
链接:点击下载PDF文件
【3】 Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
标题: 生成性语音基础模型预训练用于高质量语音提取和恢复
作者:Pin-Jui Ku,Alexander H. Liu,Roman Korostik,Sung-Feng Huang,Szu-Wei Fu,Ante Jukić
备注:5 pages, Submitted to ICASSP 2025. The implementation and configuration could be found in this https URL The audio demo page could be found in this https URL
链接:点击下载PDF文件
【4】 Scenario of Use Scheme: Threat Model Specification for Speaker Privacy Protection in the Medical Domain
标题: 使用场景方案:医疗领域演讲者隐私保护的威胁模型规范
作者:Mehtab Ur Rahman,Martha Larson,Louis ten Bosch,Cristian Tejedor-García
备注:Accepted and published at SPSC Symposium 2024 4th Symposium on Security and Privacy in Speech Communication. Interspeech 2024
链接:点击下载PDF文件
【5】 StyleSinger 2: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
标题: StyleSinger 2:具有风格转移和多层风格控制的Zero-Shot歌唱声音合成
作者:Yu Zhang,Ziyue Jiang,Ruiqi Li,Changhao Pan,Jinzheng He,Rongjie Huang,Chuxin Wang,Zhou Zhao
备注:Accepted by EMNLP 2024
链接:点击下载PDF文件
【6】 ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
标题: ESPnet-Codec:音频、音乐和语音神经编解码器的全面训练和评估
作者:Jiatong Shi,Jinchuan Tian,Yihan Wu,Jee-weon Jung,Jia Qi Yip,Yoshiki Masuyama,William Chen,Yuning Wu,Yuxun Tang,Massa Baali,Dareen Alharhi,Dong Zhang,Ruifan Deng,Tejes Srivastava,Haibin Wu,Alexander H. Liu,Bhiksha Raj,Qin Jin,Ruihua Song,Shinji Watanabe
备注:Accepted by SLT
链接:点击下载PDF文件
【7】 Interpolation filter design for sample rate independent audio effect RNNs
标题: 独立于采样率的音频效果RNN的内插过滤器设计
作者:Alistair Carson,Alec Wright,Stefan Bilbao
链接:点击下载PDF文件
【8】 Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
标题: 美杜莎耳朵里的低语:基于变形器的ASB的多头高效解码
作者:Yael Segal-Feldman,Aviv Shamsian,Aviv Navon,Gill Hetz,Joseph Keshet
备注:Under Review
链接:点击下载PDF文件
【9】 WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
标题: WeSep:一个可扩展且灵活的工具包,可推广的目标说话者提取
作者:Shuai Wang,Ke Zhang,Shaoxiong Lin,Junjie Li,Xuefei Wang,Meng Ge,Jianwei Yu,Yanmin Qian,Haizhou Li
备注:Interspeech 2024
链接:点击下载PDF文件
【10】 M-Vec: Matryoshka Speaker Embeddings with Flexible Dimensions
标题: M-Vec:具有灵活尺寸的Matryoshka扬声器嵌入式
作者:Shuai Wang,Pengcheng Zhu,Haizhou Li
备注:ICSR 2024, Shenzhen
链接:点击下载PDF文件
【11】 Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
标题: 利用随机选择策略最小化表示损失以实现高效的环境假音频检测
作者:Orchid Chetia Phukan,Girish,Mohd Mujtaba Akhtar,Swarup Ranjan Behera,Nitin Choudhury,Arun Balaji Buduru,Rajesh Sharma,S. R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【12】 Enhancing Open-Set Speaker Identification through Rapid Tuning with Speaker Reciprocal Points and Negative Sample
标题: 通过使用说话者互反点和负样本快速调整来增强开集说话者识别
作者:Zhiyong Chen,Zhiqi Ai,Xinnuo Li,Shugong Xu
Journal-ref:IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
【13】 StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
标题: StyleFusion TTC:用于Zero-Shot文本到语音合成的多模式风格控制和增强的特征融合
作者:Zhiyong Chen,Xinnuo Li,Zhiqi Ai,Shugong Xu
Journal-ref:The 7th Chinese Conference on Pattern Recognition and Computer Vision PRCV 2024
链接:点击下载PDF文件
【14】 Language-based Audio Moment Retrieval
标题: 基于百分比的音频时刻检索
作者:Hokuto Munakata,Taichi Nishimura,Shota Nakada,Tatsuya Komatsu
链接:点击下载PDF文件
【15】 Safe Guard: an LLM-agent for Real-time Voice-based Hate Speech Detection in Social Virtual Reality
标题: Safe Guard:用于社交虚拟现实中实时基于语音的仇恨言语检测的LLM代理
作者:Yiwen Xu,Qinyang Hou,Hongyu Wan,Mirjana Prpa
链接:点击下载PDF文件
【16】 Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
标题: 修改、推理和识别:通过特定描述和ASB错误纠正的基于LLM的情感识别
作者:Yuanchao Li,Yuan Gong,Chao-Han Huck Yang,Peter Bell,Catherine Lai
链接:点击下载PDF文件
【17】 Rethinking Emotion Bias in Music via Frechet Audio Distance
标题: 通过Frechet音频距离重新思考音乐中的情感偏见
作者:Yuanchao Li,Azalea Gui,Dimitra Emmanouilidou,Hannes Gamper
链接:点击下载PDF文件
【18】 Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
标题: Speech 2 rtMRI:语音引导扩散模型,用于语音期间人声实时MRI视频
作者:Hong Nguyen,Sean Foley,Kevin Huang,Xuan Shi,Tiantian Feng,Shrikanth Narayanan
备注:4 pages
链接:点击下载PDF文件
【19】 Blind Localization of Early Room Reflections with Arbitrary Microphone Array
标题: 用任意麦克风阵列实现早期房间反射的盲定位
作者:Yogev Hadadi,Vladimir Tourbabin,Zamir Ben-Hur,David Lou Alon,Boaz Rafaely
链接:点击下载PDF文件
【20】 The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
标题: 议会议事录中自动生成的语音和文本数据集的ParlaSpeech集合
作者:Nikola Ljubešić,Peter Rupnik,Danijel Koržinek
备注:Submitted to SPECOM 2024
链接:点击下载PDF文件
【21】 Toward Automated Clinical Transcriptions
标题: 迈向自动化临床Transit
作者:Mitchell A. Klusty,W. Vaiden Logan,Samuel E. Armstrong,Aaron D. Mullen,Caroline N. Leach,Jeff Talbert,V. K. Cody Bumgardner
备注:7 pages, 6 figures
链接:点击下载PDF文件
【22】 A Joint Spectro-Temporal Relational Thinking Based Acoustic Modeling Framework
标题: 基于频谱-时间关系思维的联合声学建模框架
作者:Zheng Nan,Ting Dang,Vidhyasaharan Sethu,Beena Ahmed
链接:点击下载PDF文件
【23】 TCG CREST System Description for the Second DISPLACE Challenge
标题: 第二次DISPLACE挑战赛的TCG CREST系统描述
作者:Nikhil Raghav,Subhajit Saha,Md Sahidullah,Swagatam Das
链接:点击下载PDF文件
【24】 Contextualization of ASR with LLM using phonetic retrieval-based augmentation
标题: 使用基于语音检索的增强将ASC与LLM进行上下文化
作者:Zhihong Lei,Xingyu Na,Mingbin Xu,Ernest Pusateri,Christophe Van Gysel,Yuanyuan Zhang,Shiyi Han,Zhen Huang
链接:点击下载PDF文件
【25】 A Large Dataset of Spontaneous Speech with the Accent Spoken in São Paulo for Automatic Speech Recognition Evaluation
标题: 圣保罗带有口音的自发语音大数据集用于自动语音识别评估
作者:Rodrigo Lima,Sidney Evaldo Leal,Arnaldo Candido Junior,Sandra Maria Aluísio
链接:点击下载PDF文件
【26】 WaveTransfer: A Flexible End-to-end Multi-instrument Timbre Transfer with Diffusion
标题: WaveTransfer:带扩散的灵活端到端多乐器音色传输
作者:Teysir Baoueb,Xiaoyu Bie,Hicham Janati,Gael Richard
Journal-ref:2024 IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2024), Sep 2024, London (UK), United Kingdom
链接:点击下载PDF文件
【27】 Equivariance-based self-supervised learning for audio signal recovery from clipped measurements
标题: 基于等效性的自我监督学习,从剪辑测量中恢复音频信号
作者:Victor Sechaud,Laurent Jacques,Patrice Abry,Julián Tachella
Journal-ref:EUSIPCO, Aug 2024, Lyon, France
摘要:
eess.AS
【1】 An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement标题: 用于相重建和语音增强的显式一致性保持损失函数
作者:Pin-Jui Ku,Chun-Wei Ho,Hao Yen,Sabato Marco Siniscalchi,Chin-Hui Lee
备注:5 pages, Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在这项工作中,我们提出了一种新的一致性保持损失函数恢复相位重建(PR)和语音增强(SE)的背景下的相位信息。与使用深度模型直接估计相位的传统技术不同,我们的想法是利用ad-hoc约束直接生成一致的幅度和相位对。具体地,所提出的损失迫使一组复数成为一致的短时傅立叶变换(STFT)表示,即,是真实信号的谱图。因此,我们的方法避免了估计原始相位的困难,这是高度非结构化和敏感的时移。我们提出的损失的影响,首先评估PR任务,实验证明,我们的方法是可行的。接下来,我们使用VB-DMD和WSJ 0-CHiME 3数据集展示了其在SE任务上的有效性。在VB-DMD上,我们的方法与传统解决方案相比具有竞争力。在具有挑战性的WSJ 0-CHiME 3集上,所提出的框架与显式估计相位的技术相比是有利的。摘要:In this work, we propose a novel consistency-preserving loss function for recovering the phase information in the context of phase reconstruction (PR) and speech enhancement (SE). Different from conventional techniques that directly estimate the phase using a deep model, our idea is to exploit ad-hoc constraints to directly generate a consistent pair of magnitude and phase. Specifically, the proposed loss forces a set of complex numbers to be a consistent short-time Fourier transform (STFT) representation, i.e., to be the spectrogram of a real signal. Our approach thus avoids the difficulty of estimating the original phase, which is highly unstructured and sensitive to time shift. The influence of our proposed loss is first assessed on a PR task, experimentally demonstrating that our approach is viable. Next, we show its effectiveness on an SE task, using both the VB-DMD and WSJ0-CHiME3 data sets. On VB-DMD, our approach is competitive with conventional solutions. On the challenging WSJ0-CHiME3 set, the proposed framework compares favourably over those techniques that explicitly estimate the phase.
【2】 Evaluation of state-of-the-art ASR Models in Child-Adult Interactions
标题: 儿童与成人互动中最先进的ASB模型的评估
作者:Aditya Ashvin,Rimita Lahiri,Aditya Kommineni,Somer Bishop,Catherine Lord,Sudarsana Reddy Kadiri,Shrikanth Narayanan
备注:5 pages, 3 figures, 4 tables
链接:点击下载PDF文件
摘要:在临床环境中可靠地转录儿童-成人对话的能力对于诊断和理解许多发育障碍如自闭症谱系障碍是有价值的。深度学习架构的最新进展和大规模转录数据的可用性已经导致语音基础模型的开发,这些模型在ASR性能方面表现出显着的改进。然而,这些模型的能力,以及转化为对话的儿童与成人的互动正在研究中。在这项工作中,我们使用Whisper,Wav 2 Vec 2,HuBERT和WavLM对包含自闭症诊断会话中儿童-成人互动的数据集进行了ASR性能的全面评估。我们发现,语音基础模型显示出显着的性能下降(15-20%的绝对WER)的儿童语音相比,成人语音在会话设置。然后,我们采用LoRA的最佳性能的zero shot模型(耳语大),以探讨在低资源设置微调的有效性,导致约8%的绝对WER改善儿童语音和约13%的绝对WER改善成人语音。摘要:The ability to reliably transcribe child-adult conversations in a clinical setting is valuable for diagnosis and understanding of numerous developmental disorders such as Autism Spectrum Disorder. Recent advances in deep learning architectures and availability of large scale transcribed data has led to development of speech foundation models that have shown dramatic improvements in ASR performance. However, the ability of these models to translate well to conversational child-adult interactions is under studied. In this work, we provide a comprehensive evaluation of ASR performance on a dataset containing child-adult interactions from autism diagnostic sessions, using Whisper, Wav2Vec2, HuBERT, and WavLM. We find that speech foundation models show a noticeable performance drop (15-20% absolute WER) for child speech compared to adult speech in the conversational setting. Then, we employ LoRA on the best performing zero shot model (whisper-large) to probe the effectiveness of fine-tuning in a low resource setting, resulting in ~8% absolute WER improvement for child speech and ~13% absolute WER improvement for adult speech.
【3】 Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
标题: 生成性语音基础模型预训练用于高质量语音提取和恢复
作者:Pin-Jui Ku,Alexander H. Liu,Roman Korostik,Sung-Feng Huang,Szu-Wei Fu,Ante Jukić
备注:5 pages, Submitted to ICASSP 2025. The implementation and configuration could be found in this https URL The audio demo page could be found in this https URL
链接:点击下载PDF文件
摘要:本文提出了一种高质量语音恢复任务的生成式预训练基础模型。通过直接对复值短时傅里叶变换系数进行运算,我们的模型不依赖于任何声码器进行时域信号重构。其结果是,我们的模型简化了合成过程,并删除了质量上限引入的任何梅尔频谱声码器相比,以前的工作SpeechFlow。该方法在多个语音恢复任务上进行了评估,包括语音去噪、带宽扩展、编解码器伪影去除和目标说话人提取。在所有情况下,微调我们的预训练模型都会导致优于强基线的性能。值得注意的是,在目标说话人提取任务中,我们的模型优于现有系统,包括那些利用SSL预训练编码器(如WavLM)的系统。代码和预训练的检查点在NVIDIA NeMo框架中公开提供。摘要:This paper proposes a generative pretraining foundation model for high-quality speech restoration tasks. By directly operating on complex-valued short-time Fourier transform coefficients, our model does not rely on any vocoders for time-domain signal reconstruction. As a result, our model simplifies the synthesis process and removes the quality upper-bound introduced by any mel-spectrogram vocoder compared to prior work SpeechFlow. The proposed method is evaluated on multiple speech restoration tasks, including speech denoising, bandwidth extension, codec artifact removal, and target speaker extraction. In all scenarios, finetuning our pretrained model results in superior performance over strong baselines. Notably, in the target speaker extraction task, our model outperforms existing systems, including those leveraging SSL-pretrained encoders like WavLM. The code and the pretrained checkpoints are publicly available in the NVIDIA NeMo framework.
【4】 Scenario of Use Scheme: Threat Model Specification for Speaker Privacy Protection in the Medical Domain
标题: 使用场景方案:医疗领域演讲者隐私保护的威胁模型规范
作者:Mehtab Ur Rahman,Martha Larson,Louis ten Bosch,Cristian Tejedor-García
备注:Accepted and published at SPSC Symposium 2024 4th Symposium on Security and Privacy in Speech Communication. Interspeech 2024
链接:点击下载PDF文件
摘要:语音记录越来越多地用于检测和监测疾病,导致隐私问题。除了密码学之外,还可以通过扰动、解纠缠和重新合成等方法来解决语音保护问题,这些方法可以消除说话者的敏感信息,留下医学分析所需的信息。为了制定这种隐私保护办法,有必要对有关医疗环境和医疗专业人员需求的假设作出明确和系统的说明。在本文中,我们提出了一个场景的使用计划,其中包括一个攻击者模型,其特点是对手对谁的发言人的隐私必须捍卫,和一个保护者模型,指定的防御。我们讨论了该计划与以前的工作语音隐私的连接。最后,我们提出了一个具体的例子,一个指定的使用场景和一组实验,保护说话人数据免受性别推断攻击,同时保持实用程序帕金森氏症的检测。摘要:Speech recordings are being more frequently used to detect and monitor disease, leading to privacy concerns. Beyond cryptography, protection of speech can be addressed by approaches, such as perturbation, disentanglement, and re-synthesis, that eliminate sensitive information of the speaker, leaving the information necessary for medical analysis purposes. In order for such privacy protective approaches to be developed, clear and systematic specifications of assumptions concerning medical settings and the needs of medical professionals are necessary. In this paper, we propose a Scenario of Use Scheme that incorporates an Attacker Model, which characterizes the adversary against whom the speaker's privacy must be defended, and a Protector Model, which specifies the defense. We discuss the connection of the scheme with previous work on speech privacy. Finally, we present a concrete example of a specified Scenario of Use and a set of experiments about protecting speaker data against gender inference attacks while maintaining utility for Parkinson's detection.
【5】 StyleSinger 2: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
标题: StyleSinger 2:具有风格转移和多层风格控制的Zero-Shot歌唱声音合成
作者:Yu Zhang,Ziyue Jiang,Ruiqi Li,Changhao Pan,Jinzheng He,Rongjie Huang,Chuxin Wang,Zhou Zhao
备注:Accepted by EMNLP 2024
链接:点击下载PDF文件
摘要:带风格转换和风格控制的Zero-shot歌唱声音合成(SVS)旨在从音频和文本提示中生成具有不可见音色和风格(包括演唱方法、情感、节奏、技术和发音)的高质量歌唱声音。然而,歌唱风格的多面性对有效的建模,转移和控制提出了重大挑战。此外,目前的SVS模型往往无法为看不见的歌手生成丰富的风格细微差别的歌声。为了应对这些挑战,我们引入了StyleSinger 2,这是第一个zero-shot SVS模型,用于跨语言语音和演唱风格的风格转移,以及多级风格控制。具体地说,StyleSinger 2提出了三个主要模块:1)聚类风格编码器采用聚类矢量量化模型,将风格信息稳定地压缩到一个紧凑的潜在空间中; 2)风格和时长语言模型(S &D-LM)同时预测风格信息和音素时长,这对两者都有好处; 3)风格自适应解码器使用新颖的Mel风格自适应归一化方法来生成具有增强的细节的歌声。实验结果表明,StyleSinger 2在合成质量、歌手相似度和风格可控性方面优于所有基线模型,包括zero-shot风格迁移、多级风格控制、跨语言风格迁移和语音到歌唱风格迁移。可以在https: stylesinger2.github.io 上访问唱歌的声音样本。摘要:Zero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunciation) from audio and text prompts. However, the multifaceted nature of singing styles poses a significant challenge for effective modeling, transfer, and control. Furthermore, current SVS models often fail to generate singing voices rich in stylistic nuances for unseen singers. To address these challenges, we introduce StyleSinger 2, the first zero-shot SVS model for style transfer across cross-lingual speech and singing styles, along with multi-level style control. Specifically, StyleSinger 2 proposes three primary modules: 1) the clustering style encoder employs a clustering vector quantization model to stably condense style information into a compact latent space; 2) the Style and Duration Language Model (S &D-LM) concurrently predicts style information and phoneme duration, which benefits both; 3) the style adaptive decoder uses a novel mel-style adaptive normalization method to generate singing voices with enhanced details. Experimental results show that StyleSinger 2 outperforms all baseline models in synthesis quality, singer similarity, and style controllability across various tasks, including zero-shot style transfer, multi-level style control, cross-lingual style transfer, and speech-to-singing style transfer. Singing voice samples can be accessed at https: stylesinger2.github.io .
【6】 ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
标题: ESPnet-Codec:音频、音乐和语音神经编解码器的全面训练和评估
作者:Jiatong Shi,Jinchuan Tian,Yihan Wu,Jee-weon Jung,Jia Qi Yip,Yoshiki Masuyama,William Chen,Yuning Wu,Yuxun Tang,Massa Baali,Dareen Alharhi,Dong Zhang,Ruifan Deng,Tejes Srivastava,Haibin Wu,Alexander H. Liu,Bhiksha Raj,Qin Jin,Ruihua Song,Shinji Watanabe
备注:Accepted by SLT
链接:点击下载PDF文件
摘要:神经编解码器已经成为最近的语音和音频生成研究的关键。除了信号压缩能力之外,还发现离散编解码器可以增强下游训练效率和与自回归语言模型的兼容性。然而,随着广泛的下游应用程序的调查,在确保不同应用程序之间的公平比较方面出现了挑战。为了解决这些问题,我们提出了一个新的开源平台ESPnet-Codec,该平台构建在ESPnet上,专注于神经编解码器的训练和评估。ESPnet-Codec提供音频、音乐和语音方面的各种配方,用于使用几种广泛采用的编解码器模型进行训练和评估。与ESPnet-Codec一起,我们提出了VERSA,一个独立的评估工具包,它提供了超过20个音频评估指标的编解码器性能的全面评估。值得注意的是,我们证明了ESPnet-Codec可以集成到六个ESPnet任务中,支持不同的应用程序。摘要:Neural codecs have become crucial to recent speech and audio generation research. In addition to signal compression capabilities, discrete codecs have also been found to enhance downstream training efficiency and compatibility with autoregressive language models. However, as extensive downstream applications are investigated, challenges have arisen in ensuring fair comparisons across diverse applications. To address these issues, we present a new open-source platform ESPnet-Codec, which is built on ESPnet and focuses on neural codec training and evaluation. ESPnet-Codec offers various recipes in audio, music, and speech for training and evaluation using several widely adopted codec models. Together with ESPnet-Codec, we present VERSA, a standalone evaluation toolkit, which provides a comprehensive evaluation of codec performance over 20 audio evaluation metrics. Notably, we demonstrate that ESPnet-Codec can be integrated into six ESPnet tasks, supporting diverse applications.
【7】 Interpolation filter design for sample rate independent audio effect RNNs
标题: 独立于采样率的音频效果RNN的内插过滤器设计
作者:Alistair Carson,Alec Wright,Stefan Bilbao
链接:点击下载PDF文件
摘要:递归神经网络(RNN)在模拟模拟吉他放大器的非线性、有状态行为和失真效果方面是有效的。与直接电路仿真的情况不同,RNN具有编码在其模型权重中的固定采样率,使得采样率在推理期间不可调整。最近的工作提出了通过增加样本中的反馈延迟长度来增加RNN在推理(过采样)时的采样率,使用分数延迟滤波器进行非整数转换。在这里,我们调查的任务,降低采样率的推断(欠采样),并建议使用外推滤波器来近似所需的分数信号提前。我们考虑了两种滤波器设计方法,并分析了滤波器阶数对音频质量的影响。我们的研究结果表明,滤波器的正确选择可以提供高质量的过采样和欠采样的结果,但是,在某些情况下,采样率的调整会导致输出信号中不必要的伪影。我们通过线性稳定性分析来分析这些失效情况,表明它们是由固定点周围的不稳定性引起的。这种方法能够在运行之前为给定的RNN模型提供合适的插值滤波器的知情预测。摘要:Recurrent neural networks (RNNs) are effective at emulating the non-linear, stateful behavior of analog guitar amplifiers and distortion effects. Unlike the case of direct circuit simulation, RNNs have a fixed sample rate encoded in their model weights, making the sample rate non-adjustable during inference. Recent work has proposed increasing the sample rate of RNNs at inference (oversampling) by increasing the feedback delay length in samples, using a fractional delay filter for non-integer conversions. Here, we investigate the task of lowering the sample rate at inference (undersampling), and propose using an extrapolation filter to approximate the required fractional signal advance. We consider two filter design methods and analyze the impact of filter order on audio quality. Our results show that the correct choice of filter can give high quality results for both oversampling and undersampling; however, in some cases the sample rate adjustment leads to unwanted artefacts in the output signal. We analyse these failure cases through linearised stability analysis, showing that they result from instability around a fixed point. This approach enables an informed prediction of suitable interpolation filters for a given RNN model before runtime.
【8】 Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
标题: 美杜莎耳朵里的低语:基于变形器的ASB的多头高效解码
作者:Yael Segal-Feldman,Aviv Shamsian,Aviv Navon,Gill Hetz,Joseph Keshet
备注:Under Review
链接:点击下载PDF文件
摘要:基于transformer的大型模型在语音转录和翻译方面具有巨大的潜力。它们的自我注意机制和并行处理使它们能够捕捉音频序列中的复杂模式和依赖关系。然而,这种潜力带来了挑战,因为这些大型和计算密集型模型导致推理速度缓慢。已经提出了各种优化策略来提高性能,包括有效的硬件利用率和算法增强。在本文中,我们介绍了耳语美杜莎,一种新的方法,旨在提高处理速度,最小的影响字错误率(WER)。该模型扩展了OpenAI的Whisper架构,每次迭代预测多个令牌,从而减少了50%的延迟。我们展示了Whisper-Medusa在不同学习设置和数据集上的有效性。摘要:Large transformer-based models have significant potential for speech transcription and translation. Their self-attention mechanisms and parallel processing enable them to capture complex patterns and dependencies in audio sequences. However, this potential comes with challenges, as these large and computationally intensive models lead to slow inference speeds. Various optimization strategies have been proposed to improve performance, including efficient hardware utilization and algorithmic enhancements. In this paper, we introduce Whisper-Medusa, a novel approach designed to enhance processing speed with minimal impact on Word Error Rate (WER). The proposed model extends the OpenAI's Whisper architecture by predicting multiple tokens per iteration, resulting in a 50% reduction in latency. We showcase the effectiveness of Whisper-Medusa across different learning setups and datasets.
【9】 WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
标题: WeSep:一个可扩展且灵活的工具包,可推广的目标说话者提取
作者:Shuai Wang,Ke Zhang,Shaoxiong Lin,Junjie Li,Xuefei Wang,Meng Ge,Jianwei Yu,Yanmin Qian,Haizhou Li
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:目标说话人提取(TSE)的核心是从多说话人重叠的语音中分离出特定目标说话人的语音,这是鸡尾酒会问题中的一个典型设置。近年来,TSE由于其在用户定制界面和助听器等各种应用中的潜力,或者作为语音识别和说话人识别等基本任务的关键前端处理技术,引起了越来越多的关注。然而,目前很少有开源工具包或可供现成使用的预训练模型。在这项工作中,我们介绍了WeSep,一个工具包,专为研究和实际应用在TSE。WeSep具有灵活的目标扬声器建模,可扩展的数据管理,有效的实时数据模拟,结构化配方和部署支持。该工具包在 url{https: github.com wenet-e2e WeSep.}上公开提供。摘要:Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increasing attention due to its potential for various applications such as user-customized interfaces and hearing aids, or as a crutial front-end processing technologies for subsequential tasks such as speech recognition and speaker recongtion. However, there are currently few open-source toolkits or available pre-trained models for off-the-shelf usage. In this work, we introduce WeSep, a toolkit designed for research and practical applications in TSE. WeSep is featured with flexible target speaker modeling, scalable data management, effective on-the-fly data simulation, structured recipes and deployment support. The toolkit is publicly avaliable at url{https: github.com wenet-e2e WeSep.}
【10】 M-Vec: Matryoshka Speaker Embeddings with Flexible Dimensions
标题: M-Vec:具有灵活尺寸的Matryoshka扬声器嵌入式
作者:Shuai Wang,Pengcheng Zhu,Haizhou Li
备注:ICSR 2024, Shenzhen
链接:点击下载PDF文件
摘要:固定维度的说话人嵌入已经成为说话人建模的主要方法,通常跨越数百到数千个维度。这些维度是超参数,它们不是专门挑选的,也不是按照重要性进行层次排序的。在大规模说话人表征数据库中,降低嵌入维数可以显著降低存储和计算成本。然而,直接训练低维表示通常会产生次优性能。在本文中,我们介绍了Matryoshka扬声器嵌入,一种方法,允许动态提取子维度的嵌入,同时保持性能。我们的方法在VoxCeleb数据集上进行了验证,证明它可以实现极低维的嵌入,如8维,同时保持较高的说话人验证性能。摘要:Fixed-dimensional speaker embeddings have become the dominant approach in speaker modeling, typically spanning hundreds to thousands of dimensions. These dimensions are hyperparameters that are not specifically picked, nor are they hierarchically ordered in terms of importance. In large-scale speaker representation databases, reducing the dimensionality of embeddings can significantly lower storage and computational costs. However, directly training low-dimensional representations often yields suboptimal performance. In this paper, we introduce the Matryoshka speaker embedding, a method that allows dynamic extraction of sub-dimensions from the embedding while maintaining performance. Our approach is validated on the VoxCeleb dataset, demonstrating that it can achieve extremely low-dimensional embeddings, such as 8 dimensions, while preserving high speaker verification performance.
【11】 Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
标题: 利用随机选择策略最小化表示损失以实现高效的环境假音频检测
作者:Orchid Chetia Phukan,Girish,Mohd Mujtaba Akhtar,Swarup Ranjan Behera,Nitin Choudhury,Arun Balaji Buduru,Rajesh Sharma,S. R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:基础模型的适应性极大地推进了环境音频深度伪造检测(EADD),这是一个快速增长的研究领域。这些模型通常被微调或在其冻结状态下用于下游任务。然而,其表示的维度可以实质上导致下游模型的高参数计数,从而导致更高的计算需求。因此,一般的方法是通过利用最先进的(SOTA)无监督降维技术(PCA,SVD,KPCA,GRP)来压缩这些表示,以实现有效的EADD。然而,随着这些技术的应用,我们观察到性能下降。因此,在本文中,我们证明了表示向量包含冗余信息,随机选择40-50%的表示值并在其上构建下游模型可以保留甚至有时可以提高性能。我们表明,这种随机选择保留更多的性能比SOTA降维技术,同时减少模型参数和推理时间几乎一半以上。摘要:The adaptation of foundation models has significantly advanced environmental audio deepfake detection (EADD), a rapidly growing area of research. These models are typically fine-tuned or utilized in their frozen states for downstream tasks. However, the dimensionality of their representations can substantially lead to a high parameter count of downstream models, leading to higher computational demands. So, a general way is to compress these representations by leveraging state-of-the-art (SOTA) unsupervised dimensionality reduction techniques (PCA, SVD, KPCA, GRP) for efficient EADD. However, with the application of such techniques, we observe a drop in performance. So in this paper, we show that representation vectors contain redundant information, and randomly selecting 40-50% of representation values and building downstream models on it preserves or sometimes even improves performance. We show that such random selection preserves more performance than the SOTA dimensionality reduction techniques while reducing model parameters and inference time by almost over half.
【12】 Enhancing Open-Set Speaker Identification through Rapid Tuning with Speaker Reciprocal Points and Negative Sample
标题: 通过使用说话者互反点和负样本快速调整来增强开集说话者识别
作者:Zhiyong Chen,Zhiqi Ai,Xinnuo Li,Shugong Xu
Journal-ref:IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
摘要:本文介绍了一种新的框架开放集说话人识别在家庭环境中,发挥了至关重要的作用,促进无缝的人机交互。针对当前说话人模型和分类方法的局限性,我们的工作将预训练的WavLM前端与用于注册的Few-Shot快速调谐神经网络(NN)后端集成在一起,采用任务优化的说话人倒数点学习(SRPL)来增强对多个目标说话人的区分。此外,我们提出了一个增强版的SRPL(SRPL+),它结合了负样本学习与语音合成和真正的负样本,以显着提高开集SID的准确性。我们的方法在各种多语言文本相关的说话人识别数据集进行了全面评估,证明了其在复杂的家庭多说话人识别场景中实现高可用性的有效性。与直接使用高效的WavLM基+模型相比,该系统将开集性能提高了27%.摘要:This paper introduces a novel framework for open-set speaker identification in household environments, playing a crucial role in facilitating seamless human-computer interactions. Addressing the limitations of current speaker models and classification approaches, our work integrates an pretrained WavLM frontend with a few-shot rapid tuning neural network (NN) backend for enrollment, employing task-optimized Speaker Reciprocal Points Learning (SRPL) to enhance discrimination across multiple target speakers. Furthermore, we propose an enhanced version of SRPL (SRPL+), which incorporates negative sample learning with both speech-synthesized and real negative samples to significantly improve open-set SID accuracy. Our approach is thoroughly evaluated across various multi-language text-dependent speaker recognition datasets, demonstrating its effectiveness in achieving high usability for complex household multi-speaker recognition scenarios. The proposed system enhanced open-set performance by up to 27 % over the directly use of efficient WavLM base+ model.
【13】 StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
标题: StyleFusion TTC:用于Zero-Shot文本到语音合成的多模式风格控制和增强的特征融合
作者:Zhiyong Chen,Xinnuo Li,Zhiqi Ai,Shugong Xu
Journal-ref:The 7th Chinese Conference on Pattern Recognition and Computer Vision PRCV 2024
链接:点击下载PDF文件
摘要:我们介绍StyleFusion-TTS,一个提示和 或音频参考,风格和说话人可控,zero-shot文本到语音(TTS)合成系统,旨在提高当前研究文献的可编辑性和自然性。我们提出了一个通用的前端编码器作为一个紧凑和有效的模块,利用多模态输入,包括文本提示,音频参考,扬声器音色参考在一个完全zero-shot的方式,并产生解开的风格和扬声器控制嵌入。我们的新方法还利用了一个分层的构象结构的风格和扬声器控制嵌入的融合,旨在实现最佳的功能融合在当前先进的TTS架构。StyleFusion-TTS通过主观和客观的多个指标进行评估。该系统在我们的评估中表现出良好的性能,这表明它有潜力为zero-shot文本到语音合成领域的进步做出贡献。摘要:We introduce StyleFusion-TTS, a prompt and or audio referenced, style and speaker-controllable, zero-shot text-to-speech (TTS) synthesis system designed to enhance the editability and naturalness of current research literature. We propose a general front-end encoder as a compact and effective module to utilize multimodal inputs including text prompts, audio references, and speaker timbre references in a fully zero-shot manner and produce disentangled style and speaker control embeddings. Our novel approach also leverages a hierarchical conformer structure for the fusion of style and speaker control embeddings, aiming to achieve optimal feature fusion within the current advanced TTS architecture. StyleFusion-TTS is evaluated through multiple metrics, both subjectively and objectively. The system shows promising performance across our evaluations, suggesting its potential to contribute to the advancement of the field of zero-shot text-to-speech synthesis.
【14】 Language-based Audio Moment Retrieval
标题: 基于百分比的音频时刻检索
作者:Hokuto Munakata,Taichi Nishimura,Shota Nakada,Tatsuya Komatsu
链接:点击下载PDF文件
摘要:在本文中,我们提出并设计了一个新的任务,称为音频时刻检索(AMR)。与从音频数据库中搜索短音频片段的传统基于语言的音频检索任务不同,AMR旨在基于文本查询预测未修剪长音频中的相关时刻。鉴于AMR之前缺乏相关工作,我们首先构建了一个专用数据集Clotho-Moment,由带有时刻注释的大规模模拟音频记录组成。然后,我们提出了一个基于DETR的模型,命名为音频时刻DETR(AM-DETR),AMR任务的基本框架。该模型捕获音频特征的时间依赖性,灵感来自类似的视频时刻检索任务,从而超越了传统的剪辑级音频检索方法。此外,我们提供手动注释的数据集,以正确衡量我们的方法在真实数据上的有效性和鲁棒性。实验结果表明,使用Clotho-Moment训练的AM-DETR的性能优于在所有指标上应用带有滑动窗口的剪辑级音频检索方法的基线模型,特别是将Recall1@0.7提高了9.00个点。我们的数据集和代码可在https: h-munakata.github.io Language-based-Audio-Moment-Retrieval上公开获取。摘要:In this paper, we propose and design a new task called audio moment retrieval (AMR). Unlike conventional language-based audio retrieval tasks that search for short audio clips from an audio database, AMR aims to predict relevant moments in untrimmed long audio based on a text query. Given the lack of prior work in AMR, we first build a dedicated dataset, Clotho-Moment, consisting of large-scale simulated audio recordings with moment annotations. We then propose a DETR-based model, named Audio Moment DETR (AM-DETR), as a fundamental framework for AMR tasks. This model captures temporal dependencies within audio features, inspired by similar video moment retrieval tasks, thus surpassing conventional clip-level audio retrieval methods. Additionally, we provide manually annotated datasets to properly measure the effectiveness and robustness of our methods on real data. Experimental results show that AM-DETR, trained with Clotho-Moment, outperforms a baseline model that applies a clip-level audio retrieval method with a sliding window on all metrics, particularly improving Recall1@0.7 by 9.00 points. Our datasets and code are publicly available in https: h-munakata.github.io Language-based-Audio-Moment-Retrieval.
【15】 Safe Guard: an LLM-agent for Real-time Voice-based Hate Speech Detection in Social Virtual Reality
标题: Safe Guard:用于社交虚拟现实中实时基于语音的仇恨言语检测的LLM代理
作者:Yiwen Xu,Qinyang Hou,Hongyu Wan,Mirjana Prpa
链接:点击下载PDF文件
摘要:在本文中,我们提出了安全卫士,一个LLM代理,用于检测社交VR(VRChat)中基于语音的交互中的仇恨言论。我们的系统利用Open AI GPT和音频特征提取进行实时语音交互。我们贡献了一个系统的设计和评估系统,证明了我们的方法在检测仇恨言论的能力,并减少误报相比,目前可用的方法。我们的研究结果表明,基于LLM的代理在创建更安全的虚拟环境中的潜力,并为LLM驱动的适度方法的进一步发展奠定了基础。摘要:In this paper, we present Safe Guard, an LLM-agent for the detection of hate speech in voice-based interactions in social VR (VRChat). Our system leverages Open AI GPT and audio feature extraction for real-time voice interactions. We contribute a system design and evaluation of the system that demonstrates the capability of our approach in detecting hate speech, and reducing false positives compared to currently available approaches. Our results indicate the potential of LLM-based agents in creating safer virtual environments and set the groundwork for further advancements in LLM-driven moderation approaches.
【16】 Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
标题: 修改、推理和识别:通过特定描述和ASB错误纠正的基于LLM的情感识别
作者:Yuanchao Li,Yuan Gong,Chao-Han Huck Yang,Peter Bell,Catherine Lai
链接:点击下载PDF文件
摘要:随着大语言模型(LLM)的发展,最近出现了使用提示工程来注释和识别语音情感的方法,但其有效性和可靠性仍然值得怀疑。在本文中,我们对这一主题进行了系统的研究,首先提出了新的提示,结合声学,语言学和心理学的情感特定的知识。随后,我们研究了基于LLM的提示对自动语音识别(ASR)转录的有效性,将其与地面实况转录进行对比。此外,我们提出了一个修改原因识别提示管道,用于从具有ASR错误的口语中进行鲁棒的基于LLM的情感识别。此外,上下文感知学习,在上下文学习和指令调整的实验进行检查LLM培训计划在这个方向上的有用性。最后,我们调查的灵敏度LLM轻微提示变化。实验结果表明,情感特定的提示,ASR纠错,和LLM训练计划的LLM为基础的情感识别的功效。我们的研究旨在改进LLM在情感识别和相关领域的使用。摘要:Annotating and recognizing speech emotion using prompt engineering has recently emerged with the advancement of Large Language Models (LLMs), yet its efficacy and reliability remain questionable. In this paper, we conduct a systematic study on this topic, beginning with the proposal of novel prompts that incorporate emotion-specific knowledge from acoustics, linguistics, and psychology. Subsequently, we examine the effectiveness of LLM-based prompting on Automatic Speech Recognition (ASR) transcription, contrasting it with ground-truth transcription. Furthermore, we propose a Revise-Reason-Recognize prompting pipeline for robust LLM-based emotion recognition from spoken language with ASR errors. Additionally, experiments on context-aware learning, in-context learning, and instruction tuning are performed to examine the usefulness of LLM training schemes in this direction. Finally, we investigate the sensitivity of LLMs to minor prompt variations. Experimental results demonstrate the efficacy of the emotion-specific prompts, ASR error correction, and LLM training schemes for LLM-based emotion recognition. Our study aims to refine the use of LLMs in emotion recognition and related domains.
【17】 Rethinking Emotion Bias in Music via Frechet Audio Distance
标题: 通过Frechet音频距离重新思考音乐中的情感偏见
作者:Yuanchao Li,Azalea Gui,Dimitra Emmanouilidou,Hannes Gamper
链接:点击下载PDF文件
摘要:音乐情感的主观性质在识别和生成中引入了固有的偏差,特别是当依赖于单个音频编码器、情感分类器或评估度量时。在这项工作中,我们对音乐情感识别(MER)和情感音乐生成(EMG)进行了研究,采用了不同的音频编码器以及Frechet音频距离(FAD),这是一种无参考的评估指标。我们的研究从MER的基准评估开始,强调了与使用单个音频编码器相关的限制以及在不同测量中观察到的差异。然后,我们建议使用来自多个编码器的FAD来评估MER性能,以提供更客观的音乐情感测量。此外,我们引入了一个增强的EMG方法,旨在提高所产生的音乐情感的变化和突出,从而提高现实主义。此外,我们研究了真实音乐和合成音乐中传达的情感之间的现实主义差异,将我们的EMG模型与两个基线模型进行比较。实验结果强调了MER和EMG中的情感偏差问题,并展示了使用FAD和各种音频编码器来客观评估音乐情感的潜力。摘要:The subjective nature of music emotion introduces inherent bias in both recognition and generation, especially when relying on a single audio encoder, emotion classifier, or evaluation metric. In this work, we conduct a study on Music Emotion Recognition (MER) and Emotional Music Generation (EMG), employing diverse audio encoders alongside the Frechet Audio Distance (FAD), a reference-free evaluation metric. Our study begins with a benchmark evaluation of MER, highlighting the limitations associated with using a single audio encoder and the disparities observed across different measurements. We then propose assessing MER performance using FAD from multiple encoders to provide a more objective measure of music emotion. Furthermore, we introduce an enhanced EMG approach designed to improve both the variation and prominence of generated music emotion, thus enhancing realism. Additionally, we investigate the realism disparities between the emotions conveyed in real and synthetic music, comparing our EMG model against two baseline models. Experimental results underscore the emotion bias problem in both MER and EMG and demonstrate the potential of using FAD and diverse audio encoders to evaluate music emotion objectively.
【18】 Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
标题: Speech 2 rtMRI:语音引导扩散模型,用于语音期间人声实时MRI视频
作者:Hong Nguyen,Sean Foley,Kevin Huang,Xuan Shi,Tiantian Feng,Shrikanth Narayanan
备注:4 pages
链接:点击下载PDF文件
摘要:从视觉上和运动学上理解语音产生可以为第二语言学习系统的设计提供信息,以及在视频游戏和动画中创建说话的角色。在这项工作中,我们介绍了一种数据驱动的方法来直观地表示发音关节运动的磁共振成像(MRI)视频中的人类声道在讲话的基础上,任意的音频或语音输入。我们利用嵌入先验知识的大型预训练语音模型,使用语音到视频扩散模型将视觉域推广到看不见的数据。我们的研究结果表明,视觉生成显着受益于预训练的语音表示。我们还观察到,孤立地评估音素是具有挑战性的,但在口语的背景下进行评估时变得更加简单。目前的研究结果的局限性包括舌头接触上颚时存在舌头运动不平滑和视频失真。摘要:Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven method to visually represent articulator motion in Magnetic Resonance Imaging (MRI) videos of the human vocal tract during speech based on arbitrary audio or speech input. We leverage large pre-trained speech models, which are embedded with prior knowledge, to generalize the visual domain to unseen data using a speech-to-video diffusion model. Our findings demonstrate that the visual generation significantly benefits from the pre-trained speech representations. We also observed that evaluating phonemes in isolation is challenging but becomes more straightforward when assessed within the context of spoken words. Limitations of the current results include the presence of unsmooth tongue motion and video distortion when the tongue contacts the palate.
【19】 Blind Localization of Early Room Reflections with Arbitrary Microphone Array
标题: 用任意麦克风阵列实现早期房间反射的盲定位
作者:Yogev Hadadi,Vladimir Tourbabin,Zamir Ben-Hur,David Lou Alon,Boaz Rafaely
链接:点击下载PDF文件
摘要:在没有房间脉冲响应或源信号的先验知识的情况下,盲目地估计早期房间反射的到达方向(DoA)在音频信号处理应用中是非常有价值的。FF-PHALCOR(频率聚焦相位对准相关)方法最近开发了用于此目的,扩展了原来的PHALCOR方法工作与任意阵列,而不仅仅是球形的。之前的研究仅对其性能提供了初步见解。这项研究提供了一个全面的分析方法的性能和局限性,研究反射特性,如延迟,幅度和空间密度如何影响其有效性。该研究还提出了改进措施来克服这些限制,提高检测质量并减少误报。此外,该研究还研究了使用估计的反射信息生成房间脉冲响应如何影响空间感知。研究结果表明,所提出的方法在基线上具有感知优势,当使用具有32个麦克风的球形阵列时具有特别高的感知质量。然而,当使用仅具有6个麦克风的半圆形阵列时,质量有所降低。摘要:Blindly estimating the direction of arrival (DoA) of early room reflections without prior knowledge of the room impulse response or source signal is highly valuable in audio signal processing applications. The FF-PHALCOR (Frequency Focusing PHase ALigned CORrelation) method was recently developed for this purpose, extending the original PHALCOR method to work with arbitrary arrays rather than just spherical ones. Previous studies have provided only initial insights into its performance. This study offers a comprehensive analysis of the method's performance and limitations, examining how reflection characteristics such as delay, amplitude, and spatial density affect its effectiveness. The research also proposes improvements to overcome these limitations, enhancing detection quality and reducing false alarms. Additionally, the study examined how spatial perception is affected by generating room impulse responses using estimated reflection information. The findings suggest a perceptual advantage of the proposed approach over the baseline, with particularly high perceptual quality when using the spherical array with 32 microphones. However, the quality is somewhat reduced when using a semi-circular array with only 6 microphones.
【20】 The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
标题: 议会议事录中自动生成的语音和文本数据集的ParlaSpeech集合
作者:Nikola Ljubešić,Peter Rupnik,Danijel Koržinek
备注:Submitted to SPECOM 2024
链接:点击下载PDF文件
摘要:最近语音和语言技术的重大改进来自于对原始语言数据的自我监督方法以及各种类型的显式监督。为了确保对语音数据的高质量处理,最有用的显式监督类型仍然是语音信号与其对应的文本转录之间的对齐,这是一种不适用于许多语言的数据类型。在本文中,我们提出了一种基于议会会议记录及其录音的方法来构建资源较少的语言的大型开放式语音和文本对齐数据集。我们的出发点是ParlaMint可比语料库的26个国家的欧洲议会的议会议事记录。在扩大ParlaMint语料库的试点运行中,我们将重点放在三种斯拉夫语言上,即克罗地亚语、波兰语和塞尔维亚语。我们的方法的主要挑战是ParlaMint文本和可用记录之间缺乏任何全局对齐,以及每个模式中有时不同的数据顺序,这需要一种新的方法来在大搜索空间中对齐长序列的文本和音频。这次试运行的结果是三个高质量的数据集,涵盖了超过5,000小时的演讲和附带的文本转录。虽然这些数据集已经在三种语言的口语和文本数据的可用性方面产生了巨大的差异,但我们希望强调所提出的方法在为更多语言构建类似数据集方面的潜力。摘要:Recent significant improvements in speech and language technologies come both from self-supervised approaches over raw language data as well as various types of explicit supervision. To ensure high-quality processing of spoken data, the most useful type of explicit supervision is still the alignment between the speech signal and its corresponding text transcript, which is a data type that is not available for many languages. In this paper, we present our approach to building large and open speech-and-text-aligned datasets of less-resourced languages based on transcripts of parliamentary proceedings and their recordings. Our starting point are the ParlaMint comparable corpora of transcripts of parliamentary proceedings of 26 national European parliaments. In the pilot run on expanding the ParlaMint corpora with aligned publicly available recordings, we focus on three Slavic languages, namely Croatian, Polish, and Serbian. The main challenge of our approach is the lack of any global alignment between the ParlaMint texts and the available recordings, as well as the sometimes varying data order in each of the modalities, which requires a novel approach in aligning long sequences of text and audio in a large search space. The results of this pilot run are three high-quality datasets that span more than 5,000 hours of speech and accompanying text transcripts. Although these datasets already make a huge difference in the availability of spoken and textual data for the three languages, we want to emphasize the potential of the presented approach in building similar datasets for many more languages.
【21】 Toward Automated Clinical Transcriptions
标题: 迈向自动化临床Transit
作者:Mitchell A. Klusty,W. Vaiden Logan,Samuel E. Armstrong,Aaron D. Mullen,Caroline N. Leach,Jeff Talbert,V. K. Cody Bumgardner
备注:7 pages, 6 figures
链接:点击下载PDF文件
摘要:行政文件是医疗保健成本上升的主要驱动力,并与不良后果有关,包括医生倦怠和护理质量下降。本文介绍了一种安全的系统,适用于语音到文本转录和说话人标记(日记)的最新进展,病人提供者的对话。该系统经过优化,可生成准确的transmittance并突出显示潜在的错误,以促进快速的人工验证,进一步减少必要的人工工作。应用于超过40个小时的模拟对话,该系统提供了一个有前途的基础,自动化临床transmittance。摘要:Administrative documentation is a major driver of rising healthcare costs and is linked to adverse outcomes, including physician burnout and diminished quality of care. This paper introduces a secure system that applies recent advancements in speech-to-text transcription and speaker-labeling (diarization) to patient-provider conversations. This system is optimized to produce accurate transcriptions and highlight potential errors to promote rapid human verification, further reducing the necessary manual effort. Applied to over 40 hours of simulated conversations, this system offers a promising foundation for automating clinical transcriptions.
【22】 A Joint Spectro-Temporal Relational Thinking Based Acoustic Modeling Framework
标题: 基于频谱-时间关系思维的联合声学建模框架
作者:Zheng Nan,Ting Dang,Vidhyasaharan Sethu,Beena Ahmed
链接:点击下载PDF文件
摘要:关系思维是指人类对感觉信号和先验知识之间的关系形成心理印象,并随后将其纳入其世界模型的固有能力。尽管关系思维在人类理解语音方面发挥着至关重要的作用,但它尚未在任何人工语音识别系统中得到利用。最近,已经有一些尝试来纠正这种疏忽,但这些尝试仅限于仅在时域中操作的粗略话语级模型。为了缩小人工系统和人类能力之间的差距,本文提出了一种新的基于频谱-时间关系思维的声学建模框架。具体来说,它首先生成大量概率图来对时域和频域上语音片段之间的关系进行建模。然后,将这些图中每对节点中的关系信息聚合并嵌入到可以由下游任务使用的潜在表示中。建立在这个框架上的模型优于最先进的系统,在TIMIT数据集上的音素识别任务提高了7.82%。深入的分析进一步表明,我们提出的关系思维建模主要提高了模型的识别能力,元音,这是最容易混淆的音素识别器。摘要:Relational thinking refers to the inherent ability of humans to form mental impressions about relations between sensory signals and prior knowledge, and subsequently incorporate them into their model of their world. Despite the crucial role relational thinking plays in human understanding of speech, it has yet to be leveraged in any artificial speech recognition systems. Recently, there have been some attempts to correct this oversight, but these have been limited to coarse utterance-level models that operate exclusively in the time domain. In an attempt to narrow the gap between artificial systems and human abilities, this paper presents a novel spectro-temporal relational thinking based acoustic modeling framework. Specifically, it first generates numerous probabilistic graphs to model the relationships among speech segments across both time and frequency domains. The relational information rooted in every pair of nodes within these graphs is then aggregated and embedded into latent representations that can be utilized by downstream tasks. Models built upon this framework outperform state-of-the-art systems with a 7.82 % improvement in phoneme recognition tasks over the TIMIT dataset. In-depth analyses further reveal that our proposed relational thinking modeling mainly improves the model's ability to recognize vowels, which are the most likely to be confused by phoneme recognizers.
【23】 TCG CREST System Description for the Second DISPLACE Challenge
标题: 第二次DISPLACE挑战赛的TCG CREST系统描述
作者:Nikhil Raghav,Subhajit Saha,Md Sahidullah,Swagatam Das
链接:点击下载PDF文件
摘要:在本报告中,我们描述了我们的团队为2024年第二次DISPLACE挑战赛开发的说话者日记(SD)和语言日记(LD)系统。我们的贡献致力于多语言和多发言者场景中的可持续发展轨道1和LD轨道2。我们研究了不同的语音增强技术,语音活动检测(VAD)技术,无监督域分类,神经嵌入提取架构。我们还利用了各种嵌入提取模型的融合。我们使用开源SpeechBrain工具包实现了我们的系统。我们最终的意见书使用频谱聚类的扬声器和语言日记。我们在轨道1的挑战基线上实现了约7 %$的相对改进。我们没有获得超过轨道2中挑战基线的改进。摘要:In this report, we describe the speaker diarization (SD) and language diarization (LD) systems developed by our team for the Second DISPLACE Challenge, 2024. Our contributions were dedicated to Track 1 for SD and Track 2 for LD in multilingual and multi-speaker scenarios. We investigated different speech enhancement techniques, voice activity detection (VAD) techniques, unsupervised domain categorization, and neural embedding extraction architectures. We also exploited the fusion of various embedding extraction models. We implemented our system with the open-source SpeechBrain toolkit. Our final submissions use spectral clustering for both the speaker and language diarization. We achieve about $7 %$ relative improvement over the challenge baseline in Track 1. We did not obtain improvement over the challenge baseline in Track 2.
【24】 Contextualization of ASR with LLM using phonetic retrieval-based augmentation
标题: 使用基于语音检索的增强将ASC与LLM进行上下文化
作者:Zhihong Lei,Xingyu Na,Mingbin Xu,Ernest Pusateri,Christophe Van Gysel,Yuanyuan Zhang,Shiyi Han,Zhen Huang
链接:点击下载PDF文件
摘要:大型语言模型(LLM)已经显示出对包括音频和文本在内的多模态信号进行建模的卓越能力,允许模型在给定语音输入的情况下生成口语或文本响应。然而,它仍然是一个挑战的模型,以识别个人命名的实体,如在电话簿中的联系人,当输入模态是语音。在这项工作中,我们从语音识别任务开始,并提出了一个基于检索的解决方案,以上下文的LLM:我们首先让LLM检测命名实体的语音没有任何上下文,然后使用这个命名实体作为查询检索语音相似的命名实体从个人数据库和饲料的LLM,最后运行上下文感知的LLM解码。在语音助理任务中,与没有上下文的基线系统相比,我们的解决方案实现了高达30.2%的相对单词错误率降低和73.6%的相对命名实体错误率降低。值得注意的是,我们的解决方案通过设计避免了用完整的命名实体数据库提示LLM,使其高效并适用于大型命名实体数据库。摘要:Large language models (LLMs) have shown superb capability of modeling multimodal signals including audio and text, allowing the model to generate spoken or textual response given a speech input. However, it remains a challenge for the model to recognize personal named entities, such as contacts in a phone book, when the input modality is speech. In this work, we start with a speech recognition task and propose a retrieval-based solution to contextualize the LLM: we first let the LLM detect named entities in speech without any context, then use this named entity as a query to retrieve phonetically similar named entities from a personal database and feed them to the LLM, and finally run context-aware LLM decoding. In a voice assistant task, our solution achieved up to 30.2% relative word error rate reduction and 73.6% relative named entity error rate reduction compared to a baseline system without contextualization. Notably, our solution by design avoids prompting the LLM with the full named entity database, making it highly efficient and applicable to large named entity databases.
【25】 A Large Dataset of Spontaneous Speech with the Accent Spoken in São Paulo for Automatic Speech Recognition Evaluation
标题: 圣保罗带有口音的自发语音大数据集用于自动语音识别评估
作者:Rodrigo Lima,Sidney Evaldo Leal,Arnaldo Candido Junior,Sandra Maria Aluísio
链接:点击下载PDF文件
摘要:我们提出了一个免费的自发语音语料库巴西葡萄牙语和报告初步的自动语音识别(ASR)的结果,使用Wav 2 Vec 2-XLSR-53和蒸馏耳语模型微调和训练我们的语料库。CSC-SP音频语料库包括401个不同的扬声器(204名女性,197名男性),共239.30小时的转录录音。据我们所知,这是第一个大型的Paulistano口音的自发语音语料库致力于在葡萄牙语的ASR任务。本文首先介绍了CSC-SP音频语料库的设计和开发过程,然后详细描述了四个ASR实验。实验表明,有前途的结果的适用性语料库的ASR。具体来说,我们微调了两个版本的Wav 2 Vec 2-XLSR-53模型,使用我们的数据集训练了一个Distil-Whisper模型,该数据集具有由Whisper Large-V3模型确定的标签,并使用我们的语料库微调了这个Distil-Whisper模型。我们最好的结果是Distil-Whisper在WER24.22%的WER 2 Vec 2-XLSR-53模型上进行了微调,WER为33.73%,比Distil-Whisper差了近10%。为了实现实验的可重复性,我们在Hugging-Face和Github存储库中共享了CMC-SP音频语料库数据集、预训练模型和训练配方。摘要:We present a freely available spontaneous speech corpus for the Brazilian Portuguese language and report preliminary automatic speech recognition (ASR) results, using both the Wav2Vec2-XLSR-53 and Distil-Whisper models fine-tuned and trained on our corpus. The NURC-SP Audio Corpus comprises 401 different speakers (204 females, 197 males) with a total of 239.30 hours of transcribed audio recordings. To the best of our knowledge, this is the first large Paulistano accented spontaneous speech corpus dedicated to the ASR task in Portuguese. We first present the design and development procedures of the NURC-SP Audio Corpus, and then describe four ASR experiments in detail. The experiments demonstrated promising results for the applicability of the corpus for ASR. Specifically, we fine-tuned two versions of Wav2Vec2-XLSR-53 model, trained a Distil-Whisper model using our dataset with labels determined by Whisper Large-V3 model, and fine-tuned this Distil-Whisper model with our corpus. Our best results were the Distil-Whisper fine-tuned over NURC-SP Audio Corpus with a WER of 24.22% followed by a fine-tuned versions of Wav2Vec2-XLSR-53 model with a WER of 33.73%, that is almost 10% point worse than Distil-Whisper's. To enable experiment reproducibility, we share the NURC-SP Audio Corpus dataset, pre-trained models, and training recipes in Hugging-Face and Github repositories.
【26】 WaveTransfer: A Flexible End-to-end Multi-instrument Timbre Transfer with Diffusion
标题: WaveTransfer:带扩散的灵活端到端多乐器音色传输
作者:Teysir Baoueb,Xiaoyu Bie,Hicham Janati,Gael Richard
Journal-ref:2024 IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2024), Sep 2024, London (UK), United Kingdom
链接:点击下载PDF文件
摘要:随着基于扩散的深度生成模型越来越流行,研究人员正在积极研究它们在各个领域的潜在应用,包括音乐合成和风格改变。在这项工作中,我们感兴趣的是音色转移,这是一个过程,涉及无缝地改变音乐作品的乐器特征,同时保留基本的音乐元素。本文介绍了WaveTransfer,一个为音色传递而设计的端到端扩散模型。我们特别采用双边去噪扩散模型(BDDM)的噪声调度搜索。我们的模型能够在音频混合以及单个乐器之间进行音色转移。值得注意的是,它具有多功能性,因为它在单个模型中容纳了独特乐器对之间的多种类型的音色转移,从而消除了对每个配对进行单独模型训练的需要。此外,与限制在16 kHz的最新作品不同,WaveTransfer可以以各种采样率进行训练,包括行业标准的44.1 kHz,这是音乐界特别感兴趣的一个功能。摘要:As diffusion-based deep generative models gain prevalence, researchers are actively investigating their potential applications across various domains, including music synthesis and style alteration. Within this work, we are interested in timbre transfer, a process that involves seamlessly altering the instrumental characteristics of musical pieces while preserving essential musical elements. This paper introduces WaveTransfer, an end-to-end diffusion model designed for timbre transfer. We specifically employ the bilateral denoising diffusion model (BDDM) for noise scheduling search. Our model is capable of conducting timbre transfer between audio mixtures as well as individual instruments. Notably, it exhibits versatility in that it accommodates multiple types of timbre transfer between unique instrument pairs in a single model, eliminating the need for separate model training for each pairing. Furthermore, unlike recent works limited to 16 kHz, WaveTransfer can be trained at various sampling rates, including the industry-standard 44.1 kHz, a feature of particular interest to the music community.
【27】 Equivariance-based self-supervised learning for audio signal recovery from clipped measurements
标题: 基于等效性的自我监督学习,从剪辑测量中恢复音频信号
作者:Victor Sechaud,Laurent Jacques,Patrice Abry,Julián Tachella
Journal-ref:EUSIPCO, Aug 2024, Lyon, France
链接:点击下载PDF文件
摘要:
摘要:
【28】 Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech
标题: 面部表情增强的TTC:结合面部表示和情感强度实现自适应语音
作者:Yunji Chu,Yunseob Shim,Unsang Park
备注:13 pages, 3 figures, accepted to ECCV Workshop ABAW(Affective Behavior Analysis in-the-wild)7 (to be appear)
链接:点击下载PDF文件
【29】 Leveraging Mixture of Experts for Improved Speech Deepfake Detection
标题: 利用专家混合改进语音深度伪造检测
作者:Viola Negroni,Davide Salvi,Alessandro Ilic Mezza,Paolo Bestagini,Stefano Tubaro
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【30】 Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
标题: 弥合语音和文本:通过LLM中的拼音到字符预训练增强ASB
作者:Yang Yuhang,Peng Yizhou,Eng Siong Chng,Xionghu Zhong
备注:Accepted by ISCSLP2024-Special session-Speech Processing in LLM Era
链接:点击下载PDF文件
【31】 Disentangling Age and Identity with a Mutual Information Minimization Approach for Cross-Age Speaker Verification
标题: 用互信息最小化方法解开年龄和身份,用于跨年龄段说话人验证
作者:Fengrun Zhang,Wangjin Zhou,Yiming Liu,Wang Geng,Yahui Shan,Chen Zhang
备注:Interspeech 2024
链接:点击下载PDF文件
【32】 ASD-Diffusion: Anomalous Sound Detection with Diffusion Models
标题: ASD扩散:使用扩散模型检测异常声音
作者:Fengrun Zhang,Xiang Xie,Kai Guo
备注:This paper will appear at ICPR 2024
链接:点击下载PDF文件
【33】 A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation
标题: 基于模块的语音同步翻译中梯度冲突缓解策略
作者:Xiaoqian Liu,Yangfan Du,Jianjin Wang,Yuan Ge,Chen Xu,Tong Xiao,Guocheng Chen,Jingbo Zhu
链接:点击下载PDF文件
【34】 Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
标题: 通过混合专家提高代码转换ASB增强的言语条件LLM
作者:Fengrun Zhang,Wang Geng,Hukai Huang,Cheng Yi,He Qu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【35】 On the calibration of powerset speaker diarization models
标题: 关于Powerset扬声器Dialogue模型的校准
作者:Alexis Plaquet,Hervé Bredin
Journal-ref:Interspeech 2024, Sep 2024, Kos, Greece. pp.3764-3768, ⟨10.21437Interspeech.2024-1060⟩
链接:点击下载PDF文件
【36】 NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
标题: NanoVoice:适合多个扬声器的高效扬声器自适应文本到语音
作者:Nohil Park,Heeseung Kim,Che Hyun Lee,Jooyoung Choi,Jiheum Yeom,Sungroh Yoon
备注:Submitted to ICASSP 2025, Demo Page: this https URL
链接:点击下载PDF文件
【37】 VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance
标题: VoiceGuider:通过自动引导增强参数高效扬声器自适应文本到语音的域外性能
作者:Jiheum Yeom,Heeseung Kim,Jooyoung Choi,Che Hyun Lee,Nohil Park,Sungroh Yoon
备注:Submitted to ICASSP 2025, Demo Page: this https URL
链接:点击下载PDF文件
【38】 Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
标题: 假设聚集和合并:具有说话者令牌的新型多说话者语音识别
作者:Yosuke Kashiwagi,Hayato Futami,Emiru Tsunoo,Siddhant Arora,Shinji Watanabe
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【39】 Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
标题: 超越回合制接口:同步LLM作为全速对话代理
作者:Bandhav Veluri,Benjamin N Peloquin,Bokai Yu,Hongyu Gong,Shyamnath Gollakota
备注:EMNLP Main 2024
链接:点击下载PDF文件
【40】 Generalization in birdsong classification: impact of transfer learning methods and dataset characteristics
标题: 鸟鸣分类的推广:迁移学习方法和数据集特征的影响
作者:Burooj Ghani,Vincent J. Kalkman,Bob Planqué,Willem-Pier Vellinga,Lisa Gill,Dan Stowell
备注:25 pages
链接:点击下载PDF文件
【41】 Efficient learning-based sound propagation for virtual and real-world audio processing applications
标题: 针对虚拟和现实世界音频处理应用程序的高效基于学习的声音传播
作者:Anton Jeran Ratnarajah
备注:PhD thesis
链接:点击下载PDF文件
【28】 Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech
标题: 面部表情增强的TTC:结合面部表示和情感强度实现自适应语音
作者:Yunji Chu,Yunseob Shim,Unsang Park
备注:13 pages, 3 figures, accepted to ECCV Workshop ABAW(Affective Behavior Analysis in-the-wild)7 (to be appear)
链接:点击下载PDF文件
摘要:我们提出FEIM-TTS,一个创新的zero-shot文本到语音(TTS)模型,合成情感表达的语音,与面部图像对齐,并通过情绪强度调制。利用深度学习,FEIM-TTS超越了传统的TTS系统,通过解释面部线索和调整情绪的细微差别,而不依赖于标记的数据集。为了解决稀疏的视听情感数据,该模型使用LRS 3,CREMA-D和MELD数据集进行训练,证明了其适应性。FEIM-TTS的独特能力,以产生高品质,说话者不可知的语音,使其适合创造适应性的声音,为虚拟人物。此外,FEIM-TTS大大提高了视力障碍者或视力有问题的人的可访问性。通过将情感的细微差别整合到TTS中,我们的模型为网络漫画提供了动态和引人入胜的听觉体验,使视障用户能够更充分地享受这些叙事。综合评价表明,它在调节情感和强度,提高情感语音合成和可达性方面表现出色。样品可在以下网址获得:https: feim-tts.github.io 。摘要:We propose FEIM-TTS, an innovative zero-shot text-to-speech (TTS) model that synthesizes emotionally expressive speech, aligned with facial images and modulated by emotion intensity. Leveraging deep learning, FEIM-TTS transcends traditional TTS systems by interpreting facial cues and adjusting to emotional nuances without dependence on labeled datasets. To address sparse audio-visual-emotional data, the model is trained using LRS3, CREMA-D, and MELD datasets, demonstrating its adaptability. FEIM-TTS's unique capability to produce high-quality, speaker-agnostic speech makes it suitable for creating adaptable voices for virtual characters. Moreover, FEIM-TTS significantly enhances accessibility for individuals with visual impairments or those who have trouble seeing. By integrating emotional nuances into TTS, our model enables dynamic and engaging auditory experiences for webcomics, allowing visually impaired users to enjoy these narratives more fully. Comprehensive evaluation evidences its proficiency in modulating emotion and intensity, advancing emotional speech synthesis and accessibility. Samples are available at: https: feim-tts.github.io .
【29】 Leveraging Mixture of Experts for Improved Speech Deepfake Detection
标题: 利用专家混合改进语音深度伪造检测
作者:Viola Negroni,Davide Salvi,Alessandro Ilic Mezza,Paolo Bestagini,Stefano Tubaro
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:语音深度伪造对个人安全和内容真实性构成重大威胁。文献中已经提出了几种检测器,这些系统必须面对的主要挑战之一是对看不见的数据进行泛化,以识别各种数据集中的假信号。在本文中,我们介绍了一种使用混合专家架构来增强语音深度伪造检测性能的新方法。专家混合框架非常适合语音深度伪造检测任务,因为它能够专门处理不同的输入类型并有效地处理数据变化。与传统的单一模型或集成方法相比,这种方法对未知数据具有更好的泛化能力和适应性。此外,它的模块化结构支持可扩展的更新,使其在管理深度伪造技术不断变化的复杂性方面更加灵活,同时保持高检测精度。我们提出了一种高效、轻量级的门控机制,为每个输入动态分配专家权重,从而优化检测性能。在多个数据集上的实验结果证明了我们所提出的方法的有效性和潜力。摘要:Speech deepfakes pose a significant threat to personal security and content authenticity. Several detectors have been proposed in the literature, and one of the primary challenges these systems have to face is the generalization over unseen data to identify fake signals across a wide range of datasets. In this paper, we introduce a novel approach for enhancing speech deepfake detection performance using a Mixture of Experts architecture. The Mixture of Experts framework is well-suited for the speech deepfake detection task due to its ability to specialize in different input types and handle data variability efficiently. This approach offers superior generalization and adaptability to unseen data compared to traditional single models or ensemble methods. Additionally, its modular structure supports scalable updates, making it more flexible in managing the evolving complexity of deepfake techniques while maintaining high detection accuracy. We propose an efficient, lightweight gating mechanism to dynamically assign expert weights for each input, optimizing detection performance. Experimental results across multiple datasets demonstrate the effectiveness and potential of our proposed approach.
【30】 Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
标题: 弥合语音和文本:通过LLM中的拼音到字符预训练增强ASB
作者:Yang Yuhang,Peng Yizhou,Eng Siong Chng,Xionghu Zhong
备注:Accepted by ISCSLP2024-Special session-Speech Processing in LLM Era
链接:点击下载PDF文件
摘要:大型语言模型(LLM)与预训练语音模型的集成为自动语音识别(ASR)开辟了新的途径。虽然LLM擅长多模态理解任务,但有效利用他们的ASR能力仍然是一个重大挑战。本文提出了一种新颖的培训方法来提高LLM在ASR任务中的表现。我们提出了预训练LLM拼音嵌入序列,这代表发音特征,以产生相应的汉字。该步骤使得LLM能够在遇到真实语音数据之前适应于从发音特征生成文本。此外,我们微调的LoRA参数,以提高LLM的语音模态信息的理解。在AISHELL-1语料库中,与没有拼音到字符预训练的基线相比,我们的方法在ASR任务中产生了9.5%的相对改善。此外,将辅助文本数据用于拼音到字符的预训练进一步提高了性能,实现了19.0%的相对提高。摘要:The integration of large language models (LLMs) with pre-trained speech models has opened up new avenues in automatic speech recognition (ASR). While LLMs excel in multimodal understanding tasks, effectively leveraging their capabilities for ASR remains a significant challenge. This paper presents a novel training approach to enhance LLM performance in ASR tasks. We propose pre-training LLMs on Pinyin embedding sequences, which represent pronunciation features, to generate corresponding Chinese characters. This step enables the LLM to adapt to generating text from pronunciation features before encountering real speech data. Furthermore, we fine-tune the LoRA parameters to enhance the LLM's understanding of speech modality information. In AISHELL-1 corpus, our approach yields a 9.5% relative improvement in ASR tasks compared to the baseline without Pinyi-to-Character pre-training. Additionally, incorporating auxiliary text data for Pinyi-to-Character pre-training further boosts performance, achieving a 19.0% relative improvement.
【31】 Disentangling Age and Identity with a Mutual Information Minimization Approach for Cross-Age Speaker Verification
标题: 用互信息最小化方法解开年龄和身份,用于跨年龄段说话人验证
作者:Fengrun Zhang,Wangjin Zhou,Yiming Liu,Wang Geng,Yahui Shan,Chen Zhang
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:跨年龄说话人确认(CASV)是近年来的研究热点。然而,现有的说话人确认系统在CASV中表现不佳,这是由于年龄造成的语音个体差异很大。本文提出了一种基于互信息最小化的CASV解纠缠表示学习框架。在我们的方法中,骨干模型被训练来从说话人信息中分离身份和年龄相关的嵌入,并且MI估计器被训练来通过MI最小化来最小化年龄和身份相关的嵌入之间的相关性,从而产生年龄不变的说话人嵌入。此外,通过使用积极和消极的样本之间的年龄差距,我们提出了一个老化意识的MI最小化损失函数,使骨干模型更专注于声音的变化与大的年龄差距。实验结果表明,该方法在Vox-CA的多个跨年龄测试集上的性能优于其他方法。摘要:There has been an increasing research interest in cross-age speaker verification~(CASV). However, existing speaker verification systems perform poorly in CASV due to the great individual differences in voice caused by aging. In this paper, we propose a disentangled representation learning framework for CASV based on mutual information~(MI) minimization. In our method, a backbone model is trained to disentangle the identity- and age-related embeddings from speaker information, and an MI estimator is trained to minimize the correlation between age- and identity-related embeddings via MI minimization, resulting in age-invariant speaker embeddings. Furthermore, by using the age gaps between positive and negative samples, we propose an aging-aware MI minimization loss function that allows the backbone model to focus more on the vocal changes with large age gaps. Experimental results show that the proposed method outperforms other methods on multiple Cross-Age test sets of Vox-CA.
【32】 ASD-Diffusion: Anomalous Sound Detection with Diffusion Models
标题: ASD扩散:使用扩散模型检测异常声音
作者:Fengrun Zhang,Xiang Xie,Kai Guo
备注:This paper will appear at ICPR 2024
链接:点击下载PDF文件
摘要:无监督异常声音检测(ASD)的目的是设计一种通用的方法,可以用来检测异常时,只有正常的声音。本文提出了一种基于扩散模型的异常声音检测方法(ASD-扩散),用于实际工厂中的ASD。在我们的流水线中,声学特征中的异常从其噪声损坏的特征重建成其近似正常的模式。其次,提出了一种后处理异常滤波算法,用于检测重构后与原始输入存在显著偏差的异常。此外,引入去噪扩散隐式模型,通过延长去噪过程的采样间隔,加快了推理速度。该方法是扩散模型应用的一种新方法。在DCASE 2023挑战任务2的开发集上的实验结果比基线高出7.75%,证明了该方法的有效性。摘要:Unsupervised Anomalous Sound Detection (ASD) aims to design a generalizable method that can be used to detect anomalies when only normal sounds are given. In this paper, Anomalous Sound Detection based on Diffusion Models (ASD-Diffusion) is proposed for ASD in real-world factories. In our pipeline, the anomalies in acoustic features are reconstructed from their noisy corrupted features into their approximate normal pattern. Secondly, a post-processing anomalies filter algorithm is proposed to detect anomalies that exhibit significant deviation from the original input after reconstruction. Furthermore, denoising diffusion implicit model is introduced to accelerate the inference speed by a longer sampling interval of the denoising process. The proposed method is innovative in the application of diffusion models as a new scheme. Experimental results on the development set of DCASE 2023 challenge task 2 outperform the baseline by 7.75%, demonstrating the effectiveness of the proposed method.
【33】 A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation
标题: 基于模块的语音同步翻译中梯度冲突缓解策略
作者:Xiaoqian Liu,Yangfan Du,Jianjin Wang,Yuan Ge,Chen Xu,Tong Xiao,Guocheng Chen,Jingbo Zhu
链接:点击下载PDF文件
摘要:同步语音翻译(SimulST)涉及在连续处理流式语音输入的同时生成目标语言文本,这带来了重大的实时挑战。多任务学习通常用于增强SimulST性能,但会在主任务和辅助任务之间引入优化冲突,从而可能影响整体效率。现有的模型级冲突解决方法并不适合此任务,这加剧了效率低下并导致高GPU内存消耗。为了解决这些挑战,我们提出了一个模块化梯度冲突缓解(MGCM)的策略,检测冲突在一个更细粒度的模块化水平,并解决它们利用梯度投影。实验结果表明,MGCM显着提高SimulST的性能,特别是在中等和高延迟条件下,实现了0.68 BLEU分数增益离线任务。此外,与其他冲突缓解方法相比,MGCM将GPU内存消耗减少了95%以上,使其成为SimulST任务的强大解决方案。摘要:Simultaneous Speech Translation (SimulST) involves generating target language text while continuously processing streaming speech input, presenting significant real-time challenges. Multi-task learning is often employed to enhance SimulST performance but introduces optimization conflicts between primary and auxiliary tasks, potentially compromising overall efficiency. The existing model-level conflict resolution methods are not well-suited for this task which exacerbates inefficiencies and leads to high GPU memory consumption. To address these challenges, we propose a Modular Gradient Conflict Mitigation (MGCM) strategy that detects conflicts at a finer-grained modular level and resolves them utilizing gradient projection. Experimental results demonstrate that MGCM significantly improves SimulST performance, particularly under medium and high latency conditions, achieving a 0.68 BLEU score gain in offline tasks. Additionally, MGCM reduces GPU memory consumption by over 95 % compared to other conflict mitigation methods, establishing it as a robust solution for SimulST tasks.
【34】 Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
标题: 通过混合专家提高代码转换ASB增强的言语条件LLM
作者:Fengrun Zhang,Wang Geng,Hukai Huang,Cheng Yi,He Qu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在本文中,我们介绍了一个语音条件的大语言模型(LLM)集成的混合专家(MoE)为基础的连接器,以解决自动语音识别(ASR)中的代码切换(CS)的挑战。具体来说,我们提出了一个插入和删除中断令牌(IDIT)机制,以更好地转移LLM的文本生成能力的语音识别任务。我们还提出了一个连接器与MoE架构,有效地管理多种语言。为了进一步增强多个专家的协作并利用LLM的理解能力,我们提出了一种两阶段渐进式训练策略:1)连接器被解冻并与语言专业专家一起训练,以将语音表示映射到文本空间。2)连接器和LLM LoRA适配器使用建议的IDIT机制进行培训,所有专家都被激活以学习一般表示。实验结果表明,我们的方法显着优于国家的最先进的模型,包括端到端和大规模的音频语言模型。摘要:In this paper, we introduce a speech-conditioned Large Language Model (LLM) integrated with a Mixture of Experts (MoE) based connector to address the challenge of Code-Switching (CS) in Automatic Speech Recognition (ASR). Specifically, we propose an Insertion and Deletion of Interruption Token (IDIT) mechanism for better transfer text generation ability of LLM to speech recognition task. We also present a connecter with MoE architecture that manages multiple languages efficiently. To further enhance the collaboration of multiple experts and leverage the understanding capabilities of LLM, we propose a two-stage progressive training strategy: 1) The connector is unfrozen and trained with language-specialized experts to map speech representations to the text space. 2) The connector and LLM LoRA adaptor are trained with the proposed IDIT mechanism and all experts are activated to learn general representations. Experimental results demonstrate that our method significantly outperforms state-of-the-art models, including end-to-end and large-scale audio-language models.
【35】 On the calibration of powerset speaker diarization models
标题: 关于Powerset扬声器Dialogue模型的校准
作者:Alexis Plaquet,Hervé Bredin
Journal-ref:Interspeech 2024, Sep 2024, Kos, Greece. pp.3764-3768, ⟨10.21437Interspeech.2024-1060⟩
链接:点击下载PDF文件
摘要:端到端神经日志模型通常依赖于说话人日志问题的多标签分类公式。最近,我们提出了一个powerset多类公式,它在多个数据集上击败了最先进的技术。在本文中,我们提出了研究的幂集扬声器日记模型的校准,并探讨其使用的一些。我们研究了域内和域外的校准,并探索了低置信度区域中的数据。然后在实践中测试模型置信度的可靠性:我们使用预训练模型的置信度来选择性地从未注释的数据中创建训练和验证子集,并将其与随机选择进行比较。我们发现,顶标签置信度可以用来可靠地预测高误差区域。此外,在低置信度区域上进行训练可以提供更好的校准模型,并且在低置信度区域上进行验证可以比随机区域更具注释效率。摘要:End-to-end neural diarization models have usually relied on a multilabel-classification formulation of the speaker diarization problem. Recently, we proposed a powerset multiclass formulation that has beaten the state-of-the-art on multiple datasets. In this paper, we propose to study the calibration of a powerset speaker diarization model, and explore some of its uses. We study the calibration in-domain, as well as out-of-domain, and explore the data in low-confidence regions. The reliability of model confidence is then tested in practice: we use the confidence of the pretrained model to selectively create training and validation subsets out of unannotated data, and compare this to random selection. We find that top-label confidence can be used to reliably predict high-error regions. Moreover, training on low-confidence regions provides a better calibrated model, and validating on low-confidence regions can be more annotation-efficient than random regions.
【36】 NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
标题: NanoVoice:适合多个扬声器的高效扬声器自适应文本到语音
作者:Nohil Park,Heeseung Kim,Che Hyun Lee,Jooyoung Choi,Jiheum Yeom,Sungroh Yoon
备注:Submitted to ICASSP 2025, Demo Page: this https URL
链接:点击下载PDF文件
摘要:我们提出了NanoVoice,一个个性化的文本到语音的模型,有效地构建语音适配器,同时为多个扬声器。NanoVoice引入了一种批量扬声器自适应技术,能够并行微调多个参考,显著减少训练时间。除了为每个扬声器构建单独的适配器之外,我们还提出了一种参数共享技术,该技术减少了用于扬声器自适应的参数数量。通过整合一个新的可训练的尺度矩阵,NanoVoice减轻了参数共享过程中潜在的性能下降。NanoVoice实现了与基线相当的性能,同时训练速度提高了4倍,使用40个参考语音进行扬声器自适应的参数减少了45%。广泛的消融研究和分析进一步验证了我们的模型的效率。摘要:We present NanoVoice, a personalized text-to-speech model that efficiently constructs voice adapters for multiple speakers simultaneously. NanoVoice introduces a batch-wise speaker adaptation technique capable of fine-tuning multiple references in parallel, significantly reducing training time. Beyond building separate adapters for each speaker, we also propose a parameter sharing technique that reduces the number of parameters used for speaker adaptation. By incorporating a novel trainable scale matrix, NanoVoice mitigates potential performance degradation during parameter sharing. NanoVoice achieves performance comparable to the baselines, while training 4 times faster and using 45 percent fewer parameters for speaker adaptation with 40 reference voices. Extensive ablation studies and analysis further validate the efficiency of our model.
【37】 VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance
标题: VoiceGuider:通过自动引导增强参数高效扬声器自适应文本到语音的域外性能
作者:Jiheum Yeom,Heeseung Kim,Jooyoung Choi,Che Hyun Lee,Nohil Park,Sungroh Yoon
备注:Submitted to ICASSP 2025, Demo Page: this https URL
链接:点击下载PDF文件
摘要:当通过LoRA将参数高效微调应用于说话人自适应文本到语音模型时,与完全微调的对应物相比,自适应性能可能会下降,特别是对于域外说话人。在这里,我们提出了VoiceGuider,一个参数高效的扬声器自适应文本到语音系统,通过自动指导来增强扬声器自适应性能,减少与全微调模型的差距。我们仔细探索各种加强自动导航的方法,最终找到最佳策略。结果,VoiceGuider显示了鲁棒的自适应性能,特别是在极端的域外语音数据上。我们在演示页面中提供音频样本。摘要:When applying parameter-efficient finetuning via LoRA onto speaker adaptive text-to-speech models, adaptation performance may decline compared to full-finetuned counterparts, especially for out-of-domain speakers. Here, we propose VoiceGuider, a parameter-efficient speaker adaptive text-to-speech system reinforced with autoguidance to enhance the speaker adaptation performance, reducing the gap against full-finetuned models. We carefully explore various ways of strengthening autoguidance, ultimately finding the optimal strategy. VoiceGuider as a result shows robust adaptation performance especially on extreme out-of-domain speech data. We provide audible samples in our demo page.
【38】 Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
标题: 假设聚集和合并:具有说话者令牌的新型多说话者语音识别
作者:Yosuke Kashiwagi,Hayato Futami,Emiru Tsunoo,Siddhant Arora,Shinji Watanabe
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在许多现实世界的场景中,例如会议,多个发言者与未知数量的参与者一起出现,并且他们的话语经常重叠。我们通过一种新型的基于注意力的编码器-解码器方法来解决这些多说话者挑战,该方法通过说话者聚类获得的特殊说话者类别令牌进行增强。在推理过程中,我们选择多个识别假设条件下的预测说话人集群令牌,这些假设合并凝聚层次聚类(AHC)的基础上归一化的编辑距离。聚类假设的结果在多说话人transmittance与AHC确定的适当数量的发言人。我们在LibriMix数据集上的实验表明,我们提出的方法在复杂的3混合环境中特别有效,与传统的序列化输出训练相比,在干净数据上实现了55%的相对误差减少,在嘈杂数据上实现了36%的相对误差减少。摘要:In many real-world scenarios, such as meetings, multiple speakers are present with an unknown number of participants, and their utterances often overlap. We address these multi-speaker challenges by a novel attention-based encoder-decoder method augmented with special speaker class tokens obtained by speaker clustering. During inference, we select multiple recognition hypotheses conditioned on predicted speaker cluster tokens, and these hypotheses are merged by agglomerative hierarchical clustering (AHC) based on the normalized edit distance. The clustered hypotheses result in the multi-speaker transcriptions with the appropriate number of speakers determined by AHC. Our experiments on the LibriMix dataset demonstrate that our proposed method was particularly effective in complex 3-mix environments, achieving a 55% relative error reduction on clean data and a 36% relative error reduction on noisy data compared with conventional serialized output training.
【39】 Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
标题: 超越回合制接口:同步LLM作为全速对话代理
作者:Bandhav Veluri,Benjamin N Peloquin,Bokai Yu,Hongyu Gong,Shyamnath Gollakota
备注:EMNLP Main 2024
链接:点击下载PDF文件
摘要:尽管对口语对话代理建模有广泛的兴趣,但大多数方法本质上是“半双工”的-仅限于基于回合的交互,其响应需要用户的显式提示或对中断或沉默事件的隐式跟踪。相比之下,人类的对话是“全双工”的,允许以快速和动态的话轮转换,重叠的演讲和反向引导的形式实现丰富的同步性。从技术上讲,实现与LLM的全双工对话的挑战在于对同步进行建模,因为预先训练的LLM没有“时间”感。为了弥合这一差距,我们提出了同步LLM全双工口语对话建模。我们设计了一种新的机制,将时间信息集成到Llama 3 -8b中,使它们与现实世界的时钟同步运行。我们还介绍了一个训练配方,它使用从文本对话数据生成的212 k小时的合成口语对话数据来创建一个模型,该模型可以生成有意义和自然的口语对话,而只有2k小时的真实口语对话数据。同步LLM在保持自然性的同时,在对话的意义上优于最先进的水平。最后,我们证明了模型的能力,通过模拟两个代理之间的互动,在不同的数据集上训练,同时考虑互联网规模的延迟高达240毫秒。https: syncllm.cs.washington.edu 摘要:Despite broad interest in modeling spoken dialogue agents, most approaches are inherently "half-duplex" -- restricted to turn-based interaction with responses requiring explicit prompting by the user or implicit tracking of interruption or silence events. Human dialogue, by contrast, is "full-duplex" allowing for rich synchronicity in the form of quick and dynamic turn-taking, overlapping speech, and backchanneling. Technically, the challenge of achieving full-duplex dialogue with LLMs lies in modeling synchrony as pre-trained LLMs do not have a sense of "time". To bridge this gap, we propose Synchronous LLMs for full-duplex spoken dialogue modeling. We design a novel mechanism to integrate time information into Llama3-8b so that they run synchronously with the real-world clock. We also introduce a training recipe that uses 212k hours of synthetic spoken dialogue data generated from text dialogue data to create a model that generates meaningful and natural spoken dialogue, with just 2k hours of real-world spoken dialogue data. Synchronous LLMs outperform state-of-the-art in dialogue meaningfulness while maintaining naturalness. Finally, we demonstrate the model's ability to participate in full-duplex dialogue by simulating interaction between two agents trained on different datasets, while considering Internet-scale latencies of up to 240 ms. Webpage: https: syncllm.cs.washington.edu .
【40】 Generalization in birdsong classification: impact of transfer learning methods and dataset characteristics
标题: 鸟鸣分类的推广:迁移学习方法和数据集特征的影响
作者:Burooj Ghani,Vincent J. Kalkman,Bob Planqué,Willem-Pier Vellinga,Lisa Gill,Dan Stowell
备注:25 pages
链接:点击下载PDF文件
摘要:机器学习可以自动识别动物的声音,这在生物多样性监测中发挥着重要作用。然而,尽管越来越令人印象深刻的能力,生物声学物种分类器仍然表现出跨物种和栖息地的不平衡性能,特别是在复杂的音景。在本研究中,我们探索了迁移学习在各种条件下(包括单标签和多标签场景)以及不同模型架构(例如CNN和Transformers)的大规模鸟类声音分类中的有效性。我们的实验表明,微调和知识蒸馏产生强大的性能,交叉蒸馏证明特别有效地提高在领域内的性能Xeno-canto数据。然而,当推广到音景时,与知识蒸馏相比,浅层微调表现出更优越的性能,突出了其鲁棒性和受约束的性质。我们的研究进一步探讨了如何使用多物种标签,在这些情况下,存在但不完整。我们提倡在动物声音社区内进行更全面的标记实践,包括注释背景物种和提供时间细节,以加强对强大的鸟类声音分类器的训练。这些发现为预训练模型的最佳重用提供了见解,以推进自动生物声学识别。摘要:Animal sounds can be recognised automatically by machine learning, and this has an important role to play in biodiversity monitoring. Yet despite increasingly impressive capabilities, bioacoustic species classifiers still exhibit imbalanced performance across species and habitats, especially in complex soundscapes. In this study, we explore the effectiveness of transfer learning in large-scale bird sound classification across various conditions, including single- and multi-label scenarios, and across different model architectures such as CNNs and Transformers. Our experiments demonstrate that both fine-tuning and knowledge distillation yield strong performance, with cross-distillation proving particularly effective in improving in-domain performance on Xeno-canto data. However, when generalizing to soundscapes, shallow fine-tuning exhibits superior performance compared to knowledge distillation, highlighting its robustness and constrained nature. Our study further investigates how to use multi-species labels, in cases where these are present but incomplete. We advocate for more comprehensive labeling practices within the animal sound community, including annotating background species and providing temporal details, to enhance the training of robust bird sound classifiers. These findings provide insights into the optimal reuse of pretrained models for advancing automatic bioacoustic recognition.
【41】 Efficient learning-based sound propagation for virtual and real-world audio processing applications
标题: 针对虚拟和现实世界音频处理应用程序的高效基于学习的声音传播
作者:Anton Jeran Ratnarajah
备注:PhD thesis
链接:点击下载PDF文件
摘要:声传播是声能通过介质(例如空气)以声波形式传播到周围环境的过程。房间脉冲响应(RIR)描述了这一过程,并受到声源和听众的位置、房间的几何形状及其材料的影响。几十年来,基于物理的声学模拟器一直用于计算特定声学环境的精确RIR。然而,我们遇到了现有的声学模拟器的局限性。为了解决这些问题,我们提出了三个新的解决方案。首先,我们介绍了一个基于学习的RIR生成器,它比交互式光线跟踪模拟器快两个数量级。我们的方法可以被训练为直接输入统计和传统参数,并且它可以为重建和合成的3D场景生成单声道和双耳RIR。我们生成的RIR在语音处理应用中的性能优于交互式光线跟踪模拟器,包括ASR,语音增强和语音分离。其次,我们提出了估计RIR混响语音信号和视觉线索,没有一个3D表示的环境。通过从混响语音中估计RIR,我们可以增加训练数据以匹配测试数据,从而改善ASR系统的单词错误率。在远场ASR任务中,我们估计的RIR比以前基于学习的RIR估计器提高了6.9%。我们证明,我们的视听RIR估计艾滋病的任务,如视觉声学匹配,新颖的视图声学合成,语音配音,通过感知评估验证。最后,我们引入IR-GAN来使用真实的RIR来增强准确的RIR。IR-GAN参数化控制从真实RIR中学习到的声学参数,以生成模仿不同声学环境的新RIR,在远场ASR基准上比射线追踪模拟器的性能高出8.95%。摘要:Sound propagation is the process by which sound energy travels through a medium, such as air, to the surrounding environment as sound waves. The room impulse response (RIR) describes this process and is influenced by the positions of the source and listener, the room's geometry, and its materials. Physics-based acoustic simulators have been used for decades to compute accurate RIRs for specific acoustic environments. However, we have encountered limitations with existing acoustic simulators. To address these limitations, we propose three novel solutions. First, we introduce a learning-based RIR generator that is two orders of magnitude faster than an interactive ray-tracing simulator. Our approach can be trained to input both statistical and traditional parameters directly, and it can generate both monaural and binaural RIRs for both reconstructed and synthetic 3D scenes. Our generated RIRs outperform interactive ray-tracing simulators in speech-processing applications, including ASR, Speech Enhancement, and Speech Separation. Secondly, we propose estimating RIRs from reverberant speech signals and visual cues without a 3D representation of the environment. By estimating RIRs from reverberant speech, we can augment training data to match test data, improving the word error rate of the ASR system. Our estimated RIRs achieve a 6.9% improvement over previous learning-based RIR estimators in far-field ASR tasks. We demonstrate that our audio-visual RIR estimator aids tasks like visual acoustic matching, novel-view acoustic synthesis, and voice dubbing, validated through perceptual evaluation. Finally, we introduce IR-GAN to augment accurate RIRs using real RIRs. IR-GAN parametrically controls acoustic parameters learned from real RIRs to generate new RIRs that imitate different acoustic environments, outperforming Ray-tracing simulators on the far-field ASR benchmark by 8.95%.
机器翻译,仅供参考
![]()
