本文经arXiv每日学术速递授权转载
【1】 VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
标题: VoxDirect:具有统一多语言编解码器语言建模的表达性人类指令到语音生成
作者:Yixuan Zhou,Xiaoyu Qin,Zeyu Jin,Shuoyi Zhou,Shun Lei,Songtao Zhou,Zhiyong Wu,Jia Jia
备注:Accepted by ACM Multimedia 2024
链接:点击下载PDF文件
【2】 Towards reliable respiratory disease diagnosis based on cough sounds and vision transformers
标题: 基于咳嗽声和视觉转换器实现可靠的呼吸道疾病诊断
作者:Qian Wang,Zhaoyang Bu,Jiaxuan Mao,Wenyu Zhu,Jingya Zhao,Wei Du,Guochao Shi,Min Zhou,Si Chen,Jieming Qu
链接:点击下载PDF文件
【3】 Beyond Levenshtein: Leveraging Multiple Algorithms for Robust Word Error Rate Computations And Granular Error Classifications
标题: 超越Levenshtein:利用多种算法进行稳健的误字率计算和粒度错误分类
作者:Korbinian Kuhn,Verena Kersken,Gottfried Zimmermann
备注:Accepted in INTERSPEECH 2024
链接:点击下载PDF文件
【4】 Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
标题: Whisper-PMFA:使用Whisper模型进行说话人验证的部分多尺度特征聚集
作者:Yiyang Zhao,Shuai Wang,Guangzhi Sun,Zehua Chen,Chao Zhang,Mingxing Xu,Thomas Fang Zheng
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【5】 EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
标题: MIDI攻击:利用情感语音转换对深度语音分类模型进行语音后门攻击
作者:Wenhan Yao,Zedong XingXiarun Chen,Jia Liu,yongqiang He,Weiping Wen
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【6】 Multi-modal Adversarial Training for Zero-Shot Voice Cloning
标题: Zero-Shot语音克隆的多模式对抗训练
作者:John Janiczek,Dading Chong,Dongyang Dai,Arlo Faria,Chao Wang,Tao Wang,Yuzong Liu
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【7】 Spoofing-Robust Speaker Verification Using Parallel Embedding Fusion: BTU Speech Group's Approach for ASVspoof5 Challenge
标题: 使用并行嵌入融合的欺骗稳健说话人验证:BTU语音小组应对ASVspoof 5挑战的方法
作者:Oğuzhan Kurnaz,Selim Can Demirtaş,Aykut Büker,Jagabandhu Mishra,Cemal Hanilçi
备注:Accepted in ASVspoof2024 workshop
链接:点击下载PDF文件
【8】 ModalityMirror: Improving Audio Classification in Modality Heterogeneity Federated Learning with Multimodal Distillation
标题: ModalityMirror:通过多模式蒸馏改进情态异相联邦学习中的音频分类
作者:Tiantian Feng,Tuo Zhang,Salman Avestimehr,Shrikanth S. Narayanan
链接:点击下载PDF文件
【9】 Easy, Interpretable, Effective: openSMILE for voice deepfake detection
标题: 简单、可解释、有效:用于语音深度伪造检测的openSMILE
作者:Octavian Pascu,Dan Oneata,Horia Cucu,Nicolas M. Müller
链接:点击下载PDF文件
【10】 wav2pos: Sound Source Localization using Masked Autoencoders
标题: wav 2 pos:使用掩蔽自动编码器的声音源定位
作者:Axel Berg,Jens Gulin,Mark O'Connor,Chuteng Zhou,Karl Åström,Magnus Oskarsson
备注:IPIN 2024
链接:点击下载PDF文件
【11】 A Hybrid Approach for Low-Complexity Joint Acoustic Echo and Noise Reduction
标题: 低复杂度联合声学回声和降噪的混合方法
作者:Shrishti Saha Shetu,Naveen Kumar Desiraju,Jose Miguel Martinez Aponte,Emanuël A. P. Habets,Edwin Mabande
备注:5 pages, 2 figures
链接:点击下载PDF文件
【12】 Spectral Masking with Explicit Time-Context Windowing for Neural Network-Based Monaural Speech Enhancement
标题: 基于神经网络的单耳语音增强的显式时间上下文窗口频谱掩蔽
作者:Luan Vinícius Fiorio,Boris Karanov,Bruno Defraene,Johan David,Wim van Houtum,Frans Widdershoven,Ronald M. Aarts
备注:This work has been submitted to the IEEE for possible publication
链接:点击下载PDF文件
【13】 Noise-to-mask Ratio Loss for Deep Neural Network based Audio Watermarking
标题: 基于深度神经网络的音频水印的噪屏蔽比损失
作者:Martin Moritz,Toni Olán,Tuomas Virtanen
备注:6 pages, 7 figures
链接:点击下载PDF文件
【14】 Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
标题: 放下节拍!伴奏条件说唱语音生成的Freestyler
作者:Ziqian Ning,Shuai Wang,Yuepeng Jiang,Jixun Yao,Lei He,Shifeng Pan,Jie Ding,Lei Xie
链接:点击下载PDF文件
【15】 Examining the Interplay Between Privacy and Fairness for Speech Processing: A Review and Perspective
标题: 探讨语音处理隐私与公平性之间的相互作用:回顾与展望
作者:Anna Leschanowsky,Sneha Das
链接:点击下载PDF文件
【16】 Feature Representations for Automatic Meerkat Vocalization Classification
标题: 猫鼬发声自动分类的特征表示
作者:Imen Ben Mahmoud,Eklavya Sarkar,Marta Manser,Mathew Magimai. -Doss
备注:Accepted at Interspeech 2024 satellite event (VIHAR 2024)
链接:点击下载PDF文件
标题: Zero-Shot语音克隆的多模式对抗训练
作者:John Janiczek,Dading Chong,Dongyang Dai,Arlo Faria,Chao Wang,Tao Wang,Yuzong Liu
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【2】 Spoofing-Robust Speaker Verification Using Parallel Embedding Fusion: BTU Speech Group's Approach for ASVspoof5 Challenge
标题: 使用并行嵌入融合的欺骗稳健说话人验证:BTU语音小组应对ASVspoof 5挑战的方法
作者:Oğuzhan Kurnaz,Selim Can Demirtaş,Aykut Büker,Jagabandhu Mishra,Cemal Hanilçi
备注:Accepted in ASVspoof2024 workshop
链接:点击下载PDF文件
【3】 ModalityMirror: Improving Audio Classification in Modality Heterogeneity Federated Learning with Multimodal Distillation
标题: ModalityMirror:通过多模式蒸馏改进情态异相联邦学习中的音频分类
作者:Tiantian Feng,Tuo Zhang,Salman Avestimehr,Shrikanth S. Narayanan
链接:点击下载PDF文件
【4】 Easy, Interpretable, Effective: openSMILE for voice deepfake detection
标题: 简单、可解释、有效:用于语音深度伪造检测的openSMILE
作者:Octavian Pascu,Dan Oneata,Horia Cucu,Nicolas M. Müller
链接:点击下载PDF文件
【5】 wav2pos: Sound Source Localization using Masked Autoencoders
标题: wav 2 pos:使用掩蔽自动编码器的声音源定位
作者:Axel Berg,Jens Gulin,Mark O'Connor,Chuteng Zhou,Karl Åström,Magnus Oskarsson
备注:IPIN 2024
链接:点击下载PDF文件
【6】 A Hybrid Approach for Low-Complexity Joint Acoustic Echo and Noise Reduction
标题: 低复杂度联合声学回声和降噪的混合方法
作者:Shrishti Saha Shetu,Naveen Kumar Desiraju,Jose Miguel Martinez Aponte,Emanuël A. P. Habets,Edwin Mabande
备注:5 pages, 2 figures
链接:点击下载PDF文件
【7】 Spectral Masking with Explicit Time-Context Windowing for Neural Network-Based Monaural Speech Enhancement
标题: 基于神经网络的单耳语音增强的显式时间上下文窗口频谱掩蔽
作者:Luan Vinícius Fiorio,Boris Karanov,Bruno Defraene,Johan David,Wim van Houtum,Frans Widdershoven,Ronald M. Aarts
备注:This work has been submitted to the IEEE for possible publication
链接:点击下载PDF文件
【8】 Noise-to-mask Ratio Loss for Deep Neural Network based Audio Watermarking
标题: 基于深度神经网络的音频水印的噪屏蔽比损失
作者:Martin Moritz,Toni Olán,Tuomas Virtanen
备注:6 pages, 7 figures
链接:点击下载PDF文件
【9】 Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
标题: 放下节拍!伴奏条件说唱语音生成的Freestyler
作者:Ziqian Ning,Shuai Wang,Yuepeng Jiang,Jixun Yao,Lei He,Shifeng Pan,Jie Ding,Lei Xie
链接:点击下载PDF文件
【10】 Examining the Interplay Between Privacy and Fairness for Speech Processing: A Review and Perspective
标题: 探讨语音处理隐私与公平性之间的相互作用:回顾与展望
作者:Anna Leschanowsky,Sneha Das
链接:点击下载PDF文件
【11】 YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection
标题: YOLO-Stutter:端到端区域语音流畅性检测
作者:Xuanru Zhou,Anshul Kashyap,Steve Li,Ayati Sharma,Brittany Morin,David Baquirin,Jet Vonk,Zoe Ezzes,Zachary Miller,Maria Luisa Gorno Tempini,Jiachen Lian,Gopala Krishna Anumanchipalli
备注:Interspeech 2024
链接:点击下载PDF文件
【12】 Feature Representations for Automatic Meerkat Vocalization Classification
标题: 猫鼬发声自动分类的特征表示
作者:Imen Ben Mahmoud,Eklavya Sarkar,Marta Manser,Mathew Magimai. -Doss
备注:Accepted at Interspeech 2024 satellite event (VIHAR 2024)
链接:点击下载PDF文件
【13】 VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
标题: VoxDirect:具有统一多语言编解码器语言建模的表达性人类指令到语音生成
作者:Yixuan Zhou,Xiaoyu Qin,Zeyu Jin,Shuoyi Zhou,Shun Lei,Songtao Zhou,Zhiyong Wu,Jia Jia
备注:Accepted by ACM Multimedia 2024
链接:点击下载PDF文件
【14】 Towards reliable respiratory disease diagnosis based on cough sounds and vision transformers
标题: 基于咳嗽声和视觉转换器实现可靠的呼吸道疾病诊断
作者:Qian Wang,Zhaoyang Bu,Jiaxuan Mao,Wenyu Zhu,Jingya Zhao,Wei Du,Guochao Shi,Min Zhou,Si Chen,Jieming Qu
链接:点击下载PDF文件
【15】 Beyond Levenshtein: Leveraging Multiple Algorithms for Robust Word Error Rate Computations And Granular Error Classifications
标题: 超越Levenshtein:利用多种算法进行稳健的误字率计算和粒度错误分类
作者:Korbinian Kuhn,Verena Kersken,Gottfried Zimmermann
备注:Accepted in INTERSPEECH 2024
链接:点击下载PDF文件
【16】 Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
标题: Whisper-PMFA:使用Whisper模型进行说话人验证的部分多尺度特征聚集
作者:Yiyang Zhao,Shuai Wang,Guangzhi Sun,Zehua Chen,Chao Zhang,Mingxing Xu,Thomas Fang Zheng
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【17】 EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
标题: MIDI攻击:利用情感语音转换对深度语音分类模型进行语音后门攻击
作者:Wenhan Yao,Zedong XingXiarun Chen,Jia Liu,yongqiang He,Weiping Wen
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【18】 SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
标题: SimpleSpeech 2:利用基于流的纯量潜在Transformer扩散模型实现简单有效的文本到语音
作者:Dongchao Yang,Rongjie Huang,Yuanyuan Wang,Haohan Guo,Dading Chong,Songxiang Liu,Xixin Wu,Helen Meng
备注:Submit to TASLP
链接:点击下载PDF文件
标题: VoxDirect:具有统一多语言编解码器语言建模的表达性人类指令到语音生成
作者:Yixuan Zhou,Xiaoyu Qin,Zeyu Jin,Shuoyi Zhou,Shun Lei,Songtao Zhou,Zhiyong Wu,Jia Jia
备注:Accepted by ACM Multimedia 2024
链接:点击下载PDF文件
摘要:最近的AIGC系统具有基于人类语言指令生成数字多媒体内容的能力,例如文本、图像和视频。然而,当涉及到语音时,与人类语音生成相关的现有方法表现出两个限制。首先,它们需要将输入分为内容提示(抄本)和描述提示(风格和说话人),而不是直接支持人类指令。这种划分在形式上不太自然,与其他AIGC模型不一致。其次,利用独立的描述提示来建模语音风格,而不考虑转录内容的做法,限制了在细粒度级别上控制语音的能力。为了解决这些限制,我们提出了VoxInstruct,一种新的统一的多语言编解码器语言建模框架,将传统的文本到语音的任务扩展到一个一般的人类语音转换任务。我们的方法增强了人类的表达能力的指导下的语音生成,并与其他形式的语音生成范式对齐。为了使该模型能够自动从原始文本指令中提取合成语音的内容,我们引入了语音语义令牌作为中间表示,用于解释到内容的指导。我们还将多个无分类器指导(CFG)策略纳入到我们的编解码器语言模型中,从而增强了遵循人类指令生成的语音。此外,我们的模型架构和训练策略允许同时支持结合语音提示和描述性的人类指令来进行表达性语音合成,这是第一次尝试。代码、模型和演示请访问:https: github.com thuhcsi VoxInstruct。摘要:Recent AIGC systems possess the capability to generate digital multimedia content based on human language instructions, such as text, image and video. However, when it comes to speech, existing methods related to human instruction-to-speech generation exhibit two limitations. Firstly, they require the division of inputs into content prompt (transcript) and description prompt (style and speaker), instead of directly supporting human instruction. This division is less natural in form and does not align with other AIGC models. Secondly, the practice of utilizing an independent description prompt to model speech style, without considering the transcript content, restricts the ability to control speech at a fine-grained level. To address these limitations, we propose VoxInstruct, a novel unified multilingual codec language modeling framework that extends traditional text-to-speech tasks into a general human instruction-to-speech task. Our approach enhances the expressiveness of human instruction-guided speech generation and aligns the speech generation paradigm with other modalities. To enable the model to automatically extract the content of synthesized speech from raw text instructions, we introduce speech semantic tokens as an intermediate representation for instruction-to-content guidance. We also incorporate multiple Classifier-Free Guidance (CFG) strategies into our codec language model, which strengthens the generated speech following human instructions. Furthermore, our model architecture and training strategies allow for the simultaneous support of combining speech prompt and descriptive human instruction for expressive speech synthesis, which is a first-of-its-kind attempt. Codes, models and demos are at: https: github.com thuhcsi VoxInstruct.
【2】 Towards reliable respiratory disease diagnosis based on cough sounds and vision transformers
标题: 基于咳嗽声和视觉转换器实现可靠的呼吸道疾病诊断
作者:Qian Wang,Zhaoyang Bu,Jiaxuan Mao,Wenyu Zhu,Jingya Zhao,Wei Du,Guochao Shi,Min Zhou,Si Chen,Jieming Qu
链接:点击下载PDF文件
摘要:深度学习技术的最新进展引发了各种现实应用的性能提升,包括基于多模态医疗数据的疾病诊断。基于咳嗽声数据的呼吸系统疾病(例如,COVID-19和慢性阻塞性肺疾病)诊断也备受关注。然而,现有的作品通常使用传统的机器学习或中等规模的深度模型。另一方面,开发的方法是在小规模数据上训练和评估的,因为很难大规模地管理和注释临床数据。为了解决先前工作中的这些问题,我们创建了一个统一的框架来评估来自轻量级卷积神经网络的各种深度模型(例如,ResNet 18)与现代Vision Transformers进行比较,并比较它们在呼吸系统疾病分类中的性能。基于如此广泛的实证研究的观察结果,我们提出了一种基于大规模咳嗽数据集的自监督和监督学习的基于咳嗽的疾病分类新方法。实验结果表明,我们提出的方法在COVID-19诊断的两个基准数据集和COPD 非COPD分类的专有数据集上始终优于现有技术,AUROC为92.5%。摘要:Recent advancements in deep learning techniques have sparked performance boosts in various real-world applications including disease diagnosis based on multi-modal medical data. Cough sound data-based respiratory disease (e.g., COVID-19 and Chronic Obstructive Pulmonary Disease) diagnosis has also attracted much attention. However, existing works usually utilise traditional machine learning or deep models of moderate scales. On the other hand, the developed approaches are trained and evaluated on small-scale data due to the difficulty of curating and annotating clinical data on scale. To address these issues in prior works, we create a unified framework to evaluate various deep models from lightweight Convolutional Neural Networks (e.g., ResNet18) to modern vision transformers and compare their performance in respiratory disease classification. Based on the observations from such an extensive empirical study, we propose a novel approach to cough-based disease classification based on both self-supervised and supervised learning on a large-scale cough data set. Experimental results demonstrate our proposed approach outperforms prior arts consistently on two benchmark datasets for COVID-19 diagnosis and a proprietary dataset for COPD non-COPD classification with an AUROC of 92.5%.
【3】 Beyond Levenshtein: Leveraging Multiple Algorithms for Robust Word Error Rate Computations And Granular Error Classifications
标题: 超越Levenshtein:利用多种算法进行稳健的误字率计算和粒度错误分类
作者:Korbinian Kuhn,Verena Kersken,Gottfried Zimmermann
备注:Accepted in INTERSPEECH 2024
链接:点击下载PDF文件
摘要:词错误率(WER)是自动语音识别(ASR)准确性的常用度量。转录通常通过替换特定字符来进行预处理,以解决非语义差异。由于这种标准化,标点符号或大写字母的准确性信息丢失。我们提出了一种非破坏性的,基于令牌的方法,使用扩展Levenshtein距离算法来计算一个强大的WER和额外的正交度量。转录错误也被现有的字符串相似性和语音算法更细粒度地分类。几个数据集上的评估表明,我们的方法相比,常见的WER计算的实际等效性。我们还提供了衍生用例的示例性分析,例如标点错误率,以及用于交互式使用和可视化实现的Web应用程序。该代码是开放源代码。摘要:The Word Error Rate (WER) is the common measure of accuracy for Automatic Speech Recognition (ASR). Transcripts are usually pre-processed by substituting specific characters to account for non-semantic differences. As a result of this normalisation, information on the accuracy of punctuation or capitalisation is lost. We present a non-destructive, token-based approach using an extended Levenshtein distance algorithm to compute a robust WER and additional orthographic metrics. Transcription errors are also classified more granularly by existing string similarity and phonetic algorithms. An evaluation on several datasets demonstrates the practical equivalence of our approach compared to common WER computations. We also provide an exemplary analysis of derived use cases, such as a punctuation error rate, and a web application for interactive use and visualisation of our implementation. The code is available open-source.
【4】 Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
标题: Whisper-PMFA:使用Whisper模型进行说话人验证的部分多尺度特征聚集
作者:Yiyang Zhao,Shuai Wang,Guangzhi Sun,Zehua Chen,Chao Zhang,Mingxing Xu,Thomas Fang Zheng
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在本文中,耳语,一个大规模的自动语音识别的预训练模型,提出了适用于说话人确认。提出了一种基于Whisper编码块子集的部分多尺度特征聚合(PMFA)方法,以获得高区分度的说话人信息嵌入,实验结果表明,使用Whisper编码块的中间到后面的块可以保留更多的说话人信息。在VoxCeleb 1和CN-Celeb 1数据集上,我们的系统分别实现了1.42%和8.23%的等错误率(EER),比ECAPA-TDNN基线分别减少了0.58%和1.81%,比ResNet 34基线分别减少了0.46%和0.97%。此外,我们的结果表明,使用经过多语言数据训练的Whisper模型可以有效增强模型的跨语言鲁棒性。最后,低秩自适应方法进行了评估,它减少了约45倍的可训练模型参数,而EER仅略有增加0.2%。摘要:In this paper, Whisper, a large-scale pre-trained model for automatic speech recognition, is proposed to apply to speaker verification. A partial multi-scale feature aggregation (PMFA) approach is proposed based on a subset of Whisper encoder blocks to derive highly discriminative speaker embeddings.Experimental results demonstrate that using the middle to later blocks of the Whisper encoder keeps more speaker information. On the VoxCeleb1 and CN-Celeb1 datasets, our system achieves 1.42% and 8.23% equal error rates (EERs) respectively, receiving 0.58% and 1.81% absolute EER reductions over the ECAPA-TDNN baseline, and 0.46% and 0.97% over the ResNet34 baseline. Furthermore, our results indicate that using Whisper models trained on multilingual data can effectively enhance the model's robustness across languages. Finally, the low-rank adaptation approach is evaluated, which reduces the trainable model parameters by approximately 45 times while only slightly increasing EER by 0.2%.
【5】 EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
标题: MIDI攻击:利用情感语音转换对深度语音分类模型进行语音后门攻击
作者:Wenhan Yao,Zedong XingXiarun Chen,Jia Liu,yongqiang He,Weiping Wen
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:深度语音分类任务,主要包括关键词发现和说话人确认,在基于语音的人机交互中起着至关重要的作用。最近,这些技术的安全性已被证明容易受到后门攻击。具体而言,语音样本受到噪声干扰和当前触发器中的分量修改的攻击。我们认为,语音后门攻击可以战略性地集中在情感上,这是语音中固有的更高层次的主观感知属性。此外,我们提出了情感语音转换技术可以作为语音后门攻击的触发器,这种方法被称为语音后门攻击。在此基础上,对两个语音分类任务进行了攻击实验,实验结果表明,该方法具有较好的触发效果和显著的攻击成功率及准确率差异。此外,消融实验发现,具有强烈情感的语音更适合作为攻击目标。摘要:Deep speech classification tasks, mainly including keyword spotting and speaker verification, play a crucial role in speech-based human-computer interaction. Recently, the security of these technologies has been demonstrated to be vulnerable to backdoor attacks. Specifically speaking, speech samples are attacked by noisy disruption and component modification in present triggers. We suggest that speech backdoor attacks can strategically focus on emotion, a higher-level subjective perceptual attribute inherent in speech. Furthermore, we proposed that emotional voice conversion technology can serve as the speech backdoor attack trigger, and the method is called EmoAttack. Based on this, we conducted attack experiments on two speech classification tasks, showcasing that EmoAttack method owns impactful trigger effectiveness and its remarkable attack success rate and accuracy variance. Additionally, the ablation experiments found that speech with intensive emotion is more suitable to be targeted for attacks.
【6】 Multi-modal Adversarial Training for Zero-Shot Voice Cloning
标题: Zero-Shot语音克隆的多模式对抗训练
作者:John Janiczek,Dading Chong,Dongyang Dai,Arlo Faria,Chao Wang,Tao Wang,Yuzong Liu
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:一个被训练来重建给定文本的语音的文本到语音(TTS)模型倾向于接近数据集平均特征的预测,无法对使人类语音听起来自然的变化进行建模。这个问题在zero-shot语音克隆中被放大了,这是一项需要在说话风格中具有高变化的训练数据的任务。我们建立了最近的作品,其中使用生成Advsarial网络(GAN)提出了一个Transformer编码器-解码器架构,有条件地区分真实和生成的语音特征。在训练管道中使用该训练器,该训练管道改善TTS模型的声学和韵律特征。我们通过将其应用于FastSpeech 2声学模型并在大型多说话者数据集Libriheight上进行训练来介绍我们的新型对抗训练技术,用于zero-shot语音克隆任务。我们的模型在语音质量和说话人相似性方面比基线有所改善。我们系统中的音频示例可在线获取。摘要:A text-to-speech (TTS) model trained to reconstruct speech given text tends towards predictions that are close to the average characteristics of a dataset, failing to model the variations that make human speech sound natural. This problem is magnified for zero-shot voice cloning, a task that requires training data with high variance in speaking styles. We build off of recent works which have used Generative Advsarial Networks (GAN) by proposing a Transformer encoder-decoder architecture to conditionally discriminates between real and generated speech features. The discriminator is used in a training pipeline that improves both the acoustic and prosodic features of a TTS model. We introduce our novel adversarial training technique by applying it to a FastSpeech2 acoustic model and training on Libriheavy, a large multi-speaker dataset, for the task of zero-shot voice cloning. Our model achieves improvements over the baseline in terms of speech quality and speaker similarity. Audio examples from our system are available online.
【7】 Spoofing-Robust Speaker Verification Using Parallel Embedding Fusion: BTU Speech Group's Approach for ASVspoof5 Challenge
标题: 使用并行嵌入融合的欺骗稳健说话人验证:BTU语音小组应对ASVspoof 5挑战的方法
作者:Oğuzhan Kurnaz,Selim Can Demirtaş,Aykut Büker,Jagabandhu Mishra,Cemal Hanilçi
备注:Accepted in ASVspoof2024 workshop
链接:点击下载PDF文件
摘要:本文介绍了BTU Speech Group为ASVspoof 5挑战赛开发的基于并行网络的欺骗感知说话人确认(SASV)系统。SASV系统集成了ASV和CM系统,以增强针对欺骗攻击的安全性。我们的方法采用ASV模型(ECAPA-TDNN,WavLM)和CM模型(AASIST)的得分和嵌入融合。融合嵌入使用简单的DNN结构进行处理,结合最近提出的a-DCF和BCE损失优化模型性能。我们引入了一种新的并行网络结构,其中两个相同的DNN,以不同的输入,独立地处理嵌入并产生SASV分数。最终的SASV概率是通过对这些分数求平均值而得到的,从而增强了鲁棒性和准确性。实验结果表明,所提出的并行DNN结构优于传统的单一DNN方法,提供了一个更可靠和安全的说话人确认系统对抗欺骗攻击。摘要:This paper introduces the parallel network-based spoofing-aware speaker verification (SASV) system developed by BTU Speech Group for the ASVspoof5 Challenge. The SASV system integrates ASV and CM systems to enhance security against spoofing attacks. Our approach employs score and embedding fusion from ASV models (ECAPA-TDNN, WavLM) and CM models (AASIST). The fused embeddings are processed using a simple DNN structure, optimizing model performance with a combination of recently proposed a-DCF and BCE losses. We introduce a novel parallel network structure where two identical DNNs, fed with different inputs, independently process embeddings and produce SASV scores. The final SASV probability is derived by averaging these scores, enhancing robustness and accuracy. Experimental results demonstrate that the proposed parallel DNN structure outperforms traditional single DNN methods, offering a more reliable and secure speaker verification system against spoofing attacks.
【8】 ModalityMirror: Improving Audio Classification in Modality Heterogeneity Federated Learning with Multimodal Distillation
标题: ModalityMirror:通过多模式蒸馏改进情态异相联邦学习中的音频分类
作者:Tiantian Feng,Tuo Zhang,Salman Avestimehr,Shrikanth S. Narayanan
链接:点击下载PDF文件
摘要:多模态联合学习经常遇到客户端模态异质性的挑战,导致多模态学习中的次要模态表现不理想。它在视听学习中特别普遍,音频通常被认为是识别任务中较弱的模态。为了应对这一挑战,我们引入了ModalityMirror,通过利用视听联合学习模型的知识蒸馏来提高音频模型的性能。ModalityMirror包括两个阶段:模态方面的FL阶段,以聚合单峰编码器;以及多模态客户端上的联合知识蒸馏阶段,以训练单峰学生模型。我们的研究结果表明,ModalityMirror显着提高了音频分类相比,最先进的FL方法,如和谐,特别是在视听FL面临的视频丢失。我们的方法解锁的潜力,利用固有的多模态FL的多样的模态谱。摘要:Multimodal Federated Learning frequently encounters challenges of client modality heterogeneity, leading to undesired performances for secondary modality in multimodal learning. It is particularly prevalent in audiovisual learning, with audio is often assumed to be the weaker modality in recognition tasks. To address this challenge, we introduce ModalityMirror to improve audio model performance by leveraging knowledge distillation from an audiovisual federated learning model. ModalityMirror involves two phases: a modality-wise FL stage to aggregate uni-modal encoders; and a federated knowledge distillation stage on multi-modality clients to train an unimodal student model. Our results demonstrate that ModalityMirror significantly improves the audio classification compared to the state-of-the-art FL methods such as Harmony, particularly in audiovisual FL facing video missing. Our approach unlocks the potential for exploiting the diverse modality spectrum inherent in multi-modal FL.
【9】 Easy, Interpretable, Effective: openSMILE for voice deepfake detection
标题: 简单、可解释、有效:用于语音深度伪造检测的openSMILE
作者:Octavian Pascu,Dan Oneata,Horia Cucu,Nicolas M. Müller
链接:点击下载PDF文件
摘要:在本文中,我们证明了最新ASVspoof 5数据集中的攻击-语音真实性和深度伪造检测领域的事实标准-可以使用非常简单的特征的一小部分以惊人的准确性进行识别。这些都是从openSMILE库派生的,并且是标量值的,易于计算,并且是人类可解释的。例如,攻击A10的无声段的平均长度为0.09 pm 0.02,而真正的实例的平均长度为0.18 pm 0.07。单独使用此功能,阈值分类器对攻击A10实现了10.3%的等错误率(EER)。同样,在所有攻击中,我们实现了高达0.8%的EER,总体EER为15.7 pm 6.0%。我们探讨了这些功能的泛化能力,并发现其中一些有效地转移之间的攻击,主要是当攻击来自类似的文本到语音(TTS)架构。这一发现可能表明,语音反欺骗在某种程度上是一个识别和记忆单个TTS系统的签名或指纹的问题。这有助于更好地了解反欺骗模型及其在实际应用中面临的挑战。摘要:In this paper, we demonstrate that attacks in the latest ASVspoof5 dataset -- a de facto standard in the field of voice authenticity and deepfake detection -- can be identified with surprising accuracy using a small subset of very simplistic features. These are derived from the openSMILE library, and are scalar-valued, easy to compute, and human interpretable. For example, attack A10 s unvoiced segments have a mean length of 0.09 pm 0.02, while bona fide instances have a mean length of 0.18 pm 0.07. Using this feature alone, a threshold classifier achieves an Equal Error Rate (EER) of 10.3% for attack A10. Similarly, across all attacks, we achieve up to 0.8% EER, with an overall EER of 15.7 pm 6.0%. We explore the generalization capabilities of these features and find that some of them transfer effectively between attacks, primarily when the attacks originate from similar Text-to-Speech (TTS) architectures. This finding may indicate that voice anti-spoofing is, in part, a problem of identifying and remembering signatures or fingerprints of individual TTS systems. This allows to better understand anti-spoofing models and their challenges in real-world application.
【10】 wav2pos: Sound Source Localization using Masked Autoencoders
标题: wav 2 pos:使用掩蔽自动编码器的声音源定位
作者:Axel Berg,Jens Gulin,Mark O'Connor,Chuteng Zhou,Karl Åström,Magnus Oskarsson
备注:IPIN 2024
链接:点击下载PDF文件
摘要:我们提出了一种新的方法,分布式ad-hoc麦克风阵列的三维声源定位任务,制定它作为一个集到集的回归问题。通过训练一个多模态掩蔽自动编码器模型,该模型对音频记录和麦克风坐标进行操作,我们证明了这样的公式可以通过重建输入中掩蔽的坐标来精确定位声源。我们的方法是灵活的,在这个意义上说,一个单一的模型可以使用任意数量的麦克风,即使当一个子集的音频记录和麦克风坐标丢失。我们测试我们的方法在室内环境中的音乐和语音的模拟和真实世界的录音,并展示了竞争力的表现相比,经典和其他基于学习的本地化方法。摘要:We present a novel approach to the 3D sound source localization task for distributed ad-hoc microphone arrays by formulating it as a set-to-set regression problem. By training a multi-modal masked autoencoder model that operates on audio recordings and microphone coordinates, we show that such a formulation allows for accurate localization of the sound source, by reconstructing coordinates masked in the input. Our approach is flexible in the sense that a single model can be used with an arbitrary number of microphones, even when a subset of audio recordings and microphone coordinates are missing. We test our method on simulated and real-world recordings of music and speech in indoor environments, and demonstrate competitive performance compared to both classical and other learning based localization methods.
【11】 A Hybrid Approach for Low-Complexity Joint Acoustic Echo and Noise Reduction
标题: 低复杂度联合声学回声和降噪的混合方法
作者:Shrishti Saha Shetu,Naveen Kumar Desiraju,Jose Miguel Martinez Aponte,Emanuël A. P. Habets,Edwin Mabande
备注:5 pages, 2 figures
链接:点击下载PDF文件
摘要:联合执行声学回声和降噪(AENR)任务的基于深度学习的方法通常需要较高的内存和计算资源,因此不适合在嵌入式设备等低资源平台上实时部署。我们提出了一种低复杂度的混合方法,联合AENR采用一个单一的模型来抑制残留的回声和噪声成分。具体来说,我们集成了国家的最先进的(SOTA)ULCNet模型,这是最初提出的,以实现超低复杂度的噪声抑制,在一个混合动力系统和训练它的联合AENR。我们表明,所提出的方法实现了更好的回声降低和可比的降噪性能,具有更低的计算复杂度和内存需求比所有考虑SOTA方法,在语音质量略有下降的成本。摘要:Deep learning-based methods that jointly perform the task of acoustic echo and noise reduction (AENR) often require high memory and computational resources, making them unsuitable for real-time deployment on low-resource platforms such as embedded devices. We propose a low-complexity hybrid approach for joint AENR by employing a single model to suppress both residual echo and noise components. Specifically, we integrate the state-of-the-art (SOTA) ULCNet model, which was originally proposed to achieve ultra-low complexity noise suppression, in a hybrid system and train it for joint AENR. We show that the proposed approach achieves better echo reduction and comparable noise reduction performance with much lower computational complexity and memory requirements than all considered SOTA methods, at the cost of slight degradation in speech quality.
【12】 Spectral Masking with Explicit Time-Context Windowing for Neural Network-Based Monaural Speech Enhancement
标题: 基于神经网络的单耳语音增强的显式时间上下文窗口频谱掩蔽
作者:Luan Vinícius Fiorio,Boris Karanov,Bruno Defraene,Johan David,Wim van Houtum,Frans Widdershoven,Ronald M. Aarts
备注:This work has been submitted to the IEEE for possible publication
链接:点击下载PDF文件
摘要:我们提出并分析了使用显式的时间上下文窗口基于神经网络的频谱掩蔽语音增强,以利用相邻帧之间的信号上下文相关性。特别是,我们专注于软掩蔽和损失计算重建语音的时频表示。我们表明,在神经网络模型的输入和输出的时间上下文窗口函数的应用程序,提高了软掩模的估计过程中,通过结合多个估计从不同的上下文中。所提出的方法仅在推理模式中用作后优化,不需要对神经网络模型进行额外的层或特殊训练。我们的研究结果表明,该方法一致地提高了可懂度和信号质量的去噪语音,证明了两类基于卷积的语音增强模型。重要的是,所提出的方法只需要一个微不足道的($ leq1 %$)增加模型参数的数量,使其适合于硬件受限的应用程序。摘要:We propose and analyze the use of an explicit time-context window for neural network-based spectral masking speech enhancement to leverage signal context dependencies between neighboring frames. In particular, we concentrate on soft masking and loss computed on the time-frequency representation of the reconstructed speech. We show that the application of a time-context windowing function at both input and output of the neural network model improves the soft mask estimation process by combining multiple estimates taken from different contexts. The proposed approach is only applied as post-optimization in inference mode, not requiring additional layers or special training for the neural network model. Our results show that the method consistently increases both intelligibility and signal quality of the denoised speech, as demonstrated for two classes of convolutional-based speech enhancement models. Importantly, the proposed method requires only a negligible ($ leq1 %$) increase in the number of model parameters, making it suitable for hardware-constrained applications.
【13】 Noise-to-mask Ratio Loss for Deep Neural Network based Audio Watermarking
标题: 基于深度神经网络的音频水印的噪屏蔽比损失
作者:Martin Moritz,Toni Olán,Tuomas Virtanen
备注:6 pages, 7 figures
链接:点击下载PDF文件
摘要:数字音频水印技术是在音频信号中以透明的方式嵌入一个信息,可以用来自动识别音频资料和管理版权。我们提出了一种感知损失函数,用于基于深度神经网络的音频水印系统。损失是基于噪声掩蔽比(NMR),这是一个模型的心理声学掩蔽效应特性的人耳。我们使用标记信号和宿主信号之间的NMR损失来训练深度神经模型,并使用PEAQ评估客观质量,使用MUSHRA测试评估主观质量。客观和主观测试都表明,使用NMR损失训练的模型比使用常规MSE损失训练的模型生成更多的透明水印摘要:Digital audio watermarking consists in inserting a message into audio signals in a transparent way and can be used to allow automatic recognition of audio material and management of the copyrights. We propose a perceptual loss function to be used in deep neural network based audio watermarking systems. The loss is based on the noise-to-mask ratio (NMR), which is a model of the psychoacoustic masking effect characteristic of the human ear. We use the NMR loss between marked and host signals to train the deep neural models and we evaluate the objective quality with PEAQ and the subjective quality with a MUSHRA test. Both objective and subjective tests show that models trained with NMR loss generate more transparent watermarks than models trained with the conventionally used MSE loss
【14】 Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
标题: 放下节拍!伴奏条件说唱语音生成的Freestyler
作者:Ziqian Ning,Shuai Wang,Yuepeng Jiang,Jixun Yao,Lei He,Shifeng Pan,Jie Ding,Lei Xie
链接:点击下载PDF文件
摘要:说唱是一种重要的声乐表演形式,但在声乐创作中却没有得到充分的开发。一般的声乐合成依赖于精确的音符和音长输入,要求用户具有相关的音乐知识,这限制了灵活性。相比之下,说唱通常具有更简单的旋律,其核心重点是与伴随的节拍协调的强烈节奏感。在本文中,我们提出了Freestyler,第一个系统,直接从歌词和伴奏输入生成说唱人声。Freestyler利用基于语言模型的令牌生成,然后是条件流匹配模型来生成频谱图和神经声码器来恢复音频。它允许3秒提示启用zero-shot音色控制。由于公开可用的说唱数据集的稀缺性,我们还介绍了RapBank,一个从互联网上收集的说唱歌曲数据集,以及精心设计的处理管道。实验结果表明,Freestyler产生高质量的说唱语音生成与增强的自然和强烈的对齐伴随节拍,无论是风格和节奏。摘要:Rap, a prominent genre of vocal performance, remains underexplored in vocal generation. General vocal synthesis depends on precise note and duration inputs, requiring users to have related musical knowledge, which limits flexibility. In contrast, rap typically features simpler melodies, with a core focus on a strong rhythmic sense that harmonizes with accompanying beats. In this paper, we propose Freestyler, the first system that generates rapping vocals directly from lyrics and accompaniment inputs. Freestyler utilizes language model-based token generation, followed by a conditional flow matching model to produce spectrograms and a neural vocoder to restore audio. It allows a 3-second prompt to enable zero-shot timbre control. Due to the scarcity of publicly available rap datasets, we also present RapBank, a rap song dataset collected from the internet, alongside a meticulously designed processing pipeline. Experimental results show that Freestyler produces high-quality rapping voice generation with enhanced naturalness and strong alignment with accompanying beats, both stylistically and rhythmically.
【15】 Examining the Interplay Between Privacy and Fairness for Speech Processing: A Review and Perspective
标题: 探讨语音处理隐私与公平性之间的相互作用:回顾与展望
作者:Anna Leschanowsky,Sneha Das
链接:点击下载PDF文件
摘要:语音技术已越来越多地应用于日常生活的各个领域,包括医疗保健和执法等敏感领域。为了使这些技术有效,它们必须对所有用户可靠地工作,同时保护个人隐私。虽然隐私和效用之间的权衡,以及公平和效用,已被广泛研究,在语音处理中的隐私和公平之间的具体相互作用仍然没有得到充分的探索。这篇评论和立场文件概述了整个语音处理机器学习生命周期中新兴的隐私公平权衡。通过借鉴成熟的公平和隐私框架,我们研究了语音处理模型开发过程中共存的现有偏见和隐私伤害来源。然后,我们强调了相应的隐私增强技术如何有可能无意中增加这些偏见,以及偏见缓解策略如何反过来减少隐私。通过提出开放性问题,我们主张对语音技术的隐私公平权衡进行全面评估,并在这一领域开发隐私增强和公平感知算法。摘要:Speech technology has been increasingly deployed in various areas of daily life including sensitive domains such as healthcare and law enforcement. For these technologies to be effective, they must work reliably for all users while preserving individual privacy. Although tradeoffs between privacy and utility, as well as fairness and utility, have been extensively researched, the specific interplay between privacy and fairness in speech processing remains underexplored. This review and position paper offers an overview of emerging privacy-fairness tradeoffs throughout the entire machine learning lifecycle for speech processing. By drawing on well-established frameworks on fairness and privacy, we examine existing biases and sources of privacy harm that coexist during the development of speech processing models. We then highlight how corresponding privacy-enhancing technologies have the potential to inadvertently increase these biases and how bias mitigation strategies may conversely reduce privacy. By raising open questions, we advocate for a comprehensive evaluation of privacy-fairness tradeoffs for speech technology and the development of privacy-enhancing and fairness-aware algorithms in this domain.
【16】 Feature Representations for Automatic Meerkat Vocalization Classification
标题: 猫鼬发声自动分类的特征表示
作者:Imen Ben Mahmoud,Eklavya Sarkar,Marta Manser,Mathew Magimai. -Doss
备注:Accepted at Interspeech 2024 satellite event (VIHAR 2024)
链接:点击下载PDF文件
摘要:了解群居动物声音交流的进化是一个重要的研究问题。在这种情况下,除了人类,人们对分析其他社会动物的发声感兴趣,如猫鼬,绒猴,猿。虽然现有的方法解决了某些物种的发声问题,但缺乏针对猫鼬叫声的可靠方法。在这种程度上,本文探讨了自动猫鼬发声分析的特征表示。探索了传统的基于信号处理的表示和由深度学习的进步促进的数据驱动的表示。两个数据集上进行的呼叫类型分类研究表明,人类语音处理的特征提取方法可以有效地用于自动猫鼬呼叫分析。摘要:Understanding evolution of vocal communication in social animals is an important research problem. In that context, beyond humans, there is an interest in analyzing vocalizations of other social animals such as, meerkats, marmosets, apes. While existing approaches address vocalizations of certain species, a reliable method tailored for meerkat calls is lacking. To that extent, this paper investigates feature representations for automatic meerkat vocalization analysis. Both traditional signal processing-based representations and data-driven representations facilitated by advances in deep learning are explored. Call type classification studies conducted on two data sets reveal that feature extraction methods developed for human speech processing can be effectively employed for automatic meerkat call analysis.
eess.AS音频处理
【1】 Multi-modal Adversarial Training for Zero-Shot Voice Cloning标题: Zero-Shot语音克隆的多模式对抗训练
作者:John Janiczek,Dading Chong,Dongyang Dai,Arlo Faria,Chao Wang,Tao Wang,Yuzong Liu
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:一个被训练来重建给定文本的语音的文本到语音(TTS)模型倾向于接近数据集平均特征的预测,无法对使人类语音听起来自然的变化进行建模。这个问题在zero-shot语音克隆中被放大了,这是一项需要在说话风格中具有高变化的训练数据的任务。我们建立了最近的作品,其中使用生成Advsarial网络(GAN)提出了一个Transformer编码器-解码器架构,有条件地区分真实和生成的语音特征。在训练管道中使用该训练器,该训练管道改善TTS模型的声学和韵律特征。我们通过将其应用于FastSpeech 2声学模型并在大型多说话者数据集Libriheight上进行训练来介绍我们的新型对抗训练技术,用于zero-shot语音克隆任务。我们的模型在语音质量和说话人相似性方面比基线有所改善。我们系统中的音频示例可在线获取。摘要:A text-to-speech (TTS) model trained to reconstruct speech given text tends towards predictions that are close to the average characteristics of a dataset, failing to model the variations that make human speech sound natural. This problem is magnified for zero-shot voice cloning, a task that requires training data with high variance in speaking styles. We build off of recent works which have used Generative Advsarial Networks (GAN) by proposing a Transformer encoder-decoder architecture to conditionally discriminates between real and generated speech features. The discriminator is used in a training pipeline that improves both the acoustic and prosodic features of a TTS model. We introduce our novel adversarial training technique by applying it to a FastSpeech2 acoustic model and training on Libriheavy, a large multi-speaker dataset, for the task of zero-shot voice cloning. Our model achieves improvements over the baseline in terms of speech quality and speaker similarity. Audio examples from our system are available online.
【2】 Spoofing-Robust Speaker Verification Using Parallel Embedding Fusion: BTU Speech Group's Approach for ASVspoof5 Challenge
标题: 使用并行嵌入融合的欺骗稳健说话人验证:BTU语音小组应对ASVspoof 5挑战的方法
作者:Oğuzhan Kurnaz,Selim Can Demirtaş,Aykut Büker,Jagabandhu Mishra,Cemal Hanilçi
备注:Accepted in ASVspoof2024 workshop
链接:点击下载PDF文件
摘要:本文介绍了BTU Speech Group为ASVspoof 5挑战赛开发的基于并行网络的欺骗感知说话人确认(SASV)系统。SASV系统集成了ASV和CM系统,以增强针对欺骗攻击的安全性。我们的方法采用ASV模型(ECAPA-TDNN,WavLM)和CM模型(AASIST)的得分和嵌入融合。融合嵌入使用简单的DNN结构进行处理,并结合最近提出的a-DCF和BCE损失来优化模型性能。我们引入了一种新的并行网络结构,其中两个相同的DNN,以不同的输入,独立地处理嵌入并产生SASV分数。最终的SASV概率是通过对这些分数求平均值而得到的,从而增强了鲁棒性和准确性。实验结果表明,所提出的并行DNN结构优于传统的单一DNN方法,提供了一个更可靠和安全的说话人确认系统对抗欺骗攻击。摘要:This paper introduces the parallel network-based spoofing-aware speaker verification (SASV) system developed by BTU Speech Group for the ASVspoof5 Challenge. The SASV system integrates ASV and CM systems to enhance security against spoofing attacks. Our approach employs score and embedding fusion from ASV models (ECAPA-TDNN, WavLM) and CM models (AASIST). The fused embeddings are processed using a simple DNN structure, optimizing model performance with a combination of recently proposed a-DCF and BCE losses. We introduce a novel parallel network structure where two identical DNNs, fed with different inputs, independently process embeddings and produce SASV scores. The final SASV probability is derived by averaging these scores, enhancing robustness and accuracy. Experimental results demonstrate that the proposed parallel DNN structure outperforms traditional single DNN methods, offering a more reliable and secure speaker verification system against spoofing attacks.
【3】 ModalityMirror: Improving Audio Classification in Modality Heterogeneity Federated Learning with Multimodal Distillation
标题: ModalityMirror:通过多模式蒸馏改进情态异相联邦学习中的音频分类
作者:Tiantian Feng,Tuo Zhang,Salman Avestimehr,Shrikanth S. Narayanan
链接:点击下载PDF文件
摘要:多模态联合学习经常遇到客户端模态异质性的挑战,导致多模态学习中的次要模态表现不理想。它在视听学习中特别普遍,音频通常被认为是识别任务中较弱的模态。为了应对这一挑战,我们引入了ModalityMirror,通过利用视听联合学习模型的知识蒸馏来提高音频模型的性能。ModalityMirror包括两个阶段:模态方面的FL阶段,以聚合单峰编码器;以及多模态客户端上的联合知识蒸馏阶段,以训练单峰学生模型。我们的研究结果表明,ModalityMirror显着提高了音频分类相比,最先进的FL方法,如和谐,特别是在视听FL面临的视频丢失。我们的方法解锁的潜力,利用固有的多模态FL的多样的模态谱。摘要:Multimodal Federated Learning frequently encounters challenges of client modality heterogeneity, leading to undesired performances for secondary modality in multimodal learning. It is particularly prevalent in audiovisual learning, with audio is often assumed to be the weaker modality in recognition tasks. To address this challenge, we introduce ModalityMirror to improve audio model performance by leveraging knowledge distillation from an audiovisual federated learning model. ModalityMirror involves two phases: a modality-wise FL stage to aggregate uni-modal encoders; and a federated knowledge distillation stage on multi-modality clients to train an unimodal student model. Our results demonstrate that ModalityMirror significantly improves the audio classification compared to the state-of-the-art FL methods such as Harmony, particularly in audiovisual FL facing video missing. Our approach unlocks the potential for exploiting the diverse modality spectrum inherent in multi-modal FL.
【4】 Easy, Interpretable, Effective: openSMILE for voice deepfake detection
标题: 简单、可解释、有效:用于语音深度伪造检测的openSMILE
作者:Octavian Pascu,Dan Oneata,Horia Cucu,Nicolas M. Müller
链接:点击下载PDF文件
摘要:在本文中,我们证明了最新ASVspoof 5数据集中的攻击-语音真实性和深度伪造检测领域的事实标准-可以使用非常简单的特征的一小部分以惊人的准确性进行识别。这些都是从openSMILE库派生的,并且是标量值的,易于计算,并且是人类可解释的。例如,攻击A10的无声段的平均长度为0.09 pm 0.02,而真正的实例的平均长度为0.18 pm 0.07。单独使用此功能,阈值分类器对攻击A10实现了10.3%的等错误率(EER)。同样,在所有攻击中,我们实现了高达0.8%的EER,总体EER为15.7 pm 6.0%。我们探讨了这些功能的泛化能力,并发现其中一些有效地转移之间的攻击,主要是当攻击来自类似的文本到语音(TTS)架构。这一发现可能表明,语音反欺骗在某种程度上是一个识别和记忆单个TTS系统的签名或指纹的问题。这有助于更好地了解反欺骗模型及其在实际应用中面临的挑战。摘要:In this paper, we demonstrate that attacks in the latest ASVspoof5 dataset -- a de facto standard in the field of voice authenticity and deepfake detection -- can be identified with surprising accuracy using a small subset of very simplistic features. These are derived from the openSMILE library, and are scalar-valued, easy to compute, and human interpretable. For example, attack A10 s unvoiced segments have a mean length of 0.09 pm 0.02, while bona fide instances have a mean length of 0.18 pm 0.07. Using this feature alone, a threshold classifier achieves an Equal Error Rate (EER) of 10.3% for attack A10. Similarly, across all attacks, we achieve up to 0.8% EER, with an overall EER of 15.7 pm 6.0%. We explore the generalization capabilities of these features and find that some of them transfer effectively between attacks, primarily when the attacks originate from similar Text-to-Speech (TTS) architectures. This finding may indicate that voice anti-spoofing is, in part, a problem of identifying and remembering signatures or fingerprints of individual TTS systems. This allows to better understand anti-spoofing models and their challenges in real-world application.
【5】 wav2pos: Sound Source Localization using Masked Autoencoders
标题: wav 2 pos:使用掩蔽自动编码器的声音源定位
作者:Axel Berg,Jens Gulin,Mark O'Connor,Chuteng Zhou,Karl Åström,Magnus Oskarsson
备注:IPIN 2024
链接:点击下载PDF文件
摘要:我们提出了一种新的方法,分布式ad-hoc麦克风阵列的三维声源定位任务,制定它作为一个集到集的回归问题。通过训练一个多模态掩蔽自动编码器模型,该模型对音频记录和麦克风坐标进行操作,我们证明了这样的公式可以通过重建输入中掩蔽的坐标来精确定位声源。我们的方法是灵活的,在这个意义上说,一个单一的模型可以使用任意数量的麦克风,即使当一个子集的音频记录和麦克风坐标丢失。我们测试我们的方法在室内环境中的音乐和语音的模拟和真实世界的录音,并展示了竞争力的表现相比,经典和其他基于学习的本地化方法。摘要:We present a novel approach to the 3D sound source localization task for distributed ad-hoc microphone arrays by formulating it as a set-to-set regression problem. By training a multi-modal masked autoencoder model that operates on audio recordings and microphone coordinates, we show that such a formulation allows for accurate localization of the sound source, by reconstructing coordinates masked in the input. Our approach is flexible in the sense that a single model can be used with an arbitrary number of microphones, even when a subset of audio recordings and microphone coordinates are missing. We test our method on simulated and real-world recordings of music and speech in indoor environments, and demonstrate competitive performance compared to both classical and other learning based localization methods.
【6】 A Hybrid Approach for Low-Complexity Joint Acoustic Echo and Noise Reduction
标题: 低复杂度联合声学回声和降噪的混合方法
作者:Shrishti Saha Shetu,Naveen Kumar Desiraju,Jose Miguel Martinez Aponte,Emanuël A. P. Habets,Edwin Mabande
备注:5 pages, 2 figures
链接:点击下载PDF文件
摘要:联合执行声学回声和降噪(AENR)任务的基于深度学习的方法通常需要较高的内存和计算资源,因此不适合在嵌入式设备等低资源平台上实时部署。我们提出了一种低复杂度的混合方法,联合AENR采用一个单一的模型来抑制残留的回声和噪声成分。具体来说,我们集成了国家的最先进的(SOTA)ULCNet模型,这是最初提出的,以实现超低复杂度的噪声抑制,在一个混合动力系统和训练它的联合AENR。我们表明,所提出的方法实现了更好的回声降低和可比的降噪性能,具有更低的计算复杂度和内存需求比所有考虑SOTA方法,在语音质量略有下降的成本。摘要:Deep learning-based methods that jointly perform the task of acoustic echo and noise reduction (AENR) often require high memory and computational resources, making them unsuitable for real-time deployment on low-resource platforms such as embedded devices. We propose a low-complexity hybrid approach for joint AENR by employing a single model to suppress both residual echo and noise components. Specifically, we integrate the state-of-the-art (SOTA) ULCNet model, which was originally proposed to achieve ultra-low complexity noise suppression, in a hybrid system and train it for joint AENR. We show that the proposed approach achieves better echo reduction and comparable noise reduction performance with much lower computational complexity and memory requirements than all considered SOTA methods, at the cost of slight degradation in speech quality.
【7】 Spectral Masking with Explicit Time-Context Windowing for Neural Network-Based Monaural Speech Enhancement
标题: 基于神经网络的单耳语音增强的显式时间上下文窗口频谱掩蔽
作者:Luan Vinícius Fiorio,Boris Karanov,Bruno Defraene,Johan David,Wim van Houtum,Frans Widdershoven,Ronald M. Aarts
备注:This work has been submitted to the IEEE for possible publication
链接:点击下载PDF文件
摘要:我们提出并分析了使用显式的时间上下文窗口基于神经网络的频谱掩蔽语音增强,以利用相邻帧之间的信号上下文相关性。特别是,我们专注于软掩蔽和损失计算重建语音的时频表示。我们表明,在神经网络模型的输入和输出的时间上下文窗口函数的应用程序,提高了软掩模的估计过程中,通过结合多个估计从不同的上下文中。所提出的方法仅在推理模式中用作后优化,不需要对神经网络模型进行额外的层或特殊训练。我们的研究结果表明,该方法一致地提高了可懂度和信号质量的去噪语音,证明了两类基于卷积的语音增强模型。重要的是,所提出的方法只需要一个微不足道的($ leq1 %$)增加模型参数的数量,使其适合于硬件受限的应用程序。摘要:We propose and analyze the use of an explicit time-context window for neural network-based spectral masking speech enhancement to leverage signal context dependencies between neighboring frames. In particular, we concentrate on soft masking and loss computed on the time-frequency representation of the reconstructed speech. We show that the application of a time-context windowing function at both input and output of the neural network model improves the soft mask estimation process by combining multiple estimates taken from different contexts. The proposed approach is only applied as post-optimization in inference mode, not requiring additional layers or special training for the neural network model. Our results show that the method consistently increases both intelligibility and signal quality of the denoised speech, as demonstrated for two classes of convolutional-based speech enhancement models. Importantly, the proposed method requires only a negligible ($ leq1 %$) increase in the number of model parameters, making it suitable for hardware-constrained applications.
【8】 Noise-to-mask Ratio Loss for Deep Neural Network based Audio Watermarking
标题: 基于深度神经网络的音频水印的噪屏蔽比损失
作者:Martin Moritz,Toni Olán,Tuomas Virtanen
备注:6 pages, 7 figures
链接:点击下载PDF文件
摘要:数字音频水印技术是在音频信号中以透明的方式嵌入一个信息,可以用来自动识别音频资料和管理版权。我们提出了一种感知损失函数,用于基于深度神经网络的音频水印系统。损失是基于噪声掩蔽比(NMR),这是一个模型的心理声学掩蔽效应特性的人耳。我们使用标记信号和宿主信号之间的NMR损失来训练深度神经模型,并使用PEAQ评估客观质量,使用MUSHRA测试评估主观质量。客观和主观测试都表明,使用NMR损失训练的模型比使用常规MSE损失训练的模型生成更多的透明水印摘要:Digital audio watermarking consists in inserting a message into audio signals in a transparent way and can be used to allow automatic recognition of audio material and management of the copyrights. We propose a perceptual loss function to be used in deep neural network based audio watermarking systems. The loss is based on the noise-to-mask ratio (NMR), which is a model of the psychoacoustic masking effect characteristic of the human ear. We use the NMR loss between marked and host signals to train the deep neural models and we evaluate the objective quality with PEAQ and the subjective quality with a MUSHRA test. Both objective and subjective tests show that models trained with NMR loss generate more transparent watermarks than models trained with the conventionally used MSE loss
【9】 Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
标题: 放下节拍!伴奏条件说唱语音生成的Freestyler
作者:Ziqian Ning,Shuai Wang,Yuepeng Jiang,Jixun Yao,Lei He,Shifeng Pan,Jie Ding,Lei Xie
链接:点击下载PDF文件
摘要:说唱是一种重要的声乐表演形式,但在声乐创作中却没有得到充分的开发。一般的声乐合成依赖于精确的音符和音长输入,要求用户具有相关的音乐知识,这限制了灵活性。相比之下,说唱通常具有更简单的旋律,其核心重点是与伴随的节拍协调的强烈节奏感。在本文中,我们提出了Freestyler,第一个系统,直接从歌词和伴奏输入生成说唱人声。Freestyler利用基于语言模型的令牌生成,然后是条件流匹配模型来生成频谱图和神经声码器来恢复音频。它允许3秒提示启用zero-shot音色控制。由于公开可用的说唱数据集的稀缺性,我们还介绍了RapBank,一个从互联网上收集的说唱歌曲数据集,以及精心设计的处理管道。实验结果表明,Freestyler产生高质量的说唱语音生成与增强的自然和强烈的对齐伴随节拍,无论是风格和节奏。摘要:Rap, a prominent genre of vocal performance, remains underexplored in vocal generation. General vocal synthesis depends on precise note and duration inputs, requiring users to have related musical knowledge, which limits flexibility. In contrast, rap typically features simpler melodies, with a core focus on a strong rhythmic sense that harmonizes with accompanying beats. In this paper, we propose Freestyler, the first system that generates rapping vocals directly from lyrics and accompaniment inputs. Freestyler utilizes language model-based token generation, followed by a conditional flow matching model to produce spectrograms and a neural vocoder to restore audio. It allows a 3-second prompt to enable zero-shot timbre control. Due to the scarcity of publicly available rap datasets, we also present RapBank, a rap song dataset collected from the internet, alongside a meticulously designed processing pipeline. Experimental results show that Freestyler produces high-quality rapping voice generation with enhanced naturalness and strong alignment with accompanying beats, both stylistically and rhythmically.
【10】 Examining the Interplay Between Privacy and Fairness for Speech Processing: A Review and Perspective
标题: 探讨语音处理隐私与公平性之间的相互作用:回顾与展望
作者:Anna Leschanowsky,Sneha Das
链接:点击下载PDF文件
摘要:语音技术已经越来越多地部署在日常生活的各个领域,包括医疗保健和执法等敏感领域。为了使这些技术有效,它们必须对所有用户可靠地工作,同时保护个人隐私。虽然隐私和效用之间的权衡,以及公平和效用,已被广泛研究,在语音处理中的隐私和公平之间的具体相互作用仍然没有得到充分的探索。这篇评论和立场文件概述了整个语音处理机器学习生命周期中新兴的隐私公平权衡。通过借鉴成熟的公平和隐私框架,我们研究了语音处理模型开发过程中共存的现有偏见和隐私伤害来源。然后,我们强调了相应的隐私增强技术如何有可能无意中增加这些偏见,以及偏见缓解策略如何反过来减少隐私。通过提出开放性问题,我们主张对语音技术的隐私公平权衡进行全面评估,并在这一领域开发隐私增强和公平感知算法。摘要:Speech technology has been increasingly deployed in various areas of daily life including sensitive domains such as healthcare and law enforcement. For these technologies to be effective, they must work reliably for all users while preserving individual privacy. Although tradeoffs between privacy and utility, as well as fairness and utility, have been extensively researched, the specific interplay between privacy and fairness in speech processing remains underexplored. This review and position paper offers an overview of emerging privacy-fairness tradeoffs throughout the entire machine learning lifecycle for speech processing. By drawing on well-established frameworks on fairness and privacy, we examine existing biases and sources of privacy harm that coexist during the development of speech processing models. We then highlight how corresponding privacy-enhancing technologies have the potential to inadvertently increase these biases and how bias mitigation strategies may conversely reduce privacy. By raising open questions, we advocate for a comprehensive evaluation of privacy-fairness tradeoffs for speech technology and the development of privacy-enhancing and fairness-aware algorithms in this domain.
【11】 YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection
标题: YOLO-Stutter:端到端区域语音流畅性检测
作者:Xuanru Zhou,Anshul Kashyap,Steve Li,Ayati Sharma,Brittany Morin,David Baquirin,Jet Vonk,Zoe Ezzes,Zachary Miller,Maria Luisa Gorno Tempini,Jiachen Lian,Gopala Krishna Anumanchipalli
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:不流利语音检测是不流利语音分析和口语学习的瓶颈。目前最先进的模型是由基于规则的系统,缺乏效率和鲁棒性,是敏感的模板设计。在本文中,我们提出了YOLO-Stutter:第一个端到端的方法,以时间精确的方式检测不流畅。YOLO-Stutter将不完美的语音-文本对齐作为输入,然后是空间特征聚合器和时间依赖提取器来执行区域边界和类预测。我们还介绍了两个不流利语料库,VCTK-Stutter和VCTK-TTS,模拟自然口语不流利,包括重复,块,丢失,替换和延长。我们的端到端方法在模拟数据和真实失语症语音上以最少数量的可训练参数实现了最先进的性能。代码和数据集在https: github.com rorizzz YOLO-Stutter上开源摘要:Dysfluent speech detection is the bottleneck for disordered speech analysis and spoken language learning. Current state-of-the-art models are governed by rule-based systems which lack efficiency and robustness, and are sensitive to template design. In this paper, we propose YOLO-Stutter: a first end-to-end method that detects dysfluencies in a time-accurate manner. YOLO-Stutter takes imperfect speech-text alignment as input, followed by a spatial feature aggregator, and a temporal dependency extractor to perform region-wise boundary and class predictions. We also introduce two dysfluency corpus, VCTK-Stutter and VCTK-TTS, that simulate natural spoken dysfluencies including repetition, block, missing, replacement, and prolongation. Our end-to-end method achieves state-of-the-art performance with a minimum number of trainable parameters for on both simulated data and real aphasia speech. Code and datasets are open-sourced at https: github.com rorizzz YOLO-Stutter
【12】 Feature Representations for Automatic Meerkat Vocalization Classification
标题: 猫鼬发声自动分类的特征表示
作者:Imen Ben Mahmoud,Eklavya Sarkar,Marta Manser,Mathew Magimai. -Doss
备注:Accepted at Interspeech 2024 satellite event (VIHAR 2024)
链接:点击下载PDF文件
摘要:了解群居动物声音交流的进化是一个重要的研究问题。在这种情况下,除了人类,人们对分析其他社会动物的发声感兴趣,如猫鼬,绒猴,猿。虽然现有的方法解决了某些物种的发声问题,但缺乏针对猫鼬叫声的可靠方法。在这种程度上,本文探讨了自动猫鼬发声分析的特征表示。探索了传统的基于信号处理的表示和由深度学习的进步促进的数据驱动的表示。两个数据集上进行的呼叫类型分类研究表明,人类语音处理的特征提取方法可以有效地用于自动猫鼬呼叫分析。摘要:Understanding evolution of vocal communication in social animals is an important research problem. In that context, beyond humans, there is an interest in analyzing vocalizations of other social animals such as, meerkats, marmosets, apes. While existing approaches address vocalizations of certain species, a reliable method tailored for meerkat calls is lacking. To that extent, this paper investigates feature representations for automatic meerkat vocalization analysis. Both traditional signal processing-based representations and data-driven representations facilitated by advances in deep learning are explored. Call type classification studies conducted on two data sets reveal that feature extraction methods developed for human speech processing can be effectively employed for automatic meerkat call analysis.
【13】 VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
标题: VoxDirect:具有统一多语言编解码器语言建模的表达性人类指令到语音生成
作者:Yixuan Zhou,Xiaoyu Qin,Zeyu Jin,Shuoyi Zhou,Shun Lei,Songtao Zhou,Zhiyong Wu,Jia Jia
备注:Accepted by ACM Multimedia 2024
链接:点击下载PDF文件
摘要:最近的AIGC系统具有基于人类语言指令生成数字多媒体内容的能力,例如文本、图像和视频。然而,当涉及到语音时,与人类语音生成相关的现有方法表现出两个限制。首先,它们需要将输入分为内容提示(抄本)和描述提示(风格和说话人),而不是直接支持人类指令。这种划分在形式上不太自然,与其他AIGC模型不一致。其次,利用独立的描述提示来建模语音风格,而不考虑转录内容的做法,限制了在细粒度级别上控制语音的能力。为了解决这些限制,我们提出了VoxInstruct,一种新的统一的多语言编解码器语言建模框架,将传统的文本到语音的任务扩展到一个一般的人类语音转换任务。我们的方法增强了人类的表达能力的指导下的语音生成,并与其他形式的语音生成范式对齐。为了使该模型能够自动从原始文本指令中提取合成语音的内容,我们引入了语音语义令牌作为中间表示,用于解释到内容的指导。我们还将多个无分类器指导(CFG)策略纳入到我们的编解码器语言模型中,从而增强了遵循人类指令生成的语音。此外,我们的模型架构和训练策略允许同时支持结合语音提示和描述性的人类指令来进行表达性语音合成,这是第一次尝试。代码、模型和演示请访问:https: github.com thuhcsi VoxInstruct。摘要:Recent AIGC systems possess the capability to generate digital multimedia content based on human language instructions, such as text, image and video. However, when it comes to speech, existing methods related to human instruction-to-speech generation exhibit two limitations. Firstly, they require the division of inputs into content prompt (transcript) and description prompt (style and speaker), instead of directly supporting human instruction. This division is less natural in form and does not align with other AIGC models. Secondly, the practice of utilizing an independent description prompt to model speech style, without considering the transcript content, restricts the ability to control speech at a fine-grained level. To address these limitations, we propose VoxInstruct, a novel unified multilingual codec language modeling framework that extends traditional text-to-speech tasks into a general human instruction-to-speech task. Our approach enhances the expressiveness of human instruction-guided speech generation and aligns the speech generation paradigm with other modalities. To enable the model to automatically extract the content of synthesized speech from raw text instructions, we introduce speech semantic tokens as an intermediate representation for instruction-to-content guidance. We also incorporate multiple Classifier-Free Guidance (CFG) strategies into our codec language model, which strengthens the generated speech following human instructions. Furthermore, our model architecture and training strategies allow for the simultaneous support of combining speech prompt and descriptive human instruction for expressive speech synthesis, which is a first-of-its-kind attempt. Codes, models and demos are at: https: github.com thuhcsi VoxInstruct.
【14】 Towards reliable respiratory disease diagnosis based on cough sounds and vision transformers
标题: 基于咳嗽声和视觉转换器实现可靠的呼吸道疾病诊断
作者:Qian Wang,Zhaoyang Bu,Jiaxuan Mao,Wenyu Zhu,Jingya Zhao,Wei Du,Guochao Shi,Min Zhou,Si Chen,Jieming Qu
链接:点击下载PDF文件
摘要:深度学习技术的最新进展引发了各种现实世界应用的性能提升,包括基于多模态医疗数据的疾病诊断。基于咳嗽声数据的呼吸系统疾病(例如,COVID-19和慢性阻塞性肺疾病)诊断也备受关注。然而,现有的作品通常使用传统的机器学习或中等规模的深度模型。另一方面,开发的方法是在小规模数据上训练和评估的,因为很难大规模地管理和注释临床数据。为了解决先前工作中的这些问题,我们创建了一个统一的框架来评估来自轻量级卷积神经网络的各种深度模型(例如,ResNet 18)与现代Vision Transformers进行比较,并比较它们在呼吸系统疾病分类中的性能。基于如此广泛的实证研究的观察结果,我们提出了一种基于大规模咳嗽数据集的自监督和监督学习的基于咳嗽的疾病分类新方法。实验结果表明,我们提出的方法在COVID-19诊断的两个基准数据集和COPD 非COPD分类的专有数据集上始终优于现有技术,AUROC为92.5%。摘要:Recent advancements in deep learning techniques have sparked performance boosts in various real-world applications including disease diagnosis based on multi-modal medical data. Cough sound data-based respiratory disease (e.g., COVID-19 and Chronic Obstructive Pulmonary Disease) diagnosis has also attracted much attention. However, existing works usually utilise traditional machine learning or deep models of moderate scales. On the other hand, the developed approaches are trained and evaluated on small-scale data due to the difficulty of curating and annotating clinical data on scale. To address these issues in prior works, we create a unified framework to evaluate various deep models from lightweight Convolutional Neural Networks (e.g., ResNet18) to modern vision transformers and compare their performance in respiratory disease classification. Based on the observations from such an extensive empirical study, we propose a novel approach to cough-based disease classification based on both self-supervised and supervised learning on a large-scale cough data set. Experimental results demonstrate our proposed approach outperforms prior arts consistently on two benchmark datasets for COVID-19 diagnosis and a proprietary dataset for COPD non-COPD classification with an AUROC of 92.5%.
【15】 Beyond Levenshtein: Leveraging Multiple Algorithms for Robust Word Error Rate Computations And Granular Error Classifications
标题: 超越Levenshtein:利用多种算法进行稳健的误字率计算和粒度错误分类
作者:Korbinian Kuhn,Verena Kersken,Gottfried Zimmermann
备注:Accepted in INTERSPEECH 2024
链接:点击下载PDF文件
摘要:词错误率(WER)是自动语音识别(ASR)准确性的常用度量。转录通常通过替换特定字符来进行预处理,以解决非语义差异。由于这种标准化,标点符号或大写字母的准确性信息丢失。我们提出了一种非破坏性的,基于令牌的方法,使用扩展Levenshtein距离算法来计算一个强大的WER和额外的正交度量。转录错误也被现有的字符串相似性和语音算法更细粒度地分类。几个数据集上的评估表明,我们的方法相比,常见的WER计算的实际等效性。我们还提供了衍生用例的示例性分析,例如标点错误率,以及用于交互式使用和可视化实现的Web应用程序。该代码是开放源代码。摘要:The Word Error Rate (WER) is the common measure of accuracy for Automatic Speech Recognition (ASR). Transcripts are usually pre-processed by substituting specific characters to account for non-semantic differences. As a result of this normalisation, information on the accuracy of punctuation or capitalisation is lost. We present a non-destructive, token-based approach using an extended Levenshtein distance algorithm to compute a robust WER and additional orthographic metrics. Transcription errors are also classified more granularly by existing string similarity and phonetic algorithms. An evaluation on several datasets demonstrates the practical equivalence of our approach compared to common WER computations. We also provide an exemplary analysis of derived use cases, such as a punctuation error rate, and a web application for interactive use and visualisation of our implementation. The code is available open-source.
【16】 Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
标题: Whisper-PMFA:使用Whisper模型进行说话人验证的部分多尺度特征聚集
作者:Yiyang Zhao,Shuai Wang,Guangzhi Sun,Zehua Chen,Chao Zhang,Mingxing Xu,Thomas Fang Zheng
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在本文中,耳语,一个大规模的自动语音识别的预训练模型,提出了适用于说话人确认。提出了一种基于Whisper编码块子集的部分多尺度特征聚合(PMFA)方法,以获得高区分度的说话人信息嵌入,实验结果表明,使用Whisper编码块的中间到后面的块可以保留更多的说话人信息。在VoxCeleb 1和CN-Celeb 1数据集上,我们的系统分别实现了1.42%和8.23%的等错误率(EER),比ECAPA-TDNN基线分别减少了0.58%和1.81%,比ResNet 34基线分别减少了0.46%和0.97%。此外,我们的研究结果表明,使用在多语言数据上训练的Whisper模型可以有效地增强模型在不同语言之间的鲁棒性。最后,低秩自适应方法进行了评估,它减少了约45倍的可训练模型参数,而EER仅略有增加0.2%。摘要:In this paper, Whisper, a large-scale pre-trained model for automatic speech recognition, is proposed to apply to speaker verification. A partial multi-scale feature aggregation (PMFA) approach is proposed based on a subset of Whisper encoder blocks to derive highly discriminative speaker embeddings.Experimental results demonstrate that using the middle to later blocks of the Whisper encoder keeps more speaker information. On the VoxCeleb1 and CN-Celeb1 datasets, our system achieves 1.42% and 8.23% equal error rates (EERs) respectively, receiving 0.58% and 1.81% absolute EER reductions over the ECAPA-TDNN baseline, and 0.46% and 0.97% over the ResNet34 baseline. Furthermore, our results indicate that using Whisper models trained on multilingual data can effectively enhance the model's robustness across languages. Finally, the low-rank adaptation approach is evaluated, which reduces the trainable model parameters by approximately 45 times while only slightly increasing EER by 0.2%.
【17】 EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
标题: MIDI攻击:利用情感语音转换对深度语音分类模型进行语音后门攻击
作者:Wenhan Yao,Zedong XingXiarun Chen,Jia Liu,yongqiang He,Weiping Wen
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:深度语音分类任务,主要包括关键词发现和说话人确认,在基于语音的人机交互中起着至关重要的作用。最近,这些技术的安全性已被证明容易受到后门攻击。具体而言,语音样本受到噪声干扰和当前触发器中的分量修改的攻击。我们认为,语音后门攻击可以战略性地集中在情感上,这是语音中固有的更高层次的主观感知属性。此外,我们提出了情感语音转换技术可以作为语音后门攻击的触发器,这种方法被称为语音后门攻击。在此基础上,对两个语音分类任务进行了攻击实验,实验结果表明,该方法具有较好的触发效果和显著的攻击成功率及准确率差异。此外,消融实验发现,具有强烈情感的语音更适合作为攻击目标。摘要:Deep speech classification tasks, mainly including keyword spotting and speaker verification, play a crucial role in speech-based human-computer interaction. Recently, the security of these technologies has been demonstrated to be vulnerable to backdoor attacks. Specifically speaking, speech samples are attacked by noisy disruption and component modification in present triggers. We suggest that speech backdoor attacks can strategically focus on emotion, a higher-level subjective perceptual attribute inherent in speech. Furthermore, we proposed that emotional voice conversion technology can serve as the speech backdoor attack trigger, and the method is called EmoAttack. Based on this, we conducted attack experiments on two speech classification tasks, showcasing that EmoAttack method owns impactful trigger effectiveness and its remarkable attack success rate and accuracy variance. Additionally, the ablation experiments found that speech with intensive emotion is more suitable to be targeted for attacks.
【18】 SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
标题: SimpleSpeech 2:利用基于流的纯量潜在Transformer扩散模型实现简单有效的文本到语音
作者:Dongchao Yang,Rongjie Huang,Yuanyuan Wang,Haohan Guo,Dading Chong,Songxiang Liu,Xixin Wu,Helen Meng
备注:Submit to TASLP
链接:点击下载PDF文件
摘要:将文语转换为大规模语音是提高合成语音多样性和自然度的有效方法。在高层次上,以前的大规模TTS模型可以分为基于自回归(AR)的( textit{e.g.},VALL-E)或基于非自回归(NAR)的模型(例如,NaturalSpeech 2 3)。虽然这些作品表现出良好的性能,但它们仍然存在潜在的弱点。例如,基于AR的模型存在生成质量不稳定和生成速度慢的问题;同时,一些基于NAR的模型需要音素级的时长对齐信息,从而增加了数据预处理、模型设计和损失设计的复杂性。在这项工作中,我们建立在我们以前的出版物,实现了一个简单而有效的非自回归(NAR)TTS框架,称为SimpleSpeech 2。SimpleSpeech 2有效地结合了自回归(AR)和非自回归(NAR)方法的优势,提供以下关键优势:(1)简化的数据准备;(2)简单的模型和损失设计;(3)稳定、高质量的生成性能,推理速度快。与我们以前的出版物相比,我们提出了({ romannumeral 1})语音标记器和噪声标签对TTS性能的影响的详细分析;({ romannumeral 2})四种不同类型的句子持续时间预测器;({ romannumeral 3})一种新的基于流的标量潜在Transformer扩散模型。有了这些改进,我们表现出显着的改进,在生成性能和生成速度相比,我们以前的工作和其他国家的最先进的(SOTA)大规模TTS模型。此外,我们表明,SimpleSpeech 2可以通过在多语言语音数据集上进行训练来无缝扩展到多语言TTS。演示请访问:{https: dongchaoyang.top SimpleSpeech2 _demo }。摘要:Scaling Text-to-speech (TTS) to large-scale datasets has been demonstrated as an effective method for improving the diversity and naturalness of synthesized speech. At the high level, previous large-scale TTS models can be categorized into either Auto-regressive (AR) based ( textit{e.g.}, VALL-E) or Non-auto-regressive (NAR) based models ( textit{e.g.}, NaturalSpeech 2 3). Although these works demonstrate good performance, they still have potential weaknesses. For instance, AR-based models are plagued by unstable generation quality and slow generation speed; meanwhile, some NAR-based models need phoneme-level duration alignment information, thereby increasing the complexity of data pre-processing, model design, and loss design. In this work, we build upon our previous publication by implementing a simple and efficient non-autoregressive (NAR) TTS framework, termed SimpleSpeech 2. SimpleSpeech 2 effectively combines the strengths of both autoregressive (AR) and non-autoregressive (NAR) methods, offering the following key advantages: (1) simplified data preparation; (2) straightforward model and loss design; and (3) stable, high-quality generation performance with fast inference speed. Compared to our previous publication, we present ({ romannumeral1}) a detailed analysis of the influence of speech tokenizer and noisy label for TTS performance; ({ romannumeral2}) four distinct types of sentence duration predictors; ({ romannumeral3}) a novel flow-based scalar latent transformer diffusion model. With these improvement, we show a significant improvement in generation performance and generation speed compared to our previous work and other state-of-the-art (SOTA) large-scale TTS models. Furthermore, we show that SimpleSpeech 2 can be seamlessly extended to multilingual TTS by training it on multilingual speech datasets. Demos are available on: {https: dongchaoyang.top SimpleSpeech2 _demo }.
机器翻译,仅供参考
