本文经arXiv每日学术速递授权转载
【1】 IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities
标题: IntrinsicVoice:赋予LLM固有的实时语音交互能力
作者: Xin Zhang, Xiang Lyu, Zhihao Du, Qian Chen, Dong Zhang, Hangrui Hu, Chaohong Tan, Tianyu Zhao, Yuxuan Wang, Bin Zhang, Heng Lu, Yaqian Zhou, Xipeng Qiu
链接:点击下载PDF文件
【2】 Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
标题: 不再是全等级:现代语音识别模型的低等级权重训练
作者: Adriana Fernandez-Lopez, Shiwei Liu, Lu Yin, Stavros Petridis, Maja Pantic
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【3】 Audio Explanation Synthesis with Generative Foundation Models
标题: 使用生成性基础模型的音频解释合成
作者: Alican Akman, Qiyang Sun, Björn W. Schuller
链接:点击下载PDF文件
【4】 Toward Robust Real-World Audio Deepfake Detection: Closing the Explainability Gap
标题: 迈向稳健的现实世界音频Deepfake检测:缩小可解释性差距
作者: Georgia Channing, Juil Sock, Ronald Clark, Philip Torr, Christian Schroeder de Witt
链接:点击下载PDF文件
【5】 Advocating Character Error Rate for Multilingual ASR Evaluation
标题: 倡导多语言ASB评估的字符错误率
作者: Thennal D K, Jesin James, Deepa P Gopinath, Muhammed Ashraf K
备注:8 pages
链接:点击下载PDF文件
【6】 Sound Zone Control Robust To Sound Speed Change
标题: 音区控制对音速变化具有鲁棒性
作者: Sankha Subhra Bhattacharjee, Jesper Rindom Jensen, Mads Græsbøll Christensen
备注:5 pages, 4 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【7】 Swin-BERT: A Feature Fusion System designed for Speech-based Alzheimer's Dementia Detection
标题: Swin-BERT:专为基于语音的阿尔茨海默氏症检测而设计的特征融合系统
作者: Yilin Pan, Yanpei Shi, Yijia Zhang, Mingyu Lu
链接:点击下载PDF文件
【8】 HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
标题: HALL-E:用于分钟长Zero-Shot文本到语音合成的分层神经编解码语言模型
作者: Yuto Nishimura, Takumi Hirose, Masanari Ohi, Hideki Nakayama, Nakamasa Inoue
链接:点击下载PDF文件
标题: 用于实时音乐分析的无窗口函数离散傅立叶变换,降低噪音和延迟
作者: Cai Biesinger, Hiromitsu Awano, Masanori Hashimoto
备注:5 pages, 4 figures, Submitted to ICASSP 2025
链接:点击下载PDF文件
【2】 Sound Zone Control Robust To Sound Speed Change
标题: 音区控制对音速变化具有鲁棒性
作者: Sankha Subhra Bhattacharjee, Jesper Rindom Jensen, Mads Græsbøll Christensen
备注:5 pages, 4 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【3】 Robust Fixed-Filter Sound Zone Control with Audio-Based Position Tracking
标题: 具有基于音频的位置跟踪的稳健固定过滤器声区控制
作者: Sankha Subhra Bhattacharjee, Andreas Jonas Fuglsig, Flemming Christensen, Jesper Rindom Jensen, Mads Græsbøll Christensen
备注:Equal contribution by Sankha Subhra Bhattacharjee and Andreas Jonas Fuglsig. Submitted to ICASSP 2025
链接:点击下载PDF文件
【4】 The First VoicePrivacy Attacker Challenge Evaluation Plan
标题: 首个语音隐私攻击者挑战评估计划
作者: Natalia Tomashenko, Xiaoxiao Miao, Emmanuel Vincent, Junichi Yamagishi
链接:点击下载PDF文件
【5】 Learn from Real: Reality Defender's Submission to ASVspoof5 Challenge
标题: 向真实学习:现实捍卫者提交ASVspoof 5挑战
作者: Yi Zhu, Chirag Goel, Surya Koppisetti, Trang Tran, Ankur Kumar, Gaurav Bharaj
备注:Accepted into ASVspoof5 workshop
链接:点击下载PDF文件
【6】 Swin-BERT: A Feature Fusion System designed for Speech-based Alzheimer's Dementia Detection
标题: Swin-BERT:专为基于语音的阿尔茨海默氏症检测而设计的特征融合系统
作者: Yilin Pan, Yanpei Shi, Yijia Zhang, Mingyu Lu
链接:点击下载PDF文件
【7】 Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
标题: 不再是全等级:现代语音识别模型的低等级权重训练
作者: Adriana Fernandez-Lopez, Shiwei Liu, Lu Yin, Stavros Petridis, Maja Pantic
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【8】 Audio Explanation Synthesis with Generative Foundation Models
标题: 使用生成性基础模型的音频解释合成
作者: Alican Akman, Qiyang Sun, Björn W. Schuller
链接:点击下载PDF文件
【9】 Transducer Consistency Regularization for Speech to Text Applications
标题: 语音到文本应用的传感器一致性规范化
作者: Cindy Tseng, Yun Tang, Vijendra Raj Apsingekar
备注:8 pages, 4 figures. Accepted in IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
【10】 Toward Robust Real-World Audio Deepfake Detection: Closing the Explainability Gap
标题: 迈向稳健的现实世界音频Deepfake检测:缩小可解释性差距
作者: Georgia Channing, Juil Sock, Ronald Clark, Philip Torr, Christian Schroeder de Witt
链接:点击下载PDF文件
【11】 Advocating Character Error Rate for Multilingual ASR Evaluation
标题: 倡导多语言ASB评估的字符错误率
作者: Thennal D K, Jesin James, Deepa P Gopinath, Muhammed Ashraf K
备注:8 pages
链接:点击下载PDF文件
标题: IntrinsicVoice:赋予LLM固有的实时语音交互能力
作者: Xin Zhang, Xiang Lyu, Zhihao Du, Qian Chen, Dong Zhang, Hangrui Hu, Chaohong Tan, Tianyu Zhao, Yuxuan Wang, Bin Zhang, Heng Lu, Yaqian Zhou, Xipeng Qiu
链接:点击下载PDF文件
摘要:当前构建具有语音交互能力的LLM的方法严重依赖于在语音响应生成之前或期间的显式文本自回归生成以保持内容质量,这不幸地带来了计算开销并增加了多轮交互中的延迟。为了解决这个问题,我们引入了IntrinsicVoic,这是一个具有内在实时语音交互功能的LLM。IntrinsicVoice旨在通过减轻文本和语音之间的模态差距,促进将预训练LLM的文本功能转移到语音模态。我们的新颖架构GroupFormer可以将语音序列减少到与文本序列相当的长度,同时生成高质量的音频,显著减少语音和文本之间的长度差异,加快推理速度,并缓解长文本建模问题。此外,我们构建了一个名为 method-500 k的多轮语音到语音对话数据集,其中包括近500 k轮语音到语音对话,以及一个跨模态训练策略,以增强语音和文本之间的语义对齐。实验结果表明,IntrinsicVoice在多轮对话场景下能够产生时延小于100 ms的高质量语音响应。演示可在https: instrinsicvoice.github.io 上获得。摘要:Current methods of building LLMs with voice interaction capabilities rely heavily on explicit text autoregressive generation before or during speech response generation to maintain content quality, which unfortunately brings computational overhead and increases latency in multi-turn interactions. To address this, we introduce IntrinsicVoic,e an LLM designed with intrinsic real-time voice interaction capabilities. IntrinsicVoice aims to facilitate the transfer of textual capabilities of pre-trained LLMs to the speech modality by mitigating the modality gap between text and speech. Our novelty architecture, GroupFormer, can reduce speech sequences to lengths comparable to text sequences while generating high-quality audio, significantly reducing the length difference between speech and text, speeding up inference, and alleviating long-text modeling issues. Additionally, we construct a multi-turn speech-to-speech dialogue dataset named method-500k which includes nearly 500k turns of speech-to-speech dialogues, and a cross-modality training strategy to enhance the semantic alignment between speech and text. Experimental results demonstrate that IntrinsicVoice can generate high-quality speech response with latency lower than 100ms in multi-turn dialogue scenarios. Demos are available at https: instrinsicvoice.github.io .
【2】 Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
标题: 不再是全等级:现代语音识别模型的低等级权重训练
作者: Adriana Fernandez-Lopez, Shiwei Liu, Lu Yin, Stavros Petridis, Maja Pantic
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:本文从零开始研究了大规模基于Conformer的语音识别模型的低秩权重训练的不足之处。我们的研究证明了这种训练范式对于此类模型的可行性,并得出了几个值得注意的发现。首先,我们发现将低秩结构专门应用于注意力模块可以出乎意料地提高性能,即使秩降低了12%。相比之下,前馈层提出了更大的挑战,因为它们开始表现出性能下降与适度的50%的秩减少。此外,我们发现初始化和逐层排名分配在成功的低排名训练中发挥着关键作用。具体来说,采用SVD初始化和线性逐层秩映射显著提高了低秩权重训练的效率。基于这些见解,我们引入了低秩语音模型从头开始(LR-SMS),这种方法可以实现与满秩训练的性能等同,同时大幅减少参数计数(至少2倍)和训练时间加速(ASR为1.3倍,AVSR为1.15倍)。摘要:This paper investigates the under-explored area of low-rank weight training for large-scale Conformer-based speech recognition models from scratch. Our study demonstrates the viability of this training paradigm for such models, yielding several notable findings. Firstly, we discover that applying a low-rank structure exclusively to the attention modules can unexpectedly enhance performance, even with a significant rank reduction of 12%. In contrast, feed-forward layers present greater challenges, as they begin to exhibit performance degradation with a moderate 50% rank reduction. Furthermore, we find that both initialization and layer-wise rank assignment play critical roles in successful low-rank training. Specifically, employing SVD initialization and linear layer-wise rank mapping significantly boosts the efficacy of low-rank weight training. Building on these insights, we introduce the Low-Rank Speech Model from Scratch (LR-SMS), an approach that achieves performance parity with full-rank training while delivering substantial reductions in parameters count (by at least 2x), and training time speedups (by 1.3x for ASR and 1.15x for AVSR).
【3】 Audio Explanation Synthesis with Generative Foundation Models
标题: 使用生成性基础模型的音频解释合成
作者: Alican Akman, Qiyang Sun, Björn W. Schuller
链接:点击下载PDF文件
摘要:音频基础模型在各种任务中的日益成功导致了对改进可解释性的日益增长的需求,以更好地理解其复杂的决策过程。现有的方法主要集中在解释这些模型的输入空间内的元素的重要性的基础上,他们对最终决策的影响。在本文中,我们介绍了一种新的音频解释方法,利用音频基础模型的生成能力。我们的方法通过整合已建立的特征属性技术来识别该空间中的重要特征,从而利用这些模型中嵌入空间的内在表示能力。然后,该方法通过优先考虑最重要的特征来生成可收听的音频解释。通过对标准数据集进行严格的基准测试,包括关键字识别和语音情感识别,我们的模型证明了其在生成音频解释方面的有效性。摘要:The increasing success of audio foundation models across various tasks has led to a growing need for improved interpretability to understand their intricate decision-making processes better. Existing methods primarily focus on explaining these models by attributing importance to elements within the input space based on their influence on the final decision. In this paper, we introduce a novel audio explanation method that capitalises on the generative capacity of audio foundation models. Our method leverages the intrinsic representational power of the embedding space within these models by integrating established feature attribution techniques to identify significant features in this space. The method then generates listenable audio explanations by prioritising the most important features. Through rigorous benchmarking against standard datasets, including keyword spotting and speech emotion recognition, our model demonstrates its efficacy in producing audio explanations.
【4】 Toward Robust Real-World Audio Deepfake Detection: Closing the Explainability Gap
标题: 迈向稳健的现实世界音频Deepfake检测:缩小可解释性差距
作者: Georgia Channing, Juil Sock, Ronald Clark, Philip Torr, Christian Schroeder de Witt
链接:点击下载PDF文件
摘要:人工智能操纵或生成的音频deepfake的迅速扩散对媒体完整性和选举安全构成了严重挑战。目前人工智能驱动的检测解决方案缺乏可解释性,在现实世界中表现不佳。在本文中,我们为最先进的基于transformer的音频deepfake检测器引入了新的可解释性方法,并开源了一个用于现实世界泛化的新基准。通过缩小基于transformer的音频deepfake检测器与传统方法之间的可解释性差距,我们的研究结果不仅建立了与人类专家的信任,还为释放公民智能的潜力铺平了道路,以克服音频deepfake检测中的可扩展性问题。摘要:The rapid proliferation of AI-manipulated or generated audio deepfakes poses serious challenges to media integrity and election security. Current AI-driven detection solutions lack explainability and underperform in real-world settings. In this paper, we introduce novel explainability methods for state-of-the-art transformer-based audio deepfake detectors and open-source a novel benchmark for real-world generalizability. By narrowing the explainability gap between transformer-based audio deepfake detectors and traditional methods, our results not only build trust with human experts, but also pave the way for unlocking the potential of citizen intelligence to overcome the scalability issue in audio deepfake detection.
【5】 Advocating Character Error Rate for Multilingual ASR Evaluation
标题: 倡导多语言ASB评估的字符错误率
作者: Thennal D K, Jesin James, Deepa P Gopinath, Muhammed Ashraf K
备注:8 pages
链接:点击下载PDF文件
摘要:自动语音识别(ASR)系统传统上使用英语数据集进行评估,单词错误率(WER)作为主要指标。WER的简单和易于解释有助于其广泛采用,特别是英语。然而,随着ASR系统扩展到多语言环境,WER在各种方面都失败了,特别是对于形态复杂的语言或那些没有明确单词边界的语言。我们的工作文件WER作为评估指标的局限性,并主张字符错误率(CER)作为多语言ASR评估的主要指标。我们表明,CER避免了许多WER面临的挑战,并表现出更大的一致性,跨书写系统。我们支持我们的主张,进行人类评估的ASR transmittance在三种语言:马拉雅拉姆语,英语和阿拉伯语,表现出鲜明的形态特征。我们发现,CER与人类的判断比WER更密切相关,即使是英语。为了促进进一步的研究,我们发布了我们的人类评估数据集,用于未来ASR指标的基准测试。我们的研究结果表明,CER应优先考虑,或至少是补充,在多语言ASR评估,以考虑不同语言的不同语言特点。摘要:Automatic speech recognition (ASR) systems have traditionally been evaluated using English datasets, with the word error rate (WER) serving as the predominant metric. WER's simplicity and ease of interpretation have contributed to its widespread adoption, particularly for English. However, as ASR systems expand to multilingual contexts, WER fails in various ways, particularly with morphologically complex languages or those without clear word boundaries. Our work documents the limitations of WER as an evaluation metric and advocates for the character error rate (CER) as the primary metric in multilingual ASR evaluation. We show that CER avoids many of the challenges WER faces and exhibits greater consistency across writing systems. We support our proposition by conducting human evaluations of ASR transcriptions in three languages: Malayalam, English, and Arabic, which exhibit distinct morphological characteristics. We show that CER correlates more closely with human judgments than WER, even for English. To facilitate further research, we release our human evaluation dataset for future benchmarking of ASR metrics. Our findings suggest that CER should be prioritized, or at least supplemented, in multilingual ASR evaluations to account for the varying linguistic characteristics of different languages.
【6】 Sound Zone Control Robust To Sound Speed Change
标题: 音区控制对音速变化具有鲁棒性
作者: Sankha Subhra Bhattacharjee, Jesper Rindom Jensen, Mads Græsbøll Christensen
备注:5 pages, 4 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:使用静态最优滤波器实现的声区控制(SZC)受到声学环境中的各种扰动的显著影响,其中一个重要的扰动是声速的波动,声速的波动又受到温度和湿度(TH)的变化的影响。出现这个问题是因为控制算法通常使用预先记录的静态脉冲响应(IR)来设计最佳控制滤波器。然而,由于TH变化,IR可能随时间变化,这使得导出的控制滤波器变得非最优。为了解决这一挑战,我们提出了一个简单的模型,称为sinc插值压缩 扩展恢复(SICER),它调整IR来考虑声速降低和增加。使用所提出的技术,在某一TH处测量的IR可以针对任何TH变化进行校正,并且可以重新导出控制滤波器,而无需重新测量新的IR(这在部署SZC时是不切实际的)。我们将建议的SICER IR校正方法与最近推出的可变跨度权衡(VAST)框架SZC,并提出了一个SICER校正VAST方法,是弹性的声速变化。仿真研究表明,所提出的SICER校正VAST方法显着提高声学对比度,并减少在声速变化的存在下的信号失真。摘要:Sound zone control (SZC) implemented using static optimal filters is significantly affected by various perturbations in the acoustic environment, an important one being the fluctuation in the speed of sound, which is in turn influenced by changes in temperature and humidity (TH). This issue arises because control algorithms typically use pre-recorded, static impulse responses (IRs) to design the optimal control filters. The IRs, however, may change with time due to TH changes, which renders the derived control filters to become non-optimal. To address this challenge, we propose a straightforward model called sinc interpolation-compression expansion-resampling (SICER), which adjusts the IRs to account for both sound speed reduction and increase. Using the proposed technique, IRs measured at a certain TH can be corrected for any TH change and control filters can be re-derived without the need of re-measuring the new IRs (which is impractical when SZC is deployed). We integrate the proposed SICER IR correction method with the recently introduced variable span trade-off (VAST) framework for SZC, and propose a SICER-corrected VAST method that is resilient to sound speed variations. Simulation studies show that the proposed SICER-corrected VAST approach significantly improves acoustic contrast and reduces signal distortion in the presence of sound speed changes.
【7】 Swin-BERT: A Feature Fusion System designed for Speech-based Alzheimer's Dementia Detection
标题: Swin-BERT:专为基于语音的阿尔茨海默氏症检测而设计的特征融合系统
作者: Yilin Pan, Yanpei Shi, Yijia Zhang, Mingyu Lu
链接:点击下载PDF文件
摘要:语音通常用于构建自动阿尔茨海默氏症(AD)检测系统,因为AD患者在早期阶段的声学和语言能力显示出下降。然而,言语不仅包括与AD相关的局部和全局信息,还包括与认知状态无关的其他信息,例如年龄和性别。在本文中,我们提出了一个基于语音的系统命名为Swin-BERT自动痴呆症检测。在声学部分,提出了从图像中提取局部和全局信息的移动窗口多头注意,用于设计我们的基于声学的系统。为了解耦年龄和性别对声学特征提取的影响,它们被用作所设计的声学系统的额外输入。对于语言部分,节奏相关的信息,这在患有和没有AD的人之间有很大的差异,在将音频记录转录成转录本时被删除。为了弥补被删除的节奏相关信息,字符级的成绩单被建议作为一个词级的BERT风格的系统的额外输入。最后,Swin-BERT结合了从我们提出的基于声学的系统和基于语言的系统中学习到的声学特征。这些实验基于国际痴呆症检测挑战提供的两个数据集:ADReSS和ADReSSo。结果表明,所提出的声学和语言系统可以更好地或与以前的研究在两个数据集上。在ADReSS和ADReSSo数据集上,Swin-BERT系统分别获得了85.58和87.32的F-score.摘要:Speech is usually used for constructing an automatic Alzheimer's dementia (AD) detection system, as the acoustic and linguistic abilities show a decline in people living with AD at the early stages. However, speech includes not only AD-related local and global information but also other information unrelated to cognitive status, such as age and gender. In this paper, we propose a speech-based system named Swin-BERT for automatic dementia detection. For the acoustic part, the shifted windows multi-head attention that proposed to extract local and global information from images, is used for designing our acoustic-based system. To decouple the effect of age and gender on acoustic feature extraction, they are used as an extra input of the designed acoustic system. For the linguistic part, the rhythm-related information, which varies significantly between people living with and without AD, is removed while transcribing the audio recordings into transcripts. To compensate for the removed rhythm-related information, the character-level transcripts are proposed to be used as the extra input of a word-level BERT-style system. Finally, the Swin-BERT combines the acoustic features learned from our proposed acoustic-based system with our linguistic-based system. The experiments are based on the two datasets provided by the international dementia detection challenges: the ADReSS and ADReSSo. The results show that both the proposed acoustic and linguistic systems can be better or comparable with previous research on the two datasets. Superior results are achieved by the proposed Swin-BERT system on the ADReSS and ADReSSo datasets, which are 85.58 % F-score and 87.32 % F-score respectively.
【8】 HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
标题: HALL-E:用于分钟长Zero-Shot文本到语音合成的分层神经编解码语言模型
作者: Yuto Nishimura, Takumi Hirose, Masanari Ohi, Hideki Nakayama, Nakamasa Inoue
链接:点击下载PDF文件
摘要:最近,基于将自然语言文本翻译成离散音频令牌序列的大语言模型(LLM)的文本到语音(TTS)模型获得了极大的研究关注,其中神经音频编解码器(NAC)模型使用残差矢量量化(RVQ)的进展。然而,由于高帧速率,长形式语音合成仍然是一个重大挑战,这增加了音频令牌的长度,并且使得自回归语言模型难以为甚至一分钟的语音生成音频令牌。为了应对这一挑战,本文介绍了两种新的后训练方法:1)多分辨率重新量化(MReQ)和2)HALL-E。MReQ是一个降低预训练NAC模型帧速率的框架。具体来说,它采用了多分辨率残差矢量量化(MRVQ)模块,分层重组离散的音频令牌,通过师生蒸馏。HALL-E是一种基于LLM的TTS模型,旨在预测MReQ的分层令牌。具体来说,它结合了使用MRVQ子模块的技术,并从预先训练的基于LLM的TTS模型继续训练。此外,为了促进TTS研究,我们创建了MinutesSpeech,这是一个新的基准数据集,由4万小时的过滤语音数据组成,用于训练和评估从3s到180 s的语音合成。在实验中,我们通过将我们的后训练框架应用于VALL-E来证明我们方法的有效性。我们实现了低至8 Hz的帧速率,从而在单个推理步骤中实现了稳定的小时长语音合成。音频样本、数据集、代码和预训练模型可在https: yutonishimura-v2.github.io HALL-E_DEMO 上获得。摘要:Recently, Text-to-speech (TTS) models based on large language models (LLMs) that translate natural language text into sequences of discrete audio tokens have gained great research attention, with advances in neural audio codec (NAC) models using residual vector quantization (RVQ). However, long-form speech synthesis remains a significant challenge due to the high frame rate, which increases the length of audio tokens and makes it difficult for autoregressive language models to generate audio tokens for even a minute of speech. To address this challenge, this paper introduces two novel post-training approaches: 1) Multi-Resolution Requantization (MReQ) and 2) HALL-E. MReQ is a framework to reduce the frame rate of pre-trained NAC models. Specifically, it incorporates multi-resolution residual vector quantization (MRVQ) module that hierarchically reorganizes discrete audio tokens through teacher-student distillation. HALL-E is an LLM-based TTS model designed to predict hierarchical tokens of MReQ. Specifically, it incorporates the technique of using MRVQ sub-modules and continues training from a pre-trained LLM-based TTS model. Furthermore, to promote TTS research, we create MinutesSpeech, a new benchmark dataset consisting of 40k hours of filtered speech data for training and evaluating speech synthesis ranging from 3s up to 180s. In experiments, we demonstrated the effectiveness of our approaches by applying our post-training framework to VALL-E. We achieved the frame rate down to as low as 8 Hz, enabling the stable minitue-long speech synthesis in a single inference step. Audio samples, dataset, codes and pre-trained models are available at https: yutonishimura-v2.github.io HALL-E_DEMO .
eess.AS音频处理
【1】 Window Function-less DFT with Reduced Noise and Latency for Real-Time Music Analysis标题: 用于实时音乐分析的无窗口函数离散傅立叶变换,降低噪音和延迟
作者: Cai Biesinger, Hiromitsu Awano, Masanori Hashimoto
备注:5 pages, 4 figures, Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:音乐分析应用程序需要的算法,可以提供高的时间和频率分辨率,同时最大限度地减少噪声在一个已经嘈杂的信号。实时分析还需要低延迟和低计算要求。我们提出了一种基于DFT的算法,通过扩展后处理DFT输出而不使用窗函数的方法来实现所有这些要求。我们的方法大大降低了旁瓣和噪声,并在不牺牲频率分辨率的情况下提高了时间分辨率。我们使用指数间隔的输出箱,直接映射到音乐中的音符。与现有的基于FFT和DFT的方法相比,所产生的性能改进为改进实时可视化创造了可能性,并有助于提高其他应用(如自动转录)中的分析质量。摘要:Music analysis applications demand algorithms that can provide both high time and frequency resolution while minimizing noise in an already-noisy signal. Real-time analysis additionally demands low latency and low computational requirements. We propose a DFT-based algorithm that accomplishes all these requirements by extending a method that post-processes DFT output without the use of window functions. Our approach yields greatly reduced sidelobes and noise, and improves time resolution without sacrificing frequency resolution. We use exponentially spaced output bins which directly map to notes in music. The resulting improved performance, compared to existing FFT and DFT-based approaches, creates possibilities for improved real-time visualizations, and contributes to improved analysis quality in other applications such as automatic transcription.
【2】 Sound Zone Control Robust To Sound Speed Change
标题: 音区控制对音速变化具有鲁棒性
作者: Sankha Subhra Bhattacharjee, Jesper Rindom Jensen, Mads Græsbøll Christensen
备注:5 pages, 4 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:使用静态最优滤波器实现的声区控制(SZC)受到声学环境中的各种扰动的显著影响,其中一个重要的扰动是声速的波动,声速的波动又受到温度和湿度(TH)的变化的影响。出现这个问题是因为控制算法通常使用预先记录的静态脉冲响应(IR)来设计最佳控制滤波器。然而,由于TH变化,IR可能随时间变化,这使得导出的控制滤波器变得非最优。为了解决这一挑战,我们提出了一个简单的模型,称为sinc插值压缩 扩展恢复(SICER),它调整IR来考虑声速降低和增加。使用所提出的技术,在某一TH处测量的IR可以针对任何TH变化进行校正,并且可以重新导出控制滤波器,而无需重新测量新的IR(这在部署SZC时是不切实际的)。我们将建议的SICER IR校正方法与最近推出的可变跨度权衡(VAST)框架SZC,并提出了一个SICER校正VAST方法,是弹性的声速变化。仿真研究表明,所提出的SICER校正VAST方法显着提高声学对比度,并减少在声速变化的存在下的信号失真。摘要:Sound zone control (SZC) implemented using static optimal filters is significantly affected by various perturbations in the acoustic environment, an important one being the fluctuation in the speed of sound, which is in turn influenced by changes in temperature and humidity (TH). This issue arises because control algorithms typically use pre-recorded, static impulse responses (IRs) to design the optimal control filters. The IRs, however, may change with time due to TH changes, which renders the derived control filters to become non-optimal. To address this challenge, we propose a straightforward model called sinc interpolation-compression expansion-resampling (SICER), which adjusts the IRs to account for both sound speed reduction and increase. Using the proposed technique, IRs measured at a certain TH can be corrected for any TH change and control filters can be re-derived without the need of re-measuring the new IRs (which is impractical when SZC is deployed). We integrate the proposed SICER IR correction method with the recently introduced variable span trade-off (VAST) framework for SZC, and propose a SICER-corrected VAST method that is resilient to sound speed variations. Simulation studies show that the proposed SICER-corrected VAST approach significantly improves acoustic contrast and reduces signal distortion in the presence of sound speed changes.
【3】 Robust Fixed-Filter Sound Zone Control with Audio-Based Position Tracking
标题: 具有基于音频的位置跟踪的稳健固定过滤器声区控制
作者: Sankha Subhra Bhattacharjee, Andreas Jonas Fuglsig, Flemming Christensen, Jesper Rindom Jensen, Mads Græsbøll Christensen
备注:Equal contribution by Sankha Subhra Bhattacharjee and Andreas Jonas Fuglsig. Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在实际场景中部署的声区控制(SZC)系统的性能对收听者的位置高度敏感,并且当收听者移动时会显著降低。本文提出了一个强大的SZC系统,适应动态变化,如移动的听众和不同的区域位置使用基于字典的方法。所提出的系统连续监测环境,并更新固定的控制滤波器,通过跟踪听众的位置,仅使用音频信号。为了测试所提出的SZC方法的有效性,使用实际测量的脉冲响应进行了仿真研究。这些研究表明,SZC,当与所提出的仅音频位置跟踪方案相结合时,当所有收听者位置在字典中可用时,实现最佳性能。此外,即使当不是所有收听者位置都包括在字典中时,与传统的固定滤波器SZC方案相比,该方法仍然提供良好的性能改进。摘要:Performance of sound zone control (SZC) systems deployed in practical scenarios are highly sensitive to the location of the listener(s) and can degrade significantly when listener(s) are moving. This paper presents a robust SZC system that adapts to dynamic changes such as moving listeners and varying zone locations using a dictionary-based approach. The proposed system continuously monitors the environment and updates the fixed control filters by tracking the listener position using audio signals only. To test the effectiveness of the proposed SZC method, simulation studies are carried out using practically measured impulse responses. These studies show that SZC, when incorporated with the proposed audio-only position tracking scheme, achieves optimal performance when all listener positions are available in the dictionary. Moreover, even when not all listener positions are included in the dictionary, the method still provides good performance improvement compared to a traditional fixed filter SZC scheme.
【4】 The First VoicePrivacy Attacker Challenge Evaluation Plan
标题: 首个语音隐私攻击者挑战评估计划
作者: Natalia Tomashenko, Xiaoxiao Miao, Emmanuel Vincent, Junichi Yamagishi
链接:点击下载PDF文件
摘要:第一个语音隐私攻击者挑战赛是作为语音隐私计划的一部分组织的一种新型挑战赛,由ICASSP 2025作为SP大挑战赛提供支持,它专注于开发针对语音匿名化的攻击者系统,将针对提交给语音隐私2024挑战赛的一组匿名化系统进行评估。培训、开发和评估数据集与基线攻击者系统一起提供。参赛者须自行开发以自动说话人确认系统为形式的攻击系统,并将其在开发和评估数据上的得分提交给主办方。为此,他们可以使用任何额外的训练数据和模型,只要它们是公开可用的,并在指定的截止日期之前声明。用于评估的度量是等错误率(EER)。结果将在ICASSP 2025特别会议上公布,届时将邀请5名入选的顶级参与者提交并展示他们的挑战系统。摘要:The First VoicePrivacy Attacker Challenge is a new kind of challenge organized as part of the VoicePrivacy initiative and supported by ICASSP 2025 as the SP Grand Challenge It focuses on developing attacker systems against voice anonymization, which will be evaluated against a set of anonymization systems submitted to the VoicePrivacy 2024 Challenge. Training, development, and evaluation datasets are provided along with a baseline attacker system. Participants shall develop their attacker systems in the form of automatic speaker verification systems and submit their scores on the development and evaluation data to the organizers. To do so, they can use any additional training data and models, provided that they are openly available and declared before the specified deadline. The metric for evaluation is equal error rate (EER). Results will be presented at the ICASSP 2025 special session to which 5 selected top-ranked participants will be invited to submit and present their challenge systems.
【5】 Learn from Real: Reality Defender's Submission to ASVspoof5 Challenge
标题: 向真实学习:现实捍卫者提交ASVspoof 5挑战
作者: Yi Zhu, Chirag Goel, Surya Koppisetti, Trang Tran, Ankur Kumar, Gaurav Bharaj
备注:Accepted into ASVspoof5 workshop
链接:点击下载PDF文件
摘要:音频deepfake检测对于打击恶意使用人工智能合成语音至关重要。在社区所做的许多努力中,ASVspoof挑战已成为评估检测模型通用性和鲁棒性的基准之一。在本文中,我们介绍了Reality Defender对ASVspoof5挑战的提交,强调了一种新的预训练策略,该策略显着提高了泛化能力,同时在训练过程中保持低计算成本。我们的系统SLIM使用自监督对比学习从各种类型的真实语音中学习风格语言学依赖嵌入。学习嵌入有助于区分恶搞和善意的讲话,专注于风格和语言学方面之间的关系。我们在ASVspoof5,ASV2019和野外环境中评估了我们的系统。我们提交的ASVspoof5 Track 1的最小DCF为0.1499,EER为5.5%,ASV2019和野外EER分别为7.4%和10.8%。摘要:Audio deepfake detection is crucial to combat the malicious use of AI-synthesized speech. Among many efforts undertaken by the community, the ASVspoof challenge has become one of the benchmarks to evaluate the generalizability and robustness of detection models. In this paper, we present Reality Defender's submission to the ASVspoof5 challenge, highlighting a novel pretraining strategy which significantly improves generalizability while maintaining low computational cost during training. Our system SLIM learns the style-linguistics dependency embeddings from various types of bonafide speech using self-supervised contrastive learning. The learned embeddings help to discriminate spoof from bonafide speech by focusing on the relationship between the style and linguistics aspects. We evaluated our system on ASVspoof5, ASV2019, and In-the-wild. Our submission achieved minDCF of 0.1499 and EER of 5.5% on ASVspoof5 Track 1, and EER of 7.4% and 10.8% on ASV2019 and In-the-wild respectively.
【6】 Swin-BERT: A Feature Fusion System designed for Speech-based Alzheimer's Dementia Detection
标题: Swin-BERT:专为基于语音的阿尔茨海默氏症检测而设计的特征融合系统
作者: Yilin Pan, Yanpei Shi, Yijia Zhang, Mingyu Lu
链接:点击下载PDF文件
摘要:语音通常用于构建自动阿尔茨海默氏症(AD)检测系统,因为AD患者在早期阶段的声学和语言能力显示出下降。然而,言语不仅包括与AD相关的局部和全局信息,还包括与认知状态无关的其他信息,例如年龄和性别。在本文中,我们提出了一个基于语音的系统命名为Swin-BERT自动痴呆症检测。在声学部分,提出了从图像中提取局部和全局信息的移动窗口多头注意,用于设计我们的基于声学的系统。为了解耦年龄和性别对声学特征提取的影响,它们被用作所设计的声学系统的额外输入。对于语言部分,节奏相关的信息,这在患有和没有AD的人之间有很大的差异,在将音频记录转录成转录本时被删除。为了弥补被删除的节奏相关信息,字符级的成绩单被建议作为一个词级的BERT风格的系统的额外输入。最后,Swin-BERT结合了从我们提出的基于声学的系统和基于语言的系统中学习到的声学特征。这些实验基于国际痴呆症检测挑战提供的两个数据集:ADReSS和ADReSSo。结果表明,所提出的声学和语言系统可以更好地或与以前的研究在两个数据集上。在ADReSS和ADReSSo数据集上,Swin-BERT系统分别获得了85.58和87.32的F-score.摘要:Speech is usually used for constructing an automatic Alzheimer's dementia (AD) detection system, as the acoustic and linguistic abilities show a decline in people living with AD at the early stages. However, speech includes not only AD-related local and global information but also other information unrelated to cognitive status, such as age and gender. In this paper, we propose a speech-based system named Swin-BERT for automatic dementia detection. For the acoustic part, the shifted windows multi-head attention that proposed to extract local and global information from images, is used for designing our acoustic-based system. To decouple the effect of age and gender on acoustic feature extraction, they are used as an extra input of the designed acoustic system. For the linguistic part, the rhythm-related information, which varies significantly between people living with and without AD, is removed while transcribing the audio recordings into transcripts. To compensate for the removed rhythm-related information, the character-level transcripts are proposed to be used as the extra input of a word-level BERT-style system. Finally, the Swin-BERT combines the acoustic features learned from our proposed acoustic-based system with our linguistic-based system. The experiments are based on the two datasets provided by the international dementia detection challenges: the ADReSS and ADReSSo. The results show that both the proposed acoustic and linguistic systems can be better or comparable with previous research on the two datasets. Superior results are achieved by the proposed Swin-BERT system on the ADReSS and ADReSSo datasets, which are 85.58 % F-score and 87.32 % F-score respectively.
【7】 Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
标题: 不再是全等级:现代语音识别模型的低等级权重训练
作者: Adriana Fernandez-Lopez, Shiwei Liu, Lu Yin, Stavros Petridis, Maja Pantic
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:None摘要:This paper investigates the under-explored area of low-rank weight training for large-scale Conformer-based speech recognition models from scratch. Our study demonstrates the viability of this training paradigm for such models, yielding several notable findings. Firstly, we discover that applying a low-rank structure exclusively to the attention modules can unexpectedly enhance performance, even with a significant rank reduction of 12%. In contrast, feed-forward layers present greater challenges, as they begin to exhibit performance degradation with a moderate 50% rank reduction. Furthermore, we find that both initialization and layer-wise rank assignment play critical roles in successful low-rank training. Specifically, employing SVD initialization and linear layer-wise rank mapping significantly boosts the efficacy of low-rank weight training. Building on these insights, we introduce the Low-Rank Speech Model from Scratch (LR-SMS), an approach that achieves performance parity with full-rank training while delivering substantial reductions in parameters count (by at least 2x), and training time speedups (by 1.3x for ASR and 1.15x for AVSR).
【8】 Audio Explanation Synthesis with Generative Foundation Models
标题: 使用生成性基础模型的音频解释合成
作者: Alican Akman, Qiyang Sun, Björn W. Schuller
链接:点击下载PDF文件
摘要:音频基础模型在各种任务中的日益成功导致了对改进可解释性的日益增长的需求,以更好地理解其复杂的决策过程。现有的方法主要集中在解释这些模型的输入空间内的元素的重要性的基础上,他们对最终决策的影响。在本文中,我们介绍了一种新的音频解释方法,利用音频基础模型的生成能力。我们的方法通过整合已建立的特征属性技术来识别该空间中的重要特征,从而利用这些模型中嵌入空间的内在表示能力。然后,该方法通过优先考虑最重要的特征来生成可解释的音频解释。通过对标准数据集进行严格的基准测试,包括关键字识别和语音情感识别,我们的模型证明了它在产生音频解释方面的有效性。摘要:The increasing success of audio foundation models across various tasks has led to a growing need for improved interpretability to understand their intricate decision-making processes better. Existing methods primarily focus on explaining these models by attributing importance to elements within the input space based on their influence on the final decision. In this paper, we introduce a novel audio explanation method that capitalises on the generative capacity of audio foundation models. Our method leverages the intrinsic representational power of the embedding space within these models by integrating established feature attribution techniques to identify significant features in this space. The method then generates listenable audio explanations by prioritising the most important features. Through rigorous benchmarking against standard datasets, including keyword spotting and speech emotion recognition, our model demonstrates its efficacy in producing audio explanations.
【9】 Transducer Consistency Regularization for Speech to Text Applications
标题: 语音到文本应用的传感器一致性规范化
作者: Cindy Tseng, Yun Tang, Vijendra Raj Apsingekar
备注:8 pages, 4 figures. Accepted in IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
摘要:一致性正则化是一种常用的做法,以鼓励模型从扭曲的输入特征中生成一致的表示,并提高模型的泛化能力。该算法对基于交叉熵准则优化的各种语音应用都有显著的改善。然而,这是不简单的应用一致性正则化的基于传感器的方法,这是广泛采用的语音应用程序,由于竞争的性能和流特性。主要的挑战是来自换能器优化标准的巨大对准空间,并且不是空间内的所有对准都同样有助于模型优化。在这项研究中,我们提出了传感器一致性正则化(TCR),传感器模型的一致性正则化方法。我们应用扭曲,如规格增强和辍学,以创建不同的数据视图,并最大限度地减少分布差异。我们利用职业概率对传感器输出分布赋予不同的权重,因此只有接近oracle对齐的对齐才有助于模型学习。实验结果表明,该方法优于其他一致性正则化方法,在 textsc{Librisepeech}数据集上,与强基线相比,可以有效地将单词错误率(WER)降低4.3%.摘要:Consistency regularization is a commonly used practice to encourage the model to generate consistent representation from distorted input features and improve model generalization. It shows significant improvement on various speech applications that are optimized with cross entropy criterion. However, it is not straightforward to apply consistency regularization for the transducer-based approaches, which are widely adopted for speech applications due to the competitive performance and streaming characteristic. The main challenge is from the vast alignment space of the transducer optimization criterion and not all the alignments within the space contribute to the model optimization equally. In this study, we present Transducer Consistency Regularization (TCR), a consistency regularization method for transducer models. We apply distortions such as spec augmentation and dropout to create different data views and minimize the distribution difference. We utilize occupational probabilities to give different weights on transducer output distributions, thus only alignments close to oracle alignments would contribute to the model learning. Our experiments show the proposed method is superior to other consistency regularization implementations and could effectively reduce word error rate (WER) by 4.3 % relatively comparing with a strong baseline on the textsc{Librispeech} dataset.
【10】 Toward Robust Real-World Audio Deepfake Detection: Closing the Explainability Gap
标题: 迈向稳健的现实世界音频Deepfake检测:缩小可解释性差距
作者: Georgia Channing, Juil Sock, Ronald Clark, Philip Torr, Christian Schroeder de Witt
链接:点击下载PDF文件
摘要:人工智能操纵或生成的音频deepfake的迅速扩散对媒体完整性和选举安全构成了严重挑战。目前人工智能驱动的检测解决方案缺乏可解释性,在现实世界中表现不佳。在本文中,我们为最先进的基于transformer的音频deepfake检测器引入了新的可解释性方法,并开源了一个用于现实世界泛化的新基准。通过缩小基于transformer的音频deepfake检测器与传统方法之间的可解释性差距,我们的研究结果不仅建立了与人类专家的信任,还为释放公民智能的潜力铺平了道路,以克服音频deepfake检测中的可扩展性问题。摘要:The rapid proliferation of AI-manipulated or generated audio deepfakes poses serious challenges to media integrity and election security. Current AI-driven detection solutions lack explainability and underperform in real-world settings. In this paper, we introduce novel explainability methods for state-of-the-art transformer-based audio deepfake detectors and open-source a novel benchmark for real-world generalizability. By narrowing the explainability gap between transformer-based audio deepfake detectors and traditional methods, our results not only build trust with human experts, but also pave the way for unlocking the potential of citizen intelligence to overcome the scalability issue in audio deepfake detection.
【11】 Advocating Character Error Rate for Multilingual ASR Evaluation
标题: 倡导多语言ASB评估的字符错误率
作者: Thennal D K, Jesin James, Deepa P Gopinath, Muhammed Ashraf K
备注:8 pages
链接:点击下载PDF文件
摘要:自动语音识别(ASR)系统传统上使用英语数据集进行评估,单词错误率(WER)作为主要指标。WER的简单和易于解释有助于其广泛采用,特别是英语。然而,随着ASR系统扩展到多语言环境,WER在各种方面都失败了,特别是对于形态复杂的语言或那些没有明确单词边界的语言。我们的工作文件WER作为评估指标的局限性,并主张字符错误率(CER)作为多语言ASR评估的主要指标。我们表明,CER避免了许多WER面临的挑战,并表现出更大的一致性,跨书写系统。我们支持我们的主张,进行人类评估的ASR transmittance在三种语言:马拉雅拉姆语,英语和阿拉伯语,表现出鲜明的形态特征。我们发现,CER与人类的判断比WER更密切相关,即使是英语。为了促进进一步的研究,我们发布了我们的人类评估数据集,用于未来ASR指标的基准测试。我们的研究结果表明,CER应优先考虑,或至少是补充,在多语言ASR评估,以考虑不同语言的不同语言特点。摘要:Automatic speech recognition (ASR) systems have traditionally been evaluated using English datasets, with the word error rate (WER) serving as the predominant metric. WER's simplicity and ease of interpretation have contributed to its widespread adoption, particularly for English. However, as ASR systems expand to multilingual contexts, WER fails in various ways, particularly with morphologically complex languages or those without clear word boundaries. Our work documents the limitations of WER as an evaluation metric and advocates for the character error rate (CER) as the primary metric in multilingual ASR evaluation. We show that CER avoids many of the challenges WER faces and exhibits greater consistency across writing systems. We support our proposition by conducting human evaluations of ASR transcriptions in three languages: Malayalam, English, and Arabic, which exhibit distinct morphological characteristics. We show that CER correlates more closely with human judgments than WER, even for English. To facilitate further research, we release our human evaluation dataset for future benchmarking of ASR metrics. Our findings suggest that CER should be prioritized, or at least supplemented, in multilingual ASR evaluations to account for the varying linguistic characteristics of different languages.
机器翻译,仅供参考
![]()
