今日论文合集:cs.SD语音21篇,eess.AS音频处理16篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音
【1】SpeechT: Findings of the First Mentorship in Speech Translation
标题:SpeechT:言语翻译第一任导师的调查结果
链接:https://arxiv.org/abs/2502.12050
作者:Yasmin Moslem,  Juan Julián Cea Morán,  Mariano Gonzalez-Gomez,  Muhammad Hazim Al Farouq,  Farah Abdou,  Satarupa Deb
摘要:这项工作介绍了2024年12月和2025年1月举行的第一次演讲翻译导师计划(SpeechT)的细节和结果。为了满足导师的要求,参与者参与了关键活动,包括数据编制、建模和高级研究。
摘要:This work presents the details and findings of the first mentorship in speechtranslation (SpeechT), which took place in December 2024 and January 2025. Tofulfil the requirements of the mentorship, the participants engaged in keyactivities, including data preparation, modelling, and advanced research.

【2】 Masked Latent Prediction and Classification for Self-Supervised Audio  Representation Learning
标题:自我监督音频表示学习的掩蔽潜在预测和分类
链接:https://arxiv.org/abs/2502.12031
作者:Aurian Quelennec,  Pierre Chouteau,  Geoffroy Peeters,  Slim Essid
备注:Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
摘要:最近,基于掩蔽潜在预测的自监督学习方法已被证明可以将输入数据编码为强大的表示。然而,在训练过程中,学习的潜在空间可以进一步转换,以提取更适合下游分类任务的更高级别的信息。因此,我们提出了一种新的方法:掩蔽潜伏期预测和分类(MATPAC),它是用两个联合解决的借口任务进行训练的。与以前的工作一样,第一个借口任务是一个掩蔽的潜在预测任务,确保在潜在空间中的鲁棒输入表示。第二种是无监督分类,它利用第一个借口任务的潜在表示来匹配教师和学生之间的概率分布。我们通过将其与其他最先进的建议进行比较并进行消融研究来验证MATPAC方法。MATPAC在OpenMIC、GTZAN、ESC-50和US 8 K等参考音频分类数据集上获得了最先进的自监督学习结果,并在Magna-tag-a-tune上的音乐自动标记方面优于可比的监督方法结果。
摘要:Recently, self-supervised learning methods based on masked latent predictionhave proven to encode input data into powerful representations. However, duringtraining, the learned latent space can be further transformed to extracthigher-level information that could be more suited for downstreamclassification tasks. Therefore, we propose a new method: MAsked latenTPrediction And Classification (MATPAC), which is trained with two pretext taskssolved jointly. As in previous work, the first pretext task is a masked latentprediction task, ensuring a robust input representation in the latent space.The second one is unsupervised classification, which utilises the latentrepresentations of the first pretext task to match probability distributionsbetween a teacher and a student. We validate the MATPAC method by comparing itto other state-of-the-art proposals and conducting ablations studies. MATPACreaches state-of-the-art self-supervised learning results on reference audioclassification datasets such as OpenMIC, GTZAN, ESC-50 and US8K and outperformscomparable supervised methods results for musical auto-tagging onMagna-tag-a-tune.

【3】 NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis  with Differential Digital Signal Processing
标题:NaturalL2 S:具有差异数字信号处理的端到端高质量多扬声器唇转语音合成
链接:https://arxiv.org/abs/2502.12002
作者:Yifan Liang,  Fangkun Liu,  Andong Li,  Xiaodong Li,  Chengshi Zheng
摘要:视觉语音识别(VSR)的最新进展促进了唇到语音合成的进展,其中预先训练的VSR模型通过提供有价值的语义信息来增强合成语音的可懂度。级联框架将伪VSR与伪文本到语音(TTS)相结合,或者隐式地利用转录的文本,所取得的成功突出了利用VSR模型的好处。然而,这些方法通常依赖于梅尔频谱图作为中间表示,这可能引入关键瓶颈:从固有的易错唇到语音映射生成的合成梅尔频谱图与用于训练声码器的真实梅尔频谱图之间的域间隙。这种失配不可避免地降低了合成质量。为了弥合这一差距,我们提出了自然唇到语音(NaturalL2S),一个端到端的框架集成声学感应偏差与可区分的语音生成组件。具体来说,我们引入一个基频(F0)预测捕捉合成语音的韵律变化。然后,预测的F0驱动可微分数字信号处理(DDSP)合成器以生成粗信号,该粗信号用作后续语音合成的先验信息。此外,而不是依赖于一个参考扬声器嵌入作为辅助输入,我们的方法实现了令人满意的性能扬声器相似性没有明确建模扬声器特性。客观和主观评价结果表明,NaturalL2S可以有效地提高合成语音的质量相比,国家的最先进的方法。我们的演示页面可以在https://yifan-liang.github.io/NaturalL2S/上访问。
摘要:Recent advancements in visual speech recognition (VSR) have promoted progressin lip-to-speech synthesis, where pre-trained VSR models enhance theintelligibility of synthesized speech by providing valuable semanticinformation. The success achieved by cascade frameworks, which combinepseudo-VSR with pseudo-text-to-speech (TTS) or implicitly utilize thetranscribed text, highlights the benefits of leveraging VSR models. However,these methods typically rely on mel-spectrograms as an intermediaterepresentation, which may introduce a key bottleneck: the domain gap betweensynthetic mel-spectrograms, generated from inherently error-prone lip-to-speechmappings, and real mel-spectrograms used to train vocoders. This mismatchinevitably degrades synthesis quality. To bridge this gap, we propose NaturalLip-to-Speech (NaturalL2S), an end-to-end framework integrating acousticinductive biases with differentiable speech generation components.Specifically, we introduce a fundamental frequency (F0) predictor to captureprosodic variations in synthesized speech. The predicted F0 then drives aDifferentiable Digital Signal Processing (DDSP) synthesizer to generate acoarse signal which serves as prior information for subsequent speechsynthesis. Additionally, instead of relying on a reference speaker embedding asan auxiliary input, our approach achieves satisfactory performance on speakersimilarity without explicitly modelling speaker characteristics. Both objectiveand subjective evaluation results demonstrate that NaturalL2S can effectivelyenhance the quality of the synthesized speech when compared to state-of-the-artmethods. Our demonstration page is accessible athttps://yifan-liang.github.io/NaturalL2S/.

【4】 Step-Audio: Unified Understanding and Generation in Intelligent Speech  Interaction
标题:步进音频:智能语音交互中的统一理解和生成
链接:https://arxiv.org/abs/2502.11946
作者:Ailin Huang,  Boyong Wu,  Bruce Wang,  Chao Yan,  Chen Hu,  Chengli Feng,  Fei Tian,  Feiyu Shen,  Jingbei Li,  Mingrui Chen,  Peng Liu,  Ruihang Miao,  Wang You,  Xi Chen,  Xuerui Yang,  Yechang Huang,  Yuxiang Zhang,  Zheng Gong,  Zixin Zhang,  Brian Li,  Changyi Wan,  Hanpeng Hu,  Ranchen Ming,  Song Yuan,  Xuelin Zhang,  Yu Zhou,  Bingxin Li,  Buyun Ma,  Kang An,  Wei Ji,  Wen Li,  Xuan Wen,  Yuankai Ma,  Yuanwei Liang,  Yun Mou,  Bahtiyar Ahmidi,  Bin Wang,  Bo Li,  Changxin Miao,  Chen Xu,  Chengting Feng,  Chenrun Wang,  Dapeng Shi,  Deshan Sun,  Dingyuan Hu,  Dula Sai,  Enle Liu,  Guanzhe Huang,  Gulin Yan,  Heng Wang,  Haonan Jia,  Haoyang Zhang,  Jiahao Gong,  Jianchang Wu,  Jiahong Liu,  Jianjian Sun,  Jiangjie Zhen,  Jie Feng,  Jie Wu,  Jiaoren Wu,  Jie Yang,  Jinguo Wang,  Jingyang Zhang,  Junzhe Lin,  Kaixiang Li,  Lei Xia,  Li Zhou,  Longlong Gu,  et al. (53 additional authors not shown)
摘要:实时语音交互作为人机协作的基本接口,具有巨大的潜力。然而,当前的开源模型面临语音数据收集成本高、动态控制薄弱、智能化有限等局限性。为了应对这些挑战,本文介绍了Step-Audio,这是第一个生产就绪的开源解决方案。主要贡献包括:1)130 B参数统一语音-文本多模态模型,实现统一理解和生成,Step-Audio-Chat版本开源; 2)生成语音数据引擎,建立可负担的语音克隆框架,通过提炼产生开源轻量级Step-Audio-TTS-3B模型; 3)一个由语音驱动的精细控制系统,可以在方言、情感、歌唱和RAP之间进行动态调整; 4)增强的认知架构,增强了工具调用和角色扮演能力,以有效地管理复杂的任务。基于我们新的StepEval-Audio-360评估基准,Step-Audio在人类评估方面达到了最先进的性能,特别是在指令遵循方面。在LLaMA Question等开源基准测试中,平均性能提高了9.3%,这表明我们致力于推动开源多模态语言技术的发展。我们的代码和型号可在https://github.com/stepfun-ai/Step-Audio上获得。
摘要:Real-time speech interaction, serving as a fundamental interface forhuman-machine collaboration, holds immense potential. However, currentopen-source models face limitations such as high costs in voice datacollection, weakness in dynamic control, and limited intelligence. To addressthese challenges, this paper introduces Step-Audio, the first production-readyopen-source solution. Key contributions include: 1) a 130B-parameter unifiedspeech-text multi-modal model that achieves unified understanding andgeneration, with the Step-Audio-Chat version open-sourced; 2) a generativespeech data engine that establishes an affordable voice cloning framework andproduces the open-sourced lightweight Step-Audio-TTS-3B model throughdistillation; 3) an instruction-driven fine control system enabling dynamicadjustments across dialects, emotions, singing, and RAP; 4) an enhancedcognitive architecture augmented with tool calling and role-playing abilitiesto manage complex tasks effectively. Based on our new StepEval-Audio-360evaluation benchmark, Step-Audio achieves state-of-the-art performance in humanevaluations, especially in terms of instruction following. On open-sourcebenchmarks like LLaMA Question, shows 9.3% average performance improvement,demonstrating our commitment to advancing the development of open-sourcemulti-modal language technologies. Our code and models are available athttps://github.com/stepfun-ai/Step-Audio.

【5】 Rethinking Audio-Visual Adversarial Vulnerability from Temporal and  Modality Perspectives
标题:从时间和形态角度重新思考视听对抗脆弱性
链接:https://arxiv.org/abs/2502.11858
作者:Zeliang Zhang,  Susan Liang,  Daiki Shimada,  Chenliang Xu
备注:Accepted by ICLR 2025
摘要:虽然视听学习通过利用多种感官模式为模型提供了对现实世界更丰富的理解,但这种整合也为对抗性攻击带来了新的漏洞。  在本文中,我们提出了一个全面的研究视听模型的对抗鲁棒性,同时考虑时间和模态特定的漏洞。我们提出了两种强大的对抗性攻击:1)时间不变性攻击,利用连续时间段的固有时间冗余; 2)模态错位攻击,在音频和视觉模态之间引入不一致。这些攻击旨在彻底评估视听模型对各种威胁的鲁棒性。此外,为了抵御这种攻击,我们引入了一种新的视听对抗训练框架。该框架通过结合针对多模态数据和对抗课程策略定制的高效对抗扰动工艺,解决了普通对抗训练中的关键挑战。在Kinetics-Sounds数据集中的大量实验表明,我们提出的基于时间和模态的攻击在降低模型性能方面可以实现最先进的性能,而我们的对抗性训练防御在很大程度上提高了对抗性鲁棒性和对抗性训练效率。
摘要:While audio-visual learning equips models with a richer understanding of thereal world by leveraging multiple sensory modalities, this integration alsointroduces new vulnerabilities to adversarial attacks. In this paper, we present a comprehensive study of the adversarial robustnessof audio-visual models, considering both temporal and modality-specificvulnerabilities. We propose two powerful adversarial attacks: 1) a temporalinvariance attack that exploits the inherent temporal redundancy acrossconsecutive time segments and 2) a modality misalignment attack that introducesincongruence between the audio and visual modalities. These attacks aredesigned to thoroughly assess the robustness of audio-visual models againstdiverse threats. Furthermore, to defend against such attacks, we introduce anovel audio-visual adversarial training framework. This framework addresses keychallenges in vanilla adversarial training by incorporating efficientadversarial perturbation crafting tailored to multi-modal data and anadversarial curriculum strategy. Extensive experiments in the Kinetics-Soundsdataset demonstrate that our proposed temporal and modality-based attacks indegrading model performance can achieve state-of-the-art performance, while ouradversarial training defense largely improves the adversarial robustness aswell as the adversarial training efficiency.

【6】 ChordFormer: A Conformer-Based Architecture for Large-Vocabulary Audio  Chord Recognition
标题:ChordFormer:一种用于大词汇量音频和弦识别的基于协调器的架构
链接:https://arxiv.org/abs/2502.11840
作者:Muhammad Waseem Akram,  Stefano Dettori,  Valentina Colla,  Giorgio Carlo Buttazzo
备注:13 pages, 4 figures
摘要:由于音乐分析中和弦的抽象性和描述性,和弦识别成为音乐信息检索中的一项关键任务。虽然音频和弦识别系统已经实现了对于小词汇表的显著准确性(例如,大/小调和弦),大词汇量和弦识别仍然是一个具有挑战性的问题。这种复杂性也源于和弦固有的长尾分布,其中罕见的和弦类型在大多数数据集中表现不足,导致训练样本不足。有效的和弦识别需要利用来自音频序列的上下文信息,但现有的模型,如卷积神经网络,双向长短期记忆网络和双向Transformers的组合,在捕获长期依赖性方面面临限制,并且在大词汇量和弦识别任务中表现出次优性能。这项工作提出了ChordFormer,一种新的基于一致性的架构,旨在解决结构和弦识别(例如,三和弦、低音、七度)。ChordFormer利用了将卷积神经网络与Transformers集成在一起的conformer块,从而使模型能够有效地捕获局部模式和全局依赖关系。ChordFormer通过重新加权的损失函数和结构化的和弦表示来解决类不平衡等挑战,其性能优于最先进的模型,在大词汇量和弦数据集上实现了帧精度提高2%,类精度提高6%。此外,ChordFormer擅长处理类不平衡,提供跨和弦类型的鲁棒和平衡的识别。这种方法弥合了理论音乐知识和实际应用之间的差距,推进了大词汇量和弦识别领域。
摘要:Chord recognition serves as a critical task in music information retrievaldue to the abstract and descriptive nature of chords in music analysis. Whileaudio chord recognition systems have achieved significant accuracy for smallvocabularies (e.g., major/minor chords), large-vocabulary chord recognitionremains a challenging problem. This complexity also arises from the inherentlong-tail distribution of chords, where rare chord types are underrepresentedin most datasets, leading to insufficient training samples. Effective chordrecognition requires leveraging contextual information from audio sequences,yet existing models, such as combinations of convolutional neural networks,bidirectional long short-term memory networks, and bidirectional transformers,face limitations in capturing long-term dependencies and exhibit suboptimalperformance on large-vocabulary chord recognition tasks. This work proposesChordFormer, a novel conformer-based architecture designed to tackle structuralchord recognition (e.g., triads, bass, sevenths) for large vocabularies.ChordFormer leverages conformer blocks that integrate convolutional neuralnetworks with transformers, thus enabling the model to capture both localpatterns and global dependencies effectively. By addressing challenges such asclass imbalance through a reweighted loss function and structured chordrepresentations, ChordFormer outperforms state-of-the-art models, achieving a2% improvement in frame-wise accuracy and a 6% increase in class-wise accuracyon large-vocabulary chord datasets. Furthermore, ChordFormer excels in handlingclass imbalance, providing robust and balanced recognition across chord types.This approach bridges the gap between theoretical music knowledge and practicalapplications, advancing the field of large-vocabulary chord recognition.

【7】 NablAFx: A Framework for Differentiable Black-box and Gray-box Modeling  of Audio Effects
标题:NablAFx:音效区分黑匣子和灰盒建模的框架
链接:https://arxiv.org/abs/2502.11668
作者:Marco Comunità,  Christian J. Steinmetz,  Joshua D. Reiss
摘要:我们介绍了NablAFx,这是一个开源框架,旨在支持音频效果的可区分黑盒和灰盒建模研究。NablAFx内置于PyTorch中,提供了一个多功能的生态系统来配置、训练、评估和比较各种架构方法。它包括用于管理模型架构、数据集和培训的类,以及用于计算和记录损失、指标和媒体的功能,以及用于促进详细分析的绘图功能。它集成了已建立的黑盒架构和调节方法的实现,以及可微分DSP模块和控制器,从而能够创建参数和非参数灰盒信号链。该代码可在https://github.com/mcomunita/nablafx上访问。
摘要:We present NablAFx, an open-source framework developed to support research indifferentiable black-box and gray-box modeling of audio effects. Built inPyTorch, NablAFx offers a versatile ecosystem to configure, train, evaluate,and compare various architectural approaches. It includes classes to managemodel architectures, datasets, and training, along with features to compute andlog losses, metrics and media, and plotting functions to facilitate detailedanalysis. It incorporates implementations of established black-boxarchitectures and conditioning methods, as well as differentiable DSP blocksand controllers, enabling the creation of both parametric and non-parametricgray-box signal chains. The code is accessible athttps://github.com/mcomunita/nablafx.

【8】 TAPS: Throat and Acoustic Paired Speech Dataset for Deep Learning-Based  Speech Enhancement
标题:TAPS:用于基于深度学习的语音增强的喉咙和声学配对语音数据集
链接:https://arxiv.org/abs/2502.11478
作者:Yunsik Kim,  Yonghun Song,  Yoonyoung Chung
摘要:在工厂、地铁和繁忙街道等高噪声环境中,由于背景噪声,捕获清晰的语音是一项挑战。喉部麦克风提供了一种具有噪声抑制特性的解决方案,可以在录制语音时降低噪声。然而,一个重要的限制仍然存在:高频信息衰减声波通过皮肤和组织,降低语音清晰度。最近的深度学习方法在增强喉部麦克风录音方面表现出了希望,但由于缺乏标准化数据集,进一步的进展受到限制。我们介绍了一个喉咙和声学配对语音数据集(TAPS),收集配对的话语记录从60个母语韩国人使用喉咙和声学麦克风。为了证明TAPS的实用性,我们测试了三种基线深度学习模型,并确定基于映射的方法在提高语音质量和恢复内容方面具有优势。此外,我们提出了一种最佳方法来减轻喉部和声学麦克风之间的信号失配,确保模型性能。这些结果凸显了TAPS作为标准化数据集和推进基于喉麦克风的语音增强研究的潜力。
摘要:In high-noise environments such as factories, subways, and busy streets,capturing clear speech is challenging due to background noise. Throatmicrophones provide a solution with their noise-suppressing properties,reducing the noise while recording speech. However, a significant limitationremains: high-frequency information is attenuated as sound waves pass throughskin and tissue, reducing speech clarity. Recent deep learning approaches haveshown promise in enhancing throat microphone recordings, but further progressis constrained by the absence of standardized dataset. We introduce a throatand acoustic paired speech dataset (TAPS), a collection of paired utterancesrecorded from 60 native Korean speakers using throat and acoustic microphones.To demonstrate the TAPS's utility, we tested three baseline deep learningmodels and identified the mapping-based approach as superior in improvingspeech quality and restoring content. Additionally, we propose an optimalmethod to mitigate the signal mismatch between throat and acoustic microphones,ensuring model performance. These results highlight the potential of TAPS toserve as a standardized dataset and advance research in throat microphone-basedspeech enhancement.

【9】 FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine  Flow Matching
标题:FELLE:具有令牌粗到细流匹配的自回归语音合成
链接:https://arxiv.org/abs/2502.11128
作者:Hui Wang,  Shujie Liu,  Lingwei Meng,  Jinyu Li,  Yifan Yang,  Shiwan Zhao,  Haiyang Sun,  Yanqing Liu,  Haoqin Sun,  Jiaming Zhou,  Yan Lu,  Yong Qin
摘要:为了推进连续值令牌建模和时间一致性执行,我们提出了FELLE,一个自回归模型,集成了语言建模与令牌流匹配。通过利用语言模型的自回归性质和流匹配的生成效率,FELLE有效地预测连续值标记(mel频谱图)。对于每个连续值的令牌,FELLE修改流匹配中的一般先验分布,通过合并来自前一步的信息,提高一致性和稳定性。此外,为了提高合成质量,FELLE引入了一种从粗到细的流匹配机制,以语言模型的输出为条件,分层生成连续值的令牌。实验结果证明了在自回归梅尔频谱图建模中结合流匹配技术的潜力,从而显著改善TTS生成质量,如https://aka.ms/felle中所示。
摘要:To advance continuous-valued token modeling and temporal-coherenceenforcement, we propose FELLE, an autoregressive model that integrates languagemodeling with token-wise flow matching. By leveraging the autoregressive natureof language models and the generative efficacy of flow matching, FELLEeffectively predicts continuous-valued tokens (mel-spectrograms). For eachcontinuous-valued token, FELLE modifies the general prior distribution in flowmatching by incorporating information from the previous step, improvingcoherence and stability. Furthermore, to enhance synthesis quality, FELLEintroduces a coarse-to-fine flow-matching mechanism, generatingcontinuous-valued tokens hierarchically, conditioned on the language model'soutput. Experimental results demonstrate the potential of incorporatingflow-matching techniques in autoregressive mel-spectrogram modeling, leading tosignificant improvements in TTS generation quality, as shown inhttps://aka.ms/felle.

【10】 SyncSpeech: Low-Latency and Efficient Dual-Stream Text-to-Speech based  on Temporal Masked Transformer
标题:同步语音:基于时间屏蔽Transformer的低延迟、高效的双流文本到语音
链接:https://arxiv.org/abs/2502.11094
作者:Zhengyan Sheng,  Zhihao Du,  Shiliang Zhang,  Zhijie Yan,  Yexin Yang,  Zhenhua Ling
摘要:本文提出了一种双流文本到语音(TTS)模型,同步语音,能够接收来自上游模型的流文本输入,同时生成流语音,促进与大型语言模型的无缝交互。SyncSpeech具有以下优点:低延迟,因为它在接收到第二个文本令牌时开始生成流式语音;高效率,因为它在一个步骤中解码与每个到达的文本令牌对应的所有语音令牌。为了实现这一点,我们提出了一个时间掩蔽的Transformer作为同步语音的骨干,结合令牌级的持续时间预测来预测语音令牌和持续时间的下一步。此外,我们设计了一个两阶段的训练策略,以提高训练效率和生成的语音质量。我们在英语和普通话数据集上评估了SyncSpeech。与最近的双流TTS模型相比,SyncSpeech显着降低了语音令牌的第一个数据包延迟并加速了实时因素。此外,在相同的数据规模下,SyncSpeech在语音质量和鲁棒性方面达到了与传统的基于自回归的TTS模型相当的性能。语音示例可在https://SyncSpeech.github.io/}{https://SyncSpeech.github.io/.
摘要:This paper presents a dual-stream text-to-speech (TTS) model, SyncSpeech,capable of receiving streaming text input from upstream models whilesimultaneously generating streaming speech, facilitating seamless interactionwith large language models. SyncSpeech has the following advantages: Lowlatency, as it begins generating streaming speech upon receiving the secondtext token; High efficiency, as it decodes all speech tokens corresponding tothe each arrived text token in one step. To achieve this, we propose a temporalmasked transformer as the backbone of SyncSpeech, combined with token-levelduration prediction to predict speech tokens and the duration for the nextstep. Additionally, we design a two-stage training strategy to improve trainingefficiency and the quality of generated speech. We evaluated the SyncSpeech onboth English and Mandarin datasets. Compared to the recent dual-stream TTSmodels, SyncSpeech significantly reduces the first packet delay of speechtokens and accelerates the real-time factor. Moreover, with the same datascale, SyncSpeech achieves performance comparable to that of traditionalautoregressive-based TTS models in terms of both speech quality and robustness.Speech samples are available athttps://SyncSpeech.github.io/}{https://SyncSpeech.github.io/.

【11】 Hyperdimensional Intelligent Sensing for Efficient Real-Time Audio  Processing on Extreme Edge
标题:超维智能传感在极端边缘实现高效实时音频处理
链接:https://arxiv.org/abs/2502.10718
作者:Sanggeon Yun,  Ryozo Masukawa,  Hanning Chen,  SungHeon Jeong,  Wenjun Huang,  Arghavan Rezvani,  Minhyoung Na,  Yoshiki Yamaguchi,  Mohsen Imani
备注:Accepted to IEEE Access
摘要:管理大量传感器生成的数据的挑战不断升级,特别是在音频应用中,需要创新的解决方案。当前的系统面临着巨大的计算和存储需求,特别是在枪击检测系统(GSDS)等实时应用中,边缘传感器的激增加剧了这些问题。本文提出了一种开创性的方法,采用为智能音频传感框架量身定制的近传感器模型。利用快速傅里叶变换(FFT)模块,卷积神经网络(CNN)层和超维计算(HDC),我们的模型在低能耗,快速推理和在线学习方面表现出色。它非常适合高效的ASIC设计实施,与传统的嵌入式CPU或GPU相比,具有卓越的能效,并且与缩小麦克风传感器尺寸的趋势相兼容。在软件和硬件两个层面的全面评估强调了该模型的有效性。通过详细的ROC曲线分析进行的软件评估揭示了节能和质量损失之间的微妙平衡,实现了高达82.1%的节能,而质量损失仅为1.39%。硬件评估强调了该模型通过ASIC设计实现时值得称赞的能效,特别是使用Google Edge TPU,展示了其优于主流嵌入式CPU和GPU的优势。
摘要:The escalating challenges of managing vast sensor-generated data,particularly in audio applications, necessitate innovative solutions. Currentsystems face significant computational and storage demands, especially inreal-time applications like gunshot detection systems (GSDS), and theproliferation of edge sensors exacerbates these issues. This paper proposes agroundbreaking approach with a near-sensor model tailored for intelligentaudio-sensing frameworks. Utilizing a Fast Fourier Transform (FFT) module,convolutional neural network (CNN) layers, and HyperDimensional Computing(HDC), our model excels in low-energy, rapid inference, and online learning. Itis highly adaptable for efficient ASIC design implementation, offering superiorenergy efficiency compared to conventional embedded CPUs or GPUs, and iscompatible with the trend of shrinking microphone sensor sizes. Comprehensiveevaluations at both software and hardware levels underscore the model'sefficacy. Software assessments through detailed ROC curve analysis revealed adelicate balance between energy conservation and quality loss, achieving up to82.1% energy savings with only 1.39% quality loss. Hardware evaluationshighlight the model's commendable energy efficiency when implemented via ASICdesign, especially with the Google Edge TPU, showcasing its superiority overprevalent embedded CPUs and GPUs.

【12】 Artificial intelligence-enabled detection and assessment of Parkinson's  disease using multimodal data: A survey
标题:使用多模式数据基于人工智能检测和评估帕金森病:一项调查
链接:https://arxiv.org/abs/2502.10703
作者:Aite Zhao,  Yongcan Liu,  Xinglin Yu,  Xinyue Xing
摘要:高度适应性和可重复使用的人工智能(AI)模型的快速出现将彻底改变医学领域,特别是在帕金森病(PD)的诊断和管理方面。目前,没有有效的生物标志物用于诊断PD,评估其严重程度或跟踪其进展。许多AI算法现在被用于PD诊断和治疗,能够基于多模态和异质性疾病症状数据执行各种分类任务,例如PD患者的步态,手部运动和语音模式。它们提供表达性反馈,包括预测PD的潜在可能性,评估单个或多个症状的严重程度,帮助早期检测,以及评估康复和治疗效果,从而展示先进的医学诊断能力。因此,这项工作提供了一个关于通过生物特征症状识别进行PD检测和评估的最新工作的调查汇编,重点是机器学习和深度学习方法,强调它们的好处,暴露它们的弱点,以及它们在开辟新的研究途径方面的影响。此外,它还提供了用于解决相关约束的数据集、方法和架构的分类和特征描述。此外,本文还探讨了数据驱动的人工智能技术在PD诊断中所带来的潜在机遇和挑战。
摘要:The rapid emergence of highly adaptable and reusable artificial intelligence(AI) models is set to revolutionize the medical field, particularly in thediagnosis and management of Parkinson's disease (PD). Currently, there are noeffective biomarkers for diagnosing PD, assessing its severity, or tracking itsprogression. Numerous AI algorithms are now being used for PD diagnosis andtreatment, capable of performing various classification tasks based onmultimodal and heterogeneous disease symptom data, such as gait, handmovements, and speech patterns of PD patients. They provide expressivefeedback, including predicting the potential likelihood of PD, assessing theseverity of individual or multiple symptoms, aiding in early detection, andevaluating rehabilitation and treatment effectiveness, thereby demonstratingadvanced medical diagnostic capabilities. Therefore, this work provides asurveyed compilation of recent works regarding PD detection and assessmentthrough biometric symptom recognition with a focus on machine learning and deeplearning approaches, emphasizing their benefits, and exposing their weaknesses,and their impact in opening up newer research avenues. Additionally, it alsopresents categorized and characterized descriptions of the datasets,approaches, and architectures employed to tackle associated constraints.Furthermore, the paper explores the potential opportunities and challengespresented by data-driven AI technologies in the diagnosis of PD.

【13】 F-StrIPE: Fast Structure-Informed Positional Encoding for Symbolic Music  Generation
标题:F-StrIPE:用于符号音乐生成的快速结构信息位置编码
链接:https://arxiv.org/abs/2502.10491
作者:Manvi Agarwal (IP Paris, LTCI, IDS),  Changhong Wang (LTCI),  Gael Richard (S2A, IDS)
备注:None
摘要:虽然音乐对于像Transformers这样的生成模型仍然是一个具有挑战性的领域,但最近的进展是通过利用合适的音乐信息先验来实现的。利用关于Transformers中的音乐结构的信息的一种技术是将这样的知识插入到位置编码(PE)模块中。然而,Transformers在序列长度上具有二次成本。在本文中,我们提出了F-StrIPE,一个结构通知PE计划,工作在线性复杂度。使用现有的核近似技术的基础上随机功能,我们表明,F-StrIPE是一个推广的随机位置编码(SPE)。我们说明了经验的优点F-StrIPE使用旋律协调的象征性音乐。
摘要:While music remains a challenging domain for generative models likeTransformers, recent progress has been made by exploiting suitablemusically-informed priors. One technique to leverage information about musicalstructure in Transformers is inserting such knowledge into the positionalencoding (PE) module. However, Transformers carry a quadratic cost in sequencelength. In this paper, we propose F-StrIPE, a structure-informed PE scheme thatworks in linear complexity. Using existing kernel approximation techniquesbased on random features, we show that F-StrIPE is a generalization ofStochastic Positional Encoding (SPE). We illustrate the empirical merits ofF-StrIPE using melody harmonization for symbolic music.

【14】 YNote: A Novel Music Notation for Fine-Tuning LLMs in Music Generation
标题:YNote:音乐生成中微调LLM的新颖音乐记法
链接:https://arxiv.org/abs/2502.10467
作者:Shao-Chien Lu,  Chen-Chen Yeh,  Hui-Lin Cho,  Chun-Chieh Hsu,  Tsai-Ling Hsu,  Cheng-Han Wu,  Timothy K. Shih,  Yu-Cheng Lin
摘要:使用大型语言模型(LLM)生成音乐的领域正在迅速发展,但现有的音乐符号系统,如ABC-Notation和MusicXML,仍然过于复杂,无法有效地微调LLM。由于这些格式的可变性和复杂结构,机器和人类都很难解释这些格式。为了应对这些挑战,我们引入了YNote,这是一个简化的音乐符号系统,它只使用四个字符来表示音符及其音高。YNote的固定格式确保了一致性,使其易于阅读,更适合于微调LLM。在我们的实验中,我们在YNote编码的数据集上微调了GPT-2(124 M),并分别获得了0.883和0.766的BLEU和ROUGE分数。只需两个音符作为提示,该模型就能够生成连贯且风格相关的音乐。我们相信YNote为机器学习应用程序提供了现有音乐符号的实用替代方案,并有可能显着提高使用LLM生成音乐的质量。
摘要:The field of music generation using Large Language Models (LLMs) is evolvingrapidly, yet existing music notation systems, such as MIDI, ABC Notation, andMusicXML, remain too complex for effective fine-tuning of LLMs. These formatsare difficult for both machines and humans to interpret due to theirvariability and intricate structure. To address these challenges, we introduceYNote, a simplified music notation system that uses only four characters torepresent a note and its pitch. YNote's fixed format ensures consistency,making it easy to read and more suitable for fine-tuning LLMs. In ourexperiments, we fine-tuned GPT-2 (124M) on a YNote-encoded dataset and achievedBLEU and ROUGE scores of 0.883 and 0.766, respectively. With just two notes asprompts, the model was able to generate coherent and stylistically relevantmusic. We believe YNote offers a practical alternative to existing musicnotations for machine learning applications and has the potential tosignificantly enhance the quality of music generation using LLMs.

【15】 Improving Rare-Word Recognition in Zero-Shot Settings
标题:在Zero-Shot设置中改进稀有词识别
链接:https://arxiv.org/abs/2502.11572
作者:Yash Jogi,  Vaibhav Aggarwal,  Shabari S Nair,  Yash Verma,  Aayush Kubba
备注:Accepted at IEEE SLT 2024
摘要:尽管Whisper接受了68万小时的网络规模音频数据的训练,但它在识别诸如特定领域术语等罕见单词方面面临困难,解决方案是通过提示进行上下文偏置。为了改进这种方法,本文提出了一种监督学习策略来微调Whisper以进行上下文偏置指令。我们证明,通过仅使用670小时的通用语音英语集进行微调,我们的模型可以推广到11个不同的开源英语数据集,在识别稀有单词方面实现了45.6%的改进,在识别微调过程中看不见的单词方面实现了60.8%的改进。令人惊讶的是,我们模型的上下文偏置能力甚至可以在微调过程中推广到看不见的语言。
摘要:Whisper, despite being trained on 680K hours of web-scaled audio data, facesdifficulty in recognising rare words like domain-specific terms, with asolution being contextual biasing through prompting. To improve upon thismethod, in this paper, we propose a supervised learning strategy to fine-tuneWhisper for contextual biasing instruction. We demonstrate that by using only670 hours of Common Voice English set for fine-tuning, our model generalises to11 diverse open-source English datasets, achieving a 45.6% improvement inrecognition of rare words and 60.8% improvement in recognition of words unseenduring fine-tuning over the baseline method. Surprisingly, our model'scontextual biasing ability generalises even to languages unseen duringfine-tuning.

【16】 LMFCA-Net: A Lightweight Model for Multi-Channel Speech Enhancement with  Efficient Narrow-Band and Cross-Band Attention
标题:LMFCA-Net:具有高效窄带和跨带注意力的多通道语音增强轻量级模型
链接:https://arxiv.org/abs/2502.11462
作者:Yaokai Zhang,  Hanchen Pei,  Wanqi Wang,  Gongping Huang
备注:Accepted at ICASSP 2025
摘要:基于深度学习的端到端多通道语音增强方法通过利用子带、跨带和空间信息实现了令人印象深刻的性能。然而,这些方法通常需要大量的计算资源,限制了它们在终端设备上的实用性。本文提出了一种轻量级的具有解耦全连接注意力的多通道语音增强网络(LMFCA-Net)。LMFCA-Net引入了时间轴解耦的全连接注意力(T-FCA)和频率轴解耦的全连接注意力(F-FCA)机制,有效地捕获了长距离窄带和跨带信息,而无需递归单元。实验结果表明,LMFCA-Net的性能与最先进的方法相当,同时显着降低了计算复杂性和延迟,使其成为实际应用的有希望的解决方案。
摘要:Deep learning based end-to-end multi-channel speech enhancement methods haveachieved impressive performance by leveraging sub-band, cross-band, and spatialinformation. However, these methods often demand substantial computationalresources, limiting their practicality on terminal devices. This paper presentsa lightweight multi-channel speech enhancement network with decoupled fullyconnected attention (LMFCA-Net). The proposed LMFCA-Net introduces time-axisdecoupled fully-connected attention (T-FCA) and frequency-axis decoupledfully-connected attention (F-FCA) mechanisms to effectively capture long-rangenarrow-band and cross-band information without recurrent units. Experimentalresults show that LMFCA-Net performs comparably to state-of-the-art methodswhile significantly reducing computational complexity and latency, making it apromising solution for practical applications.

【17】 AudioSpa: Spatializing Sound Events with Text
标题:AudioSpa:用文本空间化声音事件
链接:https://arxiv.org/abs/2502.11219
作者:Linfeng Feng,  Lei Zhao,  Boyu Zhu,  Xiao-Lei Zhang,  Xuelong Li
摘要:文本到音频(TTA)系统最近在从文本合成单声道音频方面表现出很强的性能。然而,从文本生成双耳空间音频的任务,通过结合空间感提供更沉浸的听觉体验,尚未被探索。在这项工作中,我们介绍了文本引导的双耳音频生成。作为早期的努力,我们专注于单声道参考音频的情况下,另外。核心问题是将特定的声音事件与它们的方向相关联,从而创建双耳空间音频。挑战在于文本描述的复杂性和单一源声音事件数据集的有限可用性。为了解决这个问题,我们提出了AudioSpa,这是一个端到端的模型,它应用大型语言模型来处理声学和文本信息。我们采用融合多头注意力(FMHA)来整合文本标记,这增强了多模态学习的生成能力。此外,我们提出了一个双耳源定位模型来评估所生成的音频的质量。最后,我们设计了一个数据增强策略来生成不同的数据集,这使得模型能够在不同的空间位置空间化声音事件。实验结果表明,我们的模型是能够把声音在指定的位置准确。它在定位精度和信号失真方面都取得了有竞争力的性能。我们的演示可在https://linfeng-feng.github.io/AudioSpa-demo上获得。
摘要:Text-to-audio (TTA) systems have recently demonstrated strong performance insynthesizing monaural audio from text. However, the task of generating binauralspatial audio from text, which provides a more immersive auditory experience byincorporating the sense of spatiality, have not been explored yet. In thiswork, we introduce text-guided binaural audio generation. As an early effort,we focus on the scenario where a monaural reference audio is givenadditionally. The core problem is to associate specific sound events with theirdirections, thereby creating binaural spatial audio. The challenge lies in thecomplexity of textual descriptions and the limited availability ofsingle-source sound event datasets. To address this, we propose AudioSpa, anend-to-end model that applies large language models to process both acousticand textual information. We employ fusion multi-head attention (FMHA) tointegrate text tokens, which enhances the generation capability of themultimodal learning. Additionally, we propose a binaural source localizationmodel to assess the quality of the generated audio. Finally, we design a dataaugmentation strategy to generate diverse datasets, which enables the model tospatialize sound events across various spatial positions. Experimental resultsdemonstrate that our model is able to put sounds at the specified locationsaccurately. It achieves competitive performance in both localization accuracyand signal distortion. Our demonstrations are available athttps://linfeng-feng.github.io/AudioSpa-demo.

【18】 Generalizable speech deepfake detection via meta-learned LoRA
标题:通过元学习LoRA进行可推广的语音深度伪造检测
链接:https://arxiv.org/abs/2502.10838
作者:Janne Laakkonen,  Ivan Kukanov,  Ville Hautamäki
备注:9 pages, 2 figures
摘要:可推广的deepfake检测可以用公式表示为一个检测问题,其中标签(真实和虚假)是固定的,但分布漂移会影响deepfake集。我们总是可以用一个选择的攻击和真实数据训练我们的检测器,但是攻击者可以通过用不同的种子重新训练他的生成器来生成新的攻击。一种合理的方法是简单地在训练时间内汇集所有不同的攻击类型。我们提出的方法是将元学习与LoRA适配器结合使用,以学习所有攻击类型所共有的训练数据中的结构。
摘要:Generalizable deepfake detection can be formulated as a detection problemwhere labels (bonafide and fake) are fixed but distributional drift affects thedeepfake set. We can always train our detector with one-selected attacks andbonafide data, but an attacker can generate new attacks by just retraining hisgenerator with a different seed. One reasonable approach is to simply pool alldifferent attack types available in training time. Our proposed approach is toutilize meta-learning in combination with LoRA adapters to learn the structurein the training data that is common to all attack types.

【19】 NeuroAMP: A Novel End-to-end General Purpose Deep Neural Amplifier for  Personalized Hearing Aids
标题:NeuroMP:一种用于个性化助听器的新型端到端通用深度神经放大器
链接:https://arxiv.org/abs/2502.10822
作者:Shafique Ahmed,  Ryandhimas E. Zezario,  Hui-Guan Yuan,  Amir Hussain,  Hsin-Min Wang,  Wei-Ho Chung,  Yu Tsao
摘要:助听器的选配方法然而,由于在传统方法中集成多个模块化组件的复杂性,优化助听器的放大过程仍然具有挑战性。为了应对这一挑战,我们提出了NeuroAMP,这是一种新型的深度神经网络,专为助听器的端到端个性化放大而设计。NeuroAMP利用频谱特征和收听者的听力图作为输入,我们研究了四种架构:卷积神经网络(CNN),长短期记忆(LSTM),卷积递归神经网络(CRNN)和Transformer。我们还介绍了Denoising NeuroAMP,这是一种扩展,它集成了降噪和放大功能,可提高真实场景中的性能。为了增强泛化能力,在对不同语音(TIMIT和TMHINT)和音乐(Cadenza Challenge MUSIC)数据集进行训练期间采用了全面的数据增强策略。使用助听器语音感知指数(HASPI)、助听器语音质量指数(HASQI)和助听器音频质量指数(HAAQI)进行的评估表明,NeuroAMP中的Transformer架构实现了最佳性能,TIMIT上的SRCC得分为0.9927(HASQI)和0.9905(HASPI),Cadenza Challenge MUSIC数据集上的SRCC得分为0.9738(HAAQI)。值得注意的是,我们的数据增强策略在看不见的数据集上保持了高性能(例如,VCTK,MUSDB18-HQ)。此外,去噪NeuroAMP优于传统的NAL-R+WDRC方法和VoiceBank+DEMAND数据集上的两阶段基线,HASPI(0.90)和HASQI(0.59)得分均提高了10%。这些结果强调了NeuroAMP和去噪NeuroAMP在个性化助听器放大方面的显着改善的潜力。
摘要:The prevalence of hearing aids is increasing. However, optimizing theamplification processes of hearing aids remains challenging due to thecomplexity of integrating multiple modular components in traditional methods.To address this challenge, we present NeuroAMP, a novel deep neural networkdesigned for end-to-end, personalized amplification in hearing aids. NeuroAMPleverages both spectral features and the listener's audiogram as inputs, and weinvestigate four architectures: Convolutional Neural Network (CNN), LongShort-Term Memory (LSTM), Convolutional Recurrent Neural Network (CRNN), andTransformer. We also introduce Denoising NeuroAMP, an extension that integratesnoise reduction along with amplification capabilities for improved performancein real-world scenarios. To enhance generalization, a comprehensive dataaugmentation strategy was employed during training on diverse speech (TIMIT andTMHINT) and music (Cadenza Challenge MUSIC) datasets. Evaluation using theHearing Aid Speech Perception Index (HASPI), Hearing Aid Speech Quality Index(HASQI), and Hearing Aid Audio Quality Index (HAAQI) demonstrates that theTransformer architecture within NeuroAMP achieves the best performance, withSRCC scores of 0.9927 (HASQI) and 0.9905 (HASPI) on TIMIT, and 0.9738 (HAAQI)on the Cadenza Challenge MUSIC dataset. Notably, our data augmentation strategymaintains high performance on unseen datasets (e.g., VCTK, MUSDB18-HQ).Furthermore, Denoising NeuroAMP outperforms both the conventional NAL-R+WDRCapproach and a two-stage baseline on the VoiceBank+DEMAND dataset, achieving a10% improvement in both HASPI (0.90) and HASQI (0.59) scores. These resultshighlight the potential of NeuroAMP and Denoising NeuroAMP to deliver notableimprovements in personalized hearing aid amplification.

【20】 Enhancing Age-Related Robustness in Children Speaker Verification
标题:增强儿童说话人验证中与语音相关的鲁棒性
链接:https://arxiv.org/abs/2502.10511
作者:Vishwas M. Shetty,  Jiusi Zheng,  Steven M. Lulich,  Abeer Alwan
备注:Accepted to ICASSP 2025
摘要:儿童说话人确认(C-SV)的主要挑战之一是儿童的声音随着他们的成长而发生显著变化。在本文中,我们提出了两种方法来提高年龄相关的鲁棒性C-SV。我们首先介绍了一个功能转换适配器(FTA)模块,将本地模式集成到更高级别的全球表示,减少过拟合特定的本地功能,提高系统的跨年度SV性能。然后,我们采用合成音频增强(SAA)来增加数据多样性和大小,从而提高对年龄相关变化的鲁棒性。由于缺乏纵向语音数据集,很难衡量与年龄相关的C-SV系统的鲁棒性,我们引入了一个纵向数据集,以评估跨年度验证C-SV系统的鲁棒性。通过整合我们提出的两种方法,平均等误差率分别降低了19.4%,13.0%和6.1%,在一年,两年和三年的差距,跨年度的评估集,与基线相比。
摘要:One of the main challenges in children's speaker verification (C-SV) is thesignificant change in children's voices as they grow. In this paper, we proposetwo approaches to improve age-related robustness in C-SV. We first introduce aFeature Transform Adapter (FTA) module that integrates local patterns intohigher-level global representations, reducing overfitting to specific localfeatures and improving the inter-year SV performance of the system. We thenemploy Synthetic Audio Augmentation (SAA) to increase data diversity and size,thereby improving robustness against age-related changes. Since the lack oflongitudinal speech datasets makes it difficult to measure age-relatedrobustness of C-SV systems, we introduce a longitudinal dataset to assessinter-year verification robustness of C-SV systems. By integrating both of ourproposed methods, the average equal error rate was reduced by 19.4%, 13.0%, and6.1% in the one-year, two-year, and three-year gap inter-year evaluation sets,respectively, compared to the baseline.

【21】 Musical Score Following using Statistical Inference
标题:使用统计推理跟踪乐谱
链接:https://arxiv.org/abs/2502.10426
作者:Josephine Cowley
摘要:乐谱跟随是将演奏实时映射到乐谱中的对应位置。乐谱跟随可用于各种应用,包括自动翻页和实时伴奏。本报告提出了一种新的分数跟踪方法,该方法受到Wilson和Adams 2013年论文的启发,该论文引入了高斯过程(GP)回归的光谱混合(SM)内核。由于SM核是从频域中的高斯混合中导出的,因此它特别适合于对音符的叠加功率谱进行建模,其中能量集中在每个音符的基频的倍数处。我们的乐谱跟随器首先使用GP来统计推断在800个样本的钢琴独奏音乐“音频帧”(~18 ms)期间演奏的音符。然后,这些预测用于持续时间相关的隐马尔可夫模型,以实时预测最有可能的得分位置。我们的两个阶段的方法实现了成功的分数以下不仅对四部分赞美诗安排键盘,但也对小提琴,双簧管和长笛作品。这展示了GP对音乐音频信号进行统计推断的强大而灵活的性质。鉴于这个项目的成功,我们贡献的文献的第一个证明的概念的应用GP在分数以下,更广泛地说,在网上音乐信息检索(MIR)的任务。该项目还提供了一个工作分数追随者产品,使用经过改编的开源用户界面实时呈现分数位置。未来的工作领域包括提高重复音符的准确性,以及在大量使用延音踏板的情况下,适应乐谱的微小偏差,以及模拟多乐器作品。
摘要:Musical score following is the real-time mapping of a performance tocorresponding locations in a musical score. Score following can be used in avariety of applications including automatic page turning and real-timeaccompaniment. This report presents a novel approach for score followingmotivated by Wilson and Adams's 2013 paper, which introduces Spectral Mixture(SM) kernels for Gaussian Process (GP) regression. Since the SM kernel isderived from a Mixture of Gaussians in the frequency domain, it is particularlysuitable for modelling the superposed power spectra of musical notes, in whichenergy is concentrated at multiples of the fundamental frequency of each note.Our score follower begins by using a GP to statistically infer the musicalnotes played during 800-sample 'audioframes' (~18 ms) of solo piano music.These predictions are then used in a duration-dependent Hidden Markov Model topredict the most likely score positions in real time. Our two-stage approachachieves successful score following not only on four-part hymns arranged forkeyboard, but also on pieces for the violin, oboe, and flute. This showcasesthe powerful and flexible nature of GPs for statistical inference on musicalaudio signals. Given the success of this project, we contribute to theliterature a first proof of concept of the application of GPs in scorefollowing, and more broadly, in online Music Information Retrieval (MIR) tasks.This project also contributes a working score follower product that rendersscore position in real time using an adapted open-source user interface. Areasfor future work include improving accuracy on repeated notes and during heavyuse of sustain pedal, adapting to minor deviations from the score, andmodelling multi-instrument works.

eess.AS音频处理

【1】 Improving Rare-Word Recognition in Zero-Shot Settings
标题:在Zero-Shot设置中改进稀有词识别
链接:https://arxiv.org/abs/2502.11572
作者:Yash Jogi,  Vaibhav Aggarwal,  Shabari S Nair,  Yash Verma,  Aayush Kubba
备注:Accepted at IEEE SLT 2024
摘要:尽管Whisper接受了68万小时的网络规模音频数据的训练,但它在识别诸如特定领域术语等罕见单词方面面临困难,解决方案是通过提示进行上下文偏置。为了改进这种方法,在本文中,我们提出了一种监督学习策略来微调耳语上下文偏置指令。我们证明,通过仅使用670小时的通用语音英语集进行微调,我们的模型可以推广到11个不同的开源英语数据集,在识别稀有单词方面实现了45.6%的改进,在识别微调过程中看不见的单词方面实现了60.8%的改进。令人惊讶的是,我们模型的上下文偏置能力甚至可以在微调过程中推广到看不见的语言。
摘要:Whisper, despite being trained on 680K hours of web-scaled audio data, facesdifficulty in recognising rare words like domain-specific terms, with asolution being contextual biasing through prompting. To improve upon thismethod, in this paper, we propose a supervised learning strategy to fine-tuneWhisper for contextual biasing instruction. We demonstrate that by using only670 hours of Common Voice English set for fine-tuning, our model generalises to11 diverse open-source English datasets, achieving a 45.6% improvement inrecognition of rare words and 60.8% improvement in recognition of words unseenduring fine-tuning over the baseline method. Surprisingly, our model'scontextual biasing ability generalises even to languages unseen duringfine-tuning.

【2】 LMFCA-Net: A Lightweight Model for Multi-Channel Speech Enhancement with  Efficient Narrow-Band and Cross-Band Attention
标题:LMFCA-Net:具有高效窄带和跨带注意力的多通道语音增强轻量级模型
链接:https://arxiv.org/abs/2502.11462
作者:Yaokai Zhang,  Hanchen Pei,  Wanqi Wang,  Gongping Huang
备注:Accepted at ICASSP 2025
摘要:基于深度学习的端到端多通道语音增强方法通过利用子带、跨带和空间信息实现了令人印象深刻的性能。然而,这些方法通常需要大量的计算资源,限制了它们在终端设备上的实用性。本文提出了一种轻量级的具有解耦全连接注意力的多通道语音增强网络(LMFCA-Net)。LMFCA-Net引入了时间轴解耦的全连接注意力(T-FCA)和频率轴解耦的全连接注意力(F-FCA)机制,有效地捕获了长距离窄带和跨带信息,而无需递归单元。实验结果表明,LMFCA-Net的性能与最先进的方法相当,同时显着降低了计算复杂性和延迟,使其成为实际应用的有希望的解决方案。
摘要:Deep learning based end-to-end multi-channel speech enhancement methods haveachieved impressive performance by leveraging sub-band, cross-band, and spatialinformation. However, these methods often demand substantial computationalresources, limiting their practicality on terminal devices. This paper presentsa lightweight multi-channel speech enhancement network with decoupled fullyconnected attention (LMFCA-Net). The proposed LMFCA-Net introduces time-axisdecoupled fully-connected attention (T-FCA) and frequency-axis decoupledfully-connected attention (F-FCA) mechanisms to effectively capture long-rangenarrow-band and cross-band information without recurrent units. Experimentalresults show that LMFCA-Net performs comparably to state-of-the-art methodswhile significantly reducing computational complexity and latency, making it apromising solution for practical applications.

【3】 AudioSpa: Spatializing Sound Events with Text
标题:AudioSpa:用文本空间化声音事件
链接:https://arxiv.org/abs/2502.11219
作者:Linfeng Feng,  Lei Zhao,  Boyu Zhu,  Xiao-Lei Zhang,  Xuelong Li
摘要:文本到音频(TTA)系统最近在从文本合成单声道音频方面表现出很强的性能。然而,从文本生成双耳空间音频的任务,通过结合空间感提供更沉浸的听觉体验,尚未被探索。在这项工作中,我们介绍了文本引导的双耳音频生成。作为早期的努力,我们专注于单声道参考音频的情况下,另外。核心问题是将特定的声音事件与它们的方向相关联,从而创建双耳空间音频。挑战在于文本描述的复杂性和单一源声音事件数据集的有限可用性。为了解决这个问题,我们提出了AudioSpa,这是一个端到端的模型,它应用大型语言模型来处理声学和文本信息。我们采用融合多头注意力(FMHA)来整合文本标记,这增强了多模态学习的生成能力。此外,我们提出了一个双耳源定位模型来评估所生成的音频的质量。最后,我们设计了一个数据增强策略来生成不同的数据集,这使得模型能够在不同的空间位置空间化声音事件。实验结果表明,我们的模型是能够把声音在指定的位置准确。它在定位精度和信号失真方面都取得了有竞争力的性能。我们的演示可在https://linfeng-feng.github.io/AudioSpa-demo上获得。
摘要:Text-to-audio (TTA) systems have recently demonstrated strong performance insynthesizing monaural audio from text. However, the task of generating binauralspatial audio from text, which provides a more immersive auditory experience byincorporating the sense of spatiality, have not been explored yet. In thiswork, we introduce text-guided binaural audio generation. As an early effort,we focus on the scenario where a monaural reference audio is givenadditionally. The core problem is to associate specific sound events with theirdirections, thereby creating binaural spatial audio. The challenge lies in thecomplexity of textual descriptions and the limited availability ofsingle-source sound event datasets. To address this, we propose AudioSpa, anend-to-end model that applies large language models to process both acousticand textual information. We employ fusion multi-head attention (FMHA) tointegrate text tokens, which enhances the generation capability of themultimodal learning. Additionally, we propose a binaural source localizationmodel to assess the quality of the generated audio. Finally, we design a dataaugmentation strategy to generate diverse datasets, which enables the model tospatialize sound events across various spatial positions. Experimental resultsdemonstrate that our model is able to put sounds at the specified locationsaccurately. It achieves competitive performance in both localization accuracyand signal distortion. Our demonstrations are available athttps://linfeng-feng.github.io/AudioSpa-demo.

【4】 SpeechT-RAG: Reliable Depression Detection in LLMs with  Retrieval-Augmented Generation Using Speech Timing Information
标题:SpeechT-RAG:使用语音定时信息进行检索增强生成,在LLM中进行可靠的抑郁检测
链接:https://arxiv.org/abs/2502.10950
作者:Xiangyu Zhang,  Hexin Liu,  Qiquan Zhang,  Beena Ahmed,  Julien Epps
摘要:大型语言模型(LLM)越来越多地被用于与健康相关的任务,但当仅仅依赖文本输入时,它们在抑郁症检测中的表现仍然有限。虽然检索增强生成(RAG)通常增强LLM功能,但我们的实验表明,传统的基于文本的RAG系统很难显着提高抑郁症检测的准确性。这一挑战部分源于编码在声学语音模式信息中的丰富的抑郁相关信息,当前仅文本的方法无法有效地捕获这些信息。为了解决这一局限性,我们进行了系统的时间言语模式的分析,比较健康的人与那些经历抑郁症。基于我们的研究结果,我们介绍了基于语音定时的检索增强生成,Speecht-RAG,一个新的系统,利用语音定时功能进行准确的抑郁症检测和可靠的置信度估计。这种集成的方法不仅优于传统的基于文本的RAG系统的检测精度,但也提高了不确定性量化,通过一个信心评分机制,自然延伸从相同的时间特征。我们的统一框架在没有额外培训的情况下实现了与微调LLM相当的结果,同时解决了心理健康评估中准确性和可信度的基本要求。
摘要:Large Language Models (LLMs) have been increasingly adopted forhealth-related tasks, yet their performance in depression detection remainslimited when relying solely on text input. While Retrieval-Augmented Generation(RAG) typically enhances LLM capabilities, our experiments indicate thattraditional text-based RAG systems struggle to significantly improve depressiondetection accuracy. This challenge stems partly from the richdepression-relevant information encoded in acoustic speech patterns informationthat current text-only approaches fail to capture effectively. To address thislimitation, we conduct a systematic analysis of temporal speech patterns,comparing healthy individuals with those experiencing depression. Based on ourfindings, we introduce Speech Timing-based Retrieval-Augmented Generation,SpeechT-RAG, a novel system that leverages speech timing features for bothaccurate depression detection and reliable confidence estimation. Thisintegrated approach not only outperforms traditional text-based RAG systems indetection accuracy but also enhances uncertainty quantification through aconfidence scoring mechanism that naturally extends from the same temporalfeatures. Our unified framework achieves comparable results to fine-tuned LLMswithout additional training while simultaneously addressing the fundamentalrequirements for both accuracy and trustworthiness in mental health assessment.

【5】 Generalizable speech deepfake detection via meta-learned LoRA
标题:通过元学习LoRA进行可推广的语音深度伪造检测
链接:https://arxiv.org/abs/2502.10838
作者:Janne Laakkonen,  Ivan Kukanov,  Ville Hautamäki
备注:9 pages, 2 figures
摘要:可推广的deepfake检测可以用公式表示为一个检测问题,其中标签(真实和虚假)是固定的,但分布漂移会影响deepfake集。我们总是可以用一个选择的攻击和真实数据训练我们的检测器,但是攻击者可以通过用不同的种子重新训练他的生成器来生成新的攻击。一种合理的方法是简单地在训练时间内汇集所有不同的攻击类型。我们提出的方法是将元学习与LoRA适配器结合使用,以学习所有攻击类型所共有的训练数据中的结构。
摘要:Generalizable deepfake detection can be formulated as a detection problemwhere labels (bonafide and fake) are fixed but distributional drift affects thedeepfake set. We can always train our detector with one-selected attacks andbonafide data, but an attacker can generate new attacks by just retraining hisgenerator with a different seed. One reasonable approach is to simply pool alldifferent attack types available in training time. Our proposed approach is toutilize meta-learning in combination with LoRA adapters to learn the structurein the training data that is common to all attack types.

【6】 NeuroAMP: A Novel End-to-end General Purpose Deep Neural Amplifier for  Personalized Hearing Aids
标题:NeuroMP:一种用于个性化助听器的新型端到端通用深度神经放大器
链接:https://arxiv.org/abs/2502.10822
作者:Shafique Ahmed,  Ryandhimas E. Zezario,  Hui-Guan Yuan,  Amir Hussain,  Hsin-Min Wang,  Wei-Ho Chung,  Yu Tsao
摘要:助听器的选配方法然而,由于在传统方法中集成多个模块化组件的复杂性,优化助听器的放大过程仍然具有挑战性。为了应对这一挑战,我们提出了NeuroAMP,这是一种新型的深度神经网络,专为助听器的端到端个性化放大而设计。NeuroAMP利用频谱特征和收听者的听力图作为输入,我们研究了四种架构:卷积神经网络(CNN),长短期记忆(LSTM),卷积递归神经网络(CRNN)和Transformer。我们还介绍了Denoising NeuroAMP,这是一种扩展,它集成了降噪和放大功能,可提高真实场景中的性能。为了增强泛化能力,在对不同语音(TIMIT和TMHINT)和音乐(Cadenza Challenge MUSIC)数据集进行训练期间采用了全面的数据增强策略。使用助听器语音感知指数(HASPI)、助听器语音质量指数(HASQI)和助听器音频质量指数(HAAQI)进行的评估表明,NeuroAMP中的Transformer架构实现了最佳性能,TIMIT上的SRCC得分为0.9927(HASQI)和0.9905(HASPI),Cadenza Challenge MUSIC数据集上的SRCC得分为0.9738(HAAQI)。值得注意的是,我们的数据增强策略在看不见的数据集上保持了高性能(例如,VCTK,MUSDB18-HQ)。此外,去噪NeuroAMP优于传统的NAL-R+WDRC方法和VoiceBank+DEMAND数据集上的两阶段基线,HASPI(0.90)和HASQI(0.59)得分均提高了10%。这些结果强调了NeuroAMP和去噪NeuroAMP在个性化助听器放大方面的显着改善的潜力。
摘要:The prevalence of hearing aids is increasing. However, optimizing theamplification processes of hearing aids remains challenging due to thecomplexity of integrating multiple modular components in traditional methods.To address this challenge, we present NeuroAMP, a novel deep neural networkdesigned for end-to-end, personalized amplification in hearing aids. NeuroAMPleverages both spectral features and the listener's audiogram as inputs, and weinvestigate four architectures: Convolutional Neural Network (CNN), LongShort-Term Memory (LSTM), Convolutional Recurrent Neural Network (CRNN), andTransformer. We also introduce Denoising NeuroAMP, an extension that integratesnoise reduction along with amplification capabilities for improved performancein real-world scenarios. To enhance generalization, a comprehensive dataaugmentation strategy was employed during training on diverse speech (TIMIT andTMHINT) and music (Cadenza Challenge MUSIC) datasets. Evaluation using theHearing Aid Speech Perception Index (HASPI), Hearing Aid Speech Quality Index(HASQI), and Hearing Aid Audio Quality Index (HAAQI) demonstrates that theTransformer architecture within NeuroAMP achieves the best performance, withSRCC scores of 0.9927 (HASQI) and 0.9905 (HASPI) on TIMIT, and 0.9738 (HAAQI)on the Cadenza Challenge MUSIC dataset. Notably, our data augmentation strategymaintains high performance on unseen datasets (e.g., VCTK, MUSDB18-HQ).Furthermore, Denoising NeuroAMP outperforms both the conventional NAL-R+WDRCapproach and a two-stage baseline on the VoiceBank+DEMAND dataset, achieving a10% improvement in both HASPI (0.90) and HASQI (0.59) scores. These resultshighlight the potential of NeuroAMP and Denoising NeuroAMP to deliver notableimprovements in personalized hearing aid amplification.

【7】 Enhancing Age-Related Robustness in Children Speaker Verification
标题:增强儿童说话人验证中与语音相关的鲁棒性
链接:https://arxiv.org/abs/2502.10511
作者:Vishwas M. Shetty,  Jiusi Zheng,  Steven M. Lulich,  Abeer Alwan
备注:Accepted to ICASSP 2025
摘要:儿童说话人确认(C-SV)的主要挑战之一是儿童的声音随着他们的成长而发生显著变化。在本文中,我们提出了两种方法来提高年龄相关的鲁棒性C-SV。我们首先介绍了一个功能转换适配器(FTA)模块,将本地模式集成到更高级别的全球表示,减少过拟合特定的本地功能,提高系统的跨年度SV性能。然后,我们采用合成音频增强(SAA)来增加数据多样性和大小,从而提高对年龄相关变化的鲁棒性。由于缺乏纵向语音数据集,很难衡量与年龄相关的鲁棒性的C-SV系统,我们引入了一个纵向数据集,以评估跨年度的C-SV系统的验证鲁棒性。通过整合我们提出的两种方法,平均等误差率分别降低了19.4%,13.0%和6.1%,在一年,两年和三年的差距,跨年度的评估集,与基线相比。
摘要:One of the main challenges in children's speaker verification (C-SV) is thesignificant change in children's voices as they grow. In this paper, we proposetwo approaches to improve age-related robustness in C-SV. We first introduce aFeature Transform Adapter (FTA) module that integrates local patterns intohigher-level global representations, reducing overfitting to specific localfeatures and improving the inter-year SV performance of the system. We thenemploy Synthetic Audio Augmentation (SAA) to increase data diversity and size,thereby improving robustness against age-related changes. Since the lack oflongitudinal speech datasets makes it difficult to measure age-relatedrobustness of C-SV systems, we introduce a longitudinal dataset to assessinter-year verification robustness of C-SV systems. By integrating both of ourproposed methods, the average equal error rate was reduced by 19.4%, 13.0%, and6.1% in the one-year, two-year, and three-year gap inter-year evaluation sets,respectively, compared to the baseline.

【8】 MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech  Recognition
标题:MoHAVE:用于稳健语音识别的分层视听专家的混合
链接:https://arxiv.org/abs/2502.10447
作者:Sungnyun Kim,  Kangwook Jang,  Sangmin Bae,  Sungwoo Cho,  Se-Young Yun
备注:Preliminary work
摘要:视听语音识别(AVSR)通过整合听觉和视觉模态,已经成为在噪声环境中增强语音识别的关键。然而,现有的AVSR系统难以在不损害计算效率的情况下扩大规模。在这项研究中,我们介绍MoHAVE(混合层次视听专家),一种新的强大的AVSR框架,旨在解决这些可扩展性的限制。通过利用混合专家(MoE)架构,MoHAVE激活特定模态的专家组,确保以最小的计算开销动态适应各种视听输入。MoHAVE的主要贡献包括:(1)稀疏MoE框架,可有效扩展AVSR模型容量;(2)分层选通机制,可基于输入上下文动态利用专家组,增强适应性和鲁棒性;(3)在鲁棒的AVSR基准测试中表现出色,包括LRS 3和MuAViC转录和翻译任务,为可扩展语音识别系统设定了新标准。
摘要:Audio-visual speech recognition (AVSR) has become critical for enhancingspeech recognition in noisy environments by integrating both auditory andvisual modalities. However, existing AVSR systems struggle to scale up withoutcompromising computational efficiency. In this study, we introduce MoHAVE(Mixture of Hierarchical Audio-Visual Experts), a novel robust AVSR frameworkdesigned to address these scalability constraints. By leveraging aMixture-of-Experts (MoE) architecture, MoHAVE activates modality-specificexpert groups, ensuring dynamic adaptation to various audio-visual inputs withminimal computational overhead. Key contributions of MoHAVE include: (1) asparse MoE framework that efficiently scales AVSR model capacity, (2) ahierarchical gating mechanism that dynamically utilizes the expert groups basedon input context, enhancing adaptability and robustness, and (3) remarkableperformance across robust AVSR benchmarks, including LRS3 and MuAViCtranscription and translation tasks, setting a new standard for scalable speechrecognition systems.

【9】 Musical Score Following using Statistical Inference
标题:使用统计推理跟踪乐谱
链接:https://arxiv.org/abs/2502.10426
作者:Josephine Cowley
摘要:乐谱跟随是将表演实时映射到乐谱中的相应位置。乐谱跟随可用于各种应用,包括自动翻页和实时伴奏。本报告提出了一种新的分数跟踪方法,该方法受到Wilson和Adams 2013年论文的启发,该论文引入了高斯过程(GP)回归的光谱混合(SM)内核。由于SM核是从频域中的高斯混合中导出的,因此它特别适合于对音符的叠加功率谱进行建模,其中能量集中在每个音符的基频的倍数处。我们的乐谱跟随器首先使用GP来统计推断在800个样本的钢琴独奏音乐“音频帧”(~18 ms)期间演奏的音符。然后,这些预测用于持续时间相关的隐马尔可夫模型,以实时预测最有可能的得分位置。我们的两个阶段的方法实现了成功的分数以下不仅对四部分赞美诗安排键盘,但也对小提琴,双簧管和长笛作品。这展示了GP对音乐音频信号进行统计推断的强大而灵活的性质。鉴于这个项目的成功,我们贡献的文献的第一个证明的概念的应用GP在分数以下,更广泛地说,在网上音乐信息检索(MIR)的任务。该项目还提供了一个工作分数追随者产品,使用经过改编的开源用户界面实时呈现分数位置。未来的工作领域包括提高重复音符和大量使用延音踏板的准确性,适应乐谱的微小偏差,以及模拟多乐器作品。
摘要:Musical score following is the real-time mapping of a performance tocorresponding locations in a musical score. Score following can be used in avariety of applications including automatic page turning and real-timeaccompaniment. This report presents a novel approach for score followingmotivated by Wilson and Adams's 2013 paper, which introduces Spectral Mixture(SM) kernels for Gaussian Process (GP) regression. Since the SM kernel isderived from a Mixture of Gaussians in the frequency domain, it is particularlysuitable for modelling the superposed power spectra of musical notes, in whichenergy is concentrated at multiples of the fundamental frequency of each note.Our score follower begins by using a GP to statistically infer the musicalnotes played during 800-sample 'audioframes' (~18 ms) of solo piano music.These predictions are then used in a duration-dependent Hidden Markov Model topredict the most likely score positions in real time. Our two-stage approachachieves successful score following not only on four-part hymns arranged forkeyboard, but also on pieces for the violin, oboe, and flute. This showcasesthe powerful and flexible nature of GPs for statistical inference on musicalaudio signals. Given the success of this project, we contribute to theliterature a first proof of concept of the application of GPs in scorefollowing, and more broadly, in online Music Information Retrieval (MIR) tasks.This project also contributes a working score follower product that rendersscore position in real time using an adapted open-source user interface. Areasfor future work include improving accuracy on repeated notes and during heavyuse of sustain pedal, adapting to minor deviations from the score, andmodelling multi-instrument works.

【10】 NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis  with Differential Digital Signal Processing
标题:NaturalL2 S:具有差异数字信号处理的端到端高质量多扬声器唇转语音合成
链接:https://arxiv.org/abs/2502.12002
作者:Yifan Liang,  Fangkun Liu,  Andong Li,  Xiaodong Li,  Chengshi Zheng
摘要:视觉语音识别(VSR)的最新进展促进了唇到语音合成的进展,其中预先训练的VSR模型通过提供有价值的语义信息来增强合成语音的可懂度。级联框架将伪VSR与伪文本到语音(TTS)相结合,或者隐式地利用转录的文本,所取得的成功突出了利用VSR模型的好处。然而,这些方法通常依赖于梅尔频谱图作为中间表示,这可能引入关键瓶颈:从固有的易错唇到语音映射生成的合成梅尔频谱图与用于训练声码器的真实梅尔频谱图之间的域间隙。这种失配不可避免地降低了合成质量。为了弥合这一差距,我们提出了自然唇到语音(NaturalL2S),一个端到端的框架集成声学感应偏差与可区分的语音生成组件。具体来说,我们引入一个基频(F0)预测捕捉合成语音的韵律变化。然后,预测的F0驱动可微分数字信号处理(DDSP)合成器以生成粗信号,该粗信号用作后续语音合成的先验信息。此外,而不是依赖于一个参考扬声器嵌入作为辅助输入,我们的方法实现了令人满意的性能扬声器相似性没有明确建模扬声器特性。客观和主观评价结果表明,NaturalL2S可以有效地提高合成语音的质量相比,国家的最先进的方法。我们的演示页面可以在https://yifan-liang.github.io/NaturalL2S/上访问。
摘要:Recent advancements in visual speech recognition (VSR) have promoted progressin lip-to-speech synthesis, where pre-trained VSR models enhance theintelligibility of synthesized speech by providing valuable semanticinformation. The success achieved by cascade frameworks, which combinepseudo-VSR with pseudo-text-to-speech (TTS) or implicitly utilize thetranscribed text, highlights the benefits of leveraging VSR models. However,these methods typically rely on mel-spectrograms as an intermediaterepresentation, which may introduce a key bottleneck: the domain gap betweensynthetic mel-spectrograms, generated from inherently error-prone lip-to-speechmappings, and real mel-spectrograms used to train vocoders. This mismatchinevitably degrades synthesis quality. To bridge this gap, we propose NaturalLip-to-Speech (NaturalL2S), an end-to-end framework integrating acousticinductive biases with differentiable speech generation components.Specifically, we introduce a fundamental frequency (F0) predictor to captureprosodic variations in synthesized speech. The predicted F0 then drives aDifferentiable Digital Signal Processing (DDSP) synthesizer to generate acoarse signal which serves as prior information for subsequent speechsynthesis. Additionally, instead of relying on a reference speaker embedding asan auxiliary input, our approach achieves satisfactory performance on speakersimilarity without explicitly modelling speaker characteristics. Both objectiveand subjective evaluation results demonstrate that NaturalL2S can effectivelyenhance the quality of the synthesized speech when compared to state-of-the-artmethods. Our demonstration page is accessible athttps://yifan-liang.github.io/NaturalL2S/.

【11】 Step-Audio: Unified Understanding and Generation in Intelligent Speech  Interaction
标题:步进音频:智能语音交互中的统一理解和生成
链接:https://arxiv.org/abs/2502.11946
作者:Ailin Huang,  Boyong Wu,  Bruce Wang,  Chao Yan,  Chen Hu,  Chengli Feng,  Fei Tian,  Feiyu Shen,  Jingbei Li,  Mingrui Chen,  Peng Liu,  Ruihang Miao,  Wang You,  Xi Chen,  Xuerui Yang,  Yechang Huang,  Yuxiang Zhang,  Zheng Gong,  Zixin Zhang,  Brian Li,  Changyi Wan,  Hanpeng Hu,  Ranchen Ming,  Song Yuan,  Xuelin Zhang,  Yu Zhou,  Bingxin Li,  Buyun Ma,  Kang An,  Wei Ji,  Wen Li,  Xuan Wen,  Yuankai Ma,  Yuanwei Liang,  Yun Mou,  Bahtiyar Ahmidi,  Bin Wang,  Bo Li,  Changxin Miao,  Chen Xu,  Chengting Feng,  Chenrun Wang,  Dapeng Shi,  Deshan Sun,  Dingyuan Hu,  Dula Sai,  Enle Liu,  Guanzhe Huang,  Gulin Yan,  Heng Wang,  Haonan Jia,  Haoyang Zhang,  Jiahao Gong,  Jianchang Wu,  Jiahong Liu,  Jianjian Sun,  Jiangjie Zhen,  Jie Feng,  Jie Wu,  Jiaoren Wu,  Jie Yang,  Jinguo Wang,  Jingyang Zhang,  Junzhe Lin,  Kaixiang Li,  Lei Xia,  Li Zhou,  Longlong Gu,  et al. (53 additional authors not shown)
摘要:实时语音交互作为人机协作的基本接口,具有巨大的潜力。然而,当前的开源模型面临语音数据收集成本高、动态控制薄弱、智能化有限等局限性。为了应对这些挑战,本文介绍了Step-Audio,这是第一个生产就绪的开源解决方案。主要贡献包括:1)130 B参数统一语音-文本多模态模型,实现统一理解和生成,Step-Audio-Chat版本开源; 2)生成语音数据引擎,建立可负担的语音克隆框架,通过提炼产生开源轻量级Step-Audio-TTS-3B模型; 3)一个由语音驱动的精细控制系统,可以在方言、情感、歌唱和RAP之间进行动态调整; 4)增强的认知架构,增强了工具调用和角色扮演能力,以有效地管理复杂的任务。基于我们新的StepEval-Audio-360评估基准,Step-Audio在人类评估方面达到了最先进的性能,特别是在指令遵循方面。在LLaMA Question等开源基准测试中,平均性能提高了9.3%,这表明我们致力于推动开源多模态语言技术的发展。我们的代码和型号可在https://github.com/stepfun-ai/Step-Audio上获得。
摘要:Real-time speech interaction, serving as a fundamental interface forhuman-machine collaboration, holds immense potential. However, currentopen-source models face limitations such as high costs in voice datacollection, weakness in dynamic control, and limited intelligence. To addressthese challenges, this paper introduces Step-Audio, the first production-readyopen-source solution. Key contributions include: 1) a 130B-parameter unifiedspeech-text multi-modal model that achieves unified understanding andgeneration, with the Step-Audio-Chat version open-sourced; 2) a generativespeech data engine that establishes an affordable voice cloning framework andproduces the open-sourced lightweight Step-Audio-TTS-3B model throughdistillation; 3) an instruction-driven fine control system enabling dynamicadjustments across dialects, emotions, singing, and RAP; 4) an enhancedcognitive architecture augmented with tool calling and role-playing abilitiesto manage complex tasks effectively. Based on our new StepEval-Audio-360evaluation benchmark, Step-Audio achieves state-of-the-art performance in humanevaluations, especially in terms of instruction following. On open-sourcebenchmarks like LLaMA Question, shows 9.3% average performance improvement,demonstrating our commitment to advancing the development of open-sourcemulti-modal language technologies. Our code and models are available athttps://github.com/stepfun-ai/Step-Audio.

【12】 TAPS: Throat and Acoustic Paired Speech Dataset for Deep Learning-Based  Speech Enhancement
标题:TAPS:用于基于深度学习的语音增强的喉咙和声学配对语音数据集
链接:https://arxiv.org/abs/2502.11478
作者:Yunsik Kim,  Yonghun Song,  Yoonyoung Chung
摘要:在工厂、地铁和繁忙街道等高噪声环境中,由于背景噪声,捕获清晰的语音是一项挑战。喉部麦克风提供了一种具有噪声抑制特性的解决方案,可以在录制语音时降低噪声。然而,一个重要的限制仍然存在:高频信息衰减声波通过皮肤和组织,降低语音清晰度。最近的深度学习方法在增强喉咙麦克风录音方面显示出了希望,但由于缺乏标准化数据集,进一步的进展受到限制。我们介绍了一个喉咙和声学配对语音数据集(TAPS),收集配对的话语记录从60个母语韩国人使用喉咙和声学麦克风。为了证明TAPS的实用性,我们测试了三种基线深度学习模型,并确定基于映射的方法在提高语音质量和恢复内容方面具有优势。此外,我们提出了一种最佳方法来减轻喉部和声学麦克风之间的信号失配,确保模型性能。这些结果凸显了TAPS作为标准化数据集和推进基于喉麦克风的语音增强研究的潜力。
摘要:In high-noise environments such as factories, subways, and busy streets,capturing clear speech is challenging due to background noise. Throatmicrophones provide a solution with their noise-suppressing properties,reducing the noise while recording speech. However, a significant limitationremains: high-frequency information is attenuated as sound waves pass throughskin and tissue, reducing speech clarity. Recent deep learning approaches haveshown promise in enhancing throat microphone recordings, but further progressis constrained by the absence of standardized dataset. We introduce a throatand acoustic paired speech dataset (TAPS), a collection of paired utterancesrecorded from 60 native Korean speakers using throat and acoustic microphones.To demonstrate the TAPS's utility, we tested three baseline deep learningmodels and identified the mapping-based approach as superior in improvingspeech quality and restoring content. Additionally, we propose an optimalmethod to mitigate the signal mismatch between throat and acoustic microphones,ensuring model performance. These results highlight the potential of TAPS toserve as a standardized dataset and advance research in throat microphone-basedspeech enhancement.

【13】 FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine  Flow Matching
标题:FELLE:具有令牌粗到细流匹配的自回归语音合成
链接:https://arxiv.org/abs/2502.11128
作者:Hui Wang,  Shujie Liu,  Lingwei Meng,  Jinyu Li,  Yifan Yang,  Shiwan Zhao,  Haiyang Sun,  Yanqing Liu,  Haoqin Sun,  Jiaming Zhou,  Yan Lu,  Yong Qin
摘要:为了推进连续值令牌建模和时间一致性执行,我们提出了FELLE,一个自回归模型,集成了语言建模与令牌流匹配。通过利用语言模型的自回归性质和流匹配的生成效率,FELLE有效地预测连续值标记(mel频谱图)。对于每个连续值的令牌,FELLE修改流匹配中的一般先验分布,通过合并来自前一步的信息,提高一致性和稳定性。此外,为了提高合成质量,FELLE引入了一种从粗到细的流匹配机制,以语言模型的输出为条件,分层生成连续值的令牌。实验结果证明了在自回归梅尔频谱图建模中结合流匹配技术的潜力,从而显著改善TTS生成质量,如https://aka.ms/felle中所示。
摘要:To advance continuous-valued token modeling and temporal-coherenceenforcement, we propose FELLE, an autoregressive model that integrates languagemodeling with token-wise flow matching. By leveraging the autoregressive natureof language models and the generative efficacy of flow matching, FELLEeffectively predicts continuous-valued tokens (mel-spectrograms). For eachcontinuous-valued token, FELLE modifies the general prior distribution in flowmatching by incorporating information from the previous step, improvingcoherence and stability. Furthermore, to enhance synthesis quality, FELLEintroduces a coarse-to-fine flow-matching mechanism, generatingcontinuous-valued tokens hierarchically, conditioned on the language model'soutput. Experimental results demonstrate the potential of incorporatingflow-matching techniques in autoregressive mel-spectrogram modeling, leading tosignificant improvements in TTS generation quality, as shown inhttps://aka.ms/felle.

【14】 Hyperdimensional Intelligent Sensing for Efficient Real-Time Audio  Processing on Extreme Edge
标题:超维智能传感在极端边缘实现高效实时音频处理
链接:https://arxiv.org/abs/2502.10718
作者:Sanggeon Yun,  Ryozo Masukawa,  Hanning Chen,  SungHeon Jeong,  Wenjun Huang,  Arghavan Rezvani,  Minhyoung Na,  Yoshiki Yamaguchi,  Mohsen Imani
备注:Accepted to IEEE Access
摘要:管理大量传感器生成的数据的挑战不断升级,特别是在音频应用中,需要创新的解决方案。当前的系统面临着巨大的计算和存储需求,特别是在枪击检测系统(GSDS)等实时应用中,边缘传感器的激增加剧了这些问题。本文提出了一种突破性的方法,为智能音频传感框架量身定制的近传感器模型。利用快速傅里叶变换(FFT)模块,卷积神经网络(CNN)层和超维计算(HDC),我们的模型在低能耗,快速推理和在线学习方面表现出色。它非常适合高效的ASIC设计实施,与传统的嵌入式CPU或GPU相比,具有卓越的能效,并且与缩小麦克风传感器尺寸的趋势相兼容。在软件和硬件两个层面的全面评估强调了该模型的有效性。通过详细的ROC曲线分析进行的软件评估揭示了节能和质量损失之间的微妙平衡,实现了高达82.1%的节能,而质量损失仅为1.39%。硬件评估强调了该模型通过ASIC设计实现时值得称赞的能效,特别是使用Google Edge TPU,展示了其优于主流嵌入式CPU和GPU的优势。
摘要:The escalating challenges of managing vast sensor-generated data,particularly in audio applications, necessitate innovative solutions. Currentsystems face significant computational and storage demands, especially inreal-time applications like gunshot detection systems (GSDS), and theproliferation of edge sensors exacerbates these issues. This paper proposes agroundbreaking approach with a near-sensor model tailored for intelligentaudio-sensing frameworks. Utilizing a Fast Fourier Transform (FFT) module,convolutional neural network (CNN) layers, and HyperDimensional Computing(HDC), our model excels in low-energy, rapid inference, and online learning. Itis highly adaptable for efficient ASIC design implementation, offering superiorenergy efficiency compared to conventional embedded CPUs or GPUs, and iscompatible with the trend of shrinking microphone sensor sizes. Comprehensiveevaluations at both software and hardware levels underscore the model'sefficacy. Software assessments through detailed ROC curve analysis revealed adelicate balance between energy conservation and quality loss, achieving up to82.1% energy savings with only 1.39% quality loss. Hardware evaluationshighlight the model's commendable energy efficiency when implemented via ASICdesign, especially with the Google Edge TPU, showcasing its superiority overprevalent embedded CPUs and GPUs.

【15】 F-StrIPE: Fast Structure-Informed Positional Encoding for Symbolic Music  Generation
标题:F-StrIPE:用于符号音乐生成的快速结构信息位置编码
链接:https://arxiv.org/abs/2502.10491
作者:Manvi Agarwal (IP Paris, LTCI, IDS),  Changhong Wang (LTCI),  Gael Richard (S2A, IDS)
备注:None
摘要:虽然音乐对于像Transformers这样的生成模型仍然是一个具有挑战性的领域,但最近的进展是通过利用合适的音乐信息先验来实现的。利用关于Transformers中的音乐结构的信息的一种技术是将这样的知识插入到位置编码(PE)模块中。然而,Transformers在序列长度上具有二次成本。在本文中,我们提出了F-StrIPE,一个结构通知PE计划,工作在线性复杂度。使用现有的核近似技术的基础上随机功能,我们表明,F-StrIPE是一个推广的随机位置编码(SPE)。我们说明了经验的优点F-StrIPE使用旋律协调的象征性音乐。
摘要:While music remains a challenging domain for generative models likeTransformers, recent progress has been made by exploiting suitablemusically-informed priors. One technique to leverage information about musicalstructure in Transformers is inserting such knowledge into the positionalencoding (PE) module. However, Transformers carry a quadratic cost in sequencelength. In this paper, we propose F-StrIPE, a structure-informed PE scheme thatworks in linear complexity. Using existing kernel approximation techniquesbased on random features, we show that F-StrIPE is a generalization ofStochastic Positional Encoding (SPE). We illustrate the empirical merits ofF-StrIPE using melody harmonization for symbolic music.

【16】 YNote: A Novel Music Notation for Fine-Tuning LLMs in Music Generation
标题:YNote:音乐生成中微调LLM的新颖音乐记法
链接:https://arxiv.org/abs/2502.10467
作者:Shao-Chien Lu,  Chen-Chen Yeh,  Hui-Lin Cho,  Chun-Chieh Hsu,  Tsai-Ling Hsu,  Cheng-Han Wu,  Timothy K. Shih,  Yu-Cheng Lin
摘要:使用大型语言模型(LLM)生成音乐的领域正在迅速发展,但现有的音乐符号系统,如ABC-Notation和MusicXML,仍然过于复杂,无法有效地微调LLM。由于这些格式的可变性和复杂结构,机器和人类都很难解释这些格式。为了应对这些挑战,我们引入了YNote,这是一个简化的音乐符号系统,它只使用四个字符来表示音符及其音高。YNote的固定格式确保了一致性,使其易于阅读,更适合于微调LLM。在我们的实验中,我们在YNote编码的数据集上微调了GPT-2(124 M),并分别获得了0.883和0.766的BLEU和ROUGE分数。只需两个音符作为提示,该模型就能够生成连贯且风格相关的音乐。我们相信YNote为机器学习应用程序提供了现有音乐符号的实用替代方案,并有可能显着提高使用LLM生成音乐的质量。
摘要:The field of music generation using Large Language Models (LLMs) is evolvingrapidly, yet existing music notation systems, such as MIDI, ABC Notation, andMusicXML, remain too complex for effective fine-tuning of LLMs. These formatsare difficult for both machines and humans to interpret due to theirvariability and intricate structure. To address these challenges, we introduceYNote, a simplified music notation system that uses only four characters torepresent a note and its pitch. YNote's fixed format ensures consistency,making it easy to read and more suitable for fine-tuning LLMs. In ourexperiments, we fine-tuned GPT-2 (124M) on a YNote-encoded dataset and achievedBLEU and ROUGE scores of 0.883 and 0.766, respectively. With just two notes asprompts, the model was able to generate coherent and stylistically relevantmusic. We believe YNote offers a practical alternative to existing musicnotations for machine learning applications and has the potential tosignificantly enhance the quality of music generation using LLMs.

机器翻译由腾讯交互翻译提供,仅供参考