微信公众号:arXiv_Daily
cs.SD语音
【1】Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba
标题:说话、编辑、重复:高保真语音编辑和Zero-ShotTTC,具有交叉注意力的曼巴
链接:https://arxiv.org/abs/2510.04738
【2】A Study on the Data Distribution Gap in Music Emotion Recognition
标题:音乐情感识别中的数据分布差距研究
链接:https://arxiv.org/abs/2510.04688
摘要:音乐情感识别(MER)是一项与人类感知密切相关的任务,在很大程度上依赖于从贡献者那里收集的主观注释。以前的研究往往集中在特定的音乐风格,而不是将各种各样的流派,如摇滚和古典,在一个单一的框架。在本文中,我们通过调查五个具有维度情感注释的数据集来解决从音频内容中识别情感的任务,这些数据集是跨各种音乐风格的音乐,包括Dummy Music,DEAM,PMEmo,WTC和WCMED。我们在一个系统的实验中证明了分布外泛化的问题。通过仔细研究多个数据和特征集,我们可以深入了解现有数据中的体裁-情感关系,并研究某些特征表示中潜在的体裁主导地位和数据集偏见。基于这些实验,我们得到了一个简单而有效的框架,它结合了从Juventure模型中提取的嵌入与色度特征,并演示了如何结合几个不同的训练集,这使我们能够训练模型,大大提高了跨数据集的泛化能力。
【3】Robustness assessment of large audio language models in multiple-choice evaluation
标题:多项选择评估中大型音频语言模型的稳健性评估
链接:https://arxiv.org/abs/2510.04584
摘要:大型音频语言模型(LALM)的最新进展主要是使用多项选择题回答(MCQA)框架进行评估的。然而,细微的变化,如改变选择的顺序,会导致实质上不同的结果。现有的MCQA框架没有考虑这种可变性,并报告每个基准或类别的单个准确性数字。我们深入了解MCQA评估框架,并对三个基准(MMAU,MMAR和MMSU)和四个模型进行了系统的研究:Audio Flamingo 2,Audio Flamingo 3,Qwen2.5-Omni-7 B-Instruct和Kimi-Audio-7 B-Instruct。我们的研究结果表明,模型不仅对选择的顺序敏感,而且对问题和选择的释义也敏感。最后,我们提出了一个更简单的评估协议和指标,占微妙的变化,并提供了更详细的评估报告的LALM MCQA框架内。
【4】Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
标题:基于语言模型的文本到音频生成:反因果对齐的协作剩余变形器
链接:https://arxiv.org/abs/2510.04577
摘要:虽然语言模型(LM)与残差矢量量化(RVQ)标记器配对在文本到音频(T2A)生成中表现出了希望,但它们仍然落后于基于扩散的模型。我们确定了一个关键的困境支撑这一差距:纳入更多的RVQ层提高音频重建保真度,但超过了传统LM的生成能力。为了解决这个问题,我们首先分析了RVQ动态并发现了两个关键限制:1)RVQ层之间特征的正交性阻碍了有效的LM训练,2)来自更深RVQ层的令牌中语义丰富度的下降加剧了自回归解码期间的暴露偏差。基于这些见解,我们提出了Siren,一种新的基于LM的框架,采用多个隔离Transformers,通过强化学习进行因果调节和反因果对齐。大量的实验表明,Siren优于现有的基于LM和基于扩散的T2A系统,实现了最先进的结果。通过桥接的代表性优势的LM与音频合成的保真度要求,我们的方法重新定位LM的竞争对手对扩散模型在T2A任务。此外,通过将音频表示与语言结构对齐,Siren为统一的多模态生成框架提供了一条有前途的途径。
【5】Evaluating Self-Supervised Speech Models via Text-Based LLMS
标题:通过基于文本的LLMS评估自我监督语音模型
链接:https://arxiv.org/abs/2510.04463
摘要:自监督学习(SSL)因其能够以低标签成本学习丰富的表示,适用于各种下游任务而获得了吸引力。然而,由于额外的培训和评估成本,评估下游任务的性能仍然具有挑战性。现有的与任务无关的评估方法也需要额外的训练或超参数调整。我们提出了一种新的评价指标,使用大型语言模型(LLM)。通过将来自SSL模型的离散令牌序列和最小域线索输入到LLM中,我们获得了平均对数似然;这些线索指导了上下文学习,使分数更可靠,而无需额外的训练或超参数调整。实验结果表明,基于LLM的分数和自动语音识别任务之间的相关性。此外,我们的研究结果表明,LLM不仅可以作为SSL评估工具,而且还提供对说话者验证任务有用的推理时间嵌入。
【6】Pitch-Conditioned Instrument Sound Synthesis From an Interactive Timbre Latent Space
标题:从交互式音色潜在空间进行音调调节乐器声音合成
链接:https://arxiv.org/abs/2510.04339
摘要:本文提出了一种新的神经乐器声音合成方法,使用两阶段半监督学习框架,能够从富有表现力的音色潜在空间生成音高准确,高质量的音乐样本。现有的实现足够质量的音乐制作的方法通常依赖于难以导航的高维潜在表示,并且提供不直观的用户体验。我们通过两个阶段的训练范例来解决这个限制:首先,我们使用变分自动编码器来训练音频样本的音高-音色分解的2D表示;其次,我们使用这种表示作为基于变换器的生成模型的条件输入。学习的2D潜在空间作为导航和探索声音景观的直观界面。我们证明,所提出的方法有效地学习一个解开的音色空间,使表达和可控的音频生成与可靠的音高调节。实验结果表明,该模型的能力,捕捉微妙的音色变化,同时保持高度的音高精度。我们的方法的可用性在一个交互式Web应用程序中得到了证明,突出了其作为未来音乐制作环境的一步的潜力,这些环境既直观又具有创造性:https://pgesam.faresschulz.com
【7】Machine Unlearning in Speech Emotion Recognition via Forget Set Alone
标题:通过Forget Set Alone实现语音情感识别的机器去学习
链接:https://arxiv.org/abs/2510.04251
摘要:语音情感识别是从语音信号中识别情感状态的一种方法,在人机交互、教育、医疗等领域有着广泛的应用。然而,由于语音数据包含丰富的敏感信息,部分数据可能会被要求删除的发言者出于隐私问题。目前的机器学习方法在很大程度上依赖于被遗忘的样本之外的数据。然而,当数据再分配受到限制并且在大数据的背景下需要大量计算资源时,这种依赖带来了挑战。我们提出了一种新的基于对抗性攻击的方法,该方法只使用要忘记的数据来微调预训练的语音情感识别模型。实验结果表明,该方法能够有效地去除模型中待遗忘数据的知识,同时保持模型在情感识别测试集上的高性能.
【8】GDiffuSE: Diffusion-based speech enhancement with noise model guidance
标题:GDffuSE:具有噪音模型指导的基于扩散的语音增强
链接:https://arxiv.org/abs/2510.04157
【9】Cross-Lingual Multi-Granularity Framework for Interpretable Parkinson's Disease Diagnosis from Speech
标题:从言语可解释帕金森病诊断的跨语言多粒度框架
链接:https://arxiv.org/abs/2510.03758
【10】Evaluating High-Resolution Piano Sustain Pedal Depth Estimation with Musically Informed Metrics
标题:用音乐信息处理评估高分辨率钢琴持续踏板深度估计
链接:https://arxiv.org/abs/2510.03750
【11】Désentrelacement Fréquentiel Doux pour les Codecs Audio Neuronaux
标题:Désentrelacement Fréquentiel Doux pour les Codecs Audio Neuronaux
链接:https://arxiv.org/abs/2510.03741
摘要:虽然基于神经的模型在音频特征提取方面取得了重大进展,但学习表示的可解释性仍然是一个关键挑战。为了解决这个问题,解纠缠技术已经被集成到离散神经音频编解码器中,以在提取的令牌上施加结构。然而,这些方法通常表现出对特定数据集或任务制定的强烈依赖性。在这项工作中,我们提出了一个解开神经音频编解码器,利用时域信号的频谱分解,以提高表示的可解释性。实验评估表明,我们的方法超越了国家的最先进的基线重建保真度和感知质量。
【12】Soft Disentanglement in Frequency Bands for Neural Audio Codecs
标题:神经音频编解码器频段的软解纠缠
链接:https://arxiv.org/abs/2510.03735
【13】Lightweight and Generalizable Acoustic Scene Representations via Contrastive Fine-Tuning and Distillation
标题:通过对比微调和蒸馏实现轻量级且可推广的声学场景表示
链接:https://arxiv.org/abs/2510.03728
【14】Audio Forensics Evaluation (SAFE) Challenge
标题:音频取证评估(SAFE)挑战
链接:https://arxiv.org/abs/2510.03387
【15】Linguistic and Audio Embedding-Based Machine Learning for Alzheimer's Dementia and Mild Cognitive Impairment Detection: Insights from the PROCESS Challenge
标题:阿尔茨海默氏痴呆症和轻度认知障碍检测的基于语言和音频嵌入的机器学习:来自Process挑战的见解
链接:https://arxiv.org/abs/2510.03336
【16】UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
标题:UniVoice:将自回归ASC和基于流匹配的TTC与大型语言模型统一起来
链接:https://arxiv.org/abs/2510.04593
【17】Differentiable physics for sound field reconstruction
标题:用于声学重建的可微物理学
链接:https://arxiv.org/abs/2510.04459
摘要:声场重建涉及从有限数量的空间分布的观测估计声场。这项工作介绍了一个可微的物理方法声场重建,其中的波动方程的初始条件近似的神经网络,和微分算子计算与可微的数值求解器。使用数值求解器可以实现稳定的网络训练,同时将物理学作为强约束,这与传统的物理学神经网络不同,传统的物理学神经网络将物理学作为损失函数中的约束。我们引入了一个额外的稀疏促进约束,以实现有意义的解决方案,即使在严重欠采样条件下。实验表明,该方法可以在极端数据稀缺的情况下重建声场,与物理信息神经网络相比,具有更高的精度和更好的收敛性。
【18】Probing Whisper for Dysarthric Speech in Detection and Assessment
标题:在检测和评估中探索耳语对发音障碍的影响
链接:https://arxiv.org/abs/2510.04219
摘要:像Whisper这样的大规模端到端模型在不同的语音任务上表现出了很强的性能,但它们对病理性语音的内部行为仍然知之甚少。了解构音障碍语音如何跨层表示对于构建可靠且可解释的临床评估工具至关重要。本研究探讨了用于构音障碍语音的Whisper-Medium模型编码器,用于检测和评估(即,严重程度分类)。我们在单任务和多任务设置下使用线性分类器评估逐层嵌入,并使用Silhouette分数和互信息补充这些结果,以提供对层信息性的看法。为了检查适应性,我们在对构音障碍语音识别任务进行微调Whisper后重复分析。在所有指标中,中级编码器层(13-15)的信息量最大,而微调只会引起适度的变化。这些发现提高了Whisper嵌入的可解释性,并突出了探测分析的潜力,以指导使用大规模的预训练模型进行病理性语音。
【19】Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning
标题:使用w2 v-BERT 2.0和知识蒸馏引导的结构化修剪增强说话者验证
链接:https://arxiv.org/abs/2510.04213
【20】Drax: Speech Recognition with Discrete Flow Matching
标题:Drax:具有离散流匹配的语音识别
链接:https://arxiv.org/abs/2510.04162
【21】MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
标题:MoME:视听语音识别Matryoshka专家的混合体
链接:https://arxiv.org/abs/2510.04136
摘要:大型语言模型(LLM)最近在视听语音识别(AVSR)中显示出强大的潜力,但其高计算需求和对令牌粒度的敏感性限制了其在资源受限环境中的实用性。令牌压缩方法可以降低推理成本,但它们需要预先固定压缩率并产生单个固定长度的输出,在推理时无法灵活地平衡信息密度和效率。Matryoshka表示学习(MRL)通过使单个模型跨多个令牌粒度运行来解决这个问题,允许动态调整压缩率。然而,目前基于MRL的方法在训练期间独立地处理每个尺度,限制了跨尺度泛化、高压缩下的鲁棒性和可解释性。为了克服这些限制,我们提出了MoME(混合Matryoshka专家),一种新的框架,将稀疏混合专家(MoE)到基于MRL的LLM AVSR。MoME通过top-k路由和共享专家来增强冻结的LLM,从而允许跨规模和模式进行动态容量分配。共享路由器可促进跨粒度的一致专家激活,使压缩序列能够受益于在较低压缩下学习的表示。在LRS 2和LRS 3上的实验表明,MoME在AVSR、ASR和VSR任务中实现了最先进的性能,同时需要的参数显著减少,并在噪声下保持了鲁棒性。MoME将MRL的适应性与MoE的效率相结合,为资源感知语音识别提供了可扩展和可解释的解决方案。
【22】A Multilingual Framework for Dysarthria: Detection, Severity Classification, Speech-to-Text, and Clean Speech Generation
标题:发音障碍的多语言框架:检测、严重性分类、语音到文本和清晰语音生成
链接:https://arxiv.org/abs/2510.03986
【23】From Qubits to Rhythm: Exploring Quantum Random Walks in Rhythmspaces
标题:从量子比特到节奏:探索节奏空间中的量子随机行走
链接:https://arxiv.org/abs/2510.03836
摘要:提出了一种用于节奏生成的量子计算算法,旨在扩展和探索量子计算在艺术,特别是音乐中的应用。该算法将量子随机行走轨迹映射到节奏空间-一个插入节奏模式的2D界面。该方法包括三个阶段。第一阶段涉及设计量子计算算法,并在量子比特空间和节奏空间之间建立映射。为了使电路深度最小化,应用将2D量子随机游走分解为两个1D量子随机游走。第二阶段的重点是通过引入经典势场来偏置量子随机行走的方向性,根据这些场内的位置梯度调整波函数的概率分布。实现了四种势场:零势、线性场、高斯势和惯性动力学下的高斯势。第三阶段通过生成磁鼓模式信息并将其传输到数字音频工作站(Digital Audio Workstation,简称DTS)来解决这些路径的发声问题。这项工作建立在现有文献的基础上,这些文献将量子计算应用于具有几个位置的更简单的量子位空间,将形式主义扩展到2D x-y平面。它作为音乐和音频应用中基于可扩展量子计算的生成随机游走算法的概念证明。此外,该方法适用于通用的多维声音空间,因为算法不严格限于节奏生成,并且可以适应不同的音乐结构。
【24】A MATLAB toolbox for Computation of Speech Transmission Index (STI)
标题:用于计算语音传输指数(STI)的LAB工具箱
链接:https://arxiv.org/abs/2510.03825
【25】Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
标题:将日记化条件耳语器适应端到端多说话者语音识别
链接:https://arxiv.org/abs/2510.03723
【26】Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
标题:使用说话者不可知活动流扩展多人ASB
链接:https://arxiv.org/abs/2510.03630
【1】MuFFIN: Multifaceted Pronunciation Feedback Model with Interactive Hierarchical Neural Modeling
标题:MuFIGN:具有交互式分层神经建模的多面发音反馈模型
链接:https://arxiv.org/abs/2510.04956
摘要:计算机辅助发音训练(CAPT)通过提供及时的指导性反馈来帮助第二语言学习者练习发音技能。为了从多个方面检查发音水平,现有的CAPT方法大致分为两类:发音错误检测和诊断(MDD)以及自动发音评估(APA)。前者旨在查明语音发音错误并提供诊断反馈,而后者则旨在量化与各个方面有关的发音熟练程度。尽管MDD和APA之间有着天然的互补性,但是研究人员和实践者经常将它们视为具有不同建模范式的独立任务。鉴于此,我们在本文中首先介绍了MuFFIN,一个多方面的发音反馈模型与交互式层次神经架构,共同解决MDD和APA的任务。为了更好地捕捉特征空间中音素之间的细微差别,然后提出了一种新的音素对比顺序正则化机制,以优化所提出的模型,以生成更多的音素区分特征,同时考虑方面分数的顺序性。此外,为了解决MDD中复杂的数据不平衡问题,我们设计了一个简单而有效的训练目标,该目标专门针对音素分类器的输出与音素特定的变化,以便更好地呈现预测音素的分布,同时考虑其错误发音特征。在Speechocean762基准数据集上进行的一系列实验证明了我们的方法在几个前沿基线上的有效性,在APA和MDD任务上都显示了最先进的性能。
【2】Perceptual Evaluation of Extrapolated Spatial Room Impulse Responses From a Mono Source
标题:从单频源推断的空间房间脉冲响应的感知评估
链接:https://arxiv.org/abs/2510.04937
摘要:沉浸在虚拟和增强现实解决方案中依赖于合理的空间音频。然而,可扩展地表示沉浸式音频的空间通常需要使用专业空间麦克风对源麦克风对进行许多单独的声学测量,这使得该过程耗时且昂贵。在这项研究中,我们使用3-替代强迫选择(3AFC)听力测试评估外推和空间化的房间脉冲响应(RIR)的可扩展性。刺激由来自三个空间的RIR与语音、管弦乐和器乐卷积组成。当被要求从一个外推刺激和两个真实刺激中选择哪些刺激是人工刺激时,20名参与者的总体准确率为38%(比预期猜测率高出5个百分点)。鉴于听力测试的结果,这项研究表明,它是可能的推断合理的空间RIR从单声道测量,减少声学测量的时间和专业设备的需要。
【3】AURA Score: A Metric For Holistic Audio Question Answering Evaluation
标题:AURA分数:整体音频问题回答评估的指标
链接:https://arxiv.org/abs/2510.04934
【4】UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
标题:UniVoice:将自回归ASC和基于流匹配的TTC与大型语言模型统一起来
链接:https://arxiv.org/abs/2510.04593
【5】Differentiable physics for sound field reconstruction
标题:用于声学重建的可微物理学
链接:https://arxiv.org/abs/2510.04459
摘要:声场重建涉及从有限数量的空间分布的观测估计声场。这项工作介绍了一个可微的物理方法声场重建,其中的波动方程的初始条件近似的神经网络,和微分算子计算与可微的数值求解器。使用数值求解器可以实现稳定的网络训练,同时将物理学作为强约束,这与传统的物理学神经网络不同,传统的物理学神经网络将物理学作为损失函数中的约束。我们引入了一个额外的稀疏促进约束,以实现有意义的解决方案,即使在严重欠采样条件下。实验表明,该方法可以在极端数据稀缺的情况下重建声场,与物理信息神经网络相比,具有更高的精度和更好的收敛性。
【6】Probing Whisper for Dysarthric Speech in Detection and Assessment
标题:在检测和评估中探索耳语对发音障碍的影响
链接:https://arxiv.org/abs/2510.04219
摘要:像Whisper这样的大规模端到端模型在不同的语音任务上表现出了很强的性能,但它们对病理性语音的内部行为仍然知之甚少。了解构音障碍语音如何跨层表示对于构建可靠且可解释的临床评估工具至关重要。本研究探讨了用于构音障碍语音的Whisper-Medium模型编码器,用于检测和评估(即,严重性分类)。我们在单任务和多任务设置下使用线性分类器评估逐层嵌入,并使用Silhouette分数和互信息补充这些结果,以提供对层信息性的看法。为了检查适应性,我们在对构音障碍语音识别任务进行微调Whisper后重复分析。在所有指标中,中级编码器层(13-15)的信息量最大,而微调只会引起适度的变化。这些发现提高了Whisper嵌入的可解释性,并突出了探测分析的潜力,以指导使用大规模的预训练模型进行病理性语音。
【7】Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning
标题:使用w2 v-BERT 2.0和知识蒸馏引导的结构化修剪增强说话者验证
链接:https://arxiv.org/abs/2510.04213
【8】Drax: Speech Recognition with Discrete Flow Matching
标题:Drax:具有离散流匹配的语音识别
链接:https://arxiv.org/abs/2510.04162
【9】MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
标题:MoME:视听语音识别Matryoshka专家的混合体
链接:https://arxiv.org/abs/2510.04136
摘要:大型语言模型(LLM)最近在视听语音识别(AVSR)中显示出强大的潜力,但其高计算需求和对令牌粒度的敏感性限制了其在资源受限环境中的实用性。令牌压缩方法可以降低推理成本,但它们需要预先固定压缩率并产生单个固定长度的输出,在推理时无法灵活地平衡信息密度和效率。Matryoshka表示学习(MRL)通过使单个模型跨多个令牌粒度运行来解决这个问题,允许动态调整压缩率。然而,目前基于MRL的方法在训练期间独立地处理每个尺度,限制了跨尺度泛化、高压缩下的鲁棒性和可解释性。为了克服这些限制,我们提出了MoME(混合Matryoshka专家),一种新的框架,将稀疏混合专家(MoE)到基于MRL的LLM AVSR。MoME通过top-k路由和共享专家来增强冻结的LLM,从而允许跨规模和模式进行动态容量分配。共享路由器可促进跨粒度的一致专家激活,使压缩序列能够受益于在较低压缩下学习的表示。在LRS 2和LRS 3上的实验表明,MoME在AVSR、ASR和VSR任务中实现了最先进的性能,同时需要的参数显著减少,并在噪声下保持了鲁棒性。MoME将MRL的适应性与MoE的效率相结合,为资源感知语音识别提供了可扩展和可解释的解决方案。
【10】A Multilingual Framework for Dysarthria: Detection, Severity Classification, Speech-to-Text, and Clean Speech Generation
标题:发音障碍的多语言框架:检测、严重性分类、语音到文本和清晰语音生成
链接:https://arxiv.org/abs/2510.03986
【11】A MATLAB toolbox for Computation of Speech Transmission Index (STI)
标题:用于计算语音传输指数(STI)的LAB工具箱
链接:https://arxiv.org/abs/2510.03825
摘要:The speech transmission index (STI) is a popular simple metric for the prediction of speech intelligibility when speech is passed through a transmission channel. Computation of STI from acoustic measurements is described in the IEC 60268-16:2020 standard. Though, reliable implementations of STI are not publicly accessible and are frequently limited to the use with a proprietary measurement hardware. We present a Matlab STI implementation of both the direct and indirect approaches according to the standard, including the shortened STIPA protocol. The suggested implementation meets prescribed requirements, as evidenced by tests on reference signals. Additionally, we conducted a verification measurement in comparison to a commercial measurement device. Our software comes with open source code.
【12】Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
标题:将日记化条件耳语器适应端到端多说话者语音识别
链接:https://arxiv.org/abs/2510.03723
摘要:We propose a speaker-attributed (SA) Whisper-based model for multi-talker speech recognition that combines target-speaker modeling with serialized output training (SOT). Our approach leverages a Diarization-Conditioned Whisper (DiCoW) encoder to extract target-speaker embeddings, which are concatenated into a single representation and passed to a shared decoder. This enables the model to transcribe overlapping speech as a serialized output stream with speaker tags and timestamps. In contrast to target-speaker ASR systems such as DiCoW, which decode each speaker separately, our approach performs joint decoding, allowing the decoder to condition on the context of all speakers simultaneously. Experiments show that the model outperforms existing SOT-based approaches and surpasses DiCoW on multi-talker mixtures (e.g., LibriMix).
【13】Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
标题:使用说话者不可知活动流扩展多人ASB
链接:https://arxiv.org/abs/2510.03630
摘要:An increasingly common training paradigm for multi-talker automatic speech recognition (ASR) is to use speaker activity signals to adapt single-speaker ASR models for overlapping speech. Although effective, these systems require running the ASR model once per speaker, resulting in inference costs that scale with the number of speakers and limiting their practicality. In this work, we propose a method that decouples the inference cost of activity-conditioned ASR systems from the number of speakers by converting speaker-specific activity outputs into two speaker-agnostic streams. A central challenge is that na\"ively merging speaker activities into streams significantly degrades recognition, since pretrained ASR models assume contiguous, single-speaker inputs. To address this, we design new heuristics aimed at preserving conversational continuity and maintaining compatibility with existing systems. We show that our approach is compatible with Diarization-Conditioned Whisper (DiCoW) to greatly reduce runtimes on the AMI and ICSI meeting datasets while retaining competitive performance.
【14】Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba
标题:说话、编辑、重复:高保真语音编辑和Zero-ShotTTC,具有交叉注意力的曼巴
链接:https://arxiv.org/abs/2510.04738
摘要:We introduce MAVE (Mamba with Cross-Attention for Voice Editing and Synthesis), a novel autoregressive architecture for text-conditioned voice editing and high-fidelity text-to-speech (TTS) synthesis, built on a cross-attentive Mamba backbone. MAVE achieves state-of-the-art performance in speech editing and very competitive results in zero-shot TTS, while not being explicitly trained on the latter task, outperforming leading autoregressive and diffusion models on diverse, real-world audio. By integrating Mamba for efficient audio sequence modeling with cross-attention for precise text-acoustic alignment, MAVE enables context-aware voice editing with exceptional naturalness and speaker consistency. In pairwise human evaluations on a random 40-sample subset of the RealEdit benchmark (400 judgments), 57.2% of listeners rated MAVE - edited speech as perceptually equal to the original, while 24.8% prefered the original and 18.0% MAVE - demonstrating that in the majority of cases edits are indistinguishable from the source. MAVE compares favorably with VoiceCraft and FluentSpeech both on pairwise comparisons and standalone mean opinion score (MOS) evaluations. For zero-shot TTS, MAVE exceeds VoiceCraft in both speaker similarity and naturalness, without requiring multiple inference runs or post-processing. Remarkably, these quality gains come with a significantly lower memory cost and approximately the same latency: MAVE requires ~6x less memory than VoiceCraft during inference on utterances from the RealEdit database (mean duration: 6.21s, A100, FP16, batch size 1). Our results demonstrate that MAVE establishes a new standard for flexible, high-fidelity voice editing and synthesis through the synergistic integration of structured state-space modeling and cross-modal attention.
【15】Robustness assessment of large audio language models in multiple-choice evaluation
标题:多项选择评估中大型音频语言模型的稳健性评估
链接:https://arxiv.org/abs/2510.04584
摘要:大型音频语言模型(LALM)的最新进展主要是使用多项选择题回答(MCQA)框架进行评估的。然而,细微的变化,如改变选择的顺序,会导致实质上不同的结果。现有的MCQA框架没有考虑这种可变性,并报告每个基准或类别的单个准确性数字。我们深入了解MCQA评估框架,并对三个基准(MMAU,MMAR和MMSU)和四个模型进行了系统的研究:Audio Flamingo 2,Audio Flamingo 3,Qwen2.5-Omni-7 B-Instruct和Kimi-Audio-7 B-Instruct。我们的研究结果表明,模型不仅对选择的顺序敏感,而且对问题和选择的释义也敏感。最后,我们提出了一个更简单的评估协议和指标,占微妙的变化,并提供了更详细的评估报告的LALM MCQA框架内。
摘要:Recent advances in large audio language models (LALMs) have primarily been assessed using a multiple-choice question answering (MCQA) framework. However, subtle changes, such as shifting the order of choices, result in substantially different results. Existing MCQA frameworks do not account for this variability and report a single accuracy number per benchmark or category. We dive into the MCQA evaluation framework and conduct a systematic study spanning three benchmarks (MMAU, MMAR and MMSU) and four models: Audio Flamingo 2, Audio Flamingo 3, Qwen2.5-Omni-7B-Instruct, and Kimi-Audio-7B-Instruct. Our findings indicate that models are sensitive not only to the ordering of choices, but also to the paraphrasing of the question and the choices. Finally, we propose a simpler evaluation protocol and metric that account for subtle variations and provide a more detailed evaluation report of LALMs within the MCQA framework.
【16】Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
标题:基于语言模型的文本到音频生成:反因果对齐的协作剩余变形器
链接:https://arxiv.org/abs/2510.04577
摘要:虽然语言模型(LM)与残差矢量量化(RVQ)标记器配对在文本到音频(T2A)生成中表现出了希望,但它们仍然落后于基于扩散的模型。我们确定了一个关键的困境支撑这一差距:纳入更多的RVQ层提高音频重建保真度,但超过了传统LM的生成能力。为了解决这个问题,我们首先分析了RVQ动态并发现了两个关键限制:1)RVQ层之间特征的正交性阻碍了有效的LM训练,2)来自更深RVQ层的令牌中语义丰富度的下降加剧了自回归解码期间的暴露偏差。基于这些见解,我们提出了Siren,一种新的基于LM的框架,采用多个隔离Transformers,通过强化学习进行因果调节和反因果对齐。大量的实验表明,Siren优于现有的基于LM和基于扩散的T2A系统,实现了最先进的结果。通过桥接的代表性优势的LM与音频合成的保真度要求,我们的方法重新定位LM的竞争对手对扩散模型在T2A任务。此外,通过将音频表示与语言结构对齐,Siren为统一的多模态生成框架提供了一条有前途的途径。
摘要:While language models (LMs) paired with residual vector quantization (RVQ) tokenizers have shown promise in text-to-audio (T2A) generation, they still lag behind diffusion-based models by a non-trivial margin. We identify a critical dilemma underpinning this gap: incorporating more RVQ layers improves audio reconstruction fidelity but exceeds the generation capacity of conventional LMs. To address this, we first analyze RVQ dynamics and uncover two key limitations: 1) orthogonality of features across RVQ layers hinders effective LMs training, and 2) descending semantic richness in tokens from deeper RVQ layers exacerbates exposure bias during autoregressive decoding. Based on these insights, we propose Siren, a novel LM-based framework that employs multiple isolated transformers with causal conditioning and anti-causal alignment via reinforcement learning. Extensive experiments demonstrate that Siren outperforms both existing LM-based and diffusion-based T2A systems, achieving state-of-the-art results. By bridging the representational strengths of LMs with the fidelity demands of audio synthesis, our approach repositions LMs as competitive contenders against diffusion models in T2A tasks. Moreover, by aligning audio representations with linguistic structures, Siren facilitates a promising pathway toward unified multi-modal generation frameworks.
【17】Evaluating Self-Supervised Speech Models via Text-Based LLMS
标题:通过基于文本的LLMS评估自我监督语音模型
链接:https://arxiv.org/abs/2510.04463
摘要:自监督学习(SSL)因其能够以低标签成本学习丰富的表示,适用于各种下游任务而获得了吸引力。然而,由于额外的培训和评估成本,评估下游任务的性能仍然具有挑战性。现有的与任务无关的评估方法也需要额外的训练或超参数调整。我们提出了一种新的评价指标,使用大型语言模型(LLM)。通过将来自SSL模型的离散令牌序列和最小域线索输入到LLM中,我们获得了平均对数似然;这些线索指导了上下文学习,使分数更可靠,而无需额外的训练或超参数调整。实验结果表明,基于LLM的分数和自动语音识别任务之间的相关性。此外,我们的研究结果表明,LLM不仅可以作为SSL评估工具,而且还提供对说话者验证任务有用的推理时间嵌入。
摘要:Self-Supervised Learning (SSL) has gained traction for its ability to learn rich representations with low labeling costs, applicable across diverse downstream tasks. However, assessing the downstream-task performance remains challenging due to the cost of extra training and evaluation. Existing methods for task-agnostic evaluation also require extra training or hyperparameter tuning. We propose a novel evaluation metric using large language models (LLMs). By inputting discrete token sequences and minimal domain cues derived from SSL models into LLMs, we obtain the mean log-likelihood; these cues guide in-context learning, rendering the score more reliable without extra training or hyperparameter tuning. Experimental results show a correlation between LLM-based scores and automatic speech recognition task. Additionally, our findings reveal that LLMs not only functions as an SSL evaluation tools but also provides inference-time embeddings that are useful for speaker verification task.
【18】Pitch-Conditioned Instrument Sound Synthesis From an Interactive Timbre Latent Space
标题:从交互式音色潜在空间进行音调调节乐器声音合成
链接:https://arxiv.org/abs/2510.04339
摘要:本文提出了一种新的神经乐器声音合成方法,使用两阶段半监督学习框架,能够从富有表现力的音色潜在空间生成音高准确,高质量的音乐样本。现有的实现足够质量的音乐制作的方法通常依赖于难以导航的高维潜在表示,并且提供不直观的用户体验。我们通过两个阶段的训练范例来解决这个限制:首先,我们使用变分自动编码器来训练音频样本的音高-音色分解的2D表示;其次,我们使用这种表示作为基于变换器的生成模型的条件输入。学习的2D潜在空间作为导航和探索声音景观的直观界面。我们证明,所提出的方法有效地学习一个解开的音色空间,使表达和可控的音频生成与可靠的音高调节。实验结果表明,该模型的能力,捕捉微妙的音色变化,同时保持高度的音高精度。我们的方法的可用性在一个交互式Web应用程序中得到了证明,突出了其作为未来音乐制作环境的一步的潜力,这些环境既直观又具有创造性:https://pgesam.faresschulz.com
摘要:This paper presents a novel approach to neural instrument sound synthesis using a two-stage semi-supervised learning framework capable of generating pitch-accurate, high-quality music samples from an expressive timbre latent space. Existing approaches that achieve sufficient quality for music production often rely on high-dimensional latent representations that are difficult to navigate and provide unintuitive user experiences. We address this limitation through a two-stage training paradigm: first, we train a pitch-timbre disentangled 2D representation of audio samples using a Variational Autoencoder; second, we use this representation as conditioning input for a Transformer-based generative model. The learned 2D latent space serves as an intuitive interface for navigating and exploring the sound landscape. We demonstrate that the proposed method effectively learns a disentangled timbre space, enabling expressive and controllable audio generation with reliable pitch conditioning. Experimental results show the model's ability to capture subtle variations in timbre while maintaining a high degree of pitch accuracy. The usability of our method is demonstrated in an interactive web application, highlighting its potential as a step towards future music production environments that are both intuitive and creatively empowering: https://pgesam.faresschulz.com
【19】Machine Unlearning in Speech Emotion Recognition via Forget Set Alone
标题:通过Forget Set Alone实现语音情感识别的机器去学习
链接:https://arxiv.org/abs/2510.04251
摘要:语音情感识别是从语音信号中识别情感状态的一种方法,在人机交互、教育、医疗等领域有着广泛的应用。然而,由于语音数据包含丰富的敏感信息,部分数据可能会被要求删除的发言者出于隐私问题。目前的机器学习方法在很大程度上依赖于被遗忘的样本之外的数据。然而,当数据再分配受到限制并且在大数据的背景下需要大量计算资源时,这种依赖带来了挑战。我们提出了一种新的基于对抗性攻击的方法,该方法只使用要忘记的数据来微调预训练的语音情感识别模型。实验结果表明,该方法能够有效地去除模型中待遗忘数据的知识,同时保持模型在情感识别测试集上的高性能.
摘要:Speech emotion recognition aims to identify emotional states from speech signals and has been widely applied in human-computer interaction, education, healthcare, and many other fields. However, since speech data contain rich sensitive information, partial data can be required to be deleted by speakers due to privacy concerns. Current machine unlearning approaches largely depend on data beyond the samples to be forgotten. However, this reliance poses challenges when data redistribution is restricted and demands substantial computational resources in the context of big data. We propose a novel adversarial-attack-based approach that fine-tunes a pre-trained speech emotion recognition model using only the data to be forgotten. The experimental results demonstrate that the proposed approach can effectively remove the knowledge of the data to be forgotten from the model, while preserving high model performance on the test set for emotion recognition.
【20】GDiffuSE: Diffusion-based speech enhancement with noise model guidance
标题:GDffuSE:具有噪音模型指导的基于扩散的语音增强
链接:https://arxiv.org/abs/2510.04157
摘要:This paper introduces a novel speech enhancement (SE) approach based on a denoising diffusion probabilistic model (DDPM), termed Guided diffusion for speech enhancement (GDiffuSE). In contrast to conventional methods that directly map noisy speech to clean speech, our method employs a lightweight helper model to estimate the noise distribution, which is then incorporated into the diffusion denoising process via a guidance mechanism. This design improves robustness by enabling seamless adaptation to unseen noise types and by leveraging large-scale DDPMs originally trained for speech generation in the context of SE. We evaluate our approach on noisy signals obtained by adding noise samples from the BBC sound effects database to LibriSpeech utterances, showing consistent improvements over state-of-the-art baselines under mismatched noise conditions. Examples are available at our project webpage.
【21】From Qubits to Rhythm: Exploring Quantum Random Walks in Rhythmspaces
标题:从量子比特到节奏:探索节奏空间中的量子随机行走
链接:https://arxiv.org/abs/2510.03836
摘要:提出了一种用于节奏生成的量子计算算法,旨在扩展和探索量子计算在艺术,特别是音乐中的应用。该算法将量子随机行走轨迹映射到节奏空间-一个插入节奏模式的2D界面。该方法包括三个阶段。第一阶段涉及设计量子计算算法,并在量子比特空间和节奏空间之间建立映射。为了使电路深度最小化,应用将2D量子随机游走分解为两个1D量子随机游走。第二阶段的重点是通过引入经典势场来偏置量子随机行走的方向性,根据这些场内的位置梯度调整波函数的概率分布。实现了四种势场:零势、线性场、高斯势和惯性动力学下的高斯势。第三阶段通过生成磁鼓模式信息并将其传输到数字音频工作站(Digital Audio Workstation,简称DTS)来解决这些路径的发声问题。这项工作建立在现有文献的基础上,这些文献将量子计算应用于具有几个位置的更简单的量子位空间,将形式主义扩展到2D x-y平面。它作为音乐和音频应用中基于可扩展量子计算的生成随机游走算法的概念证明。此外,该方法适用于通用的多维声音空间,因为算法不严格限于节奏生成,并且可以适应不同的音乐结构。
摘要:A quantum computing algorithm for rhythm generation is presented, which aims to expand and explore quantum computing applications in the arts, particularly in music. The algorithm maps quantum random walk trajectories onto a rhythmspace -- a 2D interface that interpolates rhythmic patterns. The methodology consists of three stages. The first stage involves designing quantum computing algorithms and establishing a mapping between the qubit space and the rhythmspace. To minimize circuit depth, a decomposition of a 2D quantum random walk into two 1D quantum random walks is applied. The second stage focuses on biasing the directionality of quantum random walks by introducing classical potential fields, adjusting the probability distribution of the wave function based on the position gradient within these fields. Four potential fields are implemented: a null potential, a linear field, a Gaussian potential, and a Gaussian potential under inertial dynamics. The third stage addresses the sonification of these paths by generating MIDI drum pattern messages and transmitting them to a Digital Audio Workstation (DAW). This work builds upon existing literature that applies quantum computing to simpler qubit spaces with a few positions, extending the formalism to a 2D x-y plane. It serves as a proof of concept for scalable quantum computing-based generative random walk algorithms in music and audio applications. Furthermore, the approach is applicable to generic multidimensional sound spaces, as the algorithms are not strictly constrained to rhythm generation and can be adapted to different musical structures.
【22】Cross-Lingual Multi-Granularity Framework for Interpretable Parkinson's Disease Diagnosis from Speech
标题:从言语可解释帕金森病诊断的跨语言多粒度框架
链接:https://arxiv.org/abs/2510.03758
摘要:Parkinson's Disease (PD) affects over 10 million people worldwide, with speech impairments in up to 89% of patients. Current speech-based detection systems analyze entire utterances, potentially overlooking the diagnostic value of specific phonetic elements. We developed a granularity-aware approach for multilingual PD detection using an automated pipeline that extracts time-aligned phonemes, syllables, and words from recordings. Using Italian, Spanish, and English datasets, we implemented a bidirectional LSTM with multi-head attention to compare diagnostic performance across the different granularity levels. Phoneme-level analysis achieved superior performance with AUROC of 93.78% +- 2.34% and accuracy of 92.17% +- 2.43%. This demonstrates enhanced diagnostic capability for cross-linguistic PD detection. Importantly, attention analysis revealed that the most informative speech features align with those used in established clinical protocols: sustained vowels (/a/, /e/, /o/, /i/) at phoneme level, diadochokinetic syllables (/ta/, /pa/, /la/, /ka/) at syllable level, and /pataka/ sequences at word level. Source code will be available at https://github.com/jetliqs/clearpd.
【23】Evaluating High-Resolution Piano Sustain Pedal Depth Estimation with Musically Informed Metrics
标题:用音乐信息处理评估高分辨率钢琴持续踏板深度估计
链接:https://arxiv.org/abs/2510.03750
摘要:Evaluation for continuous piano pedal depth estimation tasks remains incomplete when relying only on conventional frame-level metrics, which overlook musically important features such as direction-change boundaries and pedal curve contours. To provide more interpretable and musically meaningful insights, we propose an evaluation framework that augments standard frame-level metrics with an action-level assessment measuring direction and timing using segments of press/hold/release states and a gesture-level analysis that evaluates contour similarity of each press-release cycle. We apply this framework to compare an audio-only baseline with two variants: one incorporating symbolic information from MIDI, and another trained in a binary-valued setting, all within a unified architecture. Results show that the MIDI-informed model significantly outperforms the others at action and gesture levels, despite modest frame-level gains. These findings demonstrate that our framework captures musically relevant improvements indiscernible by traditional metrics, offering a more practical and effective approach to evaluating pedal depth estimation models.
【24】Désentrelacement Fréquentiel Doux pour les Codecs Audio Neuronaux
标题:Désentrelacement Fréquentiel Doux pour les Codecs Audio Neuronaux
链接:https://arxiv.org/abs/2510.03741
摘要:虽然基于神经的模型在音频特征提取方面取得了重大进展,但学习表示的可解释性仍然是一个关键挑战。为了解决这个问题,解纠缠技术已经被集成到离散神经音频编解码器中,以在提取的令牌上施加结构。然而,这些方法通常表现出对特定数据集或任务制定的强烈依赖性。在这项工作中,我们提出了一个解开神经音频编解码器,利用时域信号的频谱分解,以提高表示的可解释性。实验评估表明,我们的方法超越了国家的最先进的基线重建保真度和感知质量。
摘要:While neural-based models have led to significant advancements in audio feature extraction, the interpretability of the learned representations remains a critical challenge. To address this, disentanglement techniques have been integrated into discrete neural audio codecs to impose structure on the extracted tokens. However, these approaches often exhibit strong dependencies on specific datasets or task formulations. In this work, we propose a disentangled neural audio codec that leverages spectral decomposition of time-domain signals to enhance representation interpretability. Experimental evaluations demonstrate that our method surpasses a state-of-the-art baseline in both reconstruction fidelity and perceptual quality.
【25】Lightweight and Generalizable Acoustic Scene Representations via Contrastive Fine-Tuning and Distillation
标题:通过对比微调和蒸馏实现轻量级且可推广的声学场景表示
链接:https://arxiv.org/abs/2510.03728
摘要:Acoustic scene classification (ASC) models on edge devices typically operate under fixed class assumptions, lacking the transferability needed for real-world applications that require adaptation to new or refined acoustic categories. We propose ContrastASC, which learns generalizable acoustic scene representations by structuring the embedding space to preserve semantic relationships between scenes, enabling adaptation to unseen categories without retraining. Our approach combines supervised contrastive fine-tuning of pre-trained models with contrastive representation distillation to transfer this structured knowledge to compact student models. Our evaluation shows that ContrastASC demonstrates improved few-shot adaptation to unseen categories while maintaining strong closed-set performance.
【26】Audio Forensics Evaluation (SAFE) Challenge
标题:音频取证评估(SAFE)挑战
链接:https://arxiv.org/abs/2510.03387
摘要:The increasing realism of synthetic speech generated by advanced text-to-speech (TTS) models, coupled with post-processing and laundering techniques, presents a significant challenge for audio forensic detection. In this paper, we introduce the SAFE (Synthetic Audio Forensics Evaluation) Challenge, a fully blind evaluation framework designed to benchmark detection models across progressively harder scenarios: raw synthetic speech, processed audio (e.g., compression, resampling), and laundered audio intended to evade forensic analysis. The SAFE challenge consisted of a total of 90 hours of audio and 21,000 audio samples split across 21 different real sources and 17 different TTS models and 3 tasks. We present the challenge, evaluation design and tasks, dataset details, and initial insights into the strengths and limitations of current approaches, offering a foundation for advancing synthetic audio detection research. More information is available at \href{https://stresearch.github.io/SAFE/}{https://stresearch.github.io/SAFE/}.
机器翻译由腾讯交互翻译提供,仅供参考
