微信公众号:arXiv_Daily
cs.SD语音
标题: 使用谱图图像和卷积神经网络进行实时音调/F0检测
链接:https://arxiv.org/abs/2504.06165
摘要:本文提出了一种新的方法来检测F0通过卷积神经网络和图像处理技术,直接估计基音从频谱图像。我们的新方法证明了一个非常好的检测精度,预测的音高轮廓的92%有很强或中等的相关性,真正的音高轮廓。此外,我们的新方法与其他最先进的CNN方法之间的实验比较表明,我们的方法可以在各种信噪比条件下将检测率提高约5%。
摘要:This paper presents a novel approach to detect F0 through Convolutional Neural Networks and image processing techniques to directly estimate pitch from spectrogram images. Our new approach demonstrates a very good detection accuracy; a total of 92% of predicted pitch contours have strong or moderate correlations to the true pitch contours. Furthermore, the experimental comparison between our new approach and other state-of-the-art CNN methods reveals that our approach can enhance the detection rate by approximately 5% across various Signal-to-Noise Ratio conditions.
【2】 Réduire le bruit grâce à la réalité augmentée sonore -- Auditory Concealer
标题: Réduire le bruit gâce à la réalité augmentée sonore -- Auditory Concealer链接:https://arxiv.org/abs/2504.05847
备注:57 pages, in French language, 24 figures
摘要:本报告介绍了在声学/音乐研究与协调研究所(IRCAM)的音乐和声音科学与技术(STMS)实验室的声音感知和设计团队中完成的22周实习工作。作为增强现实降噪(ReNAR)项目启动的一部分,该项目旨在创建一种工具,以实时减少室内环境中令人不快或讨厌的声音的认知影响;进行了一项初步研究,以验证一种名为隐蔽器的新掩蔽方法的可行性和有效性。主要的假设是,隐藏的方法可以提供更好的结果比一个面具的方法在感知愉快。两个噪声源(通风)和五个掩蔽的声音(水声)的混合使用这两种方法在不同的水平产生。对这些混合物的感知愉悦度的评估表明,掩蔽方法仍然比隐蔽方法更有效,无论使用的噪声源、水声或水平如何。
摘要:This report presents the work done over 22 weeks of internship within the Sound Perception and Design team of the Sciences and Technologies of Music and Sound (STMS) laboratory at the Institute for Research and Coordination in Acoustics/Music (IRCAM). As part of the launch of the project Reducing Noise with Augmented Reality (ReNAR); which aims to create a tool to reduce in real-time the cognitive impact of sounds perceived as unpleasant or annoying in indoor environments; an initial study was conducted to validate the feasibility and effectiveness of a new masking approach called concealer. The main hypothesis is that the concealer approach could provide better results than a masker approach in terms of perceived pleasantness. Mixtures of two noise sources (ventilation) and five masking sounds (water sounds) were generated using both approaches at various levels. The evaluation of the perceived pleasantness of these mixtures showed that the masker approach remains more effective than the concealer approach, regardless of the noise source, water sound, or level used.
【3】 AVENet: Disentangling Features by Approximating Average Features for Voice Conversion
标题: AVENet:通过逼近平均特征来解开特征以进行语音转换链接:https://arxiv.org/abs/2504.05833
备注:Accepted by ICME 2025
摘要:语音转换(VC)在特征分解方面取得了进展,但在音色和内容信息的平衡上仍然存在困难。本文评估了语音转换中常用的预训练模型特征,并提出了一种创新的方法来解开语音特征表示。具体来说,我们首先提出了一个理想的内容功能,称为平均功能,这是通过平均帧级对齐并行语音(FAPS)数据内的功能计算。为了生成FAPS数据,我们利用了一种技术,涉及冻结的持续时间预测在文本到语音系统和操纵扬声器嵌入。为了在传统的VC数据集上拟合平均特征,我们设计了AVENet,将特征作为输入并生成紧密匹配的平均特征。在VC系统中对AVENet提取特征的性能进行了实验。实验结果表明,该方法优于现有的多种语音特征提取方法。这些发现证实了我们的解纠缠方法的有效性。
摘要:Voice conversion (VC) has made progress in feature disentanglement, but it is still difficult to balance timbre and content information. This paper evaluates the pre-trained model features commonly used in voice conversion, and proposes an innovative method for disentangling speech feature representations. Specifically, we first propose an ideal content feature, referred to as the average feature, which is calculated by averaging the features within frame-level aligned parallel speech (FAPS) data. For generating FAPS data, we utilize a technique that involves freezing the duration predictor in a Text-to-Speech system and manipulating speaker embedding. To fit the average feature on traditional VC datasets, we then design the AVENet to take features as input and generate closely matching average features. Experiments are conducted on the performance of AVENet-extracted features within a VC system. The experimental results demonstrate its superiority over multiple current speech feature disentangling methods. These findings affirm the effectiveness of our disentanglement approach.
【4】 Mass-Spring Models for Passive Keyword Spotting: A Springtronics Approach
标题: 被动关键词定位的质量弹簧模型:Springtronics方法链接:https://arxiv.org/abs/2504.05802
备注:14 pages, 8 figures
摘要:机械系统在计算历史中发挥了基础性作用,并且由于其独特的特性(如低阻尼和无需转导即可处理机械信号的能力)而重新引起人们的兴趣。然而,最近的努力主要集中在基本的计算,实现在系统的基础上预定义的水库,或在周期性的系统,如阵列的屈曲梁。在这里,我们用数字演示了一个被动的机械系统--以非线性质量弹簧模型的形式--它解决了语音信号中关键词定位的现实基准。该模型是组织在一个层次结构相结合的特征提取和连续时间卷积,与每个单独的阶段量身定制所考虑的质量弹簧系统的物理。对于计算中的每一步,通过组合一小组低阶多项式势来设计子系统。这些势能是连接质量网络的基本组成部分。类似于电子电路设计,其中复杂的功能电路是通过将基本组件组合到分层设计中来构建的,我们将这种框架称为springtronics。我们推出了具有数百个自由度的springtronic系统,实现了与现有的sub-mW电子系统相当的语音分类精度。
摘要:Mechanical systems played a foundational role in computing history, and have regained interest due to their unique properties, such as low damping and the ability to process mechanical signals without transduction. However, recent efforts have primarily focused on elementary computations, implemented in systems based on pre-defined reservoirs, or in periodic systems such as arrays of buckling beams. Here, we numerically demonstrate a passive mechanical system -- in the form of a nonlinear mass-spring model -- that tackles a real-world benchmark for keyword spotting in speech signals. The model is organized in a hierarchical architecture combining feature extraction and continuous-time convolution, with each individual stage tailored to the physics of the considered mass-spring systems. For each step in the computation, a subsystem is designed by combining a small set of low-order polynomial potentials. These potentials act as fundamental components that interconnect a network of masses. In analogy to electronic circuit design, where complex functional circuits are constructed by combining basic components into hierarchical designs, we refer to this framework as springtronics. We introduce springtronic systems with hundreds of degrees of freedom, achieving speech classification accuracy comparable to existing sub-mW electronic systems.
【5】 STAGE: Stemmed Accompaniment Generation through Prefix-Based Conditioning
标题: 阶段:通过基于前置的条件反射产生带茎伴奏链接:https://arxiv.org/abs/2504.05690
摘要:None
摘要:Recent advances in generative models have made it possible to create high-quality, coherent music, with some systems delivering production-level output.Yet, most existing models focus solely on generating music from scratch, limiting their usefulness for musicians who want to integrate such models into a human, iterative composition workflow.In this paper we introduce STAGE, our STemmed Accompaniment GEneration model, fine-tuned from the state-of-the-art MusicGen to generate single-stem instrumental accompaniments conditioned on a given mixture. Inspired by instruction-tuning methods for language models, we extend the transformer's embedding matrix with a context token, enabling the model to attend to a musical context through prefix-based conditioning.Compared to the baselines, STAGE yields accompaniments that exhibit stronger coherence with the input mixture, higher audio quality, and closer alignment with textual prompts.Moreover, by conditioning on a metronome-like track, our framework naturally supports tempo-constrained generation, achieving state-of-the-art alignment with the target rhythmic structure--all without requiring any additional tempo-specific module.As a result, STAGE offers a practical, versatile tool for interactive music creation that can be readily adopted by musicians in real-world workflows.
【6】 kNN-SVC: Robust Zero-Shot Singing Voice Conversion with Additive Synthesis and Concatenation Smoothness Optimization
标题: kNN-SRC:具有加法合成和级联平滑度优化的稳健Zero-Shot歌唱语音转换链接:https://arxiv.org/abs/2504.05686
备注:5 pages, 6 figures, 1 table, Proceedings of the International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025
摘要:鲁棒性是zero-shot歌唱声转换(SVC)的关键。本文介绍了两种新的方法,以加强鲁棒性的kNN-VC框架的SVC。首先,kNN-VC的核心表示WavLM缺乏谐波强调,导致声音沉闷和振铃伪像。为了解决这个问题,我们利用WavLM,音高轮廓和频谱图之间的双射来执行加法合成,将所得波形集成到模型中以缓解这些问题。其次,kNN-VC忽略了级联平滑,这是SVC中的关键感知因素。为了增强平滑性,我们提出了一种新的距离度量,可以过滤掉不合适的kNN候选项,并在推理过程中优化候选项的求和权重。虽然我们的技术是建立在kNN-VC框架的实现方便,他们广泛适用于一般的级联神经合成模型。实验结果验证了这些修改在实现鲁棒SVC方面的有效性。Demo:http://knnsvc.com Code:https://github.com/SmoothKen/knn-svc
摘要:Robustness is critical in zero-shot singing voice conversion (SVC). This paper introduces two novel methods to strengthen the robustness of the kNN-VC framework for SVC. First, kNN-VC's core representation, WavLM, lacks harmonic emphasis, resulting in dull sounds and ringing artifacts. To address this, we leverage the bijection between WavLM, pitch contours, and spectrograms to perform additive synthesis, integrating the resulting waveform into the model to mitigate these issues. Second, kNN-VC overlooks concatenative smoothness, a key perceptual factor in SVC. To enhance smoothness, we propose a new distance metric that filters out unsuitable kNN candidates and optimize the summing weights of the candidates during inference. Although our techniques are built on the kNN-VC framework for implementation convenience, they are broadly applicable to general concatenative neural synthesis models. Experimental results validate the effectiveness of these modifications in achieving robust SVC. Demo: http://knnsvc.com Code: https://github.com/SmoothKen/knn-svc
【7】 TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis
标题: TARO:同步视频音频合成的时步自适应表示与发生感知条件对准链接:https://arxiv.org/abs/2504.05684
备注:10 pages, 6 figures
摘要:本文介绍了时间步长自适应表示对齐与起始感知条件(TARO),高保真和时间相干的视频到音频合成的新框架。基于流的变压器,提供稳定的训练和连续的转换,以增强同步和音频质量,TARO引入了两个关键的创新:(1)时步自适应表示对齐(TRA),其通过基于噪声时间表调整对齐强度来动态地对齐潜在表示,确保平滑演进和改善的保真度,以及(2)发作感知调节(OAC),其集成了用作音频相关视觉时刻的清晰事件驱动标记的起始线索,以增强与动态视觉事件的同步。在VGGSound和Landscape数据集上的大量实验表明,TARO优于现有方法,实现了相对53%的Frechet距离(FD),29%的Frechet音频距离(FAD)和97.19%的对齐精度,突出了其卓越的音频质量和同步精度。
摘要:This paper introduces Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning (TARO), a novel framework for high-fidelity and temporally coherent video-to-audio synthesis. Built upon flow-based transformers, which offer stable training and continuous transformations for enhanced synchronization and audio quality, TARO introduces two key innovations: (1) Timestep-Adaptive Representation Alignment (TRA), which dynamically aligns latent representations by adjusting alignment strength based on the noise schedule, ensuring smooth evolution and improved fidelity, and (2) Onset-Aware Conditioning (OAC), which integrates onset cues that serve as sharp event-driven markers of audio-relevant visual moments to enhance synchronization with dynamic visual events. Extensive experiments on the VGGSound and Landscape datasets demonstrate that TARO outperforms prior methods, achieving relatively 53\% lower Frechet Distance (FD), 29% lower Frechet Audio Distance (FAD), and a 97.19% Alignment Accuracy, highlighting its superior audio quality and synchronization precision.
【8】 Contrastive Decoupled Representation Learning and Regularization for Speech-Preserving Facial Expression Manipulation
标题: 基于对比解耦表示学习和正则化的语音保持人脸表情操作链接:https://arxiv.org/abs/2504.05672
摘要:语音保留面部表情操作(SPFEM)旨在修改说话的头部以显示特定的参考情感,同时保留源口语内容的嘴部动画。因此,存在于参考和源输入中的情感和内容信息可以为SPFEM模型提供直接和准确的监督信号。然而,在谈话过程中,这些元素的内在交织对它们作为监督信号的有效性提出了挑战。在这项工作中,我们建议学习内容和情感先验知识作为对比学习的指导,通过创新的对比解耦表示学习(CDRL)算法来学习解耦的内容和情感表示。具体地,对比内容表示学习(CCRL)模块被设计为学习主要包含内容信息的音频特征,作为内容先验,以指导从源输入学习内容表示。同时,提出了对比情绪表征学习(CERL)模块,利用预先训练的视觉语言模型来学习情绪先验,然后用于指导从参考输入中学习情绪表征。我们进一步引入情感感知和情感增强的对比学习,分别训练CCRL和CERL模块,确保学习情感独立的内容表示和内容独立的情感表示。在SPFEM模型训练期间,解耦的内容和情感表示用于监督生成过程,确保更准确的情感操作以及音频-嘴唇同步。大量的实验和各种基准测试结果表明,该算法的有效性。
摘要:Speech-preserving facial expression manipulation (SPFEM) aims to modify a talking head to display a specific reference emotion while preserving the mouth animation of source spoken contents. Thus, emotion and content information existing in reference and source inputs can provide direct and accurate supervision signals for SPFEM models. However, the intrinsic intertwining of these elements during the talking process poses challenges to their effectiveness as supervisory signals. In this work, we propose to learn content and emotion priors as guidance augmented with contrastive learning to learn decoupled content and emotion representation via an innovative Contrastive Decoupled Representation Learning (CDRL) algorithm. Specifically, a Contrastive Content Representation Learning (CCRL) module is designed to learn audio feature, which primarily contains content information, as content priors to guide learning content representation from the source input. Meanwhile, a Contrastive Emotion Representation Learning (CERL) module is proposed to make use of a pre-trained visual-language model to learn emotion prior, which is then used to guide learning emotion representation from the reference input. We further introduce emotion-aware and emotion-augmented contrastive learning to train CCRL and CERL modules, respectively, ensuring learning emotion-independent content representation and content-independent emotion representation. During SPFEM model training, the decoupled content and emotion representations are used to supervise the generation process, ensuring more accurate emotion manipulation together with audio-lip synchronization. Extensive experiments and evaluations on various benchmarks show the effectiveness of the proposed algorithm.
【9】 SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding
标题: SoundVista:通过视觉-声学绑定进行新颖视角环境声音合成链接:https://arxiv.org/abs/2504.05576
备注:Highlight Accepted to CVPR 2025
摘要:我们介绍SoundVista,一种方法来产生一个任意场景的环境声音在新的观点。给定从稀疏分布的麦克风预先获取的场景录音,SoundVista可以从看不见的目标视点合成该场景的声音。该方法使用有限数量的已知记录来学习底层声学传递函数,该底层声学传递函数将在分布式麦克风处获取的信号与目标视点处的信号相关联。与现有的作品,我们的方法不需要约束或声源的细节先验知识。此外,我们的方法有效地适应不同的房间布局,参考麦克风配置和看不见的环境。为了实现这一点,我们引入了一个视觉-声学绑定模块,该模块可以从全景RGB和深度数据中学习与局部声学属性相关的视觉嵌入。我们首先利用这些嵌入来优化参考麦克风在任何给定场景中的位置。在合成过程中,我们利用从参考位置提取的多个嵌入来获得自适应权重,以目标视点为条件。我们对公开数据和现实世界的设置进行了基准测试。我们展示了对现有方法的显着改进。
摘要:We introduce SoundVista, a method to generate the ambient sound of an arbitrary scene at novel viewpoints. Given a pre-acquired recording of the scene from sparsely distributed microphones, SoundVista can synthesize the sound of that scene from an unseen target viewpoint. The method learns the underlying acoustic transfer function that relates the signals acquired at the distributed microphones to the signal at the target viewpoint, using a limited number of known recordings. Unlike existing works, our method does not require constraints or prior knowledge of sound source details. Moreover, our method efficiently adapts to diverse room layouts, reference microphone configurations and unseen environments. To enable this, we introduce a visual-acoustic binding module that learns visual embeddings linked with local acoustic properties from panoramic RGB and depth data. We first leverage these embeddings to optimize the placement of reference microphones in any given scene. During synthesis, we leverage multiple embeddings extracted from reference locations to get adaptive weights for their contribution, conditioned on target viewpoint. We benchmark the task on both publicly available data and real-world settings. We demonstrate significant improvements over existing methods.
【10】 Exploring Local Interpretable Model-Agnostic Explanations for Speech Emotion Recognition with Distribution-Shift
标题: 探索具有分布转移的语音情感识别的局部可解释模型不可知解释链接:https://arxiv.org/abs/2504.05368
备注:Published in the proceedings of ICASSP 2025
摘要:我们介绍了一个版本的本地可解释的模型不可知论解释(LIME)的黑盒语音情感识别(SER)模型。据我们所知,这是第一次尝试在SER中应用LIME。LIME生成高级可解释的解释,并确定哪些特定频率范围对确定情绪状态最有影响。该方法有助于解释复杂的高维嵌入,例如由端到端语音模型生成的嵌入。我们在三个情感语音数据集上,使用在手工制作的声学特征和Wav2Vec 2.0嵌入上训练的分类器,从定性、定量和统计上评估了Wav2Vec。我们发现,与分布变化的数据集相比,在不同的模型中,EMLIME表现出更强的鲁棒性,突出了它在数据集中SER任务中更一致的解释的潜力。
摘要:We introduce EmoLIME, a version of local interpretable model-agnostic explanations (LIME) for black-box Speech Emotion Recognition (SER) models. To the best of our knowledge, this is the first attempt to apply LIME in SER. EmoLIME generates high-level interpretable explanations and identifies which specific frequency ranges are most influential in determining emotional states. The approach aids in interpreting complex, high-dimensional embeddings such as those generated by end-to-end speech models. We evaluate EmoLIME, qualitatively, quantitatively, and statistically, across three emotional speech datasets, using classifiers trained on both hand-crafted acoustic features and Wav2Vec 2.0 embeddings. We find that EmoLIME exhibits stronger robustness across different models than across datasets with distribution shifts, highlighting its potential for more consistent explanations in SER tasks within a dataset.
【11】 Of All StrIPEs: Investigating Structure-informed Positional Encoding for Efficient Music Generation
标题: 所有StrIPE:研究结构信息位置编码以实现高效音乐生成链接:https://arxiv.org/abs/2504.05364
摘要:虽然音乐对于像Transformers这样的生成模型来说仍然是一个具有挑战性的领域,但一种双管齐下的方法最近被证明是成功的:将音乐相关的结构信息插入到位置编码(PE)模块中,并使用基于随机傅立叶特征(RFF)的核近似技术将计算成本从二次降低到线性。然而,不清楚这种基于RFF的高效PE与基于旋转矩阵的PE(诸如旋转位置编码(RoPE))相比如何。在本文中,我们提出了一个统一的框架的基础上,内核的方法来分析这两个家庭的有效PE。我们使用这个框架来开发一种新的PE方法称为RoPEPool,能够从时间序列中提取因果关系。使用RFF为基础的PE和旋转为基础的PE,我们演示了如何看似不同的PE可以共同研究,考虑他们诱导的内容上下文的相互作用。对于实证验证,我们使用一个象征性的音乐生成任务,即旋律协调。我们表明,RoPEPool,结合高信息量的结构先验,优于所有的方法。
摘要:While music remains a challenging domain for generative models like Transformers, a two-pronged approach has recently proved successful: inserting musically-relevant structural information into the positional encoding (PE) module and using kernel approximation techniques based on Random Fourier Features (RFF) to lower the computational cost from quadratic to linear. Yet, it is not clear how such RFF-based efficient PEs compare with those based on rotation matrices, such as Rotary Positional Encoding (RoPE). In this paper, we present a unified framework based on kernel methods to analyze both families of efficient PEs. We use this framework to develop a novel PE method called RoPEPool, capable of extracting causal relationships from temporal sequences. Using RFF-based PEs and rotation-based PEs, we demonstrate how seemingly disparate PEs can be jointly studied by considering the content-context interactions they induce. For empirical validation, we use a symbolic music generation task, namely, melody harmonization. We show that RoPEPool, combined with highly-informative structural priors, outperforms all methods.
【12】 Nes2Net: A Lightweight Nested Architecture for Foundation Model Driven Speech Anti-spoofing
标题: Nes 2Net:基础模型驱动语音反欺骗的轻量级嵌套架构链接:https://arxiv.org/abs/2504.05657
备注:This manuscript has been submitted for peer review
摘要:语音基础模型通过提供卓越的表示能力,显著地推进了各种与语音相关的任务。然而,它们的高维输出特征通常与下游任务模型不匹配,后者通常需要低维输入。一种常见的解决方案是应用降维(DR)层,但这种方法增加了参数开销、计算成本和丢失有价值信息的风险。为了解决这些问题,我们提出了嵌套Res 2Net(Nes 2Net),一个轻量级的后端架构,旨在直接处理高维特征,而无需DR层。嵌套结构增强了多尺度特征提取,改善了特征交互,并保留了高维信息。我们首先在CtrSVDD上验证了Nes 2Net,CtrSVDD是一个歌唱声音深度伪造检测数据集,并报告了22%的性能提升和87%的后端计算成本降低。此外,在ASVspoof 2021、ASVspoof 5、PartialSpoof和In-the-Wild四个不同的数据集上进行了广泛的测试,涵盖了完全欺骗的语音、对抗性攻击、部分欺骗和真实场景,始终突出了Nes 2Net卓越的鲁棒性和泛化能力。代码包和预训练模型可在https://github.com/Liu-Tianchi/Nes2Net上获得。
摘要:Speech foundation models have significantly advanced various speech-related tasks by providing exceptional representation capabilities. However, their high-dimensional output features often create a mismatch with downstream task models, which typically require lower-dimensional inputs. A common solution is to apply a dimensionality reduction (DR) layer, but this approach increases parameter overhead, computational costs, and risks losing valuable information. To address these issues, we propose Nested Res2Net (Nes2Net), a lightweight back-end architecture designed to directly process high-dimensional features without DR layers. The nested structure enhances multi-scale feature extraction, improves feature interaction, and preserves high-dimensional information. We first validate Nes2Net on CtrSVDD, a singing voice deepfake detection dataset, and report a 22% performance improvement and an 87% back-end computational cost reduction over the state-of-the-art baseline. Additionally, extensive testing across four diverse datasets: ASVspoof 2021, ASVspoof 5, PartialSpoof, and In-the-Wild, covering fully spoofed speech, adversarial attacks, partial spoofing, and real-world scenarios, consistently highlights Nes2Net's superior robustness and generalization capabilities. The code package and pre-trained models are available at https://github.com/Liu-Tianchi/Nes2Net.
标题: Nes 2Net:基础模型驱动语音反欺骗的轻量级嵌套架构
链接:https://arxiv.org/abs/2504.05657
备注:This manuscript has been submitted for peer review
摘要:语音基础模型通过提供卓越的表示能力,显著地推进了各种与语音相关的任务。然而,它们的高维输出特征通常与下游任务模型不匹配,后者通常需要低维输入。一种常见的解决方案是应用降维(DR)层,但这种方法增加了参数开销、计算成本和丢失有价值信息的风险。为了解决这些问题,我们提出了嵌套Res 2Net(Nes 2Net),一个轻量级的后端架构,旨在直接处理高维特征,而无需DR层。嵌套结构增强了多尺度特征提取,改善了特征交互,并保留了高维信息。我们首先在CtrSVDD上验证了Nes 2Net,CtrSVDD是一个歌唱声音深度伪造检测数据集,并报告了22%的性能提升和87%的后端计算成本降低。此外,在ASVspoof 2021、ASVspoof 5、PartialSpoof和In-the-Wild四个不同的数据集上进行了广泛的测试,涵盖了完全欺骗的语音、对抗性攻击、部分欺骗和真实场景,始终突出了Nes 2Net卓越的鲁棒性和泛化能力。代码包和预训练模型可在https://github.com/Liu-Tianchi/Nes2Net上获得。
摘要:Speech foundation models have significantly advanced various speech-related tasks by providing exceptional representation capabilities. However, their high-dimensional output features often create a mismatch with downstream task models, which typically require lower-dimensional inputs. A common solution is to apply a dimensionality reduction (DR) layer, but this approach increases parameter overhead, computational costs, and risks losing valuable information. To address these issues, we propose Nested Res2Net (Nes2Net), a lightweight back-end architecture designed to directly process high-dimensional features without DR layers. The nested structure enhances multi-scale feature extraction, improves feature interaction, and preserves high-dimensional information. We first validate Nes2Net on CtrSVDD, a singing voice deepfake detection dataset, and report a 22% performance improvement and an 87% back-end computational cost reduction over the state-of-the-art baseline. Additionally, extensive testing across four diverse datasets: ASVspoof 2021, ASVspoof 5, PartialSpoof, and In-the-Wild, covering fully spoofed speech, adversarial attacks, partial spoofing, and real-world scenarios, consistently highlights Nes2Net's superior robustness and generalization capabilities. The code package and pre-trained models are available at https://github.com/Liu-Tianchi/Nes2Net.
【2】 Réduire le bruit grâce à la réalité augmentée sonore -- Auditory Concealer
标题: Réduire le bruit gâce à la réalité augmentée sonore -- Auditory Concealer链接:https://arxiv.org/abs/2504.05847
备注:57 pages, in French language, 24 figures
摘要:本报告介绍了在声学/音乐研究与协调研究所(IRCAM)的音乐和声音科学与技术(STMS)实验室的声音感知和设计团队中完成的22周实习工作。作为增强现实降噪(ReNAR)项目启动的一部分,该项目旨在创建一种工具,以实时减少室内环境中令人不快或讨厌的声音的认知影响;进行了一项初步研究,以验证一种名为隐蔽器的新掩蔽方法的可行性和有效性。主要的假设是,隐藏的方法可以提供更好的结果比一个面具的方法在感知愉快。两个噪声源(通风)和五个掩蔽的声音(水声)的混合使用这两种方法在不同的水平产生。对这些混合物的感知愉悦度的评估表明,掩蔽方法仍然比隐蔽方法更有效,无论使用的噪声源、水声或水平如何。
摘要:This report presents the work done over 22 weeks of internship within the Sound Perception and Design team of the Sciences and Technologies of Music and Sound (STMS) laboratory at the Institute for Research and Coordination in Acoustics/Music (IRCAM). As part of the launch of the project Reducing Noise with Augmented Reality (ReNAR); which aims to create a tool to reduce in real-time the cognitive impact of sounds perceived as unpleasant or annoying in indoor environments; an initial study was conducted to validate the feasibility and effectiveness of a new masking approach called concealer. The main hypothesis is that the concealer approach could provide better results than a masker approach in terms of perceived pleasantness. Mixtures of two noise sources (ventilation) and five masking sounds (water sounds) were generated using both approaches at various levels. The evaluation of the perceived pleasantness of these mixtures showed that the masker approach remains more effective than the concealer approach, regardless of the noise source, water sound, or level used.
【3】 Exploring Local Interpretable Model-Agnostic Explanations for Speech Emotion Recognition with Distribution-Shift
标题: 探索具有分布转移的语音情感识别的局部可解释模型不可知解释链接:https://arxiv.org/abs/2504.05368
备注:Published in the proceedings of ICASSP 2025
摘要:我们介绍了一个版本的本地可解释的模型不可知论解释(LIME)的黑盒语音情感识别(SER)模型。据我们所知,这是第一次尝试在SER中应用LIME。LIME生成高级可解释的解释,并确定哪些特定频率范围对确定情绪状态最有影响。该方法有助于解释复杂的高维嵌入,例如由端到端语音模型生成的嵌入。我们在三个情感语音数据集上,使用在手工制作的声学特征和Wav2Vec 2.0嵌入上训练的分类器,从定性、定量和统计上评估了Wav2Vec。我们发现,与分布变化的数据集相比,在不同的模型中,EMLIME表现出更强的鲁棒性,突出了它在数据集中SER任务中更一致的解释的潜力。
摘要:We introduce EmoLIME, a version of local interpretable model-agnostic explanations (LIME) for black-box Speech Emotion Recognition (SER) models. To the best of our knowledge, this is the first attempt to apply LIME in SER. EmoLIME generates high-level interpretable explanations and identifies which specific frequency ranges are most influential in determining emotional states. The approach aids in interpreting complex, high-dimensional embeddings such as those generated by end-to-end speech models. We evaluate EmoLIME, qualitatively, quantitatively, and statistically, across three emotional speech datasets, using classifiers trained on both hand-crafted acoustic features and Wav2Vec 2.0 embeddings. We find that EmoLIME exhibits stronger robustness across different models than across datasets with distribution shifts, highlighting its potential for more consistent explanations in SER tasks within a dataset.
机器翻译由腾讯交互翻译提供,仅供参考
