今日论文合集:cs.SD语音11篇,eess.AS音频处理13篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音



【1】Exploring Gender Bias in Alzheimer's Disease Detection: Insights from Mandarin and Greek Speech Perception
标题:探索阿尔茨海默病检测中的性别偏见:来自普通话和希腊语言语感知的见解
链接:http://arxiv.org/pdf/2507.12356v1
作者:uanchao Li, Rui Feng, XinRan Han, Yin-Long Liu, Yuwei Yang, Zude Zhu, Jiahong Yuan

备注:12 pages, 5 figures, conference or other essential info

摘要:受性别之间基本发声差异的影响,言语感知任务中广泛存在性别偏见。本研究揭示了阿尔茨海默病(AD)言语感知中的性别偏见。在一个感知实验中,涉及16名中国听众评估中国和希腊的讲话,我们发现,男性的讲话更频繁地被确定为AD,这种偏见在中国的讲话特别明显。声学分析表明,男性言语中的闪烁值与AD感知显著相关,而言语部分与AD识别呈显著负相关。虽然语言没有显着的影响,AD的看法,我们的研究结果强调了关键作用的性别偏见在AD言语感知。这项工作强调了在开发AD检测模型时解决性别偏见的必要性,并呼吁进一步研究以验证模型在不同语言背景下的性能。
摘要:Gender bias has been widely observed in speech perception tasks, influenced by the fundamental voicing differences between genders. This study reveals a gender bias in the perception of Alzheimer's Disease (AD) speech. In a perception experiment involving 16 Chinese listeners evaluating both Chinese and Greek speech, we identified that male speech was more frequently identified as AD, with this bias being particularly pronounced in Chinese speech. Acoustic analysis showed that shimmer values in male speech were significantly associated with AD perception, while speech portion exhibited a significant negative correlation with AD identification. Although language did not have a significant impact on AD perception, our findings underscore the critical role of gender bias in AD speech perception. This work highlights the necessity of addressing gender bias when developing AD detection models and calls for further research to validate model performance across different linguistic contexts.


【2】Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations
标题:量化更多,损失更少:从剩余量化的语音表示中自回归生成
链接:http://arxiv.org/pdf/2507.12197v1

作者:n, Xiaoyang Hao, Keming Chen, Weibo Xiong, Jun He, Ruonan Zhang, Junjie Cao, Yue Liu, Bowen Li, Dongrui Zhang, Hui Xia, Huilei Fu, Kai Jia, Kaixuan Guo, Mingli Jin, Qingyun Meng, Ruidong Ma, Ruiqian Fang, Shaotong Guo, Xuhui Li, Yang Xiang, Ying Zhang, Yulong Liu, Yunfeng Li, Yuyi Zhang, Yuze Zhou, Zhen Wang, Zhaowen Chen
摘要:在离散建模范式下,文语转换(TTS)合成有了新的进展。现有的自回归方法往往依赖于单码本表示,这遭受显着的信息丢失。即使使用诸如流匹配的事后细化技术,这些方法也无法恢复细粒度的细节(例如,韵律细微差别、特定于说话者的音色),特别是在具有挑战性的场景中,如歌声或音乐合成。我们提出了QTTS,一个新的TTS框架建立在我们的新的音频编解码器,QDAC。QDAC的核心创新在于其基于ASR的自回归网络与GAN的端到端训练,该网络实现了可扩展的、近无损压缩的出色语义特征解纠缠。QTTS使用两种创新策略对这些离散代码进行建模:分层并行架构,该架构使用双AR结构对码本间依赖关系进行建模,以实现更高质量的合成,以及延迟多头方法,该方法采用具有固定延迟的并行预测来加快推理速度。我们的实验表明,该框架实现了更高的合成质量和更好地保留表达内容相比,基线。这表明,通过多码本建模来放大压缩是高保真、通用语音和音频生成的一个有前途的方向。
摘要:Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, which suffer from significant information loss. Even with post-hoc refinement techniques such as flow matching, these methods fail to recover fine-grained details (e.g., prosodic nuances, speaker-specific timbres), especially in challenging scenarios like singing voice or music synthesis. We propose QTTS, a novel TTS framework built upon our new audio codec, QDAC. The core innovation of QDAC lies in its end-to-end training of an ASR-based auto-regressive network with a GAN, which achieves superior semantic feature disentanglement for scalable, near-lossless compression. QTTS models these discrete codes using two innovative strategies: the Hierarchical Parallel architecture, which uses a dual-AR structure to model inter-codebook dependencies for higher-quality synthesis, and the Delay Multihead approach, which employs parallelized prediction with a fixed delay to accelerate inference speed. Our experiments demonstrate that the proposed framework achieves higher synthesis quality and better preserves expressive content compared to baseline. This suggests that scaling up compression via multi-codebook modeling is a promising direction for high-fidelity, general-purpose speech and audio generation.


【3】RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
标题:RUMAA:重复感知的统一音乐音频分析,用于分数性能对齐,转录和错误检测
链接:http://arxiv.org/pdf/2507.12175v1

作者:Chang, Simon Dixon, Emmanouil Benetos 备注:Accepted to WASPAA 2025
摘要:本研究介绍了RUMAA,一个基于transformer的音乐性能分析框架,它以近端到端的方式统一了分数到性能对齐,分数通知转录和错误检测。与之前单独解决这些任务的方法不同,RUMAA使用预训练的配乐和音频编码器以及通过代理任务捕获任务相互依赖性的新型三流解码器将它们集成在一起。它将人类可读的MusicXML乐谱与重复符号对齐到全长表演音频,克服了传统的基于MIDI的方法,这些方法依赖于具有预先指定的重复结构的手动展开的乐谱数据。RUMAA在非重复乐谱上匹配最先进的对齐方法,并在公共钢琴音乐数据集中的重复乐谱上优于它们,同时还提供了有希望的转录和错误检测结果。
摘要:This study introduces RUMAA, a transformer-based framework for music performance analysis that unifies score-to-performance alignment, score-informed transcription, and mistake detection in a near end-to-end manner. Unlike prior methods addressing these tasks separately, RUMAA integrates them using pre-trained score and audio encoders and a novel tri-stream decoder capturing task interdependencies through proxy tasks. It aligns human-readable MusicXML scores with repeat symbols to full-length performance audio, overcoming traditional MIDI-based methods that rely on manually unfolded score-MIDI data with pre-specified repeat structures. RUMAA matches state-of-the-art alignment methods on non-repeated scores and outperforms them on scores with repeats in a public piano music dataset, while also delivering promising transcription and mistake detection results.


【4】Room Impulse Response Generation Conditioned on Acoustic Parameters
标题:以声学参数为条件的房间脉冲响应生成
链接:http://arxiv.org/pdf/2507.12136v1

作者:ellano, Chunghsin Yeh, Gautam Bhattacharya, Daniel Arteaga
备注:4+1 pages, 2 figures; accepted in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA 2025)
摘要:使用深度神经网络生成房间脉冲响应(RIR)由于其在虚拟和增强现实、音频后期制作及相关领域的应用而引起了越来越多的研究兴趣。大多数现有的方法条件生成模型的物理描述的房间,如其大小,形状和表面材料。然而,这种对几何信息的依赖限制了它们在房间布局未知或感知现实主义(空间对听众的声音)比严格的物理准确性更重要的情况下的可用性。在这项研究中,我们提出了一种替代策略:直接在一组RIR声学参数上调节RIR生成。这些参数包括混响时间和直达声混响比的各种测量,宽带和频带。通过指定空间应该如何听起来而不是它应该如何看起来,我们的方法可以实现更灵活和感知驱动的RIR生成。我们探索了自回归和非自回归生成模型在描述音频编解码器域中运行,使用离散令牌序列或连续嵌入。具体来说,我们选择了四个模型进行评估:自回归Transformer,MaskGIT模型,流匹配模型和基于分类器的方法。进行客观和主观评价,以比较这些方法与国家的最先进的替代品。结果表明,所提出的模型匹配或优于最先进的替代方案,MaskGIT模型实现了最佳性能。
摘要:The generation of room impulse responses (RIRs) using deep neural networks has attracted growing research interest due to its applications in virtual and augmented reality, audio postproduction, and related fields. Most existing approaches condition generative models on physical descriptions of a room, such as its size, shape, and surface materials. However, this reliance on geometric information limits their usability in scenarios where the room layout is unknown or when perceptual realism (how a space sounds to a listener) is more important than strict physical accuracy. In this study, we propose an alternative strategy: conditioning RIR generation directly on a set of RIR acoustic parameters. These parameters include various measures of reverberation time and direct sound to reverberation ratio, both broadband and bandwise. By specifying how the space should sound instead of how it should look, our method enables more flexible and perceptually driven RIR generation. We explore both autoregressive and non-autoregressive generative models operating in the Descript Audio Codec domain, using either discrete token sequences or continuous embeddings. Specifically, we have selected four models to evaluate: an autoregressive transformer, the MaskGIT model, a flow matching model, and a classifier-based approach. Objective and subjective evaluations are performed to compare these methods with state-of-the-art alternatives. Results show that the proposed models match or outperform state-of-the-art alternatives, with the MaskGIT model achieving the best performance.


【5】MambaRate: Speech Quality Assessment Across Different Sampling Rates
标题:MambaSpeed:不同采样率的语音质量评估
链接:http://arxiv.org/pdf/2507.12090v1

作者:oulidis, Iakovi Alexiou, Junkwang Oh, Gunu Jho, Inchul Hwang, Pirros Tsiakoulis, Aimilios Chalamandaris

备注:Submitted to ASRU 2025 (AudioMOS Challenge 2025 Track 3)

摘要:我们提出了MambaRate,它预测平均意见分数(MOS)与有限的偏见有关的波形的采样率进行评估。它是专为AudioMOS Challenge 2025的Track 3而设计的,其重点是预测高采样频率下的语音MOS。我们的模型利用了自监督嵌入和选择性状态空间建模。目标评级通过高斯径向基函数(RBF)以连续表示进行编码。挑战的结果基于系统级斯皮尔曼等级相关系数(SRCC)度量。初始MambaRate版本(T16系统)在没有预训练的Few-Shot设置中的性能优于预训练基线(B03)约14%。T16在挑战赛中排名第四,与获胜系统相差约6%。我们在BVCC数据集上提供了额外的结果,以及使用不同表示作为输入的消融,其性能优于初始T16版本。
摘要:We propose MambaRate, which predicts Mean Opinion Scores (MOS) with limited bias regarding the sampling rate of the waveform under evaluation. It is designed for Track 3 of the AudioMOS Challenge 2025, which focuses on predicting MOS for speech in high sampling frequencies. Our model leverages self-supervised embeddings and selective state space modeling. The target ratings are encoded in a continuous representation via Gaussian radial basis functions (RBF). The results of the challenge were based on the system-level Spearman's Rank Correllation Coefficient (SRCC) metric. An initial MambaRate version (T16 system) outperformed the pre-trained baseline (B03) by ~14% in a few-shot setting without pre-training. T16 ranked fourth out of five in the challenge, differing by ~6% from the winning system. We present additional results on the BVCC dataset as well as ablations with different representations as input, which outperform the initial T16 version.


【6】Stereo Sound Event Localization and Detection with Onscreenoffscreen Classification
标题:通过屏上屏外分类进行立体声事件定位和检测
链接:http://arxiv.org/pdf/2507.12042v1
作者:imada, Archontis Politis, Iran R. Roman, Parthasaarathy Sudarsanam, David Diaz-Guerra, Ruchi Pandey, Kengo Uchida, Yuichiro Koyama, Naoya Takahashi, Takashi Shibuya, Shusuke Takahashi, Tuomas Virtanen, Yuki Mitsufuji 备注:5 pages, 2 figures
摘要:本文介绍了DCASE 2025声音事件定位和检测(SELD)挑战赛任务3的目标、数据集、基线和指标。在以前的版本中,挑战赛使用了四声道音频格式的一阶高保真度立体声(FOA)和麦克风阵列。相比之下,今年的挑战研究了具有立体声音频数据的SELD(称为立体声SELD)。此更改将重点从更专业的360度音频和视听场景分析转移到具有有限视野(FOV)的更常见的音频和媒体场景。由于立体声音频数据中固有的角度模糊性,该任务集中于方位角平面(左右轴)中的到达方向(DOA)估计以及距离估计。挑战仍然分为两个轨道:音频和视听,视听轨道引入了一个新的子任务,即有限的FOV所必需的屏幕 屏幕外事件分类。本次挑战赛介绍了DCASE 2025 Task 3 Stereo SELD数据集,其立体声音频和透视视频剪辑是从STARSS 23录音中采样和转换的。基线系统被设计为处理立体声音频和相应的视频帧作为输入。除了典型的SELD事件分类和本地化,它还集成了视听轨道的屏幕 屏幕外分类。评估指标已被修改,以引入一个可识别 屏幕外准确性指标,该指标评估模型识别哪些声源被识别的能力。在实验评估中,基线系统对立体声音频数据表现得相当好。
摘要:This paper presents the objective, dataset, baseline, and metrics of Task 3 of the DCASE2025 Challenge on sound event localization and detection (SELD). In previous editions, the challenge used four-channel audio formats of first-order Ambisonics (FOA) and microphone array. In contrast, this year's challenge investigates SELD with stereo audio data (termed stereo SELD). This change shifts the focus from more specialized 360{ deg} audio and audiovisual scene analysis to more commonplace audio and media scenarios with limited field-of-view (FOV). Due to inherent angular ambiguities in stereo audio data, the task focuses on direction-of-arrival (DOA) estimation in the azimuth plane (left-right axis) along with distance estimation. The challenge remains divided into two tracks: audio-only and audiovisual, with the audiovisual track introducing a new sub-task of onscreen offscreen event classification necessitated by the limited FOV. This challenge introduces the DCASE2025 Task3 Stereo SELD Dataset, whose stereo audio and perspective video clips are sampled and converted from the STARSS23 recordings. The baseline system is designed to process stereo audio and corresponding video frames as inputs. In addition to the typical SELD event classification and localization, it integrates onscreen offscreen classification for the audiovisual track. The evaluation metrics have been modified to introduce an onscreen offscreen accuracy metric, which assesses the models' ability to identify which sound sources are onscreen. In the experimental evaluation, the baseline system performs reasonably well with the stereo audio data.


【7】EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
标题:EME-TTS:解锁语音合成中的强调和情感链接
链接:http://arxiv.org/pdf/2507.12015v1

作者:Leyuan Qu, Jiaxi Hu, Taihao Li
备注:Accepted by INTERSPEECH 2025
摘要:近年来,情感文本到语音(TTS)合成和重点可控的语音合成取得了显着进展。然而,它们之间的相互作用仍有待探讨。我们提出强调满足情感TTS(EME-TTS),一个新的框架,旨在解决两个关键的研究问题:(1)如何有效地利用强调,以提高情感语音的表达能力,(2)如何保持知觉清晰度和稳定性的目标强调不同的情绪。EME-TTS采用具有强调伪标签和基于方差的强调特征的弱监督学习。此外,所提出的强调感知增强(EPE)块增强了情感信号和强调位置之间的相互作用。实验结果表明,EME-TTS,当与大型语言模型相结合的重点位置预测,使更自然的情感语音合成,同时保持稳定和可区分的目标强调跨情绪。可在线获得合成样品。
摘要:In recent years, emotional Text-to-Speech (TTS) synthesis and emphasis-controllable speech synthesis have advanced significantly. However, their interaction remains underexplored. We propose Emphasis Meets Emotion TTS (EME-TTS), a novel framework designed to address two key research questions: (1) how to effectively utilize emphasis to enhance the expressiveness of emotional speech, and (2) how to maintain the perceptual clarity and stability of target emphasis across different emotions. EME-TTS employs weakly supervised learning with emphasis pseudo-labels and variance-based emphasis features. Additionally, the proposed Emphasis Perception Enhancement (EPE) block enhances the interaction between emotional signals and emphasis positions. Experimental results show that EME-TTS, when combined with large language models for emphasis position prediction, enables more natural emotional speech synthesis while preserving stable and distinguishable target emphasis across emotions. Synthesized samples are available on-line.


【8】Schrödinger Bridge Consistency Trajectory Models for Speech Enhancement
标题:用于语音增强的薛定格桥一致性轨迹模型
链接:http://arxiv.org/pdf/2507.11925v1

作者:Nishigori, Koichi Saito, Naoki Murata, Masato Hirano, Shusuke Takahashi, Yuki Mitsufuji
摘要:利用扩散模型进行语音增强是一种很有前途的技术,可以提高含噪语音数据的语音质量。此外,薛定谔桥(SB)最近已被用于基于扩散的SE,以改善语音质量,通过解决的前向过程的端点和反向过程的起点之间的失配。然而,SB仍然表现出缓慢的推理,由于大量的功能评估(NFE)的推理,以获得高质量的结果的必要性。虽然一致性模型(CM)通过采用一致性训练来解决这个问题,该训练使用图像生成领域中预训练模型的蒸馏,但当步骤数量增加时,它不会提高生成质量。一致性轨迹模型(Consistency Trajectory Models,CTM)的出现,不仅加快了推理速度,而且在推理质量和推理速度之间取得了良好的平衡。此外,SoundCTM展示了CTM技术在声音生成领域的适用性。本文将CTM技术应用于SE的Schr odinger桥,提出了Schr odinger桥一致性轨迹模型(SBCTM)。此外,我们引入了一种新的辅助损失,包括感知损失,到原来的CTM的训练框架。因此,SBCTM实现了约16倍的改善,在实时因子(RTF)相比,传统的薛定谔桥的SE。此外,SBCTM中的质量和速度之间的有利权衡允许通过将多步细化限制到1步推理不足的情况来实现时间高效的推理。我们的代码、预训练模型和音频样本可在https: github.com sony sbctm 上获得。
摘要:Speech enhancement (SE) utilizing diffusion models is a promising technology that improves speech quality in noisy speech data. Furthermore, the Schr "odinger bridge (SB) has recently been used in diffusion-based SE to improve speech quality by resolving a mismatch between the endpoint of the forward process and the starting point of the reverse process. However, the SB still exhibits slow inference owing to the necessity of a large number of function evaluations (NFE) for inference to obtain high-quality results. While Consistency Models (CMs) address this issue by employing consistency training that uses distillation from pretrained models in the field of image generation, it does not improve generation quality when the number of steps increases. As a solution to this problem, Consistency Trajectory Models (CTMs) not only accelerate inference speed but also maintain a favorable trade-off between quality and speed. Furthermore, SoundCTM demonstrates the applicability of CTM techniques to the field of sound generation. In this paper, we present Schr "odinger bridge Consistency Trajectory Models (SBCTM) by applying the CTM's technique to the Schr "odinger bridge for SE. Additionally, we introduce a novel auxiliary loss, including a perceptual loss, into the original CTM's training framework. As a result, SBCTM achieves an approximately 16x improvement in the real-time factor (RTF) compared to the conventional Schr "odinger bridge for SE. Furthermore, the favorable trade-off between quality and speed in SBCTM allows for time-efficient inference by limiting multi-step refinement to cases where 1-step inference is insufficient. Our code, pretrained models, and audio samples are available at https: github.com sony sbctm .


【9】A Multimodal Data Fusion Generative Adversarial Network for Real Time Underwater Sound Speed Field Construction
标题:实时水下音速场构建的多模式数据融合生成对抗网络
链接:http://arxiv.org/pdf/2507.11812v1

作者:Yuqiang Huang, Yanan Wu, Tianhe Xu, Junting Wang, Hao Zhang
摘要:声速剖面是影响水声信号传播模式的重要参数,对水声通信的能量效率和水声定位精度有着重要影响。传统上,SSP可以通过匹配场处理(MFP)、压缩感知(CS)和深度学习(DL)方法来获得。然而,现有的方法主要依赖于现场水下声纳观测数据,这对声纳观测系统的部署提出了严格的要求。为了在没有现场水下数据测量的情况下实现给定海域声速分布的高精度估计,提出了一种用于SSP构建的带剩余注意块的多模态数据融合生成对抗网络模型(MDF-RAGAN)。为了提高模型捕捉全局空间特征相关性的能力,我们嵌入了注意机制,并使用残差模块来深度捕捉由SST变化引起的深海声速分布中的小扰动。在真实开放数据集上的实验结果表明,该模型的分类精度优于其他现有方法,误差小于0.3m s。具体而言,MDF-RAGAN不仅比卷积神经网络(CNN)和空间插值(SITP)的性能高出近2倍,而且与平均轮廓相比,均方根误差(RMSE)降低了约65.8%,充分体现了多源融合和跨模态注意对整体轮廓匹配的增强。
摘要:Sound speed profiles (SSPs) are essential parameters underwater that affects the propagation mode of underwater signals and has a critical impact on the energy efficiency of underwater acoustic communication and accuracy of underwater acoustic positioning. Traditionally, SSPs can be obtained by matching field processing (MFP), compressive sensing (CS), and deep learning (DL) methods. However, existing methods mainly rely on on-site underwater sonar observation data, which put forward strict requirements on the deployment of sonar observation systems. To achieve high-precision estimation of sound velocity distribution in a given sea area without on-site underwater data measurement, we propose a multi-modal data-fusion generative adversarial network model with residual attention block (MDF-RAGAN) for SSP construction. To improve the model's ability for capturing global spatial feature correlations, we embedded the attention mechanisms, and use residual modules for deeply capturing small disturbances in the deep ocean sound velocity distribution caused by changes of SST. Experimental results on real open dataset show that the proposed model outperforms other state-of-the-art methods, which achieves an accuracy with an error of less than 0.3m s. Specifically, MDF-RAGAN not only outperforms convolutional neural network (CNN) and spatial interpolation (SITP) by nearly a factor of two, but also achieves about 65.8 % root mean square error (RMSE) reduction compared to mean profile, which fully reflects the enhancement of overall profile matching by multi-source fusion and cross-modal attention.


【10】Towards Scalable AASIST: Refining Graph Attention for Speech Deepfake Detection
标题:迈向可扩展的AASIST:细化语音Deepfake检测的图形注意力
链接:http://arxiv.org/pdf/2507.11777v1

作者:hirev, Daniil Sirota, Aleksandr Smirnov, Kirill Borodin
摘要:语音转换和文本到语音合成的进步使得自动说话人验证(ASV)系统更容易受到欺骗攻击。这项工作探讨了适度的改进AASIST反欺骗架构。它集成了一个冻结的Wav2Vec 2.0编码器,以在有限的数据设置中保留自监督语音表示,使用异构查询投影用标准化的多头注意模块代替原始图形注意块,并用可训练的上下文感知集成层代替启发式帧段融合。在ASVspoof 5语料库上进行评估时,所提出的系统达到了7.6%的等错误率(EER),在相同训练条件下改进了重新实现的AASIST基线。消融实验表明,每个架构变化都有助于整体性能,这表明对已建立的模型进行有针对性的调整可能有助于加强实际场景中的语音深度伪造检测。该代码可在https: github.com KORALLLL AASIST_SCALING上公开获取。
摘要:Advances in voice conversion and text-to-speech synthesis have made automatic speaker verification (ASV) systems more susceptible to spoofing attacks. This work explores modest refinements to the AASIST anti-spoofing architecture. It incorporates a frozen Wav2Vec 2.0 encoder to retain self-supervised speech representations in limited-data settings, substitutes the original graph attention block with a standardized multi-head attention module using heterogeneous query projections, and replaces heuristic frame-segment fusion with a trainable, context-aware integration layer. When evaluated on the ASVspoof 5 corpus, the proposed system reaches a 7.6 % equal error rate (EER), improving on a re-implemented AASIST baseline under the same training conditions. Ablation experiments suggest that each architectural change contributes to the overall performance, indicating that targeted adjustments to established models may help strengthen speech deepfake detection in practical scenarios. The code is publicly available at https: github.com KORALLLL AASIST_SCALING.


【11】Modal Analysis of Multimode Waveguides Based on Large Step Size AdaMax from Far-Field Amplitudes
标题:基于远场放大器大步进AdaMax的多模光路模式分析
链接:http://arxiv.org/pdf/2507.12299v1

作者:Li, Dongting Huang, Minhui Xiong, Mingzhi Li
摘要:优化多模波导的性能取决于模态分析,然而,目前的方法主要集中在模态功率分布,并受实验硬件和条件的限制,表现出低精度,适应性差,计算成本高。在这项工作中,在功率归一化约束下,我们采用具有大步长策略的AdaMax优化器来根据远场振幅测量对多模波导进行模态分析。我们的方法检索模态功率分布和模态相对相位分布,我们阐明如何双图像模糊限制的能力来分析模态相对相位分布。实验结果表明,该方法对矩形波导和圆形波导都有较好的性能,在信噪比(SNR)为20 ~ 120 dB的噪声环境下仍能保持较高的精度和鲁棒性,与同类方法相比,在精度和计算成本上都有较大的提高。该方法为模态分析提供了一种新的解决方案,具有广阔的应用前景。
摘要:Optimizing multimode waveguide performance depends on modal analysis; however, current approaches focus predominantly on modal power distribution and, limited by experimental hardware and conditions, exhibit low accuracy, poor adaptability, and high computational cost. In this work, under a power-normalization constraint, we employ the AdaMax optimizer with a large-step-size strategy to perform modal analysis of multimode waveguides from far-field amplitude measurements. Our method retrieves both the modal power distribution and the modal relative-phase distribution, and we elucidate how twin-image ambiguity limits the capability to analyze modal relative-phase distributions. Experimental results demonstrate that the proposed method performs well for both rectangular and circular waveguides, maintaining high accuracy and robustness under noise with signal-to-noise ratios (SNRs) ranging from 20 to 120 dB, and achieving substantial improvements in accuracy and computational cost over comparable methods. This method provides a novel solution for modal analysis with broad application potential.



eess.AS音频处理



【1】Soft-Constrained Spatially Selective Active Noise Control for Open-fitting Hearables
标题:开放式耳机的软约束空间选择性主动噪音控制
链接:http://arxiv.org/pdf/2507.12122v1

作者:Reinhild Roden, Matthias Blau, Simon Doclo

备注:Accepted at IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2025

摘要:使用多个麦克风的空间选择性有源噪声控制(SSANC)的最新进展使得可听设备能够抑制不期望的噪声,同时保留来自特定方向的期望语音。为了实现最小的语音失真,硬约束已被用于在以前的工作中的优化问题来计算控制滤波器。在这项工作中,我们提出了一个软约束SSANC系统,使用频率无关的参数之间的权衡语音失真和降噪。我们推导出时域和频域配方,并表明,传统的有源噪声控制和硬约束SSANC代表两个极限情况下,所提出的设计。我们评估系统通过模拟使用一对开放式拟合hearables在消声环境中与一个语音源和两个噪声源。仿真结果验证了理论推导的正确性,并表明在较宽的折衷参数范围内,与硬约束设计相比,PESQ和ESTOI的信噪比、语音质量和可懂度都得到了显著改善。
摘要:Recent advances in spatially selective active noise control (SSANC) using multiple microphones have enabled hearables to suppress undesired noise while preserving desired speech from a specific direction. Aiming to achieve minimal speech distortion, a hard constraint has been used in previous work in the optimization problem to compute the control filter. In this work, we propose a soft-constrained SSANC system that uses a frequency-independent parameter to trade off between speech distortion and noise reduction. We derive both time- and frequency-domain formulations, and show that conventional active noise control and hard-constrained SSANC represent two limiting cases of the proposed design. We evaluate the system through simulations using a pair of open-fitting hearables in an anechoic environment with one speech source and two noise sources. The simulation results validate the theoretical derivations and demonstrate that for a broad range of the trade-off parameter, the signal-to-noise ratio and the speech quality and intelligibility in terms of PESQ and ESTOI can be substantially improved compared to the hard-constrained design.


【2】VoxATtack: A Multimodal Attack on Voice Anonymization Systems
标题:VoxATtack:对语音匿名系统的多模式攻击
链接:http://arxiv.org/pdf/2507.12081v1

作者:radi, Ünal Ege Gaznepoglu, Emanuël A. P. Habets, Daniel Tenbrinck
备注:5 pages, 3 figures, 3 tables, accepted at WASPAA 2025
摘要:语音匿名化系统旨在通过模糊语音特征来保护说话者隐私,同时保留与下游应用相关的语言内容。然而,由于这些语言线索保持完整,它们可以被利用来识别与特定说话者相关的语义语音模式。在这项工作中,我们提出了VoxATtack,一种新的多模态去匿名化模型,它结合了声学和文本信息来攻击匿名化系统。虽然以前的研究集中在细化从语音中提取的说话人表示,我们表明,将文本信息与标准的ECAPA-TDNN提高了攻击者的性能。我们提出的VoxATtack模型采用双分支架构,ECAPA-TDNN处理匿名语音和预训练的BERT编码的transmittance。这两个输出被投影到等维度的嵌入,然后基于在每个话语的基础上计算的置信度权重进行融合。当在语音隐私攻击者挑战(VPAC)数据集上评估我们的方法时,它在七个基准测试中的五个测试中表现优于顶级攻击者,即B3,B4,B5,T8-5和T12-5。为了进一步提高性能,我们利用匿名语音和SpecAugment作为增强技术。这一增强使VoxATtack在T10-2和T25-1上分别获得20.6%和27.2%的平均相等错误率后,在所有VPAC基准测试中达到最先进水平。我们的研究结果表明,结合文本信息和选择性数据增强揭示了当前语音匿名化方法的关键漏洞,并暴露了用于评估它们的数据集的潜在弱点。
摘要:Voice anonymization systems aim to protect speaker privacy by obscuring vocal traits while preserving the linguistic content relevant for downstream applications. However, because these linguistic cues remain intact, they can be exploited to identify semantic speech patterns associated with specific speakers. In this work, we present VoxATtack, a novel multimodal de-anonymization model that incorporates both acoustic and textual information to attack anonymization systems. While previous research has focused on refining speaker representations extracted from speech, we show that incorporating textual information with a standard ECAPA-TDNN improves the attacker's performance. Our proposed VoxATtack model employs a dual-branch architecture, with an ECAPA-TDNN processing anonymized speech and a pretrained BERT encoding the transcriptions. Both outputs are projected into embeddings of equal dimensionality and then fused based on confidence weights computed on a per-utterance basis. When evaluating our approach on the VoicePrivacy Attacker Challenge (VPAC) dataset, it outperforms the top-ranking attackers on five out of seven benchmarks, namely B3, B4, B5, T8-5, and T12-5. To further boost performance, we leverage anonymized speech and SpecAugment as augmentation techniques. This enhancement enables VoxATtack to achieve state-of-the-art on all VPAC benchmarks, after scoring 20.6% and 27.2% average equal error rate on T10-2 and T25-1, respectively. Our results demonstrate that incorporating textual information and selective data augmentation reveals critical vulnerabilities in current voice anonymization methods and exposes potential weaknesses in the datasets used to evaluate them.


【3】Self-Boosted Weight-Constrained FxLMS: A Robustness Distributed Active Noise Control Algorithm Without Internode Communication
标题:自增强加权约束FxLMS:一种无节点间通信的鲁棒分布式有源噪声控制算法
链接:http://arxiv.org/pdf/2507.12045v1
作者:Dongyuan Shi, Zhengding Luo, Boxiang Wang, Woon-Seng Gan

Journal-ref:IEEE Signal Processing Letters, pp. 1-5, 2025

摘要:与传统的集中式多通道有源噪声控制(MCANC)算法相比,它需要大量的计算资源,分散的方法表现出更高的计算效率,但通常会导致较差的降噪性能。为了提高性能,已经引入了分布式ANC方法,使得ANC节点之间能够进行信息交换;然而,由此产生的通信延迟通常会损害系统稳定性。为了克服这些局限性,我们提出了一种自提升的加权约束滤波参考最小均方(SB-WCFxLMS)算法的分布式MCANC系统没有节点间的通信。WCFxLMS算法专门设计用于缓解由节点间串扰效应引起的发散问题。自提升策略让每个ANC节点基于其局部降噪性能独立地调整其约束参数,从而确保有效的噪声消除,而无需节点间通信。在此机制的帮助下,该方法显着降低了计算复杂度和通信开销。采用真实的声学路径和压缩机噪声的数值仿真验证了所提出的系统的有效性和鲁棒性。结果表明,我们提出的方法达到了令人满意的噪声消除性能与最小的资源需求。
摘要:Compared to the conventional centralized multichannel active noise control (MCANC) algorithm, which requires substantial computational resources, decentralized approaches exhibit higher computational efficiency but typically result in inferior noise reduction performance. To enhance performance, distributed ANC methods have been introduced, enabling information exchange among ANC nodes; however, the resulting communication latency often compromises system stability. To overcome these limitations, we propose a self-boosted weight-constrained filtered-reference least mean square (SB-WCFxLMS) algorithm for the distributed MCANC system without internode communication. The WCFxLMS algorithm is specifically designed to mitigate divergence issues caused by the internode cross-talk effect. The self-boosted strategy lets each ANC node independently adapt its constraint parameters based on its local noise reduction performance, thus ensuring effective noise cancellation without the need for inter-node communication. With the assistance of this mechanism, this approach significantly reduces both computational complexity and communication overhead. Numerical simulations employing real acoustic paths and compressor noise validate the effectiveness and robustness of the proposed system. The results demonstrate that our proposed method achieves satisfactory noise cancellation performance with minimal resource requirements.


【4】VoJSQA: Speech Quality Assessment with Perceptually-Inspired Contrastive Pretraining Based on JND Audio PairsxATtack: A Multimodal Attack on Voice Anonymization Systems
标题:JSQA:基于JND音频对的感知对比预训练的语音质量评估
链接:http://arxiv.org/pdf/2507.11636v1
作者:Donald Williamson
备注:5 pages, 3 figures, 3 tables, accepted at WASPAA 2025
摘要:语音质量评估(SQA)通常用于学习从高维输入空间到表示感知语音质量的平均意见得分(MOS)的标量的映射。学习这样的映射是具有挑战性的原因有很多,但主要是因为MOS表现出高水平的固有差异,由于感知和实验设计的差异。已经提出了许多解决方案,但许多方法没有将感知因素正确地纳入其学习算法(超出MOS标签),这可能导致不满意的结果。为此,我们提出了JSQA,这是一个两阶段的框架,它使用感知引导的对比学习对JND对音频编码器进行预训练,然后对MOS预测进行微调。我们首先在JND级别内生成音频数据对,然后将其用于预训练编码器,以利用感知质量相似性信息并将其映射到嵌入空间中。JND对来自与来自CHiME-3的背景噪声以不同信噪比(SNR)混合的干净LibriSpeech话语。编码器稍后使用来自NISQA数据集的音频样本进行微调以进行MOS预测。实验结果表明,感知启发的对比预训练显着提高了模型的性能评估的各种指标相比,从零开始训练没有预训练的相同的网络。这些发现表明,将知觉因素纳入预训练大大有助于提高性能的SQA。
摘要:

Speech quality assessment (SQA) is often used to learn a mapping from a high-dimensional input space to a scalar that represents the mean opinion score (MOS) of the perceptual speech quality. Learning such a mapping is challenging for many reasons, but largely because MOS exhibits high levels of inherent variance due to perceptual and experimental-design differences. Many solutions have been proposed, but many approaches do not properly incorporate perceptual factors into their learning algorithms (beyond the MOS label), which could lead to unsatisfactory results. To this end, we propose JSQA, a two-stage framework that pretrains an audio encoder using perceptually-guided contrastive learning on just noticeable difference (JND) pairs, followed by fine-tuning for MOS prediction. We first generate pairs of audio data within JND levels, which are then used to pretrain an encoder to leverage perceptual quality similarity information and map it into an embedding space. The JND pairs come from clean LibriSpeech utterances that are mixed with background noise from CHiME-3, at different signal-to-noise ratios (SNRs). The encoder is later fine-tuned with audio samples from the NISQA dataset for MOS prediction. Experimental results suggest that perceptually-inspired contrastive pretraining significantly improves the model performance evaluated by various metrics when compared against the same network trained from scratch without pretraining. These findings suggest that incorporating perceptual factors into pretraining greatly contributes to the improvement in performance for SQA.


【5】Towards few-shot isolated word reading assessment
标题:走向几次孤立的单词阅读评估
链接:http://arxiv.org/pdf/2507.12217v1
作者:it, Retief Louw, Herman Kamper

备注:Accepted to SLaTE 2025

摘要:我们探索了一种无ASR的方法,用于低资源环境中的孤立词阅读评估。我们的Few-Shot方法将输入的儿童语音与成人提供的一小组参考模板进行比较。输入和模板使用来自大型自监督学习(SSL)模型的中间层进行编码。使用南非荷兰语儿童语音基准,我们调查的设计方案,如离散化SSL功能和重心平均的模板。理想化的实验表明,合理的性能为成人,但大幅下降的儿童语音输入,即使与儿童模板。尽管成功地采用SSL表示在低资源的语音任务,我们的工作突出了SSL表示用于处理子数据时,在一个Few-Shot分类系统的局限性。
摘要:We explore an ASR-free method for isolated word reading assessment in low-resource settings. Our few-shot approach compares input child speech to a small set of adult-provided reference templates. Inputs and templates are encoded using intermediate layers from large self-supervised learned (SSL) models. Using an Afrikaans child speech benchmark, we investigate design options such as discretising SSL features and barycentre averaging of the templates. Idealised experiments show reasonable performance for adults, but a substantial drop for child speech input, even with child templates. Despite the success of employing SSL representations in low-resource speech tasks, our work highlights the limitations of SSL representations for processing child data when used in a few-shot classification system.


【6】RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
标题:RUMAA:重复感知的统一音乐音频分析,用于分数性能对齐,转录和错误检测
链接:http://arxiv.org/pdf/2507.12175v1

作者:Chang, Simon Dixon, Emmanouil Benetos 备注:Accepted to WASPAA 2025
摘要:本研究介绍了RUMAA,一个基于transformer的音乐性能分析框架,它以近端到端的方式统一了分数到性能对齐,分数通知转录和错误检测。与之前单独解决这些任务的方法不同,RUMAA使用预训练的配乐和音频编码器以及通过代理任务捕获任务相互依赖性的新型三流解码器将它们集成在一起。它将人类可读的MusicXML乐谱与重复符号对齐到全长表演音频,克服了传统的基于MIDI的方法,这些方法依赖于具有预先指定的重复结构的手动展开的乐谱数据。RUMAA在非重复乐谱上匹配最先进的对齐方法,并在公共钢琴音乐数据集中的重复乐谱上优于它们,同时还提供了有希望的转录和错误检测结果。
摘要:This study introduces RUMAA, a transformer-based framework for music performance analysis that unifies score-to-performance alignment, score-informed transcription, and mistake detection in a near end-to-end manner. Unlike prior methods addressing these tasks separately, RUMAA integrates them using pre-trained score and audio encoders and a novel tri-stream decoder capturing task interdependencies through proxy tasks. It aligns human-readable MusicXML scores with repeat symbols to full-length performance audio, overcoming traditional MIDI-based methods that rely on manually unfolded score-MIDI data with pre-specified repeat structures. RUMAA matches state-of-the-art alignment methods on non-repeated scores and outperforms them on scores with repeats in a public piano music dataset, while also delivering promising transcription and mistake detection results.


【7】Room Impulse Response Generation Conditioned on Acoustic Parameters
标题:以声学参数为条件的房间脉冲响应生成
链接:http://arxiv.org/pdf/2507.12136v1

作者:ellano, Chunghsin Yeh, Gautam Bhattacharya, Daniel Arteaga
备注:4+1 pages, 2 figures; accepted in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA 2025)
摘要:使用深度神经网络生成房间脉冲响应(RIR)由于其在虚拟和增强现实、音频后期制作及相关领域的应用而引起了越来越多的研究兴趣。大多数现有的方法条件生成模型的物理描述的房间,如其大小,形状和表面材料。然而,这种对几何信息的依赖限制了它们在房间布局未知或感知现实主义(空间对听众的声音)比严格的物理准确性更重要的情况下的可用性。在这项研究中,我们提出了一种替代策略:直接在一组RIR声学参数上调节RIR生成。这些参数包括混响时间和直达声混响比的各种测量,宽带和频带。通过指定空间应该如何听起来而不是它应该如何看起来,我们的方法可以实现更灵活和感知驱动的RIR生成。我们探索了自回归和非自回归生成模型在描述音频编解码器域中运行,使用离散令牌序列或连续嵌入。具体来说,我们选择了四个模型进行评估:自回归Transformer,MaskGIT模型,流匹配模型和基于分类器的方法。进行客观和主观评价,以比较这些方法与国家的最先进的替代品。结果表明,所提出的模型匹配或优于最先进的替代方案,MaskGIT模型实现了最佳性能。
摘要:The generation of room impulse responses (RIRs) using deep neural networks has attracted growing research interest due to its applications in virtual and augmented reality, audio postproduction, and related fields. Most existing approaches condition generative models on physical descriptions of a room, such as its size, shape, and surface materials. However, this reliance on geometric information limits their usability in scenarios where the room layout is unknown or when perceptual realism (how a space sounds to a listener) is more important than strict physical accuracy. In this study, we propose an alternative strategy: conditioning RIR generation directly on a set of RIR acoustic parameters. These parameters include various measures of reverberation time and direct sound to reverberation ratio, both broadband and bandwise. By specifying how the space should sound instead of how it should look, our method enables more flexible and perceptually driven RIR generation. We explore both autoregressive and non-autoregressive generative models operating in the Descript Audio Codec domain, using either discrete token sequences or continuous embeddings. Specifically, we have selected four models to evaluate: an autoregressive transformer, the MaskGIT model, a flow matching model, and a classifier-based approach. Objective and subjective evaluations are performed to compare these methods with state-of-the-art alternatives. Results show that the proposed models match or outperform state-of-the-art alternatives, with the MaskGIT model achieving the best performance.


【8】MambaRate: Speech Quality Assessment Across Different Sampling Rates
标题:MambaSpeed:不同采样率的语音质量评估
链接:http://arxiv.org/pdf/2507.12090v1

作者:oulidis, Iakovi Alexiou, Junkwang Oh, Gunu Jho, Inchul Hwang, Pirros Tsiakoulis, Aimilios Chalamandaris

备注:Submitted to ASRU 2025 (AudioMOS Challenge 2025 Track 3)

摘要:我们提出了MambaRate,它预测平均意见分数(MOS)与有限的偏见有关的波形的采样率进行评估。它是专为AudioMOS Challenge 2025的Track 3而设计的,其重点是预测高采样频率下的语音MOS。我们的模型利用了自监督嵌入和选择性状态空间建模。目标评级通过高斯径向基函数(RBF)以连续表示进行编码。挑战的结果基于系统级斯皮尔曼等级相关系数(SRCC)度量。初始MambaRate版本(T16系统)在没有预训练的Few-Shot设置中的性能优于预训练基线(B03)约14%。T16在挑战赛中排名第四,与获胜系统相差约6%。我们在BVCC数据集上提供了额外的结果,以及使用不同表示作为输入的消融,其性能优于初始T16版本。
摘要:We propose MambaRate, which predicts Mean Opinion Scores (MOS) with limited bias regarding the sampling rate of the waveform under evaluation. It is designed for Track 3 of the AudioMOS Challenge 2025, which focuses on predicting MOS for speech in high sampling frequencies. Our model leverages self-supervised embeddings and selective state space modeling. The target ratings are encoded in a continuous representation via Gaussian radial basis functions (RBF). The results of the challenge were based on the system-level Spearman's Rank Correllation Coefficient (SRCC) metric. An initial MambaRate version (T16 system) outperformed the pre-trained baseline (B03) by ~14% in a few-shot setting without pre-training. T16 ranked fourth out of five in the challenge, differing by ~6% from the winning system. We present additional results on the BVCC dataset as well as ablations with different representations as input, which outperform the initial T16 version.


【9】Stereo Sound Event Localization and Detection with Onscreenoffscreen Classification
标题:通过屏上屏外分类进行立体声事件定位和检测
链接:http://arxiv.org/pdf/2507.12042v1
作者:imada, Archontis Politis, Iran R. Roman, Parthasaarathy Sudarsanam, David Diaz-Guerra, Ruchi Pandey, Kengo Uchida, Yuichiro Koyama, Naoya Takahashi, Takashi Shibuya, Shusuke Takahashi, Tuomas Virtanen, Yuki Mitsufuji 备注:5 pages, 2 figures
摘要:本文介绍了DCASE 2025声音事件定位和检测(SELD)挑战赛任务3的目标、数据集、基线和指标。在以前的版本中,挑战赛使用了四声道音频格式的一阶高保真度立体声(FOA)和麦克风阵列。相比之下,今年的挑战研究了具有立体声音频数据的SELD(称为立体声SELD)。此更改将重点从更专业的360度音频和视听场景分析转移到具有有限视野(FOV)的更常见的音频和媒体场景。由于立体声音频数据中固有的角度模糊性,该任务集中于方位角平面(左右轴)中的到达方向(DOA)估计以及距离估计。挑战仍然分为两个轨道:音频和视听,视听轨道引入了一个新的子任务,即有限的FOV所必需的屏幕 屏幕外事件分类。本次挑战赛介绍了DCASE 2025 Task 3 Stereo SELD数据集,其立体声音频和透视视频剪辑是从STARSS 23录音中采样和转换的。基线系统被设计为处理立体声音频和相应的视频帧作为输入。除了典型的SELD事件分类和本地化,它还集成了视听轨道的屏幕 屏幕外分类。评估指标已被修改,以引入一个可识别 屏幕外准确性指标,该指标评估模型识别哪些声源被识别的能力。在实验评估中,基线系统对立体声音频数据表现得相当好。
摘要:This paper presents the objective, dataset, baseline, and metrics of Task 3 of the DCASE2025 Challenge on sound event localization and detection (SELD). In previous editions, the challenge used four-channel audio formats of first-order Ambisonics (FOA) and microphone array. In contrast, this year's challenge investigates SELD with stereo audio data (termed stereo SELD). This change shifts the focus from more specialized 360{ deg} audio and audiovisual scene analysis to more commonplace audio and media scenarios with limited field-of-view (FOV). Due to inherent angular ambiguities in stereo audio data, the task focuses on direction-of-arrival (DOA) estimation in the azimuth plane (left-right axis) along with distance estimation. The challenge remains divided into two tracks: audio-only and audiovisual, with the audiovisual track introducing a new sub-task of onscreen offscreen event classification necessitated by the limited FOV. This challenge introduces the DCASE2025 Task3 Stereo SELD Dataset, whose stereo audio and perspective video clips are sampled and converted from the STARSS23 recordings. The baseline system is designed to process stereo audio and corresponding video frames as inputs. In addition to the typical SELD event classification and localization, it integrates onscreen offscreen classification for the audiovisual track. The evaluation metrics have been modified to introduce an onscreen offscreen accuracy metric, which assesses the models' ability to identify which sound sources are onscreen. In the experimental evaluation, the baseline system performs reasonably well with the stereo audio data.


【10】Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
标题:Vox视频引导对比视听掩蔽自动编码器,从视频中自动生成视听文本三重组ATtack:对语音匿名系统的多模式攻击
链接:http://arxiv.org/pdf/2507.11967v1
作者:ikawa, Shota Nakada, Hokuto Munakata, Kazuhiro Saito, Tatsuya Komatsu, Yoshimitsu Aoki
备注:Interspeech 2025
摘要:在本文中,我们提出了一种基于图像引导的对比视听掩蔽自编码器(LG-CAV-MAE)来改进视听表示学习。LG-CAV-MAE将预训练的文本编码器集成到对比视听掩蔽自动编码器中,使模型能够跨音频,视觉和文本模式进行学习。为了训练LG-CAV-MAE,我们引入了一种从未标记的视频中自动生成视听文本三元组的方法。我们首先使用图像字幕模型生成帧级字幕,然后应用基于CLAP的过滤,以确保音频和字幕之间的强对齐。这种方法产生高质量的视听文本三元组,而不需要手动注释。我们评估LG-CAV-MAE的视听检索任务,以及视听分类任务。我们的方法显著优于现有的方法,检索任务的recall@10提高了5.6%,分类任务提高了3.2%。
摘要:In this paper, we propose Language-Guided Contrastive Audio-Visual Masked Autoencoders (LG-CAV-MAE) to improve audio-visual representation learning. LG-CAV-MAE integrates a pretrained text encoder into contrastive audio-visual masked autoencoders, enabling the model to learn across audio, visual and text modalities. To train LG-CAV-MAE, we introduce an automatic method to generate audio-visual-text triplets from unlabeled videos. We first generate frame-level captions using an image captioning model and then apply CLAP-based filtering to ensure strong alignment between audio and captions. This approach yields high-quality audio-visual-text triplets without requiring manual annotations. We evaluate LG-CAV-MAE on audio-visual retrieval tasks, as well as an audio-visual classification task. Our method significantly outperforms existing approaches, achieving up to a 5.6% improvement in recall@10 for retrieval tasks and a 3.2% improvement for the classification task.


【11】Schrödinger Bridge Consistency Trajectory Models for Speech Enhancement
标题:用于语音增强的薛定格桥一致性轨迹模型
链接:http://arxiv.org/pdf/2507.11925v1

作者:Nishigori, Koichi Saito, Naoki Murata, Masato Hirano, Shusuke Takahashi, Yuki Mitsufuji
摘要:利用扩散模型进行语音增强是一种很有前途的技术,可以提高含噪语音数据的语音质量。此外,薛定谔桥(SB)最近已被用于基于扩散的SE,以改善语音质量,通过解决的前向过程的端点和反向过程的起点之间的失配。然而,SB仍然表现出缓慢的推理,由于大量的功能评估(NFE)的推理,以获得高质量的结果的必要性。虽然一致性模型(CM)通过采用一致性训练来解决这个问题,该训练使用图像生成领域中预训练模型的蒸馏,但当步骤数量增加时,它不会提高生成质量。一致性轨迹模型(Consistency Trajectory Models,CTM)的出现,不仅加快了推理速度,而且在推理质量和推理速度之间取得了良好的平衡。此外,SoundCTM展示了CTM技术在声音生成领域的适用性。本文将CTM技术应用于SE的Schr odinger桥,提出了Schr odinger桥一致性轨迹模型(SBCTM)。此外,我们引入了一种新的辅助损失,包括感知损失,到原来的CTM的训练框架。因此,SBCTM实现了约16倍的改善,在实时因子(RTF)相比,传统的薛定谔桥的SE。此外,SBCTM中的质量和速度之间的有利权衡允许通过将多步细化限制到1步推理不足的情况来实现时间高效的推理。我们的代码、预训练模型和音频样本可在https: github.com sony sbctm 上获得。
摘要:Speech enhancement (SE) utilizing diffusion models is a promising technology that improves speech quality in noisy speech data. Furthermore, the Schr "odinger bridge (SB) has recently been used in diffusion-based SE to improve speech quality by resolving a mismatch between the endpoint of the forward process and the starting point of the reverse process. However, the SB still exhibits slow inference owing to the necessity of a large number of function evaluations (NFE) for inference to obtain high-quality results. While Consistency Models (CMs) address this issue by employing consistency training that uses distillation from pretrained models in the field of image generation, it does not improve generation quality when the number of steps increases. As a solution to this problem, Consistency Trajectory Models (CTMs) not only accelerate inference speed but also maintain a favorable trade-off between quality and speed. Furthermore, SoundCTM demonstrates the applicability of CTM techniques to the field of sound generation. In this paper, we present Schr "odinger bridge Consistency Trajectory Models (SBCTM) by applying the CTM's technique to the Schr "odinger bridge for SE. Additionally, we introduce a novel auxiliary loss, including a perceptual loss, into the original CTM's training framework. As a result, SBCTM achieves an approximately 16x improvement in the real-time factor (RTF) compared to the conventional Schr "odinger bridge for SE. Furthermore, the favorable trade-off between quality and speed in SBCTM allows for time-efficient inference by limiting multi-step refinement to cases where 1-step inference is insufficient. Our code, pretrained models, and audio samples are available at https: github.com sony sbctm .


【12】A Multimodal Data Fusion Generative Adversarial Network for Real Time Underwater Sound Speed Field Construction
标题:实时水下音速场构建的多模式数据融合生成对抗网络
链接:http://arxiv.org/pdf/2507.11812v1

作者:Yuqiang Huang, Yanan Wu, Tianhe Xu, Junting Wang, Hao Zhang
摘要:声速剖面是影响水声信号传播模式的重要参数,对水声通信的能量效率和水声定位精度有着重要影响。传统上,SSP可以通过匹配场处理(MFP)、压缩感知(CS)和深度学习(DL)方法来获得。然而,现有的方法主要依赖于现场水下声纳观测数据,这对声纳观测系统的部署提出了严格的要求。为了在没有现场水下数据测量的情况下实现给定海域声速分布的高精度估计,提出了一种用于SSP构建的带剩余注意块的多模态数据融合生成对抗网络模型(MDF-RAGAN)。为了提高模型捕捉全局空间特征相关性的能力,我们嵌入了注意机制,并使用残差模块来深度捕捉由SST变化引起的深海声速分布中的小扰动。在真实开放数据集上的实验结果表明,该模型的分类精度优于其他现有方法,误差小于0.3m s。具体而言,MDF-RAGAN不仅比卷积神经网络(CNN)和空间插值(SITP)的性能高出近2倍,而且与平均轮廓相比,均方根误差(RMSE)降低了约65.8%,充分体现了多源融合和跨模态注意对整体轮廓匹配的增强。
摘要:Sound speed profiles (SSPs) are essential parameters underwater that affects the propagation mode of underwater signals and has a critical impact on the energy efficiency of underwater acoustic communication and accuracy of underwater acoustic positioning. Traditionally, SSPs can be obtained by matching field processing (MFP), compressive sensing (CS), and deep learning (DL) methods. However, existing methods mainly rely on on-site underwater sonar observation data, which put forward strict requirements on the deployment of sonar observation systems. To achieve high-precision estimation of sound velocity distribution in a given sea area without on-site underwater data measurement, we propose a multi-modal data-fusion generative adversarial network model with residual attention block (MDF-RAGAN) for SSP construction. To improve the model's ability for capturing global spatial feature correlations, we embedded the attention mechanisms, and use residual modules for deeply capturing small disturbances in the deep ocean sound velocity distribution caused by changes of SST. Experimental results on real open dataset show that the proposed model outperforms other state-of-the-art methods, which achieves an accuracy with an error of less than 0.3m s. Specifically, MDF-RAGAN not only outperforms convolutional neural network (CNN) and spatial interpolation (SITP) by nearly a factor of two, but also achieves about 65.8 % root mean square error (RMSE) reduction compared to mean profile, which fully reflects the enhancement of overall profile matching by multi-source fusion and cross-modal attention.


【13】Towards Scalable AASIST: Refining Graph Attention for Speech Deepfake Detection
标题:迈向可扩展的AASIST:细化语音Deepfake检测的图形注意力
链接:http://arxiv.org/pdf/2507.11777v1

作者:hirev, Daniil Sirota, Aleksandr Smirnov, Kirill Borodin
摘要:语音转换和文本到语音合成的进步使得自动说话人验证(ASV)系统更容易受到欺骗攻击。这项工作探讨了适度的改进AASIST反欺骗架构。它集成了一个冻结的Wav2Vec 2.0编码器,以在有限的数据设置中保留自监督语音表示,使用异构查询投影用标准化的多头注意模块代替原始图形注意块,并用可训练的上下文感知集成层代替启发式帧段融合。在ASVspoof 5语料库上进行评估时,所提出的系统达到了7.6%的等错误率(EER),在相同训练条件下改进了重新实现的AASIST基线。消融实验表明,每个架构变化都有助于整体性能,这表明对已建立的模型进行有针对性的调整可能有助于加强实际场景中的语音深度伪造检测。该代码可在https: github.com KORALLLL AASIST_SCALING上公开获取。
摘要:Advances in voice conversion and text-to-speech synthesis have made automatic speaker verification (ASV) systems more susceptible to spoofing attacks. This work explores modest refinements to the AASIST anti-spoofing architecture. It incorporates a frozen Wav2Vec 2.0 encoder to retain self-supervised speech representations in limited-data settings, substitutes the original graph attention block with a standardized multi-head attention module using heterogeneous query projections, and replaces heuristic frame-segment fusion with a trainable, context-aware integration layer. When evaluated on the ASVspoof 5 corpus, the proposed system reaches a 7.6 % equal error rate (EER), improving on a re-implemented AASIST baseline under the same training conditions. Ablation experiments suggest that each architectural change contributes to the overall performance, indicating that targeted adjustments to established models may help strengthen speech deepfake detection in practical scenarios. The code is publicly available at https: github.com KORALLLL AASIST_SCALING.


机器翻译由腾讯交互翻译提供,仅供参考