今日论文合集:cs.SD语音3篇,eess.AS音频处理4篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 Differentiable Room Acoustic Rendering with Multi-View Vision Priors
标题: 利用多视图视觉先验的区分房间声学渲染
链接:https://arxiv.org/abs/2504.21847

作者: Derong Jin,  Ruohan Gao (University of Maryland, College Park) 
备注:Project Page: this https URL
摘要:在创建逼真的虚拟环境时,由空间音频实现的沉浸式声学体验与视觉方面一样重要。然而,用于房间脉冲响应估计的现有方法依赖于数据要求高的基于学习的模型或计算上昂贵的基于物理的建模。在这项工作中,我们介绍了视听区分房间声学渲染(AV-DAR),一个框架,利用视觉线索提取的多视图图像和声学波束跟踪基于物理的房间声学渲染。来自两个数据集的六个真实世界环境的实验表明,我们的多模态,基于物理的方法是有效的,可解释的,准确的,显着优于一系列先前的方法。值得注意的是,在Real Acoustic Field数据集上,AV-DAR实现了与使用10倍以上数据训练的模型相当的性能,同时在相同规模下训练时提供了16.6%至50.9%的相对增益。
摘要:An immersive acoustic experience enabled by spatial audio is just as crucial as the visual aspect in creating realistic virtual environments. However, existing methods for room impulse response estimation rely either on data-demanding learning-based models or computationally expensive physics-based modeling. In this work, we introduce Audio-Visual Differentiable Room Acoustic Rendering (AV-DAR), a framework that leverages visual cues extracted from multi-view images and acoustic beam tracing for physics-based room acoustic rendering. Experiments across six real-world environments from two datasets demonstrate that our multimodal, physics-based approach is efficient, interpretable, and accurate, significantly outperforming a series of prior methods. Notably, on the Real Acoustic Field dataset, AV-DAR achieves comparable performance to models trained on 10 times more data while delivering relative gains ranging from 16.6% to 50.9% when trained at the same scale.


【2】 DGFNet: End-to-End Audio-Visual Source Separation Based on Dynamic  Gating Fusion
标题: DGFNet:基于动态门控融合的端到端视听源分离
链接:https://arxiv.org/abs/2504.21366

作者: Yinfeng Yu,  Shiyu Sun 
备注:Main paper (9 pages). Accepted for publication by ICMR(International Conference on Multimedia Retrieval) 2025
摘要:现有的音视频源分离方法主要采用两种设计策略。第一种策略涉及在编码器的瓶颈层融合音频和视觉特征,然后通过解码器处理融合的特征。然而,当两种模式之间存在显著差异时,这种方法可能会导致关键信息的丢失。第二种策略避免了直接融合,而是依赖于解码器来处理音频和视觉特征之间的交互。尽管如此,如果编码器未能充分整合模态之间的信息,则解码器可能无法有效地捕获它们之间的复杂关系。为了解决这些问题,本文提出了一种基于门控机制的动态融合方法,该方法动态调整模态融合度。这种方法减轻了仅依赖于解码器的限制,并促进了音频和视觉特征之间的有效协作。此外,引入音频注意模块,增强音频特征的表达能力,从而进一步提高模型的性能。实验结果表明,该方法在两个基准数据集上取得了显著的性能改善,验证了其在视听源分离任务中的有效性和优势。
摘要:Current Audio-Visual Source Separation methods primarily adopt two design strategies. The first strategy involves fusing audio and visual features at the bottleneck layer of the encoder, followed by processing the fused features through the decoder. However, when there is a significant disparity between the two modalities, this approach may lead to the loss of critical information. The second strategy avoids direct fusion and instead relies on the decoder to handle the interaction between audio and visual features. Nonetheless, if the encoder fails to integrate information across modalities adequately, the decoder may be unable to effectively capture the complex relationships between them. To address these issues, this paper proposes a dynamic fusion method based on a gating mechanism that dynamically adjusts the modality fusion degree. This approach mitigates the limitations of solely relying on the decoder and facilitates efficient collaboration between audio and visual features. Additionally, an audio attention module is introduced to enhance the expressive capacity of audio features, thereby further improving model performance. Experimental results demonstrate that our method achieves significant performance improvements on two benchmark datasets, validating its effectiveness and advantages in Audio-Visual Source Separation tasks.


【3】 Design, analysis, and experimental validation of a stepped plate  parametric array loudspeaker
标题: 阶梯板参数阵列扬声器的设计、分析和实验验证
链接:https://arxiv.org/abs/2504.21171

作者: Woongji Kim,  Beomseok Oh,  Chayeong Kim,  Wonkyu Moon 
备注:51 pages, 18 figures, arXiv:YYMM.NNNN(N) format preferred, submitted to The Journal of the Acoustical Society of America (AIP)
摘要:本研究探讨阶梯式平板参数阵列扬声器(SPPAL)的设计与分析,以取代传统的阵列式参数扬声器。SPPAL利用单个Langevin型超声换能器与弯曲阶梯板耦合,通过非线性声学相互作用产生窄波束可听声音。为了评估和优化SPPAL的性能,开发了一个集成的建模框架,包括换能器动力学的近似解析3D模型,与阶梯板和刚性活塞行为相关的当量比公式,以及用于非线性声场模拟的球面波展开方法。通过多目标分析优化换能器的双共振行为,以提高低频音频性能。实验验证包括换能器的频率响应和模态分析,以及声场测量。通过与实验数据的比较,进一步验证了分析方法的正确性。此外,组合谐振-一个意外的结构激励互调产生的-被确定为一个固有的现象SPPAL操作。研究结果提供了实用的指导,有效的,紧凑的,可制造的参数阵列扬声器采用板基弯曲振动的发展。
摘要:This study investigates the design and analysis of a stepped plate parametric array loudspeaker (SPPAL) as an alternative to conventional array-based parametric loudspeakers. The SPPAL utilizes a single Langevin-type ultrasonic transducer coupled with a flexural stepped plate to generate narrow-beam audible sound via nonlinear acoustic interaction. To evaluate and optimize the performance of the SPPAL, an integrated modeling framework is developed, consisting of an approximate analytical 3D model for transducer dynamics, an equivalence ratio formulation to relate stepped plate and rigid piston behavior, and a spherical wave expansion method for nonlinear sound field simulation. The dual-resonance behavior of the transducer is optimized through multi-objective analysis to enhance low-frequency audio performance. Experimental validation includes frequency response and modal analysis of the transducer, as well as sound field measurements. The analytical methods are further verified through comparison with experimental data. Furthermore, combination resonance--an unintended structural excitation resulting from intermodulation--is identified as an inherent phenomenon in SPPAL operation. The findings offer practical guidance for the development of efficient, compact, and manufacturable parametric array loudspeakers employing plate-based flexural vibration.


eess.AS音频处理


【1】 From Aesthetics to Human Preferences: Comparative Perspectives of  Evaluating Text-to-Music Systems
标题: 从美学到人类偏好:评估文本到音乐系统的比较视角
链接:https://arxiv.org/abs/2504.21815

作者: Huan Zhang,  Jinhua Liang,  Huy Phan,  Wenwu Wang,  Emmanouil Benetos 
摘要:评估生成模型仍然是一个根本性的挑战,特别是当目标是反映人类偏好时。在本文中,我们使用音乐生成作为一个案例研究,调查自动评估指标和人类偏好之间的差距。我们在五个国家的最先进的音乐生成方法进行比较实验,评估感知质量和分布的相似性,人类组成的音乐。具体来说,我们从各种感知维度评估合成音乐,并检查基于参考的指标,如Mauve音频发散度(MAD)和内核音频距离(KAD)。我们的研究结果显示,在不同的指标显着不一致,突出了目前的评价实践的局限性。为了支持进一步的研究,我们发布了一个基准数据集,包括来自多个模型的样本。这项研究提供了一个更广泛的角度对人类偏好的生成建模对齐,倡导更以人为本的跨领域的评估策略。
摘要:Evaluating generative models remains a fundamental challenge, particularly when the goal is to reflect human preferences. In this paper, we use music generation as a case study to investigate the gap between automatic evaluation metrics and human preferences. We conduct comparative experiments across five state-of-the-art music generation approaches, assessing both perceptual quality and distributional similarity to human-composed music. Specifically, we evaluate synthesis music from various perceptual dimensions and examine reference-based metrics such as Mauve Audio Divergence (MAD) and Kernel Audio Distance (KAD). Our findings reveal significant inconsistencies across the different metrics, highlighting the limitation of the current evaluation practice. To support further research, we release a benchmark dataset comprising samples from multiple models. This study provides a broader perspective on the alignment of human preference in generative modeling, advocating for more human-centered evaluation strategies across domains.


【2】 Impairments are Clustered in Latents of Deep Neural Network-based Speech  Quality Models
标题: 基于深度神经网络的语音质量模型的潜在损害被忽视
链接:https://arxiv.org/abs/2504.21528

作者: Fredrik Cumlin,  Xinyu Liang,  Victor Ungureanu,  Chandan K. A. Reddy,  Christian Schüldt,  Saikat Chatterjee 
摘要:在这篇文章中,我们提供了一个实验观察:基于深度神经网络(DNN)的语音质量评估(SQA)模型具有固有的潜在表示,其中许多类型的损伤被聚集在一起。虽然基于DNN的SQA模型没有经过损伤分类的训练,但我们的实验表明,在适当的SQA潜在表示中,损伤分类结果良好。我们调查使用各种音频退化,包括不同类型的噪声,波形剪裁,增益转换,音调偏移,压缩,混响等集群的损伤的聚类,我们执行的损伤的分类在SQA-潜在的表示域使用标准的k-最近邻(kNN)分类。我们还开发了一个新的基于DNN的SQA模型,名为DNSMOS+,以检查SQA的改进是否会导致损伤分类的改进。LibriAugmented数据集的分类准确率为94%,具有16种类型的损伤,ESC-50数据集的分类准确率为54%,具有50种类型的真实噪声。
摘要:In this article, we provide an experimental observation: Deep neural network (DNN) based speech quality assessment (SQA) models have inherent latent representations where many types of impairments are clustered. While DNN-based SQA models are not trained for impairment classification, our experiments show good impairment classification results in an appropriate SQA latent representation. We investigate the clustering of impairments using various kinds of audio degradations that include different types of noises, waveform clipping, gain transition, pitch shift, compression, reverberation, etc. To visualize the clusters we perform classification of impairments in the SQA-latent representation domain using a standard k-nearest neighbor (kNN) classifier. We also develop a new DNN-based SQA model, named DNSMOS+, to examine whether an improvement in SQA leads to an improvement in impairment classification. The classification accuracy is 94% for LibriAugmented dataset with 16 types of impairments and 54% for ESC-50 dataset with 50 types of real noises.


【3】 Pretraining Large Brain Language Model for Active BCI: Silent Speech
标题: 主动BCI预训练大大脑语言模型:无声语音
链接:https://arxiv.org/abs/2504.21214

作者: Jinzhao Zhou,  Zehong Cao,  Yiqun Duan,  Connor Barkley,  Daniel Leong,  Xiaowei Jiang,  Quoc-Toan Nguyen,  Ziyi Zhao,  Thomas Do,  Yu-Cheng Chang,  Sheng-Fu Liang,  Chin-teng Lin 
摘要:本文探讨了主动脑机接口(BCI)系统中的无声语音解码,它提供了比传统BCI应用更自然和灵活的通信。我们从12名受试者中收集了超过120小时的脑电图(EEG)记录,捕获了24个常用的英语单词,用于语言模型预训练和解码。继最近成功地用自监督范式预训练大型模型以提高EEG分类性能之后,我们提出了经过预训练的大型脑语言模型(LBLM),以解码主动BCI的无声语音。为了预训练LBLM,我们提出了未来谱-时预测(FSTP)预训练范式,从未标记的EEG数据中学习有效的表示。与现有的EEG预训练方法主要遵循掩蔽重建范式不同,我们提出的FSTP方法在时域和频域中采用自回归建模来捕获EEG信号的时间和频谱依赖性。在预训练之后,我们在下游任务上微调LBLM,包括单词级和语义级分类。大量的实验证明了LBLM在完全监督和预训练的基线模型上的显着性能增益。例如,在困难的跨会话环境中,我们的模型在语义级分类上达到了47.0%的准确率,在词级分类上达到了39.6%,分别比基线方法高出5.4%和7.3%.我们的研究推进了主动BCI系统中的无声语音解码,为EEG语言模型预训练提供了创新的解决方案,并为基础研究提供了新的数据集。
摘要:This paper explores silent speech decoding in active brain-computer interface (BCI) systems, which offer more natural and flexible communication than traditional BCI applications. We collected a new silent speech dataset of over 120 hours of electroencephalogram (EEG) recordings from 12 subjects, capturing 24 commonly used English words for language model pretraining and decoding. Following the recent success of pretraining large models with self-supervised paradigms to enhance EEG classification performance, we propose Large Brain Language Model (LBLM) pretrained to decode silent speech for active BCI. To pretrain LBLM, we propose Future Spectro-Temporal Prediction (FSTP) pretraining paradigm to learn effective representations from unlabeled EEG data. Unlike existing EEG pretraining methods that mainly follow a masked-reconstruction paradigm, our proposed FSTP method employs autoregressive modeling in temporal and frequency domains to capture both temporal and spectral dependencies from EEG signals. After pretraining, we finetune our LBLM on downstream tasks, including word-level and semantic-level classification. Extensive experiments demonstrate significant performance gains of the LBLM over fully-supervised and pretrained baseline models. For instance, in the difficult cross-session setting, our model achieves 47.0\% accuracy on semantic-level classification and 39.6\% in word-level classification, outperforming baseline methods by 5.4\% and 7.3\%, respectively. Our research advances silent speech decoding in active BCI systems, offering an innovative solution for EEG language model pretraining and a new dataset for fundamental research.


【4】 Design, analysis, and experimental validation of a stepped plate  parametric array loudspeaker
标题: 阶梯板参数阵列扬声器的设计、分析和实验验证
链接:https://arxiv.org/abs/2504.21171

作者: Woongji Kim,  Beomseok Oh,  Chayeong Kim,  Wonkyu Moon 
备注:51 pages, 18 figures, arXiv:YYMM.NNNN(N) format preferred, submitted to The Journal of the Acoustical Society of America (AIP)
摘要:本研究探讨阶梯式平板参数阵列扬声器(SPPAL)的设计与分析,以取代传统的阵列式参数扬声器。SPPAL利用单个Langevin型超声换能器与弯曲阶梯板耦合,通过非线性声学相互作用产生窄波束可听声音。为了评估和优化SPPAL的性能,开发了一个集成的建模框架,包括换能器动力学的近似解析3D模型,与阶梯板和刚性活塞行为相关的当量比公式,以及用于非线性声场模拟的球面波展开方法。通过多目标分析优化换能器的双共振行为,以提高低频音频性能。实验验证包括换能器的频率响应和模态分析,以及声场测量。通过与实验数据的比较,进一步验证了分析方法的正确性。此外,组合谐振-一个意外的结构激励互调产生的-被确定为一个固有的现象SPPAL操作。研究结果提供了实用的指导,有效的,紧凑的,可制造的参数阵列扬声器采用板基弯曲振动的发展。
摘要:This study investigates the design and analysis of a stepped plate parametric array loudspeaker (SPPAL) as an alternative to conventional array-based parametric loudspeakers. The SPPAL utilizes a single Langevin-type ultrasonic transducer coupled with a flexural stepped plate to generate narrow-beam audible sound via nonlinear acoustic interaction. To evaluate and optimize the performance of the SPPAL, an integrated modeling framework is developed, consisting of an approximate analytical 3D model for transducer dynamics, an equivalence ratio formulation to relate stepped plate and rigid piston behavior, and a spherical wave expansion method for nonlinear sound field simulation. The dual-resonance behavior of the transducer is optimized through multi-objective analysis to enhance low-frequency audio performance. Experimental validation includes frequency response and modal analysis of the transducer, as well as sound field measurements. The analytical methods are further verified through comparison with experimental data. Furthermore, combination resonance--an unintended structural excitation resulting from intermodulation--is identified as an inherent phenomenon in SPPAL operation. The findings offer practical guidance for the development of efficient, compact, and manufacturable parametric array loudspeakers employing plate-based flexural vibration.


机器翻译由腾讯交互翻译提供,仅供参考