今日论文合集:cs.SD语音8篇,eess.AS音频处理13篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
标题:SpA 2V:利用空间听觉线索实现音频驱动的空间感知视频生成
链接:http://arxiv.org/pdf/2508.00782v1

作者:Kien T. Pham, Yingqing He, Yazhou Xing, Qifeng Chen, Long Chen

备注:The 33rd ACM Multimedia Conference (MM '25)

摘要:音频驱动的视频生成旨在合成与输入音频记录一致的逼真视频,类似于人类从听觉输入中可视化场景的能力。然而,现有的方法主要集中于探索语义信息,例如音频中存在的声音源的类别,限制了它们生成具有准确内容和空间组成的视频的能力。相比之下,我们人类不仅可以自然地识别声源的语义类别,而且还可以确定其深度编码的空间属性,包括位置和运动方向。这种有用的信息可以通过考虑从声音的固有物理特性(如响度或频率)导出的特定空间指标来阐明。由于先前的方法在很大程度上忽略了这个因素,我们提出了SpA2V,第一个框架明确地利用这些空间听觉线索从音频生成具有高语义和空间对应性的视频。SpA2V将生成过程分解为两个阶段:1)音频引导的视频规划:我们精心调整了最先进的MLLM,用于利用来自输入音频的空间和语义线索来构建视频场景布局(VSL)的新任务。这用作中间表示,以弥合音频和视频模态之间的差距。2)基于布局的视频生成:我们开发了一种高效有效的方法,将VSL作为条件指导无缝集成到预先训练的扩散模型中,从而以免训练的方式实现基于VSL的视频生成。大量的实验表明,SpA2V擅长生成逼真的视频与语义和空间对齐的输入音频。
摘要:Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring semantic information, such as the classes of sounding sources present in the audio, limiting their ability to generate videos with accurate content and spatial composition. In contrast, we humans can not only naturally identify the semantic categories of sounding sources but also determine their deeply encoded spatial attributes, including locations and movement directions. This useful information can be elucidated by considering specific spatial indicators derived from the inherent physical properties of sound, such as loudness or frequency. As prior methods largely ignore this factor, we present SpA2V, the first framework explicitly exploits these spatial auditory cues from audios to generate videos with high semantic and spatial correspondence. SpA2V decomposes the generation process into two stages: 1) Audio-guided Video Planning: We meticulously adapt a state-of-the-art MLLM for a novel task of harnessing spatial and semantic cues from input audio to construct Video Scene Layouts (VSLs). This serves as an intermediate representation to bridge the gap between the audio and video modalities. 2) Layout-grounded Video Generation: We develop an efficient and effective approach to seamlessly integrate VSLs as conditional guidance into pre-trained diffusion models, enabling VSL-grounded video generation in a training-free manner. Extensive experiments demonstrate that SpA2V excels in generating realistic videos with semantic and spatial alignment to the input audios.


【2】AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

标题:AudioGen-Omni:用于视频同步音频、语音和歌曲生成的统一多模式扩散Transformer

链接:http://arxiv.org/pdf/2508.00733v1
作者:Le Wang, Jun Wang, Feng Deng, Chen Zhang, Kun Gai, Di Zhang

备注:12 pages, 2 figures

摘要:我们提出了AudioGen-Omni -一种基于多模态扩散Transformers(MMDit)的统一方法,能够生成与输入视频同步的高保真音频,语音和歌曲。AudioGen-Omni引入了一种新型的联合训练范式,该范式无缝集成了大规模的视频-文本-音频语料库,使模型能够在多模态输入的条件下生成语义丰富、声学多样的音频,并适应各种音频生成任务。AudioGen-Omni采用统一的歌词转录编码器,将来自演唱和口语输入的字素和音素编码为密集的帧级表示。密集的帧级表示融合使用基于AdaLN的联合注意力机制增强与相位对齐的各向异性位置注入(PAAPI),其中RoPE被选择性地应用于时间结构化的模态,以确保精确和鲁棒的跨模态对齐。通过解冻所有模态并掩蔽缺失的输入,AudioGen-Omni减轻了文本冻结范式的语义限制,从而实现有效的跨模态条件反射。这种联合训练方法提高了音频质量、语义对齐和嘴唇同步准确性,同时还在文本到音频 语音 歌曲任务上实现了最先进的结果。对于8秒的音频,推理时间为1.91秒,它在效率和通用性方面都有很大的提高。
摘要:We present AudioGen-Omni - a unified approach based on multimodal diffusion transformers (MMDit), capable of generating high-fidelity audio, speech, and songs coherently synchronized with the input video. AudioGen-Omni introduces a novel joint training paradigm that seamlessly integrates large-scale video-text-audio corpora, enabling a model capable of generating semantically rich, acoustically diverse audio conditioned on multimodal inputs and adaptable to a wide range of audio generation tasks. AudioGen-Omni employs a unified lyrics-transcription encoder that encodes graphemes and phonemes from both sung and spoken inputs into dense frame-level representations. Dense frame-level representations are fused using an AdaLN-based joint attention mechanism enhanced with phase-aligned anisotropic positional infusion (PAAPI), wherein RoPE is selectively applied to temporally structured modalities to ensure precise and robust cross-modal alignment. By unfreezing all modalities and masking missing inputs, AudioGen-Omni mitigates the semantic constraints of text-frozen paradigms, enabling effective cross-modal conditioning. This joint training approach enhances audio quality, semantic alignment, and lip-sync accuracy, while also achieving state-of-the-art results on Text-to-Audio Speech Song tasks. With an inference time of 1.91 seconds for 8 seconds of audio, it offers substantial improvements in both efficiency and generality.


【3】Advancing Speech Quality Assessment Through Scientific Challenges and Open-source Activities
标题:通过科学挑战和开源活动推进语音质量评估
链接:http://arxiv.org/pdf/2508.00317v1

作者:Wen-Chin Huang
备注:APSIPA ASC 2025 perspective paper
摘要:语音质量评估(SQA)是指对语音质量的评估,为了跟上生成式人工智能的热潮,开发一种准确的自动SQA方法来反映人类的感知变得越来越重要。近年来,SQA已经发展到一个点,研究人员开始忠实地使用自动SQA在研究论文作为一个严格的衡量良好的语音生成系统。我们认为,最近的科学挑战和开放源码活动刺激了这一领域的增长。在本文中,我们回顾了最近的挑战以及SQA的开源实现和工具包,并强调了维护这些活动的重要性,以促进SQA本身以及语音生成AI的发展。
摘要:Speech quality assessment (SQA) refers to the evaluation of speech quality, and developing an accurate automatic SQA method that reflects human perception has become increasingly important, in order to keep up with the generative AI boom. In recent years, SQA has progressed to a point that researchers started to faithfully use automatic SQA in research papers as a rigorous measurement of goodness for speech generation systems. We believe that the scientific challenges and open-source activities of late have stimulated the growth in this field. In this paper, we review recent challenges as well as open-source implementations and toolkits for SQA, and highlight the importance of maintaining such activities to facilitate the development of not only SQA itself but also generative AI for speech.


【4】DeformTune: A Deformable XAI Music Prototype for Non-Musicians
标题:变形曲调:适合非音乐家的可变形XAI音乐原型
链接:http://arxiv.org/pdf/2508.00160v1

作者:Ziqing Xu, Nick Bryan-Kinns

备注:In Proceedings of Explainable AI for the Arts Workshop 2025 (XAIxArts 2025) arXiv:2406.14485

摘要:许多现有的人工智能音乐生成工具依赖于文本提示、复杂的界面或类似乐器的控件,这可能需要非音乐家所不具备的音乐或技术知识。本文介绍了DeformTune,这是一个原型系统,它将触觉可变形界面与MeasureVAE模型相结合,以探索更直观,更具体,更可解释的AI交互。我们对11名没有接受过正式音乐训练的成年参与者进行了初步研究,以调查他们在人工智能辅助音乐创作方面的经验。对他们反馈的主题分析揭示了反复出现的挑战-包括不清楚的控制映射,有限的表达范围,以及在整个使用过程中需要指导。我们讨论了几个增强AI可解释性的设计机会,包括多模态反馈和渐进式交互支持。这些发现为使AI音乐系统更易于解释和授权新手用户提供了早期见解。
摘要:Many existing AI music generation tools rely on text prompts, complex interfaces, or instrument-like controls, which may require musical or technical knowledge that non-musicians do not possess. This paper introduces DeformTune, a prototype system that combines a tactile deformable interface with the MeasureVAE model to explore more intuitive, embodied, and explainable AI interaction. We conducted a preliminary study with 11 adult participants without formal musical training to investigate their experience with AI-assisted music creation. Thematic analysis of their feedback revealed recurring challenge--including unclear control mappings, limited expressive range, and the need for guidance throughout use. We discuss several design opportunities for enhancing explainability of AI, including multimodal feedback and progressive interaction support. These findings contribute early insights toward making AI music systems more explainable and empowering for novice users.


【5】VR-PTOLEMAIC: A Virtual Environment for the Perceptual Testing of Spatial Audio Algorithms
标题:空间音频算法感知测试的虚拟环境VR-PPEMAIC
链接:http://arxiv.org/pdf/2508.00501v1

作者:Paolo Ostan, Francesca Del Gaudio, Federico Miotello, Mirco Pezzoli, Fabio Antonacci

备注:to appear in EAA Forum Acusticum 2025

摘要:空间音频算法的感知评估是沉浸式音频应用开发的重要一步,因为它确保合成声场在聆听体验、空间感知和听觉真实性方面符合质量标准。为了支持这些评估,虚拟现实可以通过提供沉浸式和交互式测试环境来提供强大的平台。在本文中,我们提出了VR-PADMAIC,一个虚拟现实评估系统,旨在评估空间音频算法。该系统将MUSHRA(具有隐藏参考和锚点的多刺激测试)评估方法实施到虚拟环境中。特别地,用户可以将他们自己定位在虚拟重建的研讨会房间的25个模拟收听位置中的每一个中,并且相对于实际记录的二阶立体混响房间脉冲响应来评估模拟声学响应,所有这些都与各种源信号卷积。我们通过广泛的测试活动,评估人员被要求比较各种声场重建算法的重建能力,评估了所提出的框架的可用性。结果表明,虚拟现实平台有效地支持空间音频算法的评估,用户体验和沉浸感普遍积极的反馈。
摘要:The perceptual evaluation of spatial audio algorithms is an important step in the development of immersive audio applications, as it ensures that synthesized sound fields meet quality standards in terms of listening experience, spatial perception and auditory realism. To support these evaluations, virtual reality can offer a powerful platform by providing immersive and interactive testing environments. In this paper, we present VR-PTOLEMAIC, a virtual reality evaluation system designed for assessing spatial audio algorithms. The system implements the MUSHRA (MUlti-Stimulus test with Hidden Reference and Anchor) evaluation methodology into a virtual environment. In particular, users can position themselves in each of the 25 simulated listening positions of a virtually recreated seminar room and evaluate simulated acoustic responses with respect to the actually recorded second-order ambisonic room impulse responses, all convolved with various source signals. We evaluated the usability of the proposed framework through an extensive testing campaign in which assessors were asked to compare the reconstruction capabilities of various sound field reconstruction algorithms. Results show that the VR platform effectively supports the assessment of spatial audio algorithms, with generally positive feedback on user experience and immersivity.


【6】Wavelet-Based Time-Frequency Fingerprinting for Feature Extraction of Traditional Irish Music
标题:基于子波的时频指纹识别用于爱尔兰传统音乐特征提取
链接:http://arxiv.org/pdf/2508.00479v1

备注:Master's thesis. The focus of the thesis is on the underlying techniques for signal fingerprinting
摘要:这项工作提出了一种基于小波的时间序列特征提取的时频指纹识别方法,重点是从现场录音的传统爱尔兰曲调的音频识别。识别时间序列数据中的特征的挑战是通过采用连续小波变换来提取频谱特征和小波相干分析来比较记录的音频频谱图合成生成的曲调。合成曲调来自ABC记谱法,这是爱尔兰音乐的一种常见符号表示。实验结果表明,基于小波变换的方法能够准确、有效地识别录音曲调。本研究还详细介绍了小波相干模型的性能,突出了其优于其他时频分解方法的优势。此外,我们还讨论了该模型,并将其部署在音乐以外的几个应用程序中,包括EEG信号分析和金融时间序列预测。
摘要:This work presents a wavelet-based approach to time-frequency fingerprinting for time series feature extraction, with a focus on audio identification from live recordings of traditional Irish tunes. The challenges of identifying features in time-series data are addressed by employing a continuous wavelet transform to extract spectral features and wavelet coherence analysis is used to compare recorded audio spectrograms to synthetically generated tunes. The synthetic tunes are derived from ABC notation, which is a common symbolic representation for Irish music. Experimental results demonstrate that the wavelet-based method can accurately and efficiently identify recorded tunes. This research study also details the performance of the wavelet coherence model, highlighting its strengths over other methods of time-frequency decomposition. Additionally, we discuss and deploy the model on several applications beyond music, including in EEG signal analysis and financial time series forecasting.


【7】Beamformed 360° Sound Maps: U-Net-Driven Acoustic Source Segmentation and Localization
标题:束成形360°声音地图:U-Net驱动的光源分割和定位
链接:http://arxiv.org/pdf/2508.00307v1

作者:Belman J. Rodriguez, Sergio F. Chevtchenko, Marcelo Herrera Martinez, Yeshwant Bethy, Saeed Afshar
摘要:我们介绍了一个U-网络模型的360{ deg}声源定位制定为一个球形的语义分割任务。我们的模型不是回归离散的到达方向(DoA)角度,而是将波束成形的音频地图(方位角和仰角)分割成活动声音存在的区域。在定制的24麦克风阵列上使用延迟求和(DAS)波束成形,我们生成与无人机GPS遥测对齐的信号,以创建二进制监督掩码。修改后的U-Net在这些地图的频域表示上进行训练,学习识别空间分布的源区域,同时通过Tversky损失解决类别不平衡问题。由于该网络在波束成形能量图上运行,因此该方法本质上与阵列无关,并且可以适应不同的麦克风配置,而无需从头开始重新训练。通过计算激活区域上的质心对分割输出进行后处理,从而实现稳健的DoA估计。我们的数据集包括DJI Air 3无人机的真实野外记录,与多个日期和地点的360度视频和飞行日志同步。实验结果表明,U-net可以跨环境进行推广,提供更高的角度精度,为传统声源定位(SSL)之外的密集空间音频理解提供了新的范例。
摘要:We introduce a U-net model for 360{ deg} acoustic source localization formulated as a spherical semantic segmentation task. Rather than regressing discrete direction-of-arrival (DoA) angles, our model segments beamformed audio maps (azimuth and elevation) into regions of active sound presence. Using delay-and-sum (DAS) beamforming on a custom 24-microphone array, we generate signals aligned with drone GPS telemetry to create binary supervision masks. A modified U-Net, trained on frequency-domain representations of these maps, learns to identify spatially distributed source regions while addressing class imbalance via the Tversky loss. Because the network operates on beamformed energy maps, the approach is inherently array-independent and can adapt to different microphone configurations without retraining from scratch. The segmentation outputs are post-processed by computing centroids over activated regions, enabling robust DoA estimates. Our dataset includes real-world open-field recordings of a DJI Air 3 drone, synchronized with 360{ deg} video and flight logs across multiple dates and locations. Experimental results show that U-net generalizes across environments, providing improved angular precision, offering a new paradigm for dense spatial audio understanding beyond traditional Sound Source Localization (SSL).


【8】Ambisonics Super-Resolution Using A Waveform-Domain Neural Network
标题:使用波域神经网络的立体声立体声超分辨率
链接:http://arxiv.org/pdf/2508.00240v1

作者:Ismael Nawfal, Symeon Delikaris Manias, Mehrez Souden, Juha Merimaa, Joshua Atkins, Elisabeth McMullin, Shadi Pirhosseinloo, Daniel Phillips
摘要:高保真度立体声是描述声场的空间音频格式。一阶高保真度立体声(FOA)是一种流行的格式,仅包括四个通道。这种有限的通道数是以牺牲空间精度为代价的。理想情况下,人们将能够采取FOA格式的效率,而不受其限制。我们设计了一种数据驱动的空间音频解决方案,它保留了FOA格式的效率,但质量超过了传统的渲染器。利用全卷积时域音频神经网络(Conv-TasNet),我们创建了一个解决方案,该解决方案采用FOA输入并提供高阶高保真度立体声(HOA)输出。这种数据驱动的方法是新颖的,相比典型的物理和心理声学的渲染。定量评估显示预测和实际三阶HOA之间的平均位置均方误差差为0.6dB。中位数的定性评级显示了80%的改善,在感知质量比传统的渲染方法。
摘要:Ambisonics is a spatial audio format describing a sound field. First-order Ambisonics (FOA) is a popular format comprising only four channels. This limited channel count comes at the expense of spatial accuracy. Ideally one would be able to take the efficiency of a FOA format without its limitations. We have devised a data-driven spatial audio solution that retains the efficiency of the FOA format but achieves quality that surpasses conventional renderers. Utilizing a fully convolutional time-domain audio neural network (Conv-TasNet), we created a solution that takes a FOA input and provides a higher order Ambisonics (HOA) output. This data driven approach is novel when compared to typical physics and psychoacoustic based renderers. Quantitative evaluations showed a 0.6dB average positional mean squared error difference between predicted and actual 3rd order HOA. The median qualitative rating showed an 80% improvement in perceived quality over the traditional rendering approach.


eess.AS音频处理


【1】Subband Architecture Aided Selective Fixed-Filter Active Noise Control
标题:子带架构辅助选择性固定过滤器主动噪音控制
链接:http://arxiv.org/pdf/2508.00603v1

作者:Hong-Cheng Liang, Man-Wai Mak, Kong Aik Lee
摘要:前馈选择性固定滤波器方法根据检测到的参考信号的频谱特征选择最合适的预训练控制滤波器,有效避免了传统自适应算法收敛速度慢的问题。然而,它只能处理有限类型的噪声,并且当输入噪声呈现非均匀功率谱密度时,性能会下降。为了解决这些限制,本文设计了一种新的选择性固定滤波器方案的基础上的无延迟子带结构。在离线训练阶段,子带控制滤波器被预先训练用于不同的频率范围,并存储在专用的子滤波器数据库中。在在线控制阶段,使用多相FFT滤波器组分解输入噪声,并且频带匹配机制为每个子带信号分配最合适的控制滤波器。随后,采用权重堆叠技术将所有子带权重组合成全带滤波器,从而实现实时噪声抑制。实验结果表明,该方法具有收敛速度快、降噪效果好、鲁棒性强等优点。
摘要:The feedforward selective fixed-filter method selects the most suitable pre-trained control filter based on the spectral features of the detected reference signal, effectively avoiding slow convergence in conventional adaptive algorithms. However, it can only handle limited types of noises, and the performance degrades when the input noise exhibits non-uniform power spectral density. To address these limitations, this paper devises a novel selective fixed-filter scheme based on a delayless subband structure. In the off-line training stage, subband control filters are pre-trained for different frequency ranges and stored in a dedicated sub-filter database. During the on-line control stage, the incoming noise is decomposed using a polyphase FFT filter bank, and a frequency-band-matching mechanism assigns each subband signal the most appropriate control filter. Subsequently, a weight stacking technique is employed to combine all subband weights into a fullband filter, enabling real-time noise suppression. Experimental results demonstrate that the proposed scheme provides fast convergence, effective noise reduction, and strong robustness in handling more complicated noisy environments.


【2】Dynamic Real-Time Ambisonics Order Adaptation for Immersive Networked Music Performances
标题:沉浸式网络音乐表演的动态实时立体声立体声顺序自适应
链接:http://arxiv.org/pdf/2508.00509v1

作者:Paolo Ostan, Carlo Centofanti, Mirco Pezzoli, Alberto Bernardini, Claudia Rinaldi, Fabio Antonacci
备注:to appear in EUSIPCO 2025
摘要:网络音乐表演(NMP)等高级远程应用需要解决方案来保证用户之间的沉浸式真实世界般的交互。因此,采用空间音频格式(如Ambisonics)是让用户体验沉浸式声学场景的基础。声音场景再现的准确度随着高保真度立体声编码的阶数而增加,从而以更大数量的音频通道为代价而导致改善的沉浸感,这又升级了带宽要求和对网络损伤(例如,延迟、抖动和分组丢失)。这些因素对要求高空间保真度和低端到端延迟的交互式音乐会话提出了重大挑战。我们提出了一种实时自适应高阶高保真度立体声响复制策略,连续监控网络吞吐量和动态缩放的高保真度立体声响复制顺序。当可用带宽下降到预设阈值以下时,会降低阶数以防止音频丢失;然后,一旦条件恢复,它会恢复到更高的阶数,从而平衡沉浸感和可靠性。基于MUSHRA的评估表明,这种自适应的方法是有希望的,以保证在带宽有限的NMP的情况下的用户体验。
摘要:Advanced remote applications such as Networked Music Performance (NMP) require solutions to guarantee immersive real-world-like interaction among users. Therefore, the adoption of spatial audio formats, such as Ambisonics, is fundamental to let the user experience an immersive acoustic scene. The accuracy of the sound scene reproduction increases with the order of the Ambisonics enconding, resulting in an improved immersivity at the cost of a greater number of audio channels, which in turn escalates both bandwidth requirements and susceptibility to network impairments (e.g., latency, jitter, and packet loss). These factors pose a significant challenge for interactive music sessions, which demand high spatial fidelity and low end-to-end delay. We propose a real-time adaptive higher-order Ambisonics strategy that continuously monitors network throughput and dynamically scales the Ambisonics order. When available bandwidth drops below a preset threshold, the order is lowered to prevent audio dropouts; it then reverts to higher orders once conditions recover, thus balancing immersion and reliability. A MUSHRA-based evaluation indicates that this adaptive approach is promising to guarantee user experience in bandwidth-limited NMP scenarios.


【3】VR-PTOLEMAIC: A Virtual Environment for the Perceptual Testing of Spatial Audio Algorithms
标题:空间音频算法感知测试的虚拟环境VR-PPEMAIC
链接:http://arxiv.org/pdf/2508.00501v1

作者:Paolo Ostan, Francesca Del Gaudio, Federico Miotello, Mirco Pezzoli, Fabio Antonacci

备注:to appear in EAA Forum Acusticum 2025

摘要:空间音频算法的感知评估是沉浸式音频应用开发的重要一步,因为它确保合成声场在聆听体验、空间感知和听觉真实性方面符合质量标准。为了支持这些评估,虚拟现实可以通过提供沉浸式和交互式测试环境来提供强大的平台。在本文中,我们提出了VR-PADMAIC,一个虚拟现实评估系统,旨在评估空间音频算法。该系统将MUSHRA(具有隐藏参考和锚点的多刺激测试)评估方法实施到虚拟环境中。特别地,用户可以将他们自己定位在虚拟重建的研讨会房间的25个模拟收听位置中的每一个中,并且相对于实际记录的二阶立体混响房间脉冲响应来评估模拟声学响应,所有这些都与各种源信号卷积。我们通过广泛的测试活动,评估人员被要求比较各种声场重建算法的重建能力,评估了所提出的框架的可用性。结果表明,虚拟现实平台有效地支持空间音频算法的评估,用户体验和沉浸感普遍积极的反馈。
摘要:The perceptual evaluation of spatial audio algorithms is an important step in the development of immersive audio applications, as it ensures that synthesized sound fields meet quality standards in terms of listening experience, spatial perception and auditory realism. To support these evaluations, virtual reality can offer a powerful platform by providing immersive and interactive testing environments. In this paper, we present VR-PTOLEMAIC, a virtual reality evaluation system designed for assessing spatial audio algorithms. The system implements the MUSHRA (MUlti-Stimulus test with Hidden Reference and Anchor) evaluation methodology into a virtual environment. In particular, users can position themselves in each of the 25 simulated listening positions of a virtually recreated seminar room and evaluate simulated acoustic responses with respect to the actually recorded second-order ambisonic room impulse responses, all convolved with various source signals. We evaluated the usability of the proposed framework through an extensive testing campaign in which assessors were asked to compare the reconstruction capabilities of various sound field reconstruction algorithms. Results show that the VR platform effectively supports the assessment of spatial audio algorithms, with generally positive feedback on user experience and immersivity.


【4】Wavelet-Based Time-Frequency Fingerprinting for Feature Extraction of Traditional Irish Music
标题:基于子波的时频指纹识别用于爱尔兰传统音乐特征提取
链接:http://arxiv.org/pdf/2508.00479v1

备注:Master's thesis. The focus of the thesis is on the underlying techniques for signal fingerprinting
摘要:这项工作提出了一种基于小波的时间序列特征提取的时频指纹识别方法,重点是从现场录音的传统爱尔兰曲调的音频识别。识别时间序列数据中的特征的挑战是通过采用连续小波变换来提取频谱特征和小波相干分析来比较记录的音频频谱图合成生成的曲调。合成曲调来自ABC记谱法,这是爱尔兰音乐的一种常见符号表示。实验结果表明,基于小波变换的方法能够准确、有效地识别录音曲调。本研究还详细介绍了小波相干模型的性能,突出了其优于其他时频分解方法的优势。此外,我们还讨论了该模型,并将其部署在音乐以外的几个应用程序中,包括EEG信号分析和金融时间序列预测。
摘要:This work presents a wavelet-based approach to time-frequency fingerprinting for time series feature extraction, with a focus on audio identification from live recordings of traditional Irish tunes. The challenges of identifying features in time-series data are addressed by employing a continuous wavelet transform to extract spectral features and wavelet coherence analysis is used to compare recorded audio spectrograms to synthetically generated tunes. The synthetic tunes are derived from ABC notation, which is a common symbolic representation for Irish music. Experimental results demonstrate that the wavelet-based method can accurately and efficiently identify recorded tunes. This research study also details the performance of the wavelet coherence model, highlighting its strengths over other methods of time-frequency decomposition. Additionally, we discuss and deploy the model on several applications beyond music, including in EEG signal analysis and financial time series forecasting.


【5】Beamformed 360° Sound Maps: U-Net-Driven Acoustic Source Segmentation and Localization
标题:束成形360°声音地图:U-Net驱动的光源分割和定位
链接:http://arxiv.org/pdf/2508.00307v1

作者:Belman J. Rodriguez, Sergio F. Chevtchenko, Marcelo Herrera Martinez, Yeshwant Bethy, Saeed Afshar
摘要:我们介绍了一个U-网络模型的360{ deg}声源定位制定为一个球形的语义分割任务。我们的模型不是回归离散的到达方向(DoA)角度,而是将波束成形的音频地图(方位角和仰角)分割成活动声音存在的区域。在定制的24麦克风阵列上使用延迟求和(DAS)波束成形,我们生成与无人机GPS遥测对齐的信号,以创建二进制监督掩码。修改后的U-Net在这些地图的频域表示上进行训练,学习识别空间分布的源区域,同时通过Tversky损失解决类别不平衡问题。由于该网络在波束成形能量图上运行,因此该方法本质上与阵列无关,并且可以适应不同的麦克风配置,而无需从头开始重新训练。通过计算激活区域上的质心对分割输出进行后处理,从而实现稳健的DoA估计。我们的数据集包括DJI Air 3无人机的真实野外记录,与多个日期和地点的360度视频和飞行日志同步。实验结果表明,U-net可以跨环境进行推广,提供更高的角度精度,为传统声源定位(SSL)之外的密集空间音频理解提供了新的范例。
摘要:We introduce a U-net model for 360{ deg} acoustic source localization formulated as a spherical semantic segmentation task. Rather than regressing discrete direction-of-arrival (DoA) angles, our model segments beamformed audio maps (azimuth and elevation) into regions of active sound presence. Using delay-and-sum (DAS) beamforming on a custom 24-microphone array, we generate signals aligned with drone GPS telemetry to create binary supervision masks. A modified U-Net, trained on frequency-domain representations of these maps, learns to identify spatially distributed source regions while addressing class imbalance via the Tversky loss. Because the network operates on beamformed energy maps, the approach is inherently array-independent and can adapt to different microphone configurations without retraining from scratch. The segmentation outputs are post-processed by computing centroids over activated regions, enabling robust DoA estimates. Our dataset includes real-world open-field recordings of a DJI Air 3 drone, synchronized with 360{ deg} video and flight logs across multiple dates and locations. Experimental results show that U-net generalizes across environments, providing improved angular precision, offering a new paradigm for dense spatial audio understanding beyond traditional Sound Source Localization (SSL).


【6】Ambisonics Super-Resolution Using A Waveform-Domain Neural Network
标题:使用波域神经网络的立体声立体声超分辨率
链接:http://arxiv.org/pdf/2508.00240v1

作者:Ismael Nawfal, Symeon Delikaris Manias, Mehrez Souden, Juha Merimaa, Joshua Atkins, Elisabeth McMullin, Shadi Pirhosseinloo, Daniel Phillips
摘要:高保真度立体声是描述声场的空间音频格式。一阶高保真度立体声(FOA)是一种流行的格式,仅包括四个通道。这种有限的通道数是以牺牲空间精度为代价的。理想情况下,人们将能够采取FOA格式的效率,而不受其限制。我们设计了一种数据驱动的空间音频解决方案,它保留了FOA格式的效率,但质量超过了传统的渲染器。利用全卷积时域音频神经网络(Conv-TasNet),我们创建了一个解决方案,该解决方案采用FOA输入并提供高阶高保真度立体声(HOA)输出。这种数据驱动的方法是新颖的,相比典型的物理和心理声学的渲染。定量评估显示预测和实际三阶HOA之间的平均位置均方误差差为0.6dB。中位数的定性评级显示了80%的改善,在感知质量比传统的渲染方法。
摘要:Ambisonics is a spatial audio format describing a sound field. First-order Ambisonics (FOA) is a popular format comprising only four channels. This limited channel count comes at the expense of spatial accuracy. Ideally one would be able to take the efficiency of a FOA format without its limitations. We have devised a data-driven spatial audio solution that retains the efficiency of the FOA format but achieves quality that surpasses conventional renderers. Utilizing a fully convolutional time-domain audio neural network (Conv-TasNet), we created a solution that takes a FOA input and provides a higher order Ambisonics (HOA) output. This data driven approach is novel when compared to typical physics and psychoacoustic based renderers. Quantitative evaluations showed a 0.6dB average positional mean squared error difference between predicted and actual 3rd order HOA. The median qualitative rating showed an 80% improvement in perceived quality over the traditional rendering approach.


【7】Melody-Lyrics Matching with Contrastive Alignment Loss
标题:旋律-歌词匹配与对比对齐损失
链接:http://arxiv.org/pdf/2508.00123v1

作者:Changhong Wang, Michel Olvera, Gaël Richard
备注:10 pages, 7 figures, 3 tables. This work has been submitted to the IEEE for possible publication
摘要:音乐和歌词之间的联系远远超出了语义纽带。节奏与韵律、音长与音节重音、结构对应等概念对是音乐信息检索领域中一个引人注目而又很少探索的方向。在本文中,我们提出了旋律歌词匹配(MLM),一个新的任务,检索潜在的歌词为一个给定的符号旋律从文本源。而不是从零开始产生歌词,传销基本上利用旋律和歌词之间的关系。我们提出了一个自我监督的表示学习框架与对比对齐损失的旋律和歌词。这有可能利用丰富的现有歌曲与配对的旋律和歌词。不需要路线注释。此外,我们介绍了sylphone,一种新的表示歌词在音节水平激活的音素身份和元音重音。我们证明,我们的方法可以匹配旋律与连贯和可唱的歌词与经验的结果和直观的例子。我们开放源代码,并在配套网页上提供匹配的示例:https: github.com changhongw mlm。
摘要:The connection between music and lyrics is far beyond semantic bonds. Conceptual pairs in the two modalities such as rhythm and rhyme, note duration and syllabic stress, and structure correspondence, raise a compelling yet seldom-explored direction in the field of music information retrieval. In this paper, we present melody-lyrics matching (MLM), a new task which retrieves potential lyrics for a given symbolic melody from text sources. Rather than generating lyrics from scratch, MLM essentially exploits the relationships between melody and lyrics. We propose a self-supervised representation learning framework with contrastive alignment loss for melody and lyrics. This has the potential to leverage the abundance of existing songs with paired melody and lyrics. No alignment annotations are required. Additionally, we introduce sylphone, a novel representation for lyrics at syllable-level activated by phoneme identity and vowel stress. We demonstrate that our method can match melody with coherent and singable lyrics with empirical results and intuitive examples. We open source code and provide matching examples on the companion webpage: https: github.com changhongw mlm.


【8】SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
标题:SpA 2V:利用空间听觉线索实现音频驱动的空间感知视频生成
链接:http://arxiv.org/pdf/2508.00782v1

作者:Kien T. Pham, Yingqing He, Yazhou Xing, Qifeng Chen, Long Chen

备注:The 33rd ACM Multimedia Conference (MM '25)

摘要:音频驱动的视频生成旨在合成与输入音频记录一致的逼真视频,类似于人类从听觉输入中可视化场景的能力。然而,现有的方法主要集中于探索语义信息,例如音频中存在的声音源的类别,限制了它们生成具有准确内容和空间组成的视频的能力。相比之下,我们人类不仅可以自然地识别声源的语义类别,而且还可以确定其深度编码的空间属性,包括位置和运动方向。这种有用的信息可以通过考虑从声音的固有物理特性(如响度或频率)导出的特定空间指标来阐明。由于先前的方法在很大程度上忽略了这个因素,我们提出了SpA2V,第一个框架明确地利用这些空间听觉线索从音频生成具有高语义和空间对应性的视频。SpA2V将生成过程分解为两个阶段:1)音频引导的视频规划:我们精心调整了最先进的MLLM,用于利用来自输入音频的空间和语义线索来构建视频场景布局(VSL)的新任务。这用作中间表示,以弥合音频和视频模态之间的差距。2)基于布局的视频生成:我们开发了一种高效有效的方法,将VSL作为条件指导无缝集成到预先训练的扩散模型中,从而以免训练的方式实现基于VSL的视频生成。大量的实验表明,SpA2V擅长生成逼真的视频与语义和空间对齐的输入音频。
摘要:Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring semantic information, such as the classes of sounding sources present in the audio, limiting their ability to generate videos with accurate content and spatial composition. In contrast, we humans can not only naturally identify the semantic categories of sounding sources but also determine their deeply encoded spatial attributes, including locations and movement directions. This useful information can be elucidated by considering specific spatial indicators derived from the inherent physical properties of sound, such as loudness or frequency. As prior methods largely ignore this factor, we present SpA2V, the first framework explicitly exploits these spatial auditory cues from audios to generate videos with high semantic and spatial correspondence. SpA2V decomposes the generation process into two stages: 1) Audio-guided Video Planning: We meticulously adapt a state-of-the-art MLLM for a novel task of harnessing spatial and semantic cues from input audio to construct Video Scene Layouts (VSLs). This serves as an intermediate representation to bridge the gap between the audio and video modalities. 2) Layout-grounded Video Generation: We develop an efficient and effective approach to seamlessly integrate VSLs as conditional guidance into pre-trained diffusion models, enabling VSL-grounded video generation in a training-free manner. Extensive experiments demonstrate that SpA2V excels in generating realistic videos with semantic and spatial alignment to the input audios.


【9】AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

标题:AudioGen-Omni:用于视频同步音频、语音和歌曲生成的统一多模式扩散Transformer

链接:http://arxiv.org/pdf/2508.00733v1
作者:Le Wang,Jun Wang, Feng Deng, Chen Zhang, Kun Gai, Di Zhang

备注:12 pages, 2 figures

摘要:我们提出了AudioGen-Omni -一种基于多模态扩散Transformers(MMDit)的统一方法,能够生成与输入视频同步的高保真音频,语音和歌曲。AudioGen-Omni引入了一种新型的联合训练范式,该范式无缝集成了大规模的视频-文本-音频语料库,使模型能够在多模态输入的条件下生成语义丰富、声学多样的音频,并适应各种音频生成任务。AudioGen-Omni采用统一的歌词转录编码器,将来自演唱和口语输入的字素和音素编码为密集的帧级表示。密集的帧级表示融合使用基于AdaLN的联合注意力机制增强与相位对齐的各向异性位置注入(PAAPI),其中RoPE被选择性地应用于时间结构化的模态,以确保精确和鲁棒的跨模态对齐。通过解冻所有模态并掩蔽缺失的输入,AudioGen-Omni减轻了文本冻结范式的语义限制,从而实现有效的跨模态条件反射。这种联合训练方法提高了音频质量、语义对齐和嘴唇同步准确性,同时还在文本到音频 语音 歌曲任务上实现了最先进的结果。对于8秒的音频,推理时间为1.91秒,它在效率和通用性方面都有很大的提高。
摘要:We present AudioGen-Omni - a unified approach based on multimodal diffusion transformers (MMDit), capable of generating high-fidelity audio, speech, and songs coherently synchronized with the input video. AudioGen-Omni introduces a novel joint training paradigm that seamlessly integrates large-scale video-text-audio corpora, enabling a model capable of generating semantically rich, acoustically diverse audio conditioned on multimodal inputs and adaptable to a wide range of audio generation tasks. AudioGen-Omni employs a unified lyrics-transcription encoder that encodes graphemes and phonemes from both sung and spoken inputs into dense frame-level representations. Dense frame-level representations are fused using an AdaLN-based joint attention mechanism enhanced with phase-aligned anisotropic positional infusion (PAAPI), wherein RoPE is selectively applied to temporally structured modalities to ensure precise and robust cross-modal alignment. By unfreezing all modalities and masking missing inputs, AudioGen-Omni mitigates the semantic constraints of text-frozen paradigms, enabling effective cross-modal conditioning. This joint training approach enhances audio quality, semantic alignment, and lip-sync accuracy, while also achieving state-of-the-art results on Text-to-Audio Speech Song tasks. With an inference time of 1.91 seconds for 8 seconds of audio, it offers substantial improvements in both efficiency and generality.


【10】Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition
标题:提示代理:一种用于自动提示语音识别的协作多代理系统
链接:http://arxiv.org/pdf/2508.00391v1

作者:Guanjie Huang, Danny H.K. Tsang, Shan Yang, Guangzhi Lei, Li Liu
备注:9 pages
摘要:提示语音(CS)是一种视觉交流系统,结合唇读和手势编码,以促进听力障碍者的交流。自动CS识别(ACSR)旨在通过AI驱动的方法将CS手势和嘴唇运动转换为文本。传统上,手和嘴唇运动之间的时间相关性需要设计复杂的模块来促进有效的多模态融合。然而,有限的数据可用性的约束,目前的方法表现出足够的能力,充分训练这些融合机制,导致次优性能。最近,多智能体系统在处理有限数据可用性的复杂任务方面表现出了很好的能力。为此,我们提出了第一个合作的多智能体系统ACSR,命名为线索代理。它集成了四个专门的子代理:采用关键帧筛选和CS专家提示策略来解码手部运动的基于多模态大语言模型的手部识别代理,从输入视频中提取唇部特征的预训练的基于变换器的唇部识别代理,在推理期间以无训练方式动态地将手部提示与唇部特征集成的手部提示解码代理,以及自校正音素到单词代理,其首次通过语义细化实现从音素序列到自然语言句子的后处理和端到端转换。为了支持这项研究,我们扩展了现有的普通话CS数据集,收集数据从八个听力受损的cuers,建立一个混合数据集的十四个主题。大量的实验表明,我们的线索代理表现出色,在正常和听力受损的情况下相比,国家的最先进的方法。该实现可在https: github.com DennisHgj Cued-Agent上获得。
摘要:Cued Speech (CS) is a visual communication system that combines lip-reading with hand coding to facilitate communication for individuals with hearing impairments. Automatic CS Recognition (ACSR) aims to convert CS hand gestures and lip movements into text via AI-driven methods. Traditionally, the temporal asynchrony between hand and lip movements requires the design of complex modules to facilitate effective multimodal fusion. However, constrained by limited data availability, current methods demonstrate insufficient capacity for adequately training these fusion mechanisms, resulting in suboptimal performance. Recently, multi-agent systems have shown promising capabilities in handling complex tasks with limited data availability. To this end, we propose the first collaborative multi-agent system for ACSR, named Cued-Agent. It integrates four specialized sub-agents: a Multimodal Large Language Model-based Hand Recognition agent that employs keyframe screening and CS expert prompt strategies to decode hand movements, a pretrained Transformer-based Lip Recognition agent that extracts lip features from the input video, a Hand Prompt Decoding agent that dynamically integrates hand prompts with lip features during inference in a training-free manner, and a Self-Correction Phoneme-to-Word agent that enables post-process and end-to-end conversion from phoneme sequences to natural language sentences for the first time through semantic refinement. To support this study, we expand the existing Mandarin CS dataset by collecting data from eight hearing-impaired cuers, establishing a mixed dataset of fourteen subjects. Extensive experiments demonstrate that our Cued-Agent performs superbly in both normal and hearing-impaired scenarios compared with state-of-the-art methods. The implementation is available at https: github.com DennisHgj Cued-Agent.


【11】Advancing Speech Quality Assessment Through Scientific Challenges and Open-source Activities
标题:通过科学挑战和开源活动推进语音质量评估
链接:http://arxiv.org/pdf/2508.00317v1

作者:Wen-Chin Huang
备注:APSIPA ASC 2025 perspective paper
摘要:语音质量评估(SQA)是指对语音质量的评估,为了跟上生成式人工智能的热潮,开发一种准确的自动SQA方法来反映人类的感知变得越来越重要。近年来,SQA已经发展到一个点,研究人员开始忠实地使用自动SQA在研究论文作为一个严格的衡量良好的语音生成系统。我们认为,最近的科学挑战和开放源码活动刺激了这一领域的增长。在本文中,我们回顾了最近的挑战以及SQA的开源实现和工具包,并强调了维护这些活动的重要性,以促进SQA本身以及语音生成AI的发展。
摘要:Speech quality assessment (SQA) refers to the evaluation of speech quality, and developing an accurate automatic SQA method that reflects human perception has become increasingly important, in order to keep up with the generative AI boom. In recent years, SQA has progressed to a point that researchers started to faithfully use automatic SQA in research papers as a rigorous measurement of goodness for speech generation systems. We believe that the scientific challenges and open-source activities of late have stimulated the growth in this field. In this paper, we review recent challenges as well as open-source implementations and toolkits for SQA, and highlight the importance of maintaining such activities to facilitate the development of not only SQA itself but also generative AI for speech.


【12】Audio Prototypical Network For Controllable Music Recommendation
标题:可控音乐推荐的音频原型网络
链接:http://arxiv.org/pdf/2508.00194v1

作者:Fırat Oncel, Emiliano Penaloza, Haolun Wu, Shubham Gupta, Mirco Ravanelli, Laurent Charlin, Cem Subakan
备注:Accepted to MLSP2025
摘要:传统的推荐系统通过黑盒编码器模型获得密集表示来表示用户偏好。虽然这些模型通常提供强大的推荐性能,但它们缺乏对用户的可解释性,使用户无法理解或控制系统对其偏好的建模。这种限制在音乐推荐中尤其具有挑战性,其中用户偏好是高度个人化的,并且通常基于情绪,流派,节奏或乐器等细微差别的质量而变化。在本文中,我们提出了一个音频原型网络的可控音乐推荐。这个网络表示用户的喜好,在原型的语义有意义的功能有关的音乐品质的代表。我们表明,该模型获得竞争力的推荐性能相比,流行的基线模型,同时还提供可解释和可控的用户配置文件。
摘要:Traditional recommendation systems represent user preferences in dense representations obtained through black-box encoder models. While these models often provide strong recommendation performance, they lack interpretability for users, leaving users unable to understand or control the system's modeling of their preferences. This limitation is especially challenging in music recommendation, where user preferences are highly personal and often evolve based on nuanced qualities like mood, genre, tempo, or instrumentation. In this paper, we propose an audio prototypical network for controllable music recommendation. This network expresses user preferences in terms of prototypes representative of semantically meaningful features pertaining to musical qualities. We show that the model obtains competitive recommendation performance compared to popular baseline models while also providing interpretable and controllable user profiles.


【13】DeformTune: A Deformable XAI Music Prototype for Non-Musicians
标题:变形曲调:适合非音乐家的可变形XAI音乐原型
链接:http://arxiv.org/pdf/2508.00160v1

作者:Ziqing Xu, Nick Bryan-Kinns

备注:In Proceedings of Explainable AI for the Arts Workshop 2025 (XAIxArts 2025) arXiv:2406.14485

摘要:许多现有的人工智能音乐生成工具依赖于文本提示、复杂的界面或类似乐器的控件,这可能需要非音乐家所不具备的音乐或技术知识。本文介绍了DeformTune,这是一个原型系统,它将触觉可变形界面与MeasureVAE模型相结合,以探索更直观,更具体,更可解释的AI交互。我们对11名没有接受过正式音乐训练的成年参与者进行了初步研究,以调查他们在人工智能辅助音乐创作方面的经验。对他们反馈的主题分析揭示了反复出现的挑战-包括不清楚的控制映射,有限的表达范围,以及在整个使用过程中需要指导。我们讨论了几个增强AI可解释性的设计机会,包括多模态反馈和渐进式交互支持。这些发现为使AI音乐系统更易于解释和授权新手用户提供了早期见解。
摘要:Many existing AI music generation tools rely on text prompts, complex interfaces, or instrument-like controls, which may require musical or technical knowledge that non-musicians do not possess. This paper introduces DeformTune, a prototype system that combines a tactile deformable interface with the MeasureVAE model to explore more intuitive, embodied, and explainable AI interaction. We conducted a preliminary study with 11 adult participants without formal musical training to investigate their experience with AI-assisted music creation. Thematic analysis of their feedback revealed recurring challenge--including unclear control mappings, limited expressive range, and the need for guidance throughout use. We discuss several design opportunities for enhancing explainability of AI, including multimodal feedback and progressive interaction support. These findings contribute early insights toward making AI music systems more explainable and empowering for novice users.


机器翻译由腾讯交互翻译提供,仅供参考