微信公众号:arXiv_Daily
cs.SD语音
【1】Representing Classical Compositions through Implication-Realization Temporal-Gestalt Graphs
标题:通过含义实现时间完形图表示古典作品
链接:https://arxiv.org/abs/2510.27530
备注:8 pages, 11 figures
摘要:理解音乐作品的结构和认知基础仍然是音乐理论和计算音乐学的关键挑战。虽然传统的方法侧重于和声和节奏,但认知模型,如含义-实现(I-R)模型和时间完形理论,可以深入了解听众如何感知和预测音乐结构。本研究提出了一种基于图的计算方法,通过将旋律分割成感知单元并用I-R模式对其进行注释来实现这些模型。这些段使用动态时间规整进行比较,并组织成k-最近邻图,以模拟段内和段间的关系。 每个片段在图中被表示为节点,并且节点进一步用来自Schellenberg的双因素I-R模型的旋律期望值来标记-在片段水平上量化音高接近度和音高反转。这种标记使图形能够编码结构和认知信息,反映听众如何体验音乐的张力和解决方案。 为了评估这些图的表现力,我们应用Weisfeiler-Lehman图内核来测量组合物之间和组合物内的相似性。结果显示,统计上显着的区别内和图间结构。通过多维尺度的段级分析证实,在图形级别的结构相似性反映在段级的感知相似性。Graph 2 vec嵌入和聚类表明,这些表示捕捉到了超越作曲家身份的风格和结构特征。 这些研究结果突出了基于图形的方法作为计算音乐分析的结构化,认知信息框架的潜力,通过听众感知的镜头,使音乐结构和风格的理解更加细致入微。
摘要:Understanding the structural and cognitive underpinnings of musical compositions remains a key challenge in music theory and computational musicology. While traditional methods focus on harmony and rhythm, cognitive models such as the Implication-Realization (I-R) model and Temporal Gestalt theory offer insight into how listeners perceive and anticipate musical structure. This study presents a graph-based computational approach that operationalizes these models by segmenting melodies into perceptual units and annotating them with I-R patterns. These segments are compared using Dynamic Time Warping and organized into k-nearest neighbors graphs to model intra- and inter-segment relationships. Each segment is represented as a node in the graph, and nodes are further labeled with melodic expectancy values derived from Schellenberg's two-factor I-R model-quantifying pitch proximity and pitch reversal at the segment level. This labeling enables the graphs to encode both structural and cognitive information, reflecting how listeners experience musical tension and resolution. To evaluate the expressiveness of these graphs, we apply the Weisfeiler-Lehman graph kernel to measure similarity between and within compositions. Results reveal statistically significant distinctions between intra- and inter-graph structures. Segment-level analysis via multidimensional scaling confirms that structural similarity at the graph level reflects perceptual similarity at the segment level. Graph2vec embeddings and clustering demonstrate that these representations capture stylistic and structural features that extend beyond composer identity. These findings highlight the potential of graph-based methods as a structured, cognitively informed framework for computational music analysis, enabling a more nuanced understanding of musical structure and style through the lens of listener perception.
【2】Expressive Range Characterization of Open Text-to-Audio Models
标题:开放文本到音频模型的表达范围特征
链接:https://arxiv.org/abs/2510.27102
备注:Accepted at the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE 2025)
摘要:文本到音频模型是一种生成模型,其响应于给定的文本提示而产生音频输出。虽然级别生成器和它们创建的功能内容的属性(例如,可玩性)主导程序生成内容(PCG)中的大多数话语,但在情感上与玩家产生共鸣的游戏倾向于将一系列创造性和多模态内容编织在一起(例如,音乐、声音、视觉、叙述语气),并且多模态模型已经开始看到至少用于该目的的实验性使用。然而,目前还不清楚这些模型究竟生成什么,以及具有多大程度的可变性和保真度:音频是生成系统的目标输出的一个非常广泛的类别。 在PCG社区,表达范围分析(ERA)已被用作一种定量的方法来表征发电机的输出空间,特别是水平发电机。本文适应ERA的文本到音频模型,通过查看特定的,固定的提示输出的表达范围,使分析易于处理。实验进行了提示模型与几个标准化的提示来自环境声音分类(ESC-50)数据集。沿着关键声学维度(例如,音高、响度和音色)。更广泛地说,本文提供了一个基于ERA的生成音频模型的探索性评估框架。
摘要:Text-to-audio models are a type of generative model that produces audio output in response to a given textual prompt. Although level generators and the properties of the functional content that they create (e.g., playability) dominate most discourse in procedurally generated content (PCG), games that emotionally resonate with players tend to weave together a range of creative and multimodal content (e.g., music, sounds, visuals, narrative tone), and multimodal models have begun seeing at least experimental use for this purpose. However, it remains unclear what exactly such models generate, and with what degree of variability and fidelity: audio is an extremely broad class of output for a generative system to target. Within the PCG community, expressive range analysis (ERA) has been used as a quantitative way to characterize generators' output space, especially for level generators. This paper adapts ERA to text-to-audio models, making the analysis tractable by looking at the expressive range of outputs for specific, fixed prompts. Experiments are conducted by prompting the models with several standardized prompts derived from the Environmental Sound Classification (ESC-50) dataset. The resulting audio is analyzed along key acoustic dimensions (e.g., pitch, loudness, and timbre). More broadly, this paper offers a framework for ERA-based exploratory evaluation of generative audio models.
【3】Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling
标题:复杂场景下分离与去混响联合建模的音视频语音增强
链接:https://arxiv.org/abs/2510.26825
摘要:视听语音增强(AVSE)是利用视觉辅助信息从混合音频中提取目标说话人的语音。在现实生活中,往往存在着复杂的声学环境,伴随着各种干扰声和混响。大多数以前的方法都难以应付这种复杂的条件,导致提取的语音的感知质量差。在本文中,我们提出了一个有效的AVSE系统,在复杂的声学环境中表现良好。具体来说,我们设计了一个“去混响前分离”的管道,可以扩展到其他AVSE网络。第四届COGMHEAR视听语音增强挑战赛(AVSEC)旨在探索多模态复杂环境中语音处理的新方法。我们在AVSEC-4中验证了我们系统的性能:我们在竞争排行榜上的三个客观指标中取得了优异的成绩,并最终在人类主观听力测试中获得了第一名。
摘要:Audio-visual speech enhancement (AVSE) is a task that uses visual auxiliary information to extract a target speaker's speech from mixed audio. In real-world scenarios, there often exist complex acoustic environments, accompanied by various interfering sounds and reverberation. Most previous methods struggle to cope with such complex conditions, resulting in poor perceptual quality of the extracted speech. In this paper, we propose an effective AVSE system that performs well in complex acoustic environments. Specifically, we design a "separation before dereverberation" pipeline that can be extended to other AVSE networks. The 4th COGMHEAR Audio-Visual Speech Enhancement Challenge (AVSEC) aims to explore new approaches to speech processing in multimodal complex environments. We validated the performance of our system in AVSEC-4: we achieved excellent results in the three objective metrics on the competition leaderboard, and ultimately secured first place in the human subjective listening test.
【4】Cross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features
标题:使用领域知识声学特征进行乌尔都语语音情感识别的跨Corpus验证
链接:https://arxiv.org/abs/2510.26823
备注:Conference paper, 4 pages, including 3 figures and 3 tables
摘要:语音情感识别(SER)是实现情感智能人工智能的关键情感计算技术。虽然SER总体上具有挑战性,但对于乌尔都语等低资源语言来说尤其困难。本研究探讨乌尔都语SER在跨语料库设置,一个领域,在很大程度上仍然是未开发的。我们采用跨语料库的评估框架,在三个不同的乌尔都语情感语音数据集测试模型的泛化。两个标准的领域知识为基础的声学特征集,eGeMAPS和ComparE,用于表示语音信号的特征向量,然后通过逻辑回归和多层感知器分类。分类性能评估使用未加权平均召回(UAR),同时考虑类标签的不平衡。结果表明,自语料库验证往往高估了性能,UAR超过跨语料库评估高达13%,强调跨语料库评估提供了一个更现实的衡量模型的鲁棒性。总体而言,这项工作强调了跨语料库验证的重要性,乌尔都语SER及其影响有助于推进情感计算研究的代表性不足的语言社区。
摘要:Speech Emotion Recognition (SER) is a key affective computing technology that enables emotionally intelligent artificial intelligence. While SER is challenging in general, it is particularly difficult for low-resource languages such as Urdu. This study investigates Urdu SER in a cross-corpus setting, an area that has remained largely unexplored. We employ a cross-corpus evaluation framework across three different Urdu emotional speech datasets to test model generalization. Two standard domain-knowledge based acoustic feature sets, eGeMAPS and ComParE, are used to represent speech signals as feature vectors which are then passed to Logistic Regression and Multilayer Perceptron classifiers. Classification performance is assessed using unweighted average recall (UAR) whilst considering class-label imbalance. Results show that Self-corpus validation often overestimates performance, with UAR exceeding cross-corpus evaluation by up to 13%, underscoring that cross-corpus evaluation offers a more realistic measure of model robustness. Overall, this work emphasizes the importance of cross-corpus validation for Urdu SER and its implications contribute to advancing affective computing research for underrepresented language communities.
【5】GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
标题:GACA-DiT:基于扩散的舞蹈到音乐生成,具有流派自适应节奏和上下文感知对齐
链接:https://arxiv.org/abs/2510.26818
备注:5 pages, 3 figures, submitted to ICASSP 2026
摘要:舞蹈到音乐(D2 M)生成的目的是自动创作音乐,节奏和时间与舞蹈动作一致。现有的方法通常依赖于粗糙的节奏嵌入,例如全局运动特征或二进制化的基于关节的节奏值,其丢弃细粒度的运动线索并导致弱节奏对齐。此外,由特征下采样引入的时间失配进一步阻碍了舞蹈和音乐之间的精确同步。为了解决这些问题,我们提出了\textbf{GACA-DiT},一个基于扩散变换的框架,具有两个新颖的模块,用于节奏一致和时间对齐的音乐生成。首先,一个\textbf{genre-adaptive rhythm extraction}模块结合了多尺度时间小波分析和空间相位直方图与自适应联合加权,以捕获细粒度的,特定于流派的节奏模式。第二,一个\textbf{上下文感知的时间对齐}模块解决时间不匹配使用可学习的上下文查询对齐音乐潜伏期与相关的舞蹈节奏功能。在AIST++和TikTok数据集上进行的大量实验表明,GACA-DiT在客观指标和人工评估方面都优于最先进的方法。项目页面:https://beria-moon.github.io/GACA-DiT/。
摘要:Dance-to-music (D2M) generation aims to automatically compose music that is rhythmically and temporally aligned with dance movements. Existing methods typically rely on coarse rhythm embeddings, such as global motion features or binarized joint-based rhythm values, which discard fine-grained motion cues and result in weak rhythmic alignment. Moreover, temporal mismatches introduced by feature downsampling further hinder precise synchronization between dance and music. To address these problems, we propose \textbf{GACA-DiT}, a diffusion transformer-based framework with two novel modules for rhythmically consistent and temporally aligned music generation. First, a \textbf{genre-adaptive rhythm extraction} module combines multi-scale temporal wavelet analysis and spatial phase histograms with adaptive joint weighting to capture fine-grained, genre-specific rhythm patterns. Second, a \textbf{context-aware temporal alignment} module resolves temporal mismatches using learnable context queries to align music latents with relevant dance rhythm features. Extensive experiments on the AIST++ and TikTok datasets demonstrate that GACA-DiT outperforms state-of-the-art methods in both objective metrics and human evaluation. Project page: https://beria-moon.github.io/GACA-DiT/.
【6】Oral Tradition-Encoded NanyinHGNN: Integrating Nanyin Music Preservation and Generation through a Pipa-Centric Dataset
标题:口述编码南音HGNN:通过以Pipa为中心的数据集整合南音音乐保存和生成
链接:https://arxiv.org/abs/2510.26817
备注:10 pages, 2 figures
摘要:我们提出了南音HGNN,一个异构的图网络模型,用于生成南音器乐。作为联合国教科文组织认可的非物质文化遗产,南音遵循以琵琶为中心的异声传统,其中核心旋律用传统记谱法记谱,而表达则通过口头传递,这对保护和当代创新都提出了挑战。为了解决这个问题,我们构建了一个以Pipa为中心的数据集,开发了NanyinTok作为一种专门的标记化方法,并使用Graph Converter将符号序列转换为图形结构,以确保保留关键的音乐特征。我们的关键创新重新表述为异构图内的表示节点的创建表示生成。首先,一个图形神经网络生成优化的旋律轮廓。然后,由南音演奏实践告知的规则引导系统将这些轮廓细化为完整的表达,而不需要在训练期间进行显式的表达注释。实验结果表明,我们的模型成功地产生真实的四个传统乐器的杂音合奏。这些发现验证了将特定领域的知识整合到模型架构中可以有效地缓解计算民族音乐学中的数据稀缺挑战。
摘要:We propose NanyinHGNN, a heterogeneous graph network model for generating Nanyin instrumental music. As a UNESCO-recognized intangible cultural heritage, Nanyin follows a heterophonic tradition centered around the pipa, where core melodies are notated in traditional notation while ornamentations are passed down orally, presenting challenges for both preservation and contemporary innovation. To address this, we construct a Pipa-Centric MIDI dataset, develop NanyinTok as a specialized tokenization method, and convert symbolic sequences into graph structures using a Graph Converter to ensure that key musical features are preserved. Our key innovation reformulates ornamentation generation as the creation of ornamentation nodes within a heterogeneous graph. First, a graph neural network generates melodic outlines optimized for ornamentations. Then, a rule-guided system informed by Nanyin performance practices refines these outlines into complete ornamentations without requiring explicit ornamentation annotations during training. Experimental results demonstrate that our model successfully generates authentic heterophonic ensembles featuring four traditional instruments. These findings validate that integrating domain-specific knowledge into model architecture can effectively mitigate data scarcity challenges in computational ethnomusicology.
【7】Reference Microphone Selection for Guided Source Separation based on the Normalized L-p Norm
标题:基于标准化L-p规范的引导源分离参考麦克风选择
链接:https://arxiv.org/abs/2510.27198
摘要:引导源分离(GSS)是使用空间分布麦克风的远程自动语音识别(ASR)系统的流行前端。当考虑空间分布的麦克风时,参考麦克风的选择可能对输出信号的质量和下游ASR性能具有很大影响。在基于GSS的语音增强中,通常使用信噪比(SNR)来执行参考麦克风选择,SNR对于降噪是最佳的,但可能忽略麦克风之间的早到晚混响比(ELR)的差异。在本文中,我们提出了两个参考麦克风选择基于高斯的语音增强方法是基于归一化的$\ell_p $-范数,或者只使用归一化的$\ell_p $-范数或结合归一化的$\ell_p$-范数和SNR来考虑两个麦克风的SNR和ELR的差异。使用CHiME-8远程ASR系统的实验评估表明,所提出的基于$\ell_p$范数的方法优于基线方法,降低了宏观平均单词错误率。
摘要:Guided Source Separation (GSS) is a popular front-end for distant automatic speech recognition (ASR) systems using spatially distributed microphones. When considering spatially distributed microphones, the choice of reference microphone may have a large influence on the quality of the output signal and the downstream ASR performance. In GSS-based speech enhancement, reference microphone selection is typically performed using the signal-to-noise ratio (SNR), which is optimal for noise reduction but may neglect differences in early-to-late-reverberant ratio (ELR) across microphones. In this paper, we propose two reference microphone selection methods for GSS-based speech enhancement that are based on the normalized $\ell_p$-norm, either using only the normalized $\ell_p$-norm or combining the normalized $\ell_p$-norm and the SNR to account for both differences in SNR and ELR across microphones. Experimental evaluation using a CHiME-8 distant ASR system shows that the proposed $\ell_p$-norm-based methods outperform the baseline method, reducing the macro-average word error rate.
【8】Beamforming in the Reproducing Kernel Domain Based on Spatial Differentiation
标题:基于空间分化的再生核域的束形成
链接:https://arxiv.org/abs/2510.27143
摘要:本文提出了一种新的波束形成框架中的再生核域,来自一个统一的解释方向响应的声场的空间微分。通过使用多项式微分算子表示方向响应,所提出的方法能够制定任意波束方向图,包括非轴对称波束方向图。与内场相关的再生核的推导在数学上由霍布森定理支持,该定理允许简洁的解析表达式。此外,所提出的框架概括了传统的球谐域波束形成器,通过重新解释它们作为空间微分算子,从而澄清其理论结构和可扩展性。在二维空间中进行的三个数值模拟证实了该方法的有效性。
摘要:This paper proposes a novel beamforming framework in the reproducing kernel domain, derived from a unified interpretation of directional response as spatial differentiation of the sound field. By representing directional response using polynomial differential operators, the proposed method enables the formulation of arbitrary beam patterns including non-axisymmetric. The derivation of the reproducing kernel associated with the interior fields is mathematically supported by Hobson's theorem, which allows concise analytical expressions. Furthermore, the proposed framework generalizes conventional spherical harmonic domain beamformers by reinterpreting them as spatial differential operators, thereby clarifying their theoretical structure and extensibility. Three numerical simulations conducted in two-dimensional space confirm the validity of the method.
【9】Multi-Representation Attention Framework for Underwater Bioacoustic Denoising and Recognition
标题:基于多表征注意框架的水下生物声信号去噪与识别
链接:https://arxiv.org/abs/2510.26838
摘要:在圣劳伦斯河口的海洋哺乳动物的自动化监测面临着极端的挑战:调用跨越低频呻吟声超声波点击,往往重叠,并嵌入在可变的人为和环境噪声。我们引入了一个多步骤,注意力引导的框架,首先分割频谱图以生成生物相关能量的软掩模,然后将这些掩模与多波段去噪分类的原始输入融合。图像和掩模嵌入通过中级融合进行集成,使模型能够专注于显着的频谱图区域,同时保留全局上下文。使用来自加拿大Saguenay St. Lawrence海洋公园研究站的真实记录,我们证明了分割驱动的注意力和中级融合可以提高信号辨别力,减少假阳性检测,并为不同环境条件和信噪比的海洋哺乳动物监测提供可靠的表示。除了在分布评估,我们进一步评估的泛化面具引导分类(MGC)下分布的变化,通过测试与替代声学变换生成的频谱图。虽然高容量基线模型在这种分布外(OOD)设置中失去了准确性,但MGC保持了稳定的性能,即使是简单的融合机制(门控,concat)也可以在不同的分布中获得可比的结果。这种鲁棒性突出了MGC学习可转移表示的能力,而不是过度拟合特定的转换,从而加强了其对大规模,真实世界的生物多样性监测的适用性。我们表明,在所有的实验设置中,MGC框架始终优于基线架构,在分布和OOD数据的准确性方面都有很大的提高。
摘要:Automated monitoring of marine mammals in the St. Lawrence Estuary faces extreme challenges: calls span low-frequency moans to ultrasonic clicks, often overlap, and are embedded in variable anthropogenic and environmental noise. We introduce a multi-step, attention-guided framework that first segments spectrograms to generate soft masks of biologically relevant energy and then fuses these masks with the raw inputs for multi-band, denoised classification. Image and mask embeddings are integrated via mid-level fusion, enabling the model to focus on salient spectrogram regions while preserving global context. Using real-world recordings from the Saguenay St. Lawrence Marine Park Research Station in Canada, we demonstrate that segmentation-driven attention and mid-level fusion improve signal discrimination, reduce false positive detections, and produce reliable representations for operational marine mammal monitoring across diverse environmental conditions and signal-to-noise ratios. Beyond in-distribution evaluation, we further assess the generalization of Mask-Guided Classification (MGC) under distributional shifts by testing on spectrograms generated with alternative acoustic transformations. While high-capacity baseline models lose accuracy in this Out-of-distribution (OOD) setting, MGC maintains stable performance, with even simple fusion mechanisms (gated, concat) achieving comparable results across distributions. This robustness highlights the capacity of MGC to learn transferable representations rather than overfitting to a specific transformation, thereby reinforcing its suitability for large-scale, real-world biodiversity monitoring. We show that in all experimental settings, the MGC framework consistently outperforms baseline architectures, yielding substantial gains in accuracy on both in-distribution and OOD data.
【10】See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region Refinement
标题:见演讲者:在事先指导和区域细化的情况下从演讲中制作高分辨率的说话面孔
链接:https://arxiv.org/abs/2510.26819
备注:16 pages,15 figures, accepted by TASLP
摘要:与现有的依赖于源图像作为外观参考并使用源语音来生成运动的方法不同,这项工作提出了一种新的方法,该方法直接从语音中提取信息,解决了语音到说话人脸的关键挑战。具体来说,我们首先采用语音到人脸的肖像生成阶段,利用语音条件扩散模型结合统计面部先验和样本自适应加权模块,以实现高质量的肖像生成。在随后的语音驱动的说话人脸生成阶段,我们嵌入表达的动态,如嘴唇运动,面部表情,眼球运动到扩散模型的潜在空间,并进一步优化唇同步使用区域增强模块。为了生成高分辨率的输出,我们将预先训练的基于transformer的离散码本与图像渲染网络集成在一起,以端到端的方式增强视频帧细节。实验结果表明,我们的方法优于现有的方法在HDTF,VoxCeleb和AVSpeech数据集。值得注意的是,这是第一种能够仅从单个语音输入生成高分辨率、高质量的说话面部视频的方法。
摘要:Unlike existing methods that rely on source images as appearance references and use source speech to generate motion, this work proposes a novel approach that directly extracts information from the speech, addressing key challenges in speech-to-talking face. Specifically, we first employ a speech-to-face portrait generation stage, utilizing a speech-conditioned diffusion model combined with statistical facial prior and a sample-adaptive weighting module to achieve high-quality portrait generation. In the subsequent speech-driven talking face generation stage, we embed expressive dynamics such as lip movement, facial expressions, and eye movements into the latent space of the diffusion model and further optimize lip synchronization using a region-enhancement module. To generate high-resolution outputs, we integrate a pre-trained Transformer-based discrete codebook with an image rendering network, enhancing video frame details in an end-to-end manner. Experimental results demonstrate that our method outperforms existing approaches on the HDTF, VoxCeleb, and AVSpeech datasets. Notably, this is the first method capable of generating high-resolution, high-quality talking face videos exclusively from a single speech input.
【1】Reference Microphone Selection for Guided Source Separation based on the Normalized L-p Norm
标题:基于标准化L-p规范的引导源分离参考麦克风选择
链接:https://arxiv.org/abs/2510.27198
摘要:引导源分离(GSS)是使用空间分布麦克风的远程自动语音识别(ASR)系统的流行前端。当考虑空间分布的麦克风时,参考麦克风的选择可能对输出信号的质量和下游ASR性能具有很大影响。在基于GSS的语音增强中,通常使用信噪比(SNR)来执行参考麦克风选择,SNR对于降噪是最佳的,但可能忽略麦克风之间的早到晚混响比(ELR)的差异。在本文中,我们提出了两个参考麦克风选择基于高斯的语音增强方法是基于归一化的$\ell_p $-范数,或者只使用归一化的$\ell_p $-范数或结合归一化的$\ell_p$-范数和SNR来考虑两个麦克风的SNR和ELR的差异。使用CHiME-8远程ASR系统的实验评估表明,所提出的基于$\ell_p$范数的方法优于基线方法,降低了宏观平均单词错误率。
摘要:Guided Source Separation (GSS) is a popular front-end for distant automatic speech recognition (ASR) systems using spatially distributed microphones. When considering spatially distributed microphones, the choice of reference microphone may have a large influence on the quality of the output signal and the downstream ASR performance. In GSS-based speech enhancement, reference microphone selection is typically performed using the signal-to-noise ratio (SNR), which is optimal for noise reduction but may neglect differences in early-to-late-reverberant ratio (ELR) across microphones. In this paper, we propose two reference microphone selection methods for GSS-based speech enhancement that are based on the normalized $\ell_p$-norm, either using only the normalized $\ell_p$-norm or combining the normalized $\ell_p$-norm and the SNR to account for both differences in SNR and ELR across microphones. Experimental evaluation using a CHiME-8 distant ASR system shows that the proposed $\ell_p$-norm-based methods outperform the baseline method, reducing the macro-average word error rate.
【2】Beamforming in the Reproducing Kernel Domain Based on Spatial Differentiation
标题:基于空间分化的再生核域的束形成
链接:https://arxiv.org/abs/2510.27143
摘要:本文提出了一种新的波束形成框架中的再生核域,来自一个统一的解释方向响应的声场的空间微分。通过使用多项式微分算子表示方向响应,所提出的方法能够制定任意波束方向图,包括非轴对称波束方向图。与内场相关的再生核的推导在数学上由霍布森定理支持,该定理允许简洁的解析表达式。此外,所提出的框架概括了传统的球谐域波束形成器,通过重新解释它们作为空间微分算子,从而澄清其理论结构和可扩展性。在二维空间中进行的三个数值模拟证实了该方法的有效性。
摘要:This paper proposes a novel beamforming framework in the reproducing kernel domain, derived from a unified interpretation of directional response as spatial differentiation of the sound field. By representing directional response using polynomial differential operators, the proposed method enables the formulation of arbitrary beam patterns including non-axisymmetric. The derivation of the reproducing kernel associated with the interior fields is mathematically supported by Hobson's theorem, which allows concise analytical expressions. Furthermore, the proposed framework generalizes conventional spherical harmonic domain beamformers by reinterpreting them as spatial differential operators, thereby clarifying their theoretical structure and extensibility. Three numerical simulations conducted in two-dimensional space confirm the validity of the method.
【3】Multi-Representation Attention Framework for Underwater Bioacoustic Denoising and Recognition
标题:基于多表征注意框架的水下生物声信号去噪与识别
链接:https://arxiv.org/abs/2510.26838
摘要:在圣劳伦斯河口的海洋哺乳动物的自动化监测面临着极端的挑战:调用跨越低频呻吟声超声波点击,往往重叠,并嵌入在可变的人为和环境噪声。我们引入了一个多步骤,注意力引导的框架,首先分割频谱图以生成生物相关能量的软掩模,然后将这些掩模与多波段去噪分类的原始输入融合。图像和掩模嵌入通过中级融合进行集成,使模型能够专注于显着的频谱图区域,同时保留全局上下文。使用来自加拿大Saguenay St. Lawrence海洋公园研究站的真实记录,我们证明了分割驱动的注意力和中级融合可以提高信号辨别力,减少假阳性检测,并为不同环境条件和信噪比的海洋哺乳动物监测提供可靠的表示。除了在分布评估,我们进一步评估的泛化面具引导分类(MGC)下分布的变化,通过测试与替代声学变换生成的频谱图。虽然高容量基线模型在这种分布外(OOD)设置中失去了准确性,但MGC保持了稳定的性能,即使是简单的融合机制(门控,concat)也可以在不同的分布中获得可比的结果。这种鲁棒性突出了MGC学习可转移表示的能力,而不是过度拟合特定的转换,从而加强了其对大规模,真实世界的生物多样性监测的适用性。我们表明,在所有的实验设置中,MGC框架始终优于基线架构,在分布和OOD数据的准确性方面都有很大的提高。
摘要:Automated monitoring of marine mammals in the St. Lawrence Estuary faces extreme challenges: calls span low-frequency moans to ultrasonic clicks, often overlap, and are embedded in variable anthropogenic and environmental noise. We introduce a multi-step, attention-guided framework that first segments spectrograms to generate soft masks of biologically relevant energy and then fuses these masks with the raw inputs for multi-band, denoised classification. Image and mask embeddings are integrated via mid-level fusion, enabling the model to focus on salient spectrogram regions while preserving global context. Using real-world recordings from the Saguenay St. Lawrence Marine Park Research Station in Canada, we demonstrate that segmentation-driven attention and mid-level fusion improve signal discrimination, reduce false positive detections, and produce reliable representations for operational marine mammal monitoring across diverse environmental conditions and signal-to-noise ratios. Beyond in-distribution evaluation, we further assess the generalization of Mask-Guided Classification (MGC) under distributional shifts by testing on spectrograms generated with alternative acoustic transformations. While high-capacity baseline models lose accuracy in this Out-of-distribution (OOD) setting, MGC maintains stable performance, with even simple fusion mechanisms (gated, concat) achieving comparable results across distributions. This robustness highlights the capacity of MGC to learn transferable representations rather than overfitting to a specific transformation, thereby reinforcing its suitability for large-scale, real-world biodiversity monitoring. We show that in all experimental settings, the MGC framework consistently outperforms baseline architectures, yielding substantial gains in accuracy on both in-distribution and OOD data.
【4】See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region Refinement
标题:见演讲者:在事先指导和区域细化的情况下从演讲中制作高分辨率的说话面孔
链接:https://arxiv.org/abs/2510.26819
备注:16 pages,15 figures, accepted by TASLP
摘要:与现有的依赖于源图像作为外观参考并使用源语音来生成运动的方法不同,这项工作提出了一种新的方法,该方法直接从语音中提取信息,解决了语音到说话人脸的关键挑战。具体来说,我们首先采用语音到人脸的肖像生成阶段,利用语音条件扩散模型结合统计面部先验和样本自适应加权模块,以实现高质量的肖像生成。在随后的语音驱动的说话人脸生成阶段,我们嵌入表达的动态,如嘴唇运动,面部表情,眼球运动到扩散模型的潜在空间,并进一步优化唇同步使用区域增强模块。为了生成高分辨率的输出,我们将预先训练的基于transformer的离散码本与图像渲染网络集成在一起,以端到端的方式增强视频帧细节。实验结果表明,我们的方法优于现有的方法在HDTF,VoxCeleb和AVSpeech数据集。值得注意的是,这是第一种能够仅从单个语音输入生成高分辨率、高质量的说话面部视频的方法。
摘要:Unlike existing methods that rely on source images as appearance references and use source speech to generate motion, this work proposes a novel approach that directly extracts information from the speech, addressing key challenges in speech-to-talking face. Specifically, we first employ a speech-to-face portrait generation stage, utilizing a speech-conditioned diffusion model combined with statistical facial prior and a sample-adaptive weighting module to achieve high-quality portrait generation. In the subsequent speech-driven talking face generation stage, we embed expressive dynamics such as lip movement, facial expressions, and eye movements into the latent space of the diffusion model and further optimize lip synchronization using a region-enhancement module. To generate high-resolution outputs, we integrate a pre-trained Transformer-based discrete codebook with an image rendering network, enhancing video frame details in an end-to-end manner. Experimental results demonstrate that our method outperforms existing approaches on the HDTF, VoxCeleb, and AVSpeech datasets. Notably, this is the first method capable of generating high-resolution, high-quality talking face videos exclusively from a single speech input.
【5】Inferring trust in recommendation systems from brain, behavioural, and physiological data
标题:从大脑、行为和生理数据推断对推荐系统的信任度
链接:https://arxiv.org/abs/2510.27272
摘要:随着人们越来越依赖人工智能(AI)来管理信息和做出决策,在自动化智能系统中分配适当的信任变得越来越重要。然而,目前对自动化的信任度的测量在很大程度上仍然依赖于对用户主观和破坏性的自我报告。在这里,我们以音乐推荐为模型来研究自动化中信任的神经和认知过程。我们观察到,系统的准确性直接关系到用户的信任和调制的音乐偏好的推荐线索的影响。使用强化学习模型对用户的奖励编码过程进行建模,进一步揭示了系统准确性,预期奖励和预测误差与通过EEG记录的振荡神经活动和瞳孔直径的变化有关。我们的研究结果提供了一个基于神经的自动化信任校准的解释,并强调了开发可信赖的AI系统的多模式方法的承诺。
摘要:As people nowadays increasingly rely on artificial intelligence (AI) to curate information and make decisions, assigning the appropriate amount of trust in automated intelligent systems has become ever more important. However, current measurements of trust in automation still largely rely on self-reports that are subjective and disruptive to the user. Here, we take music recommendation as a model to investigate the neural and cognitive processes underlying trust in automation. We observed that system accuracy was directly related to users' trust and modulated the influence of recommendation cues on music preference. Modelling users' reward encoding process with a reinforcement learning model further revealed that system accuracy, expected reward, and prediction error were related to oscillatory neural activity recorded via EEG and changes in pupil diameter. Our results provide a neurally grounded account of calibrating trust in automation and highlight the promises of a multimodal approach towards developing trustable AI systems.
【6】Expressive Range Characterization of Open Text-to-Audio Models
标题:开放文本到音频模型的表达范围特征
链接:https://arxiv.org/abs/2510.27102
备注:Accepted at the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE 2025)
摘要:文本到音频模型是一种生成模型,其响应于给定的文本提示而产生音频输出。虽然级别生成器和它们创建的功能内容的属性(例如,可玩性)主导程序生成内容(PCG)中的大多数话语,但在情感上与玩家产生共鸣的游戏倾向于将一系列创造性和多模态内容编织在一起(例如,音乐、声音、视觉、叙述语气),并且多模态模型已经开始看到至少用于该目的的实验性使用。然而,目前还不清楚这些模型究竟生成什么,以及具有多大程度的可变性和保真度:音频是生成系统的目标输出的一个非常广泛的类别。 在PCG社区,表达范围分析(ERA)已被用作一种定量的方法来表征发电机的输出空间,特别是水平发电机。本文适应ERA的文本到音频模型,通过查看特定的,固定的提示输出的表达范围,使分析易于处理。实验进行了提示模型与几个标准化的提示来自环境声音分类(ESC-50)数据集。沿着关键声学维度(例如,音高、响度和音色)。更广泛地说,本文提供了一个基于ERA的生成音频模型的探索性评估框架。
摘要:Text-to-audio models are a type of generative model that produces audio output in response to a given textual prompt. Although level generators and the properties of the functional content that they create (e.g., playability) dominate most discourse in procedurally generated content (PCG), games that emotionally resonate with players tend to weave together a range of creative and multimodal content (e.g., music, sounds, visuals, narrative tone), and multimodal models have begun seeing at least experimental use for this purpose. However, it remains unclear what exactly such models generate, and with what degree of variability and fidelity: audio is an extremely broad class of output for a generative system to target. Within the PCG community, expressive range analysis (ERA) has been used as a quantitative way to characterize generators' output space, especially for level generators. This paper adapts ERA to text-to-audio models, making the analysis tractable by looking at the expressive range of outputs for specific, fixed prompts. Experiments are conducted by prompting the models with several standardized prompts derived from the Environmental Sound Classification (ESC-50) dataset. The resulting audio is analyzed along key acoustic dimensions (e.g., pitch, loudness, and timbre). More broadly, this paper offers a framework for ERA-based exploratory evaluation of generative audio models.
【7】Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling
标题:复杂场景下分离与去混响联合建模的音视频语音增强
链接:https://arxiv.org/abs/2510.26825
摘要:视听语音增强(AVSE)是利用视觉辅助信息从混合音频中提取目标说话人的语音。在现实生活中,往往存在着复杂的声学环境,伴随着各种干扰声和混响。大多数以前的方法都难以应付这种复杂的条件,导致提取的语音的感知质量差。在本文中,我们提出了一个有效的AVSE系统,在复杂的声学环境中表现良好。具体来说,我们设计了一个“去混响前分离”的管道,可以扩展到其他AVSE网络。第四届COGMHEAR视听语音增强挑战赛(AVSEC)旨在探索多模态复杂环境中语音处理的新方法。我们在AVSEC-4中验证了我们系统的性能:我们在竞争排行榜上的三个客观指标中取得了优异的成绩,并最终在人类主观听力测试中获得了第一名。
摘要:Audio-visual speech enhancement (AVSE) is a task that uses visual auxiliary information to extract a target speaker's speech from mixed audio. In real-world scenarios, there often exist complex acoustic environments, accompanied by various interfering sounds and reverberation. Most previous methods struggle to cope with such complex conditions, resulting in poor perceptual quality of the extracted speech. In this paper, we propose an effective AVSE system that performs well in complex acoustic environments. Specifically, we design a "separation before dereverberation" pipeline that can be extended to other AVSE networks. The 4th COGMHEAR Audio-Visual Speech Enhancement Challenge (AVSEC) aims to explore new approaches to speech processing in multimodal complex environments. We validated the performance of our system in AVSEC-4: we achieved excellent results in the three objective metrics on the competition leaderboard, and ultimately secured first place in the human subjective listening test.
【8】Cross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features
标题:使用领域知识声学特征进行乌尔都语语音情感识别的跨Corpus验证
链接:https://arxiv.org/abs/2510.26823
备注:Conference paper, 4 pages, including 3 figures and 3 tables
摘要:语音情感识别(SER)是实现情感智能人工智能的关键情感计算技术。虽然SER总体上具有挑战性,但对于乌尔都语等低资源语言来说尤其困难。本研究探讨乌尔都语SER在跨语料库设置,一个领域,在很大程度上仍然是未开发的。我们采用跨语料库的评估框架,在三个不同的乌尔都语情感语音数据集测试模型的泛化。两个标准的领域知识为基础的声学特征集,eGeMAPS和ComparE,用于表示语音信号的特征向量,然后通过逻辑回归和多层感知器分类。分类性能评估使用未加权平均召回(UAR),同时考虑类标签的不平衡。结果表明,自语料库验证往往高估了性能,UAR超过跨语料库评估高达13%,强调跨语料库评估提供了一个更现实的衡量模型的鲁棒性。总体而言,这项工作强调了跨语料库验证的重要性,乌尔都语SER及其影响有助于推进情感计算研究的代表性不足的语言社区。
摘要:Speech Emotion Recognition (SER) is a key affective computing technology that enables emotionally intelligent artificial intelligence. While SER is challenging in general, it is particularly difficult for low-resource languages such as Urdu. This study investigates Urdu SER in a cross-corpus setting, an area that has remained largely unexplored. We employ a cross-corpus evaluation framework across three different Urdu emotional speech datasets to test model generalization. Two standard domain-knowledge based acoustic feature sets, eGeMAPS and ComParE, are used to represent speech signals as feature vectors which are then passed to Logistic Regression and Multilayer Perceptron classifiers. Classification performance is assessed using unweighted average recall (UAR) whilst considering class-label imbalance. Results show that Self-corpus validation often overestimates performance, with UAR exceeding cross-corpus evaluation by up to 13%, underscoring that cross-corpus evaluation offers a more realistic measure of model robustness. Overall, this work emphasizes the importance of cross-corpus validation for Urdu SER and its implications contribute to advancing affective computing research for underrepresented language communities.
【9】GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
标题:GACA-DiT:基于扩散的舞蹈到音乐生成,具有流派自适应节奏和上下文感知对齐
链接:https://arxiv.org/abs/2510.26818
备注:5 pages, 3 figures, submitted to ICASSP 2026
摘要:舞蹈到音乐(D2 M)生成的目的是自动创作音乐,节奏和时间与舞蹈动作一致。现有的方法通常依赖于粗糙的节奏嵌入,例如全局运动特征或二进制化的基于关节的节奏值,其丢弃细粒度的运动线索并导致弱节奏对齐。此外,由特征下采样引入的时间失配进一步阻碍了舞蹈和音乐之间的精确同步。为了解决这些问题,我们提出了\textbf{GACA-DiT},一个基于扩散变换的框架,具有两个新颖的模块,用于节奏一致和时间对齐的音乐生成。首先,一个\textbf{genre-adaptive rhythm extraction}模块结合了多尺度时间小波分析和空间相位直方图与自适应联合加权,以捕获细粒度的,特定于流派的节奏模式。第二,一个\textbf{上下文感知的时间对齐}模块解决时间不匹配使用可学习的上下文查询对齐音乐潜伏期与相关的舞蹈节奏功能。在AIST++和TikTok数据集上进行的大量实验表明,GACA-DiT在客观指标和人工评估方面都优于最先进的方法。项目页面:https://beria-moon.github.io/GACA-DiT/。
摘要:Dance-to-music (D2M) generation aims to automatically compose music that is rhythmically and temporally aligned with dance movements. Existing methods typically rely on coarse rhythm embeddings, such as global motion features or binarized joint-based rhythm values, which discard fine-grained motion cues and result in weak rhythmic alignment. Moreover, temporal mismatches introduced by feature downsampling further hinder precise synchronization between dance and music. To address these problems, we propose \textbf{GACA-DiT}, a diffusion transformer-based framework with two novel modules for rhythmically consistent and temporally aligned music generation. First, a \textbf{genre-adaptive rhythm extraction} module combines multi-scale temporal wavelet analysis and spatial phase histograms with adaptive joint weighting to capture fine-grained, genre-specific rhythm patterns. Second, a \textbf{context-aware temporal alignment} module resolves temporal mismatches using learnable context queries to align music latents with relevant dance rhythm features. Extensive experiments on the AIST++ and TikTok datasets demonstrate that GACA-DiT outperforms state-of-the-art methods in both objective metrics and human evaluation. Project page: https://beria-moon.github.io/GACA-DiT/.
【10】Oral Tradition-Encoded NanyinHGNN: Integrating Nanyin Music Preservation and Generation through a Pipa-Centric Dataset
标题:口述编码南音HGNN:通过以Pipa为中心的数据集整合南音音乐保存和生成
链接:https://arxiv.org/abs/2510.26817
备注:10 pages, 2 figures
摘要:我们提出了南音HGNN,一个异构的图网络模型,用于生成南音器乐。作为联合国教科文组织认可的非物质文化遗产,南音遵循以琵琶为中心的异声传统,其中核心旋律用传统记谱法记谱,而表达则通过口头传递,这对保护和当代创新都提出了挑战。为了解决这个问题,我们构建了一个以Pipa为中心的数据集,开发了NanyinTok作为一种专门的标记化方法,并使用Graph Converter将符号序列转换为图形结构,以确保保留关键的音乐特征。我们的关键创新重新表述为异构图内的表示节点的创建表示生成。首先,一个图形神经网络生成优化的旋律轮廓。然后,由南音演奏实践告知的规则引导系统将这些轮廓细化为完整的表达,而不需要在训练期间进行显式的表达注释。实验结果表明,我们的模型成功地产生真实的四个传统乐器的杂音合奏。这些发现验证了将特定领域的知识整合到模型架构中可以有效地缓解计算民族音乐学中的数据稀缺挑战。
摘要:We propose NanyinHGNN, a heterogeneous graph network model for generating Nanyin instrumental music. As a UNESCO-recognized intangible cultural heritage, Nanyin follows a heterophonic tradition centered around the pipa, where core melodies are notated in traditional notation while ornamentations are passed down orally, presenting challenges for both preservation and contemporary innovation. To address this, we construct a Pipa-Centric MIDI dataset, develop NanyinTok as a specialized tokenization method, and convert symbolic sequences into graph structures using a Graph Converter to ensure that key musical features are preserved. Our key innovation reformulates ornamentation generation as the creation of ornamentation nodes within a heterogeneous graph. First, a graph neural network generates melodic outlines optimized for ornamentations. Then, a rule-guided system informed by Nanyin performance practices refines these outlines into complete ornamentations without requiring explicit ornamentation annotations during training. Experimental results demonstrate that our model successfully generates authentic heterophonic ensembles featuring four traditional instruments. These findings validate that integrating domain-specific knowledge into model architecture can effectively mitigate data scarcity challenges in computational ethnomusicology.
机器翻译由腾讯交互翻译提供,仅供参考
