今日论文合集:cs.SD语音14篇,eess.AS音频处理6篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Frequency-Invariant Beamforming in Elevation and Azimuth via Autograd and Concentric Circular Microphone Arrays
标题:通过Autograd和同心圆形麦克风阵列在高度和方位角进行频率不变的射束形成
链接:https://arxiv.org/pdf/2511.19403v1

作者:Jorge Ortigoso-Narro,Jose A. Belloch,Maximo Morales-Cespedes,Maximo Cobos
摘要:平面和同心圆麦克风阵列在波束形成中的使用由于其优化方位角和仰角的能力而受到关注,使其成为声源定位和噪声抑制等空间音频任务的理想选择。与将转向限制在单轴的线性阵列不同,2D阵列提供双轴优化,尽管仰角控制仍然具有挑战性。本研究探讨整合autograd,自动微分工具,同心圆阵列施加波束宽度和频率不变的约束。这使得能够在两个角度上持续优化,同时在宽频率范围内保持性能。我们评估我们的方法,通过模拟的波束宽度,白噪声增益,并在多个频率的方向性。对标准和先进的波束形成器,包括延迟和总和,修改后的延迟和总和,雅可比-安格展开为基础的方法,和高斯窗口为基础的梯度下降方法的比较分析。我们的方法实现了优越的空间选择性和更窄的主瓣,特别是在较低频率的仰角轴。这些结果强调了我们的方法在提高声学传感和空间音频应用需要精确的双轴控制的波束成形性能的有效性。摘要:The use of planar and concentric circular microphone arrays in beamforming has gained attention due to their ability to optimize both azimuth and elevation angles, making them ideal for spatial audio tasks like sound source localization and noise suppression. Unlike linear arrays, which restrict steering to a single axis, 2D arrays offer dual-axis optimization, although elevation control remains challenging. This study explores the integration of autograd, an automatic differentiation tool, with concentric circular arrays to impose beamwidth and frequency invariance constraints. This enables continuous optimization over both angles while maintaining performance across a wide frequency range. We evaluate our method through simulations of beamwidth, white noise gain, and directivity across multiple frequencies. A comparative analysis is presented against standard and advanced beamformers, including delay-and-sum, modified delay-and-sum, a Jacobi-Anger expansion-based method, and a Gaussian window-based gradient descent approach. Our method achieves superior spatial selectivity and narrower mainlobes, particularly in the elevation axis at lower frequencies. These results underscore the effectiveness of our approach in enhancing beamforming performance for acoustic sensing and spatial audio applications requiring precise dual-axis control.


【2】Real-Time Object Tracking with On-Device Deep Learning for Adaptive Beamforming in Dynamic Acoustic Environments
标题:动态声学环境中基于设备深度学习的自适应波束形成实时目标跟踪
链接:https://arxiv.org/pdf/2511.19396v1

作者:Jorge Ortigoso-Narro,Jose A. Belloch,Adrian Amor-Martin,Sandra Roger,Maximo Cobos
摘要:目标跟踪和声学波束形成的进步正在推动监控、人机交互和机器人技术的新功能。这项工作提出了一种嵌入式系统,该系统集成了基于深度学习的跟踪与波束成形,以实现动态环境中精确的声源定位和定向音频捕获。该方法结合了单摄像机深度估计和立体视觉,以实现移动物体的精确3D定位。利用MEMS麦克风构造的平面同心圆形麦克风阵列提供了一种紧凑的、节能的平台,该平台支持跨方位角和仰角的2D波束转向。实时跟踪输出不断调整阵列的焦点,使声学响应与目标的位置同步。通过将学习的空间感知与动态转向相结合,系统在存在多个或移动源的情况下保持稳健的性能。实验评估表明,信号干扰比的显着增益,使设计非常适合电话会议,智能家居设备和辅助技术。摘要:Advances in object tracking and acoustic beamforming are driving new capabilities in surveillance, human-computer interaction, and robotics. This work presents an embedded system that integrates deep learning-based tracking with beamforming to achieve precise sound source localization and directional audio capture in dynamic environments. The approach combines single-camera depth estimation and stereo vision to enable accurate 3D localization of moving objects. A planar concentric circular microphone array constructed with MEMS microphones provides a compact, energy-efficient platform supporting 2D beam steering across azimuth and elevation. Real-time tracking outputs continuously adapt the array's focus, synchronizing the acoustic response with the target's position. By uniting learned spatial awareness with dynamic steering, the system maintains robust performance in the presence of multiple or moving sources. Experimental evaluation demonstrates significant gains in signal-to-interference ratio, making the design well-suited for teleconferencing, smart home devices, and assistive technologies.


【3】Explicit Tonal Tension Conditioning via Dual-Level Beam Search for Symbolic Music Generation
标题:通过双级梁搜索显式调性张力条件反射以产生象征性音乐
链接:https://arxiv.org/pdf/2511.19342v1

作者:Maral Ebrahimzadeh,Gilberto Bernardes,Sebastian Stober

备注:12 pages, 2 Figures, Accepted at the 17th International Symposium on Computer Music Multidisciplinary Research (CMMR) 2025

摘要:国家的最先进的符号音乐生成模型最近取得了显着的输出质量,但明确的控制组成的特点,如音调的张力,仍然具有挑战性。我们提出了一种新的方法,集成了一个计算音调张力模型,音调间隔矢量分析的基础上,到一个Transformer框架。我们的方法在推理过程中采用了两级波束搜索策略。在令牌级别,使用模型概率和多样性度量对生成的候选者进行重新排名,以保持整体质量。在小节级别,应用基于张力的重新排序,以确保生成的音乐与期望的张力曲线对齐。客观的评估表明,我们的方法有效地调节音调的紧张,和主观听力测试证实,该系统产生的输出,符合目标紧张。这些结果表明,通过双层梁搜索的显式张力调节提供了一个强大而直观的工具来指导人工智能生成的音乐。此外,我们的实验表明,我们的方法可以产生多个不同的音乐解释在相同的张力条件下。摘要:State-of-the-art symbolic music generation models have recently achieved remarkable output quality, yet explicit control over compositional features, such as tonal tension, remains challenging. We propose a novel approach that integrates a computational tonal tension model, based on tonal interval vector analysis, into a Transformer framework. Our method employs a two-level beam search strategy during inference. At the token level, generated candidates are re-ranked using model probability and diversity metrics to maintain overall quality. At the bar level, a tension-based re-ranking is applied to ensure that the generated music aligns with a desired tension curve. Objective evaluations indicate that our approach effectively modulates tonal tension, and subjective listening tests confirm that the system produces outputs that align with the target tension. These results demonstrate that explicit tension conditioning through a dual-level beam search provides a powerful and intuitive tool to guide AI-generated music. Furthermore, our experiments demonstrate that our method can generate multiple distinct musical interpretations under the same tension condition.


【4】Dynamic Multi-Species Bird Soundscape Generation with Acoustic Patterning and 3D Spatialization
标题:利用声学图案化和3D空间化生成动态多物种鸟类声景
链接:https://arxiv.org/pdf/2511.19275v1

作者:Ellie L. Zhang,Duoduo Liao,Callie C. Liao

备注:Accepted by IEEE Big Data 2025

摘要:生成动态的、可扩展的多物种鸟类音景仍然是计算机音乐和算法声音设计中的一个重大挑战。鸟鸣声涉及快速调频啁啾,复杂的振幅包络,独特的声学模式,重叠的呼叫和动态的鸟间互动,所有这些都需要在3D环境中精确的时间和空间控制。现有的方法,无论是基于数字信号处理(DSP)还是数据驱动的,通常仅关注单物种建模、静态调用结构或直接从记录合成,并且通常遭受噪声、有限的灵活性或大数据需求。为了解决这些挑战,我们提出了一种新颖的,完全算法驱动的框架,使用基于DSP的啁啾生成和3D空间化生成动态多物种鸟类音景,而不依赖于录音或训练数据。我们的方法模拟多个独立移动的鸟类,每个物种沿着不同的移动3D轨迹,支持可控的啁啾序列,重叠的合唱,和现实的3D运动在可扩展的音景,同时保留物种特定的声学模式。一个可视化界面提供了鸟类的轨迹,光谱图,活动时间表和声波的分析和创造性的目的。视觉和音频评估都证明了该系统能够生成密集的,身临其境的和生态启发的音景,突出了其在计算机音乐,交互式虚拟环境和计算生物声学研究中的潜力。摘要:Generation of dynamic, scalable multi-species bird soundscapes remains a significant challenge in computer music and algorithmic sound design. Birdsongs involve rapid frequency-modulated chirps, complex amplitude envelopes, distinctive acoustic patterns, overlapping calls, and dynamic inter-bird interactions, all of which require precise temporal and spatial control in 3D environments. Existing approaches, whether Digital Signal Processing (DSP)-based or data-driven, typically focus only on single species modeling, static call structures, or synthesis directly from recordings, and often suffer from noise, limited flexibility, or large data needs. To address these challenges, we present a novel, fully algorithm-driven framework that generates dynamic multi-species bird soundscapes using DSP-based chirp generation and 3D spatialization, without relying on recordings or training data. Our approach simulates multiple independently-moving birds per species along different moving 3D trajectories, supporting controllable chirp sequences, overlapping choruses, and realistic 3D motion in scalable soundscapes while preserving species-specific acoustic patterns. A visualization interface provides bird trajectories, spectrograms, activity timelines, and sound waves for analytical and creative purposes. Both visual and audio evaluations demonstrate the ability of the system to generate dense, immersive, and ecologically inspired soundscapes, highlighting its potential for computer music, interactive virtual environments, and computational bioacoustics research.


【5】Multidimensional Music Aesthetic Evaluation via Semantically Consistent C-Mixup Augmentation
标题:基于语义一致性C-Mixup增强的多维音乐审美评价
链接:https://arxiv.org/pdf/2511.18869v1

作者:Shuyang Liu,Yuan Jin,Rui Lin,Shizhe Chen,Junyu Dai,Tao Jiang
摘要:由于音乐感知的多维性,评估生成歌曲的美学质量具有挑战性。我们提出了一个强大的音乐美学评估框架,该框架结合了(1)多源多尺度特征提取以获得互补的片段和曲目级表示,(2)分层音频增强策略以丰富训练数据,以及(3)集成回归和排名损失的混合训练目标,以获得准确的评分和可靠的热门歌曲识别。在ICASSP 2026 SongEval基准测试上的实验表明,我们的方法在相关性和顶级指标上始终优于基线方法。摘要:Evaluating the aesthetic quality of generated songs is challenging due to the multi-dimensional nature of musical perception. We propose a robust music aesthetic evaluation framework that combines (1) multi-source multi-scale feature extraction to obtain complementary segment- and track-level representations, (2) a hierarchical audio augmentation strategy to enrich training data, and (3) a hybrid training objective that integrates regression and ranking losses for accurate scoring and reliable top-song identification. Experiments on the ICASSP 2026 SongEval benchmark demonstrate that our approach consistently outperforms baseline methods across correlation and top-tier metrics.


【6】PrismAudio: Decomposed Chain-of-Thoughts and Multi-dimensional Rewards for Video-to-Audio Generation
标题:棱镜音频:分解思想链和视频到音频生成的多维奖励
链接:https://arxiv.org/pdf/2511.18833v1

作者:Huadai Liu,Kaicheng Luo,Wen Wang,Qian Chen,Peiwen Sun,Rongjie Huang,Xiangang Li,Jieping Ye,Wei Xue
备注:Preprint
摘要:视频到音频(V2 A)的生成需要平衡四个关键的感知维度:语义一致性,视听时间同步性,美学质量和空间准确性;然而现有的方法存在客观纠缠,将竞争目标合并在单一损失函数中,并且缺乏人类偏好对齐。我们介绍了PrismAudio,这是第一个将强化学习集成到V2 A生成中的框架,具有专门的思想链(CoT)规划。我们的方法将整体推理分解为四个专门的CoT模块(语义,时间,美学和空间CoT),每个模块都与有针对性的奖励函数配对。这种CoT-奖励对应关系实现了多维RL优化,引导模型在所有角度上联合生成更好的推理,解决了客观纠缠问题,同时保持了可解释性。为了使这种优化在计算上切实可行,我们提出了快速GRPO,它采用混合ODE-GRPO采样,大大减少了训练开销相比,现有的GRPO实现。我们还引入了AudioCanvas,这是一个严格的基准测试,它在分布上更加平衡,比现有的数据集覆盖了更现实的多样化和更具挑战性的场景,具有300个单事件类和501个多事件样本。实验结果表明,PrismAudio在域内VGGSound测试集和域外AudioCanvas基准测试中的所有四个感知维度上都达到了最先进的性能。项目页面可在https: PrismAudio-Project.github.io上找到。摘要:Video-to-Audio (V2A) generation requires balancing four critical perceptual dimensions: semantic consistency, audio-visual temporal synchrony, aesthetic quality, and spatial accuracy; yet existing methods suffer from objective entanglement that conflates competing goals in single loss functions and lack human preference alignment. We introduce PrismAudio, the first framework to integrate Reinforcement Learning into V2A generation with specialized Chain-of-Thought (CoT) planning. Our approach decomposes monolithic reasoning into four specialized CoT modules (Semantic, Temporal, Aesthetic, and Spatial CoT), each paired with targeted reward functions. This CoT-reward correspondence enables multidimensional RL optimization that guides the model to jointly generate better reasoning across all perspectives, solving the objective entanglement problem while preserving interpretability. To make this optimization computationally practical, we propose Fast-GRPO, which employs hybrid ODE-SDE sampling that dramatically reduces the training overhead compared to existing GRPO implementations. We also introduce AudioCanvas, a rigorous benchmark that is more distributionally balanced and covers more realistically diverse and challenging scenarios than existing datasets, with 300 single-event classes and 501 multi-event samples. Experimental results demonstrate that PrismAudio achieves state-of-the-art performance across all four perceptual dimensions on both the in-domain VGGSound test set and out-of-domain AudioCanvas benchmark. The project page is available at https: PrismAudio-Project.github.io.


【7】Multimodal Real-Time Anomaly Detection and Industrial Applications
标题:多模式实时异常检测和工业应用
链接:https://arxiv.org/pdf/2511.18698v1

作者:Aman Verma,Keshav Samdani,Mohd. Samiuddin Shafi
备注:Accepted to ASRU 2025
摘要:本文介绍了综合多模式房间监控系统的设计、实现和发展,该系统集成了同步视频和音频处理,用于实时活动识别和异常检测。我们描述了系统的两个迭代:使用YOLOv8,ByteTrack和Audio Spectrogram Transformer(AST)的初始轻量级实现,以及包含多模型音频合奏,混合对象检测,双向跨模态注意力和多方法异常检测的高级版本。这一演变在准确性、鲁棒性和工业适用性方面都有了显著的改进。先进的系统结合了三种音频模型(AST,Wav2Vec2和HuBERT),用于全面的音频理解,双对象检测器(YOLO和DETR)用于提高准确性,以及用于增强跨模态学习的复杂融合机制。实验评估表明,该系统在一般监控场景以及专业工业安全应用中的有效性,在标准硬件上实现实时性能,同时保持高精度。摘要:This paper presents the design, implementation, and evolution of a comprehensive multimodal room-monitoring system that integrates synchronized video and audio processing for real-time activity recognition and anomaly detection. We describe two iterations of the system: an initial lightweight implementation using YOLOv8, ByteTrack, and the Audio Spectrogram Transformer (AST), and an advanced version that incorporates multi-model audio ensembles, hybrid object detection, bidirectional cross-modal attention, and multi-method anomaly detection. The evolution demonstrates significant improvements in accuracy, robustness, and industrial applicability. The advanced system combines three audio models (AST, Wav2Vec2, and HuBERT) for comprehensive audio understanding, dual object detectors (YOLO and DETR) for improved accuracy, and sophisticated fusion mechanisms for enhanced cross-modal learning. Experimental evaluation shows the system's effectiveness in general monitoring scenarios as well as specialized industrial safety applications, achieving real-time performance on standard hardware while maintaining high accuracy.


【8】DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation
标题:DHAuDS:一个动态异构的测试时自适应音频基准
链接:https://arxiv.org/pdf/2511.18421v1

作者:Weichuang Shao,Iman Yi Liao,Tomas Henrique Bode Maul,Tissa Chandesa
摘要:当在一个数据集上训练的模型在声学上不同的条件下记录的数据上失去准确性时,音频分类器经常面临域偏移。以前的语音和声音分析中的测试时自适应(TTA)研究通常在固定或不匹配的噪声设置下评估模型,无法模拟真实世界的变化。为了克服这些局限性,本文提出了DHAuDS(动态和异构音频域偏移),一个基准,旨在评估TTA方法下更现实和多样化的声学变化。DHAuDS包括四个标准化基准:UrbanSound8K-C,SpeechCommandsV2-C,VocalSound-C和ReefSet-C,每个基准都使用动态损坏严重级别和异构噪声类型构建,以模拟真实的音频降级场景。该框架为每个基准测试定义了14个评估标准(UrbanSound8K-C为8个),产生了50个不重复的标准(124个实验),这些标准共同实现了TTA算法的公平,可重复和跨域比较。通过包含动态和混合域噪声设置,DHAuDS提供了一个一致且可公开复制的测试平台,以支持正在进行的鲁棒和自适应音频建模研究。摘要:Audio classifiers frequently face domain shift, when models trained on one dataset lose accuracy on data recorded in acoustically different conditions. Previous Test-Time Adaptation (TTA) research in speech and sound analysis often evaluates models under fixed or mismatched noise settings, that fail to mimic real-world variability. To overcome these limitations, this paper presents DHAuDS (Dynamic and Heterogeneous Audio Domain Shift), a benchmark designed to assess TTA approaches under more realistic and diverse acoustic shifts. DHAuDS comprises four standardized benchmarks: UrbanSound8K-C, SpeechCommandsV2-C, VocalSound-C, and ReefSet-C, each constructed with dynamic corruption severity levels and heterogeneous noise types to simulate authentic audio degradation scenarios. The framework defines 14 evaluation criteria for each benchmark (8 for UrbanSound8K-C), resulting in 50 unrepeated criteria (124 experiments) that collectively enable fair, reproducible, and cross-domain comparison of TTA algorithms. Through the inclusion of dynamic and mixed-domain noise settings, DHAuDS offers a consistent and publicly reproducible testbed to support ongoing studies in robust and adaptive audio modeling.


【9】NSTR: Neural Spectral Transport Representation for Space-Varying Frequency Fields
标题:NTR:空变频场的神经谱传输表示
链接:https://arxiv.org/pdf/2511.18384v1

作者:Plein Versace
摘要:隐式神经表征(INRs)已经成为一种强大的范式,用于表示图像,音频和3D场景等信号。然而,现有的INR框架--包括具有傅立叶特征的MLP、SIREN和多分辨率哈希网格--隐含地假设了一个 textit{全局和固定}谱基。这种假设从根本上与真实世界的信号不一致,真实世界的信号的频率特性在空间上变化很大,表现出局部高频纹理,平滑区域和频率漂移现象。我们提出了 textbf{神经频谱传输表示(NSTR)},第一个INR框架, textbf{显式地模拟空间变化的局部频率场}。NSTR引入了一个可学习的频率输运方程,这是一个PDE,它控制着局部光谱成分如何在空间中演变。给定一个可学习的局部谱场$S(x)$和一个频率传输网络$F_θ$,NSTR通过空间调制一组紧凑的全局正弦基来重构信号。该公式具有很强的局部适应性,并通过可视化频率流提供了一个新的水平的可解释性。在2D图像回归、音频重建和隐式3D几何上的实验表明,NSTR比SIREN、傅立叶特征MLP和Instant-NGP实现了更好的精度-参数权衡。NSTR需要更少的全局频率,收敛更快,并且通过频谱传输场自然地解释信号结构。我们相信,NSTR通过引入空间变化谱的显式建模,为INR研究开辟了一个新的方向。摘要:Implicit Neural Representations (INRs) have emerged as a powerful paradigm for representing signals such as images, audio, and 3D scenes. However, existing INR frameworks -- including MLPs with Fourier features, SIREN, and multiresolution hash grids -- implicitly assume a textit{global and stationary} spectral basis. This assumption is fundamentally misaligned with real-world signals whose frequency characteristics vary significantly across space, exhibiting local high-frequency textures, smooth regions, and frequency drift phenomena. We propose textbf{Neural Spectral Transport Representation (NSTR)}, the first INR framework that textbf{explicitly models a spatially varying local frequency field}. NSTR introduces a learnable emph{frequency transport equation}, a PDE that governs how local spectral compositions evolve across space. Given a learnable local spectrum field $S(x)$ and a frequency transport network $F_θ$ enforcing $ nabla S(x) approx F_θ(x, S(x))$, NSTR reconstructs signals by spatially modulating a compact set of global sinusoidal bases. This formulation enables strong local adaptivity and offers a new level of interpretability via visualizing frequency flows. Experiments on 2D image regression, audio reconstruction, and implicit 3D geometry show that NSTR achieves significantly better accuracy-parameter trade-offs than SIREN, Fourier-feature MLPs, and Instant-NGP. NSTR requires fewer global frequencies, converges faster, and naturally explains signal structure through spectral transport fields. We believe NSTR opens a new direction in INR research by introducing explicit modeling of space-varying spectrum.


【10】Diffusion-based Surrogate Model for Time-varying Underwater Acoustic Channels
标题:时变水下声通道的基于扩散的代理模型
链接:https://arxiv.org/pdf/2511.18078v1

作者:Kexin Li,Mandar Chitre
摘要:时变水声信道的精确建模对于可靠的水下通信系统的设计、评估和部署至关重要。传统的物理模型需要详细的环境知识,而随机重放方法受到测量通道的有限多样性的限制,并且通常无法推广到看不见的场景,从而降低了它们的实用性。为了应对这些挑战,我们提出了StableUASim,这是一种预训练的条件潜在扩散代理模型,可以捕获水声通信信道的随机动态。利用生成式建模,StableUASim可以生成多样化和统计上真实的信道实现,同时支持从特定测量样本中进行条件生成。预训练可以使用最少的额外数据快速适应新环境,自动编码器的潜在表示有助于有效的信道分析和压缩。实验结果表明,StableUASim准确地再现了关键的信道特性和通信性能,为系统设计和机器学习驱动的水下应用提供了一个可扩展的,数据高效的,物理上一致的代理模型。摘要:Accurate modeling of time-varying underwater acoustic channels is essential for the design, evaluation, and deployment of reliable underwater communication systems. Conventional physics models require detailed environmental knowledge, while stochastic replay methods are constrained by the limited diversity of measured channels and often fail to generalize to unseen scenarios, reducing their practical applicability. To address these challenges, we propose StableUASim, a pre-trained conditional latent diffusion surrogate model that captures the stochastic dynamics of underwater acoustic communication channels. Leveraging generative modeling, StableUASim produces diverse and statistically realistic channel realizations, while supporting conditional generation from specific measurement samples. Pre-training enables rapid adaptation to new environments using minimal additional data, and the autoencoder latent representation facilitates efficient channel analysis and compression. Experimental results demonstrate that StableUASim accurately reproduces key channel characteristics and communication performance, providing a scalable, data-efficient, and physically consistent surrogate model for both system design and machine learning-driven underwater applications.


【11】Three-Class Emotion Classification for Audiovisual Scenes Based on Ensemble Learning Scheme
标题:基于整体学习方案的视听场景三级情感分类
链接:https://arxiv.org/pdf/2511.17926v1

作者:Xiangrui Xiong,Zhou Zhou,Guocai Nong,Junlin Deng,Ning Wu
摘要:情感识别在增强人机交互中起着关键作用,特别是在理解情感内容至关重要的电影推荐系统中。虽然结合音频和视频的多模式方法已经证明了有效性,但它们对高性能图形计算的依赖限制了在资源受限设备(如个人计算机或家庭视听系统)上的部署。为了解决这个问题,本研究提出了一种新的音频集成学习框架,能够将电影场景分为三个情感类别:好,中性和坏。该模型集成了10个支持向量机和6个神经网络的堆叠集成架构,以提高分类性能。一个定制的数据预处理管道,包括特征提取,离群值处理和特征工程,旨在优化音频输入的情感信息。在模拟数据集上进行的实验达到了67%的准确率,而从15部不同电影中收集的真实数据集的准确率达到了令人印象深刻的86%。这些结果强调了基于音频的轻量级情感识别方法在更广泛的消费级应用中的潜力,同时提供了计算效率和强大的分类能力。摘要:Emotion recognition plays a pivotal role in enhancing human-computer interaction, particularly in movie recommendation systems where understanding emotional content is essential. While multimodal approaches combining audio and video have demonstrated effectiveness, their reliance on high-performance graphical computing limits deployment on resource-constrained devices such as personal computers or home audiovisual systems. To address this limitation, this study proposes a novel audio-only ensemble learning framework capable of classifying movie scenes into three emotional categories: Good, Neutral, and Bad. The model integrates ten support vector machines and six neural networks within a stacking ensemble architecture to enhance classification performance. A tailored data preprocessing pipeline, including feature extraction, outlier handling, and feature engineering, is designed to optimize emotional information from audio inputs. Experiments on a simulated dataset achieve 67% accuracy, while a real-world dataset collected from 15 diverse films yields an impressive 86% accuracy. These results underscore the potential of audio-based, lightweight emotion recognition methods for broader consumer-level applications, offering both computational efficiency and robust classification capabilities.


【12】Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
标题:生成性对抗训练后缓解实时人机音乐互动中的奖励黑客行为
链接:https://arxiv.org/pdf/2511.17879v1

作者:Yusong Wu,Stephen Brade,Teng Ma,Tia-Jane Fowler,Enning Yang,Berker Banar,Aaron Courville,Natasha Jaques,Cheng-Zhi Anna Huang
摘要:生成式人工智能的大多数应用都涉及顺序交互,其中一个人输入提示并等待响应,反应时间和适应性不是重要因素。相比之下,现场干扰是一种协作互动,需要实时协调和适应,而不需要访问其他玩家的未来动作,同时保留多样性以维持创造性的流动。强化学习后训练通过策略上的交互实现了有效的适应,但它往往通过利用基于一致性的奖励来减少输出多样性。这种崩溃被称为“奖励黑客”,影响了许多RL后训练管道,但在现场干扰中尤其有害,因为音乐创造力依赖于动态变化和相互响应。在本文中,我们提出了一种基于策略生成轨迹的新型对抗性训练方法,以减轻旋律到和弦伴奏的RL后训练中的奖励黑客攻击。一个共同进化的策略将策略轨迹与数据分布分开,而策略除了一致性奖励之外还最大化策略输出,以防止崩溃为微不足道的输出。我们评估伴奏质量和输出的多样性,在模拟与固定的测试旋律和学习旋律代理,我们进行了用户研究与专家音乐家的实时交互系统中部署的模型。定量评价和用户反馈表明,产出多样性、协调一致性、适应速度和用户机构得到了改善。我们的研究结果证明了一种简单而有效的方法,可以在生成序列模型的RL后训练中减轻奖励黑客攻击。摘要:Most applications of generative AI involve a sequential interaction in which a person inputs a prompt and waits for a response, and where reaction time and adaptivity are not important factors. In contrast, live jamming is a collaborative interaction that requires real-time coordination and adaptation without access to the other player's future moves, while preserving diversity to sustain a creative flow. Reinforcement learning post-training enables effective adaptation through on-policy interaction, yet it often reduces output diversity by exploiting coherence-based rewards. This collapse, known as reward hacking'', affects many RL post-training pipelines, but is especially harmful in live jamming, where musical creativity relies on dynamic variation and mutual responsiveness. In this paper, we propose a novel adversarial training method on policy-generated trajectories to mitigate reward hacking in RL post-training for melody-to-chord accompaniment. A co-evolving discriminator separates policy trajectories from the data distribution, while the policy maximizes the discriminator output in addition to coherence rewards to prevent collapse to trivial outputs. We evaluate accompaniment quality and output diversity in simulation with both fixed test melodies and learned melody agents, and we conduct a user study with the model deployed in a real-time interactive system with expert musicians. Quantitative evaluation and user feedback demonstrate improved output diversity, harmonic coherence, adaptation speed and user agency. Our results demonstrate a simple yet effective method to mitigate reward hacking in RL post-training of generative sequence models.


【13】Point of Order: Action-Aware LLM Persona Modeling for Realistic Civic Simulation
标题:秩序问题:面向现实市政模拟的感知LLM角色建模
链接:https://arxiv.org/pdf/2511.17813v1

作者:Scott Merrill,Shashank Srivastava

备注:8 pages (29 pages including appendix), 18 figures. Code and datasets are available at https:github.comsmerrilluncaction-aware-llms. Submitted to ACL 2026

摘要:大型语言模型提供了模拟多方审议的机会,但现实的建模仍然受到缺乏说话者属性数据的限制。通过自动语音识别(ASR)产生的转录本分配匿名说话者标签(例如,Speaker_1),阻止模型捕捉一致的人类行为。这项工作引入了一个可复制的管道,将公共Zoom录音转换为具有元数据(如人物配置文件和实用动作标签)的说话者属性转录本(例如,[propose_motion])。我们发布了三个地方政府审议数据集:上诉法院听证会,学校董事会会议和市议会会议。微调LLM使用这种“动作感知”数据对特定参与者进行建模,可以减少67%的困惑,并将基于分类器的性能指标提高近一倍,以实现扬声器保真度和真实性。图灵风格的人类评估表明,我们的模拟通常与真实的审议难以区分,为复杂的现实城市模拟提供了一种实用且可扩展的方法。摘要:Large language models offer opportunities to simulate multi-party deliberation, but realistic modeling remains limited by a lack of speaker-attributed data. Transcripts produced via automatic speech recognition (ASR) assign anonymous speaker labels (e.g., Speaker_1), preventing models from capturing consistent human behavior. This work introduces a reproducible pipeline to transform public Zoom recordings into speaker-attributed transcripts with metadata like persona profiles and pragmatic action tags (e.g., [propose_motion]). We release three local government deliberation datasets: Appellate Court hearings, School Board meetings, and Municipal Council sessions. Fine-tuning LLMs to model specific participants using this "action-aware" data produces a 67% reduction in perplexity and nearly doubles classifier-based performance metrics for speaker fidelity and realism. Turing-style human evaluations show our simulations are often indistinguishable from real deliberations, providing a practical and scalable method for complex realistic civic simulations.


【14】InstructAudio: Unified speech and music generation with natural language instruction
标题:DirectAudio:通过自然语言教学统一生成语音和音乐
链接:https://arxiv.org/pdf/2511.18487v1

作者:Chunyu Qiang,Kang Yin,Xiaopeng Wang,Yuzhe Liang,Jiahui Zhao,Ruibo Fu,Tianrui Wang,Cheng Gong,Chen Zhang,Longbiao Wang,Jianwu Dang
摘要:文本到语音(TTS)和文本到音乐(TTM)模型在基于语音的控制中面临显著的限制。TTS系统通常依赖于参考音频的音色,只提供有限的文本级属性控制,很少支持对话生成。TTM系统受到依赖于专家知识注释的输入条件要求的约束。这些输入控制条件的高度异质性使得它们难以与语音合成联合建模。尽管共享共同的声学建模特征,但这两个任务长期以来一直是独立开发的,因此通过自然语言指令实现统一建模的挑战仍然存在。我们介绍了InstructAudio,一个统一的框架,使基于语音(自然语言描述)控制的声学属性,包括音色(性别,年龄),语言(情感,风格,口音),和音乐(流派,乐器,节奏,气氛)。它支持表达语音,音乐和英语和中文对话生成。该模型采用联合和单扩散Transformer层,具有标准化的音素输入格式,在50K小时的语音和20K小时的音乐数据上进行训练,从而实现多任务学习和跨模态对齐。图1显示了与主流TTS和TTM模型的性能比较,表明InstructAudio在大多数指标上都达到了最佳效果。据我们所知,InstructAudio代表了第一个统一语音和音乐生成的语音控制框架。音频样本可在以下网址获得:www.example.com摘要:Text-to-speech (TTS) and text-to-music (TTM) models face significant limitations in instruction-based control. TTS systems usually depend on reference audio for timbre, offer only limited text-level attribute control, and rarely support dialogue generation. TTM systems are constrained by input conditioning requirements that depend on expert knowledge annotations. The high heterogeneity of these input control conditions makes them difficult to joint modeling with speech synthesis. Despite sharing common acoustic modeling characteristics, these two tasks have long been developed independently, leaving open the challenge of achieving unified modeling through natural language instructions. We introduce InstructAudio, a unified framework that enables instruction-based (natural language descriptions) control of acoustic attributes including timbre (gender, age), paralinguistic (emotion, style, accent), and musical (genre, instrument, rhythm, atmosphere). It supports expressive speech, music, and dialogue generation in English and Chinese. The model employs joint and single diffusion transformer layers with a standardized instruction-phoneme input format, trained on 50K hours of speech and 20K hours of music data, enabling multi-task learning and cross-modal alignment. Fig. 1 visualizes performance comparisons with mainstream TTS and TTM models, demonstrating that InstructAudio achieves optimal results on most metrics. To our best knowledge, InstructAudio represents the first instruction-controlled framework unifying speech and music generation. Audio samples are available at: https: qiangchunyu.github.io InstructAudio


eess.AS音频处理


【1】First Deep Learning Approach to Hammering Acoustics for Stem Stability Assessment in Total Hip Arthroplasty
标题:第一种深度学习方法,通过敲击声学来评估全髋关节置换术中的柄稳定性
链接:https://arxiv.org/pdf/2511.18725v1

作者:Dongqi Zhu,Zhuwen Xu,Youyuan Chen,Minghao Jin,Wan Zheng,Yi Zhou,Huiwu Li,Yongyun Chang,Feng Hong,Zanjing Zhai
摘要:音频事件分类最近在医学应用中成为一种有前途的方法。在全髋关节置换术(THA)中,术中锤击声为评估股骨柄的初始稳定性提供了关键线索,但由于股骨形态、植入物尺寸和手术技术的差异性限制了传统的评估方法。我们为此任务提出了第一个深度学习框架,采用了一个TimeMIL模型,该模型在Log-Mel Spectrogram特征上进行了训练,并使用伪标记进行了增强。在术中记录中,该方法达到了91.17%+ -2.79%的准确度,证明了股骨柄稳定性的可靠估计。对比实验进一步表明,减少股骨柄品牌的多样性提高了模型性能,尽管有限的数据集大小仍然是一个瓶颈。这些结果确立了基于深度学习的音频事件分类作为THA术中稳定性评估的可行方法。摘要:Audio event classification has recently emerged as a promising approach in medical applications. In total hip arthroplasty (THA), intra-operative hammering acoustics provide critical cues for assessing the initial stability of the femoral stem, yet variability due to femoral morphology, implant size, and surgical technique constrains conventional assessment methods. We propose the first deep learning framework for this task, employing a TimeMIL model trained on Log-Mel Spectrogram features and enhanced with pseudo-labeling. On intra-operative recordings, the method achieved 91.17 % + - 2.79 % accuracy, demonstrating reliable estimation of stem stability. Comparative experiments further show that reducing the diversity of femoral stem brands improves model performance, although limited dataset size remains a bottleneck. These results establish deep learning-based audio event classification as a feasible approach for intra-operative stability assessment in THA.


【2】InstructAudio: Unified speech and music generation with natural language instruction
标题:DirectAudio:通过自然语言教学统一生成语音和音乐
链接:https://arxiv.org/pdf/2511.18487v1

作者:Chunyu Qiang,Kang Yin,Xiaopeng Wang,Yuzhe Liang,Jiahui Zhao,Ruibo Fu,Tianrui Wang,Cheng Gong,Chen Zhang,Longbiao Wang,Jianwu Dang
摘要:文本到语音(TTS)和文本到音乐(TTM)模型在基于语音的控制中面临显著的限制。TTS系统通常依赖于参考音频的音色,只提供有限的文本级属性控制,很少支持对话生成。TTM系统受到依赖于专家知识注释的输入条件要求的约束。这些输入控制条件的高度异质性使得它们难以与语音合成联合建模。尽管共享共同的声学建模特征,但这两个任务长期以来一直是独立开发的,因此通过自然语言指令实现统一建模的挑战仍然存在。我们介绍了InstructAudio,一个统一的框架,使基于语音(自然语言描述)控制的声学属性,包括音色(性别,年龄),语言(情感,风格,口音),和音乐(流派,乐器,节奏,气氛)。它支持表达语音,音乐和英语和中文对话生成。该模型采用联合和单扩散Transformer层,具有标准化的音素输入格式,在50K小时的语音和20K小时的音乐数据上进行训练,从而实现多任务学习和跨模态对齐。图1显示了与主流TTS和TTM模型的性能比较,表明InstructAudio在大多数指标上都达到了最佳效果。据我们所知,InstructAudio代表了第一个统一语音和音乐生成的语音控制框架。音频样本可在以下网址获得:www.example.com摘要:Text-to-speech (TTS) and text-to-music (TTM) models face significant limitations in instruction-based control. TTS systems usually depend on reference audio for timbre, offer only limited text-level attribute control, and rarely support dialogue generation. TTM systems are constrained by input conditioning requirements that depend on expert knowledge annotations. The high heterogeneity of these input control conditions makes them difficult to joint modeling with speech synthesis. Despite sharing common acoustic modeling characteristics, these two tasks have long been developed independently, leaving open the challenge of achieving unified modeling through natural language instructions. We introduce InstructAudio, a unified framework that enables instruction-based (natural language descriptions) control of acoustic attributes including timbre (gender, age), paralinguistic (emotion, style, accent), and musical (genre, instrument, rhythm, atmosphere). It supports expressive speech, music, and dialogue generation in English and Chinese. The model employs joint and single diffusion transformer layers with a standardized instruction-phoneme input format, trained on 50K hours of speech and 20K hours of music data, enabling multi-task learning and cross-modal alignment. Fig. 1 visualizes performance comparisons with mainstream TTS and TTM models, demonstrating that InstructAudio achieves optimal results on most metrics. To our best knowledge, InstructAudio represents the first instruction-controlled framework unifying speech and music generation. Audio samples are available at: https: qiangchunyu.github.io InstructAudio


【3】Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
标题:语音识别模型使用细粒度奖励改进文本到语音合成
链接:https://arxiv.org/pdf/2511.17555v1

作者:Guansu Wang,Peijie Sun

备注:The paper makes an important contribution to the very challenging problem of training TTS models, with a novel application of reinforcement learning and demonstrating convincing improvements

摘要:文本到语音(TTS)的最新进展使模型能够克隆任意看不见的说话者,并合成高质量,听起来自然的语音。然而,评估方法落后:典型的平均意见得分(MOS)估计执行整个话语回归,而失败通常发生在几个有问题的话。我们观察到编码器-解码器ASR模型(例如,Whisper)通过交叉注意力(cross-attention)来显示语音和文本之间的单词级别的不匹配,从而提供细粒度的奖励信号。在此基础上,我们引入了由ASR驱动的注意奖励(W3 AR)的单词级TTS对齐。在没有显式奖励注释的情况下,W3 AR使用来自预训练的ASR模型的注意力来驱动TTS模型预测的序列的细粒度对齐和优化。实验结果表明,W3 AR提高了现有TTS系统的质量,增强了对未知说话人的zero-shot鲁棒性。更广泛地说,我们的研究结果为生成式建模提供了一个简单的方法:理解模型可以充当评估者,为优化提供信息丰富的细粒度反馈。摘要:Recent advances in text-to-speech (TTS) have enabled models to clone arbitrary unseen speakers and synthesize high-quality, natural-sounding speech. However, evaluation methods lag behind: typical mean opinion score (MOS) estimators perform regression over entire utterances, while failures usually occur in a few problematic words. We observe that encoder-decoder ASR models (e.g., Whisper) surface word-level mismatches between speech and text via cross-attention, providing a fine-grained reward signal. Building on this, we introduce Word-level TTS Alignment by ASR-driven Attentive Reward (W3AR). Without explicit reward annotations, W3AR uses attention from a pre-trained ASR model to drive finer-grained alignment and optimization of sequences predicted by a TTS model. Experiments show that W3AR improves the quality of existing TTS systems and strengthens zero-shot robustness on unseen speakers. More broadly, our results suggest a simple recipe for generative modeling: understanding models can act as evaluators, delivering informative, fine-grained feedback for optimization.


【4】Dynamic Multi-Species Bird Soundscape Generation with Acoustic Patterning and 3D Spatialization
标题:利用声学图案化和3D空间化生成动态多物种鸟类声景
链接:https://arxiv.org/pdf/2511.19275v1

作者:Ellie L. Zhang,Duoduo Liao,Callie C. Liao

备注:Accepted by IEEE Big Data 2025

摘要:生成动态的、可扩展的多物种鸟类音景仍然是计算机音乐和算法声音设计中的一个重大挑战。鸟鸣声涉及快速调频啁啾,复杂的振幅包络,独特的声学模式,重叠的呼叫和动态的鸟间互动,所有这些都需要在3D环境中精确的时间和空间控制。现有的方法,无论是基于数字信号处理(DSP)还是数据驱动的,通常仅关注单物种建模、静态调用结构或直接从记录合成,并且通常遭受噪声、有限的灵活性或大数据需求。为了解决这些挑战,我们提出了一种新颖的,完全算法驱动的框架,使用基于DSP的啁啾生成和3D空间化生成动态多物种鸟类音景,而不依赖于录音或训练数据。我们的方法模拟多个独立移动的鸟类,每个物种沿着不同的移动3D轨迹,支持可控的啁啾序列,重叠的合唱,和现实的3D运动在可扩展的音景,同时保留物种特定的声学模式。一个可视化界面提供了鸟类的轨迹,光谱图,活动时间表和声波的分析和创造性的目的。视觉和音频评估都证明了该系统能够生成密集的,身临其境的和生态启发的音景,突出了其在计算机音乐,交互式虚拟环境和计算生物声学研究中的潜力。摘要:Generation of dynamic, scalable multi-species bird soundscapes remains a significant challenge in computer music and algorithmic sound design. Birdsongs involve rapid frequency-modulated chirps, complex amplitude envelopes, distinctive acoustic patterns, overlapping calls, and dynamic inter-bird interactions, all of which require precise temporal and spatial control in 3D environments. Existing approaches, whether Digital Signal Processing (DSP)-based or data-driven, typically focus only on single species modeling, static call structures, or synthesis directly from recordings, and often suffer from noise, limited flexibility, or large data needs. To address these challenges, we present a novel, fully algorithm-driven framework that generates dynamic multi-species bird soundscapes using DSP-based chirp generation and 3D spatialization, without relying on recordings or training data. Our approach simulates multiple independently-moving birds per species along different moving 3D trajectories, supporting controllable chirp sequences, overlapping choruses, and realistic 3D motion in scalable soundscapes while preserving species-specific acoustic patterns. A visualization interface provides bird trajectories, spectrograms, activity timelines, and sound waves for analytical and creative purposes. Both visual and audio evaluations demonstrate the ability of the system to generate dense, immersive, and ecologically inspired soundscapes, highlighting its potential for computer music, interactive virtual environments, and computational bioacoustics research.


【5】Multidimensional Music Aesthetic Evaluation via Semantically Consistent C-Mixup Augmentation
标题:基于语义一致性C-Mixup增强的多维音乐审美评价
链接:https://arxiv.org/pdf/2511.18869v1

作者:Shuyang Liu,Yuan Jin,Rui Lin,Shizhe Chen,Junyu Dai,Tao Jiang
摘要:由于音乐感知的多维性,评估生成歌曲的美学质量具有挑战性。我们提出了一个强大的音乐美学评估框架,该框架结合了(1)多源多尺度特征提取以获得互补的片段和曲目级表示,(2)分层音频增强策略以丰富训练数据,以及(3)集成回归和排名损失的混合训练目标,以获得准确的评分和可靠的热门歌曲识别。在ICASSP 2026 SongEval基准测试上的实验表明,我们的方法在相关性和顶级指标上始终优于基线方法。摘要:Evaluating the aesthetic quality of generated songs is challenging due to the multi-dimensional nature of musical perception. We propose a robust music aesthetic evaluation framework that combines (1) multi-source multi-scale feature extraction to obtain complementary segment- and track-level representations, (2) a hierarchical audio augmentation strategy to enrich training data, and (3) a hybrid training objective that integrates regression and ranking losses for accurate scoring and reliable top-song identification. Experiments on the ICASSP 2026 SongEval benchmark demonstrate that our approach consistently outperforms baseline methods across correlation and top-tier metrics.


【6】PrismAudio: Decomposed Chain-of-Thoughts and Multi-dimensional Rewards for Video-to-Audio Generation
标题:棱镜音频:分解思想链和视频到音频生成的多维奖励
链接:https://arxiv.org/pdf/2511.18833v1

作者:Huadai Liu,Kaicheng Luo,Wen Wang,Qian Chen,Peiwen Sun,Rongjie Huang,Xiangang Li,Jieping Ye,Wei Xue
备注:Preprint
摘要:视频到音频(V2 A)的生成需要平衡四个关键的感知维度:语义一致性,视听时间同步性,美学质量和空间准确性;然而现有的方法存在客观纠缠,将竞争目标合并在单一损失函数中,并且缺乏人类偏好对齐。我们介绍了PrismAudio,这是第一个将强化学习集成到V2 A生成中的框架,具有专门的思想链(CoT)规划。我们的方法将整体推理分解为四个专门的CoT模块(语义,时间,美学和空间CoT),每个模块都与有针对性的奖励函数配对。这种CoT-奖励对应关系实现了多维RL优化,引导模型在所有角度上联合生成更好的推理,解决了客观纠缠问题,同时保持了可解释性。为了使这种优化在计算上切实可行,我们提出了快速GRPO,它采用混合ODE-GRPO采样,大大减少了训练开销相比,现有的GRPO实现。我们还引入了AudioCanvas,这是一个严格的基准测试,它在分布上更加平衡,比现有的数据集覆盖了更现实的多样化和更具挑战性的场景,具有300个单事件类和501个多事件样本。实验结果表明,PrismAudio在域内VGGSound测试集和域外AudioCanvas基准测试中的所有四个感知维度上都达到了最先进的性能。项目页面可在https: PrismAudio-Project.github.io上找到。摘要:Video-to-Audio (V2A) generation requires balancing four critical perceptual dimensions: semantic consistency, audio-visual temporal synchrony, aesthetic quality, and spatial accuracy; yet existing methods suffer from objective entanglement that conflates competing goals in single loss functions and lack human preference alignment. We introduce PrismAudio, the first framework to integrate Reinforcement Learning into V2A generation with specialized Chain-of-Thought (CoT) planning. Our approach decomposes monolithic reasoning into four specialized CoT modules (Semantic, Temporal, Aesthetic, and Spatial CoT), each paired with targeted reward functions. This CoT-reward correspondence enables multidimensional RL optimization that guides the model to jointly generate better reasoning across all perspectives, solving the objective entanglement problem while preserving interpretability. To make this optimization computationally practical, we propose Fast-GRPO, which employs hybrid ODE-SDE sampling that dramatically reduces the training overhead compared to existing GRPO implementations. We also introduce AudioCanvas, a rigorous benchmark that is more distributionally balanced and covers more realistically diverse and challenging scenarios than existing datasets, with 300 single-event classes and 501 multi-event samples. Experimental results demonstrate that PrismAudio achieves state-of-the-art performance across all four perceptual dimensions on both the in-domain VGGSound test set and out-of-domain AudioCanvas benchmark. The project page is available at https: PrismAudio-Project.github.io.


机器翻译由腾讯交互翻译提供,仅供参考