今日论文合集:cs.SD语音16篇,eess.AS音频处理18篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Recomposer: Event-roll-guided generative audio editing
标题:重新组合器:事件滚动引导生成音频编辑
链接:https://arxiv.org/abs/2509.05256

作者:Daniel P. W. Ellis, Eduardo Fonseca, Ron J. Weiss, Kevin Wilson, Scott Wisdom, Hakan Erdogan, John R. Hershey, Aren Jansen, R. Channing Moore, Manoj Plakal
备注:5 pages, 5 figures
摘要:编辑复杂的真实世界声音场景是困难的,因为各个声源在时间上重叠。生成模型可以根据其对数据域的强大先验理解来填充缺失或损坏的细节。我们提出了一种用于在复杂场景内编辑各个声音事件的系统,该系统能够基于文本编辑描述(例如,"增强门“)和从"事件卷”转录导出的事件定时的图形表示。我们提出了一个编码器-解码器Transformer,它工作在SoundStream表示上,在合成(输入,所需输出)音频示例对上进行训练,这些音频示例对是通过将孤立的声音事件添加到密集的真实世界背景中而形成的。评估揭示了编辑描述的每个部分的重要性--动作、类、时间。我们的工作表明“重组”是一个重要且实际的应用。
摘要:Editing complex real-world sound scenes is difficult because individual sound sources overlap in time. Generative models can fill-in missing or corrupted details based on their strong prior understanding of the data domain. We present a system for editing individual sound events within complex scenes able to delete, insert, and enhance individual sound events based on textual edit descriptions (e.g., ``enhance Door'') and a graphical representation of the event timing derived from an ``event roll'' transcription. We present an encoder-decoder transformer working on SoundStream representations, trained on synthetic (input, desired output) audio example pairs formed by adding isolated sound events to dense, real-world backgrounds. Evaluation reveals the importance of each part of the edit descriptions -- action, class, timing. Our work demonstrates ``recomposition'' is an important and practical application.


【2】Exploring Situated Stabilities of a Rhythm Generation System through Variational Cross-Examination
标题:通过变分盘问探索节奏生成系统的情境稳定性
链接:https://arxiv.org/abs/2509.05145

作者:Błazej Kotowski, Nicholas Evans, Behzad Haki, Frederic Font, Sergi Jordà
备注:AI Music Creativity 2025
摘要:本文通过变分交叉检验(VCE)的后现象学框架,研究了一个实时节奏生成系统Groupy Transformer。通过反思其部署在三个不同的艺术背景下,我们确定了三个稳定性:一个自主的鼓伴奏发生器,节奏控制电压音序器在Eurorack格式,和一个和谐的伴奏系统的节奏驱动程序。其应用的多功能性从项目一开始就不是一个明确的目标。因此,我们要问:这种多重稳定性是如何出现的?通过VCE,我们确定了其出现的三个关键因素:系统不变量的启示,跨学科的合作,以及其发展的定位性质。最后,我们反思VCE的可行性作为一种描述和分析方法,数字乐器(MIDI)设计,强调其价值,揭示技术如何调解,共同塑造,并共同塑造用户和上下文。
摘要:This paper investigates GrooveTransformer, a real-time rhythm generation system, through the postphenomenological framework of Variational Cross-Examination (VCE). By reflecting on its deployment across three distinct artistic contexts, we identify three stabilities: an autonomous drum accompaniment generator, a rhythmic control voltage sequencer in Eurorack format, and a rhythm driver for a harmonic accompaniment system. The versatility of its applications was not an explicit goal from the outset of the project. Thus, we ask: how did this multistability emerge? Through VCE, we identify three key contributors to its emergence: the affordances of system invariants, the interdisciplinary collaboration, and the situated nature of its development. We conclude by reflecting on the viability of VCE as a descriptive and analytical method for Digital Musical Instrument (DMI) design, emphasizing its value in uncovering how technologies mediate, co-shape, and are co-shaped by users and contexts.


【3】Training a Perceptual Model for Evaluating Auditory Similarity in Music Adversarial Attack
标题:训练用于评估音乐对抗性攻击中听觉相似性的感知模型
链接:https://arxiv.org/abs/2509.04985

作者:Yuxuan Liu, Rui Sang, Peihong Zhang, Zhixin Li, Shengchen Li
摘要:音乐信息检索(MIR)系统非常容易受到对抗性攻击,这些攻击通常是人类无法感知的,主要是由于模型特征空间和人类听觉感知之间的不一致。现有的防御措施和感知指标经常无法充分捕捉这些听觉细微差别,我们最初的听力测试支持这一限制,显示通用指标与人类判断之间的相关性较低。为了弥合这一差距,我们引入了感知对齐的MERT Transformer(PAMT),这是一种用于学习鲁棒的感知对齐音乐表示的新框架。我们的核心创新在于心理声学调节的顺序对比Transformer,这是一种建立在冻结MERT编码器上的轻量级投影头。PAMT与主观评分的斯皮尔曼相关系数为0.65,优于现有的感知指标。我们的方法在具有挑战性的MIR任务上的鲁棒准确性平均提高了9.15%,包括在各种感知对抗攻击下的封面歌曲识别和音乐流派分类。这项工作开创了架构集成的心理声学条件反射,产生的表示更符合人类的感知,并对音乐对抗性攻击具有鲁棒性。
摘要:Music Information Retrieval (MIR) systems are highly vulnerable to adversarial attacks that are often imperceptible to humans, primarily due to a misalignment between model feature spaces and human auditory perception. Existing defenses and perceptual metrics frequently fail to adequately capture these auditory nuances, a limitation supported by our initial listening tests showing low correlation between common metrics and human judgments. To bridge this gap, we introduce Perceptually-Aligned MERT Transformer (PAMT), a novel framework for learning robust, perceptually-aligned music representations. Our core innovation lies in the psychoacoustically-conditioned sequential contrastive transformer, a lightweight projection head built atop a frozen MERT encoder. PAMT achieves a Spearman correlation coefficient of 0.65 with subjective scores, outperforming existing perceptual metrics. Our approach also achieves an average of 9.15\% improvement in robust accuracy on challenging MIR tasks, including Cover Song Identification and Music Genre Classification, under diverse perceptual adversarial attacks. This work pioneers architecturally-integrated psychoacoustic conditioning, yielding representations significantly more aligned with human perception and robust against music adversarial attacks.


【4】MAIA: An Inpainting-Based Approach for Music Adversarial Attacks
标题:MAIA:一种基于修补的音乐对抗攻击方法
链接:https://arxiv.org/abs/2509.04980

作者:Yuxuan Liu, Peihong Zhang, Rui Sang, Zhixin Li, Shengchen Li
备注:Accepted at ISMIR2025
摘要:音乐对抗性攻击在音乐信息检索(MIR)领域引起了人们的极大兴趣。在本文中,我们提出了音乐对抗性修复攻击(MAIA),一种新的对抗性攻击框架,支持白盒和黑盒攻击场景。MAIA从重要性分析开始,以识别关键的音频片段,然后将其作为修改的目标。利用生成修复模型,这些片段在受攻击模型输出的指导下进行重建,确保微妙而有效的对抗性扰动。我们在多个MIR任务上评估MAIA,在白盒和黑盒设置中展示了高攻击成功率,同时保持最小的感知失真。此外,主观听力测试证实了对抗样本的高音频保真度。我们的研究结果突出了当前MIR系统的漏洞,并强调需要更强大和安全的模型。
摘要:Music adversarial attacks have garnered significant interest in the field of Music Information Retrieval (MIR). In this paper, we present Music Adversarial Inpainting Attack (MAIA), a novel adversarial attack framework that supports both white-box and black-box attack scenarios. MAIA begins with an importance analysis to identify critical audio segments, which are then targeted for modification. Utilizing generative inpainting models, these segments are reconstructed with guidance from the output of the attacked model, ensuring subtle and effective adversarial perturbations. We evaluate MAIA on multiple MIR tasks, demonstrating high attack success rates in both white-box and black-box settings while maintaining minimal perceptual distortion. Additionally, subjective listening tests confirm the high audio fidelity of the adversarial samples. Our findings highlight vulnerabilities in current MIR systems and emphasize the need for more robust and secure models.


【5】Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper
标题:通过多个基础模型映射器高效生成视频到音频
链接:https://arxiv.org/abs/2509.04957

作者:Gehui Chen, Guan'an Wang, Xiaowen Huang, Jitao Sang
摘要:最近的视频到音频(V2 A)生成依赖于从视频中提取语义和时间特征到条件生成模型。从头开始训练这些模型是资源密集型的。因此,利用基础模型(FM)由于其跨模态知识转移和泛化能力而获得了吸引力。之前的一项工作探索了微调轻量级映射器网络,以将预训练的视觉编码器与V2 A的文本到音频生成模型连接起来。受此启发,我们引入了多基础模型映射器(MFM-Mapper)。与以前的映射器方法相比,MFM-Mapper通过融合来自双视觉编码器的特征来获得更丰富的语义和时间信息。此外,通过用GPT-2替换线性映射器,MFM-Mapper改进了特征对齐,在跨模态特征映射和自回归翻译任务之间绘制了平行线。我们的MFM映射器具有显着的训练效率。它在语义和时间一致性方面取得了更好的性能,训练消耗更少,与以前的基于映射器的工作相比,只需要16%的训练规模,但在更大的规模上训练的模型却取得了有竞争力的性能。
摘要:Recent Video-to-Audio (V2A) generation relies on extracting semantic and temporal features from video to condition generative models. Training these models from scratch is resource intensive. Consequently, leveraging foundation models (FMs) has gained traction due to their cross-modal knowledge transfer and generalization capabilities. One prior work has explored fine-tuning a lightweight mapper network to connect a pre-trained visual encoder with a text-to-audio generation model for V2A. Inspired by this, we introduce the Multiple Foundation Model Mapper (MFM-Mapper). Compared to the previous mapper approach, MFM-Mapper benefits from richer semantic and temporal information by fusing features from dual visual encoders. Furthermore, by replacing a linear mapper with GPT-2, MFM-Mapper improves feature alignment, drawing parallels between cross-modal features mapping and autoregressive translation tasks. Our MFM-Mapper exhibits remarkable training efficiency. It achieves better performance in semantic and temporal consistency with fewer training consuming, requiring only 16\% of the training scale compared to previous mapper-based work, yet achieves competitive performance with models trained on a much larger scale.


【6】Learning and composing of classical music using restricted Boltzmann machines
标题:使用受限制的Boltzmann机器学习和创作古典音乐
链接:https://arxiv.org/abs/2509.04899

作者:Mutsumi Kobayashi, Hiroshi Watanabe
备注:19 pages, 10 figures
摘要:最近,已经开发出软件,利用机器学习来模仿特定作曲家的风格,例如J. S。巴赫.但是,由于这类软件往往采用结构复杂的机器学习模型,因此很难分析软件如何理解作曲家音乐的特征。在这项研究中,我们采用了J。巴赫的音乐用于限制玻尔兹曼机(RBM)的训练。由于RBM的结构简单,它允许我们研究学习后的内部状态。我们发现,学习RBM是能够创作音乐。
摘要:Recently, software has been developed that uses machine learning to mimic the style of a particular composer, such as J. S. Bach. However, since such software often adopts machine learning models with complex structures, it is difficult to analyze how the software understands the characteristics of the composer's music. In this study, we adopted J. S. Bach's music for training of a restricted Boltzmann machine (RBM). Since the structure of RBMs is simple, it allows us to investigate the internal states after learning. We found that the learned RBM is able to compose music.


【7】Quantum Fourier Transform Based Denoising: Unitary Filtering for Enhanced Speech Clarity
标题:基于量子傅里叶变换的去噪:用于增强语音清晰度的元化过滤
链接:https://arxiv.org/abs/2509.04851

作者:Rajeshwar Tripathi, Sahil Tomar, Sandeep Kumar, Monika Aggarwal
备注:8 pages
摘要:本文介绍了一种量子启发的去噪框架,该框架将量子傅立叶变换(QFT)集成到经典的音频增强管道中。与传统的基于快速傅立叶变换(FFT)的方法不同,QFT提供了具有全局相位相干性和能量保存的酉变换,使得能够改进语音和噪声之间的区分。所提出的方法用QFT算子取代Wiener和谱减法滤波器中的FFT,确保一致的超参数设置以进行公平比较。在不同信噪比(SNR)条件下对干净语音、合成音调和嘈杂混合物进行的实验表明,SNR在统计上有显著的提高,最高可达15 dB的改善,并减少了伪影的产生。结果证实,基于QFT的去噪在低SNR和非平稳噪声场景下提供了鲁棒性,而无需额外的计算开销,突出了其作为量子增强语音处理的可扩展途径的潜力。
摘要:This paper introduces a quantum-inspired denoising framework that integrates the Quantum Fourier Transform (QFT) into classical audio enhancement pipelines. Unlike conventional Fast Fourier Transform (FFT) based methods, QFT provides a unitary transformation with global phase coherence and energy preservation, enabling improved discrimination between speech and noise. The proposed approach replaces FFT in Wiener and spectral subtraction filters with a QFT operator, ensuring consistent hyperparameter settings for fair comparison. Experiments on clean speech, synthetic tones, and noisy mixtures across diverse signal to noise ratio (SNR) conditions, demonstrate statistically significant gains in SNR, with up to 15 dB improvement and reduced artifact generation. Results confirm that QFT based denoising offers robustness under low SNR and nonstationary noise scenarios without additional computational overhead, highlighting its potential as a scalable pathway toward quantum-enhanced speech processing.


【8】WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
标题:WildScore:狂野符号音乐推理中的MLLM基准
链接:https://arxiv.org/abs/2509.04744

作者:Gagan Mundada, Yash Vishe, Amit Namburi, Xin Xu, Zachary Novack, Julian McAuley, Junda Wu
摘要:多模态大型语言模型(MLLM)的最新进展已经在各种视觉语言任务中展示了令人印象深刻的能力。然而,他们在多模态符号音乐领域的推理能力仍然在很大程度上未被探索。我们介绍WildScore,第一个在野外多模态符号音乐推理和分析基准,旨在评估MLLM的能力,解释现实世界的乐谱和回答复杂的音乐学查询。WildScore中的每个实例都来自真实的音乐作品,并伴随着真实的用户生成的问题和讨论,捕捉实际音乐分析的复杂性。为了便于系统的评估,我们提出了一个系统的分类,包括高层次和细粒度的音乐本体。此外,我们框架复杂的音乐推理作为多项选择题的回答,使控制和可扩展的评估MLLM的符号音乐的理解。在WildScore上对最先进的MLLM进行的经验基准测试揭示了他们的视觉符号推理中的有趣模式,揭示了MLLM在符号音乐推理和分析中的有前途的方向和持续的挑战。我们发布数据集和代码。
摘要:Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, their reasoning abilities in the multimodal symbolic music domain remain largely unexplored. We introduce WildScore, the first in-the-wild multimodal symbolic music reasoning and analysis benchmark, designed to evaluate MLLMs' capacity to interpret real-world music scores and answer complex musicological queries. Each instance in WildScore is sourced from genuine musical compositions and accompanied by authentic user-generated questions and discussions, capturing the intricacies of practical music analysis. To facilitate systematic evaluation, we propose a systematic taxonomy, comprising both high-level and fine-grained musicological ontologies. Furthermore, we frame complex music reasoning as multiple-choice question answering, enabling controlled and scalable assessment of MLLMs' symbolic music understanding. Empirical benchmarking of state-of-the-art MLLMs on WildScore reveals intriguing patterns in their visual-symbolic reasoning, uncovering both promising directions and persistent challenges for MLLMs in symbolic music reasoning and analysis. We release the dataset and code.


【9】A Multiclass Acoustic Dataset and Interactive Tool for Analyzing Drone Signatures in Real-World Environments
标题:用于分析现实环境中无人机签名的多类声学数据集和交互工具
链接:https://arxiv.org/abs/2509.04715

作者:Mia Y. Wang, Mackenzie Linn, Andrew P. Berg, Qian Zhang
备注:This article extends our previous work presented in the 2024 Artificial Intelligence x Humanities, Education, and Art (2024 AIxHeart) Conference
摘要:无人机在各个行业的迅速普及带来了与隐私、安全和噪音污染相关的重大挑战。目前的无人机探测系统主要基于视觉和雷达技术,在某些条件下面临局限性,这突出表明需要有效的基于声学的探测方法。本文提供了一个独特而全面的无人机声学特征数据集,包括按品牌和型号区分的32个不同类别。该数据集包括每架无人机的原始音频记录、声谱图和梅尔频率倒谱系数(MFCC)图。此外,我们还介绍了一个交互式Web应用程序,允许用户通过选择特定的无人机类别、收听相关音频以及查看相应的频谱图和MFCC图来探索此数据集。该工具旨在促进无人机检测,分类和声学分析的研究,支持技术进步和教育计划。本文详细介绍了数据集的创建过程,Web应用程序的设计和实现,并提供了实验结果和用户反馈。最后,我们讨论了潜在的应用和未来的工作,以扩大和加强该项目。
摘要:The rapid proliferation of drones across various industries has introduced significant challenges related to privacy, security, and noise pollution. Current drone detection systems, primarily based on visual and radar technologies, face limitations under certain conditions, highlighting the need for effective acoustic-based detection methods. This paper presents a unique and comprehensive dataset of drone acoustic signatures, encompassing 32 different categories differentiated by brand and model. The dataset includes raw audio recordings, spectrogram plots, and Mel-frequency cepstral coefficient (MFCC) plots for each drone. Additionally, we introduce an interactive web application that allows users to explore this dataset by selecting specific drone categories, listening to the associated audio, and viewing the corresponding spectrogram and MFCC plots. This tool aims to facilitate research in drone detection, classification, and acoustic analysis, supporting both technological advancements and educational initiatives. The paper details the dataset creation process, the design and implementation of the web application, and provides experimental results and user feedback. Finally, we discuss potential applications and future work to expand and enhance the project.


【10】Ecologically Valid Benchmarking and Adaptive Attention: Scalable Marine Bioacoustic Monitoring
标题:公共有效的基准和适应性注意力:可扩展的海洋生物声学监测
链接:https://arxiv.org/abs/2509.04682

作者:Nicholas R. Rasmussen, Rodrigue Rizk, Longwei Wang, KC Santosh
备注:Under review as an anonymous submission to IEEETAI - We are allowed an archive submission. Final formatting is yet to be determined
摘要:水下被动声监测(UPAM)为长期生态分析提供了丰富的时空数据,但固有噪声和复杂的信号依赖性阻碍了模型的稳定性和通用性。多层窗口提高了目标声音定位,但变化的环境噪声,不同的传播效果,以及混合的生物和人为来源需要强大的架构和严格的评估。我们介绍GetNetUPAM,一个分层嵌套的交叉验证框架,旨在量化生态现实的变化下的模型稳定性。数据被划分为不同的站点年段,保留记录异质性,并确保每个验证折叠反映了一个独特的环境子集,减少局部噪声和传感器伪影的过拟合。站点年阻塞强制执行对真正的环境多样性的评估,而标准的交叉验证随机子集的措施,在UPAM的完整的信号分布,一个维度缺席目前的基准泛化。使用GetNetUPAM作为评估骨干,我们提出了自适应分辨率池和注意力网络(ARPA-N),这是一种用于不规则频谱图维度的神经架构。具有空间注意力的自适应池扩展了感受野,在没有过多参数的情况下捕获全局上下文。在GetNetUPAM下,ARPA-N在DenseNet基线上实现了14.4%的平均精度增益,并且在所有指标上的变异性都下降了log 2量级,从而实现了跨站点年折叠的一致检测,并推进了可扩展的准确生物声学监测。
摘要:Underwater Passive Acoustic Monitoring (UPAM) provides rich spatiotemporal data for long-term ecological analysis, but intrinsic noise and complex signal dependencies hinder model stability and generalization. Multilayered windowing has improved target sound localization, yet variability from shifting ambient noise, diverse propagation effects, and mixed biological and anthropogenic sources demands robust architectures and rigorous evaluation. We introduce GetNetUPAM, a hierarchical nested cross-validation framework designed to quantify model stability under ecologically realistic variability. Data are partitioned into distinct site-year segments, preserving recording heterogeneity and ensuring each validation fold reflects a unique environmental subset, reducing overfitting to localized noise and sensor artifacts. Site-year blocking enforces evaluation against genuine environmental diversity, while standard cross-validation on random subsets measures generalization across UPAM's full signal distribution, a dimension absent from current benchmarks. Using GetNetUPAM as the evaluation backbone, we propose the Adaptive Resolution Pooling and Attention Network (ARPA-N), a neural architecture for irregular spectrogram dimensions. Adaptive pooling with spatial attention extends the receptive field, capturing global context without excessive parameters. Under GetNetUPAM, ARPA-N achieves a 14.4% gain in average precision over DenseNet baselines and a log2-scale order-of-magnitude drop in variability across all metrics, enabling consistent detection across site-year folds and advancing scalable, accurate bioacoustic monitoring.


【11】Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
标题:基于大语言模型的多说话者语音识别的序列化输出预处理
链接:https://arxiv.org/abs/2509.04488

作者:Hao Shi, Yusuke Fujita, Tomoya Mizumoto, Lianbo Liu, Atsushi Kojima, Yui Sudo
摘要:摘要对于任务定义和提高基于大型语言模型(LLM)的系统的性能至关重要。然而,现有的基于LLM的多说话者(MT)自动语音识别(ASR)系统要么省略提示,要么依赖于简单的任务定义提示,没有先前的工作探索提示的设计,以提高性能。在本文中,我们提出了提取序列化的输出提示(SOP),并明确指导LLM使用结构化提示,以提高系统性能(SOP-MT-ASR)。在语音编码器之后插入分离器和序列化的连接主义时间分类(CTC)层,以按照先说先出的方式从混合语音编码中分离和提取MT内容。随后,SOP,这是作为一个提示LLM,通过使用贪婪搜索解码的串行CTC输出。为了有效地训练模型,我们设计了一个三阶段的训练策略,包括序列化输出训练(SOT)微调,序列化语音信息提取,和基于SOP的适应。LibriMix数据集上的实验结果表明,尽管基于LLM的SOT模型在两个谈话者场景中表现良好,但在更复杂的条件下,如三个谈话者场景,它无法充分利用LLM。所提出的SOP方法显着提高了性能,在两个和三个谈话者的条件下。
摘要:Prompts are crucial for task definition and for improving the performance of large language models (LLM)-based systems. However, existing LLM-based multi-talker (MT) automatic speech recognition (ASR) systems either omit prompts or rely on simple task-definition prompts, with no prior work exploring the design of prompts to enhance performance. In this paper, we propose extracting serialized output prompts (SOP) and explicitly guiding the LLM using structured prompts to improve system performance (SOP-MT-ASR). A Separator and serialized Connectionist Temporal Classification (CTC) layers are inserted after the speech encoder to separate and extract MT content from the mixed speech encoding in a first-speaking-first-out manner. Subsequently, the SOP, which serves as a prompt for LLMs, is obtained by decoding the serialized CTC outputs using greedy search. To train the model effectively, we design a three-stage training strategy, consisting of serialized output training (SOT) fine-tuning, serialized speech information extraction, and SOP-based adaptation. Experimental results on the LibriMix dataset show that, although the LLM-based SOT model performs well in the two-talker scenario, it fails to fully leverage LLMs under more complex conditions, such as the three-talker scenario. The proposed SOP approach significantly improved performance under both two- and three-talker conditions.


【12】MEAN-RIR: Multi-Modal Environment-Aware Network for Robust Room Impulse Response Estimation
标题:MEAN-RIR:用于稳健房间脉冲响应估计的多模式环境感知网络
链接:https://arxiv.org/abs/2509.05205

作者:Jiajian Chen, Jiakang Chen, Hang Chen, Qing Wang, Yu Gao, Jun Du
备注:Accepted by ASRU 2025
摘要:本文提出了一种多模态环境感知网络(MEAN-RIR),它使用编码器-解码器框架来预测房间脉冲响应(RIR)的基础上,从音频,视频和文本源的多层次环境信息。具体地,混响语音捕获房间声学特性作为主要输入,其与作为补充输入的全景图像和文本描述相结合。每个输入由其各自的编码器处理,输出被馈送到交叉注意模块,以实现不同模态之间的有效交互。MEAN-RIR解码器生成两个不同的分量:第一个分量捕获直达声和早期反射,而第二个分量产生调制可学习的滤波噪声以合成后期混响的掩模。将这两个分量混合以重建最终的RIR。结果表明,MEAN-RIR显着改善RIR估计,与声学参数的显着增益。
摘要:This paper presents a Multi-Modal Environment-Aware Network (MEAN-RIR), which uses an encoder-decoder framework to predict room impulse response (RIR) based on multi-level environmental information from audio, visual, and textual sources. Specifically, reverberant speech capturing room acoustic properties serves as the primary input, which is combined with panoramic images and text descriptions as supplementary inputs. Each input is processed by its respective encoder, and the outputs are fed into cross-attention modules to enable effective interaction between different modalities. The MEAN-RIR decoder generates two distinct components: the first component captures the direct sound and early reflections, while the second produces masks that modulate learnable filtered noise to synthesize the late reverberation. These two components are mixed to reconstruct the final RIR. The results show that MEAN-RIR significantly improves RIR estimation, with notable gains in acoustic parameters.


【13】Lightweight DNN for Full-Band Speech Denoising on Mobile Devices: Exploiting Long and Short Temporal Patterns
标题:用于移动设备上全频段语音去噪的轻量级DNN:利用长时间和短时间模式
链接:https://arxiv.org/abs/2509.05079

作者:Konstantinos Drossos, Mikko Heikkinen, Paschalis Tsiaflakis
备注:Accepted for publication in Proceedings of the 2025 IEEE 27th International Workshop on Multimedia Signal Processing (MMSP)
摘要:语音去噪(SD)是许多(如果不是全部)用于设备和日常生活应用的现代信号处理链的重要任务。虽然有许多已发布的强大的基于深度神经网络(DNN)的SD方法,但很少有针对移动设备等资源受限平台进行优化。此外,用于SD的大多数基于DNN的方法并不关注全频带(FB)信号,即具有48 kHz采样率和/或低延迟情况。在本文中,我们提出了一种因果关系,低延迟,轻量级的基于DNN的方法,用于全频带SD,利用短期和长期的时间模式。该方法是基于一个修改后的UNet架构,采用回看帧,卷积核的时间跨度,和循环神经网络,利用短期和长期的信号和估计的去噪掩模的时间模式。DNN以STFT幅度作为输入,在因果逐帧的基础上进行操作,利用MobileNet启发的反向瓶颈,采用因果实例归一化进行信道归一化,并在部署在现代移动电话上时实现低于0.02的实时因子。使用已建立的语音去噪指标和公开可用的数据集对所提出的方法进行评估,证明其在实现优于现有FB和低延迟SD方法的(SI-)SDR值方面的有效性。
摘要:Speech denoising (SD) is an important task of many, if not all, modern signal processing chains used in devices and for everyday-life applications. While there are many published and powerful deep neural network (DNN)-based methods for SD, few are optimized for resource-constrained platforms such as mobile devices. Additionally, most DNN-based methods for SD are not focusing on full-band (FB) signals, i.e. having 48 kHz sampling rate, and/or low latency cases. In this paper we present a causal, low latency, and lightweight DNN-based method for full-band SD, leveraging both short and long temporal patterns. The method is based on a modified UNet architecture employing look-back frames, temporal spanning of convolutional kernels, and recurrent neural networks for exploiting short and long temporal patterns in the signal and estimated denoising mask. The DNN operates on a causal frame-by-frame basis taking as an input the STFT magnitude, utilizes inverted bottlenecks inspired by MobileNet, employs causal instance normalization for channel-wise normalization, and achieves a real-time factor below 0.02 when deployed on a modern mobile phone. The proposed method is evaluated using established speech denoising metrics and publicly available datasets, demonstrating its effectiveness in achieving an (SI-)SDR value that outperforms existing FB and low latency SD methods.


【14】Layer-wise Analysis for Quality of Multilingual Synthesized Speech
标题:多语言合成语音质量的分层分析
链接:https://arxiv.org/abs/2509.04830

作者:Erica Cooper, Takuma Okamoto, Yamato Ohtani, Tomoki Toda, Hisashi Kawai
备注:Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
摘要:虽然合成语音的监督质量预测器已经证明与人类评级具有很强的相关性,但它们对域内标记训练数据的要求阻碍了它们对新领域的泛化能力。基于预训练的自监督学习(SSL)模型和自动语音识别(ASR)模型的无监督方法是一种很有前途的选择;然而,人们对这些模型如何编码语音质量信息知之甚少。为了更好地理解语音质量的不同方面是如何在多语言环境中编码的,我们提出了一种基于参考建模的多语言预训练语音模型的分层分析。我们发现,从早期的SSL层提取的特征显示出与人类对合成语音的评级相关,而ASR模型的后期层可以预测非神经系统的质量以及可懂度。我们还证明了使用匹配良好的参考数据的重要性。
摘要:While supervised quality predictors for synthesized speech have demonstrated strong correlations with human ratings, their requirement for in-domain labeled training data hinders their generalization ability to new domains. Unsupervised approaches based on pretrained self-supervised learning (SSL) based models and automatic speech recognition (ASR) models are a promising alternative; however, little is known about how these models encode information about speech quality. Towards the goal of better understanding how different aspects of speech quality are encoded in a multilingual setting, we present a layer-wise analysis of multilingual pretrained speech models based on reference modeling. We find that features extracted from early SSL layers show correlations with human ratings of synthesized speech, and later layers of ASR models can predict quality of non-neural systems as well as intelligibility. We also demonstrate the importance of using well-matched reference data.


【15】Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
标题:少花钱多说:通过自适应集群和隐式持续时间编码的可变帧率语音令牌化
链接:https://arxiv.org/abs/2509.04685

作者:Rui-Chen Zheng, Wenrui Liu, Hui-Peng Du, Qinglin Zhang, Chong Deng, Qian Chen, Wen Wang, Yang Ai, Zhen-Hua Ling
摘要:现有的语音标记器通常每秒分配固定数量的标记,而不管语音信号中的变化的信息密度或时间波动。这种统一的令牌分配与语音的内在结构不匹配,其中信息随时间不均匀地分布。为了解决这个问题,我们提出了VARSTok,可变帧速率语音令牌,适应令牌分配的基础上,本地特征相似性。VARSTok引入了两项关键创新:(1)时间感知密度峰值聚类算法,可自适应地将语音分段为可变长度单元,以及(2)新颖的隐式持续时间编码方案,将内容和时间跨度嵌入单个令牌索引,无需辅助持续时间预测器。大量实验表明,VARSTok显著优于强固定速率基线。值得注意的是,它实现了卓越的重建自然度,同时使用的令牌比40 Hz固定帧速率基线少23%。VARSTok进一步降低了字错误率,并提高了zero-shot文本到语音合成的自然度。据我们所知,这是第一个证明完全动态,可变帧速率声学语音标记器可以无缝集成到下游语音语言模型的工作。语音样本可在https://zhengrachel.github.io/VARSTok上获得。
摘要:Existing speech tokenizers typically assign a fixed number of tokens per second, regardless of the varying information density or temporal fluctuations in the speech signal. This uniform token allocation mismatches the intrinsic structure of speech, where information is distributed unevenly over time. To address this, we propose VARSTok, a VAriable-frame-Rate Speech Tokenizer that adapts token allocation based on local feature similarity. VARSTok introduces two key innovations: (1) a temporal-aware density peak clustering algorithm that adaptively segments speech into variable-length units, and (2) a novel implicit duration coding scheme that embeds both content and temporal span into a single token index, eliminating the need for auxiliary duration predictors. Extensive experiments show that VARSTok significantly outperforms strong fixed-rate baselines. Notably, it achieves superior reconstruction naturalness while using up to 23% fewer tokens than a 40 Hz fixed-frame-rate baseline. VARSTok further yields lower word error rates and improved naturalness in zero-shot text-to-speech synthesis. To the best of our knowledge, this is the first work to demonstrate that a fully dynamic, variable-frame-rate acoustic speech tokenizer can be seamlessly integrated into downstream speech language models. Speech samples are available at https://zhengrachel.github.io/VARSTok.


【16】On Time Delay Interpolation for Improved Acoustic Reflector Localization
标题:改进声反射器定位的延时插值
链接:https://arxiv.org/abs/2509.04629

作者:Hannes Rosseel, Toon van Waterschoot
备注:20 pages, 13 figures, 2 tables, submitted to J. Acoust. Soc. Am
摘要:声反射体的定位是室内声学分析、声源定位和声场景分析等应用中的一个基本组成部分。时间延迟估计(TDE)是确定反射器相对于传感器阵列的位置的关键。传统的TDE算法通常产生的时间延迟是操作采样周期的整数倍,可能缺乏足够的时间分辨率。为了实现子采样TDE精度,已经提出了各种插值方法,包括抛物线、高斯、频率和sinc插值。本文提出了一个全面的研究时间延迟插值,以实现混响条件下的声反射定位的子样本精度。我们推导出惠特克-香农插值公式,从先前提出的sinc插值的上下文中的短时间加窗时域估计声反射定位。仿真表明,sinc和Whittaker-Shannon插值在临界采样和带限反射的时延误差和位置误差方面优于现有方法。在MYRiAD数据集的真实测量结果上评估了性能,结果表明sinc和Whittaker-Shannon插值在不同的传感器-源对和扬声器位置上始终提供可靠的性能。这些结果可以提高声反射器定位系统的精度,对于诸如室内声学分析、声源定位和声学场景分析等应用至关重要。
摘要:The localization of acoustic reflectors is a fundamental component in various applications, including room acoustics analysis, sound source localization, and acoustic scene analysis. Time Delay Estimation (TDE) is essential for determining the position of reflectors relative to a sensor array. Traditional TDE algorithms generally yield time delays that are integer multiples of the operating sampling period, potentially lacking sufficient time resolution. To achieve subsample TDE accuracy, various interpolation methods, including parabolic, Gaussian, frequency, and sinc interpolation, have been proposed. This paper presents a comprehensive study on time delay interpolation to achieve subsample accuracy for acoustic reflector localization in reverberant conditions. We derive the Whittaker-Shannon interpolation formula from the previously proposed sinc interpolation in the context of short-time windowed TDE for acoustic reflector localization. Simulations show that sinc and Whittaker-Shannon interpolation outperform existing methods in terms of time delay error and positional error for critically sampled and band-limited reflections. Performance is evaluated on real-world measurements from the MYRiAD dataset, showing that sinc and Whittaker-Shannon interpolation consistently provide reliable performance across different sensor-source pairs and loudspeaker positions. These results can enhance the precision of acoustic reflector localization systems, vital for applications such as room acoustics analysis, sound source localization, and acoustic scene analysis.


eess.AS音频处理


【1】MEAN-RIR: Multi-Modal Environment-Aware Network for Robust Room Impulse Response Estimation
标题:MEAN-RIR:用于稳健房间脉冲响应估计的多模式环境感知网络
链接:https://arxiv.org/abs/2509.05205

作者:Jiajian Chen, Jiakang Chen, Hang Chen, Qing Wang, Yu Gao, Jun Du
备注:Accepted by ASRU 2025
摘要:本文提出了一种多模态环境感知网络(MEAN-RIR),它使用编码器-解码器框架来预测房间脉冲响应(RIR)的基础上,从音频,视频和文本源的多层次环境信息。具体地,混响语音捕获房间声学特性作为主要输入,其与作为补充输入的全景图像和文本描述相结合。每个输入由其各自的编码器处理,输出被馈送到交叉注意模块,以实现不同模态之间的有效交互。MEAN-RIR解码器生成两个不同的分量:第一个分量捕获直达声和早期反射,而第二个分量产生调制可学习的滤波噪声以合成后期混响的掩模。将这两个分量混合以重建最终的RIR。结果表明,MEAN-RIR显着改善RIR估计,与声学参数的显着增益。
摘要:This paper presents a Multi-Modal Environment-Aware Network (MEAN-RIR), which uses an encoder-decoder framework to predict room impulse response (RIR) based on multi-level environmental information from audio, visual, and textual sources. Specifically, reverberant speech capturing room acoustic properties serves as the primary input, which is combined with panoramic images and text descriptions as supplementary inputs. Each input is processed by its respective encoder, and the outputs are fed into cross-attention modules to enable effective interaction between different modalities. The MEAN-RIR decoder generates two distinct components: the first component captures the direct sound and early reflections, while the second produces masks that modulate learnable filtered noise to synthesize the late reverberation. These two components are mixed to reconstruct the final RIR. The results show that MEAN-RIR significantly improves RIR estimation, with notable gains in acoustic parameters.


【2】Room-acoustic simulations as an alternative to measurements for audio-algorithm evaluation
标题:室内声学模拟作为音频算法评估测量的替代方案
链接:https://arxiv.org/abs/2509.05175

作者:Georg Götz, Daniel Gert Nielsen, Steinar Guðjónsson, Finnur Pind
摘要:音频信号处理和音频机器学习(ASP/AML)算法在智能设备、可穿戴设备和娱乐系统等现代技术中无处不在。此类算法和模型的开发通常涉及正式评估,以证明其有效性和超越最新技术水平的进展。理想情况下,全面评估应涵盖许多不同的应用场景和室内声学条件。然而,在实践中,评价数据集的规模和多样性往往有限,因为它们依赖于昂贵和耗时的测量。本文探讨了如何室内声学模拟可用于评估ASP/AML算法。为此,我们评估三个ASP/AML算法与房间声学测量和来自不同模拟引擎的数据,并评估从测量和模拟获得的评估结果之间的匹配。所提出的调查比较数值波为基础的求解器与两个几何声学模拟器。虽然基于数值波的模拟产生了与所有三种评估的ASP/AML算法的测量结果相似的评估结果,但几何声学模拟无法可靠地复制测量的评估结果。
摘要:Audio-signal-processing and audio-machine-learning (ASP/AML) algorithms are ubiquitous in modern technology like smart devices, wearables, and entertainment systems. Development of such algorithms and models typically involves a formal evaluation to demonstrate their effectiveness and progress beyond the state-of-the-art. Ideally, a thorough evaluation should cover many diverse application scenarios and room-acoustic conditions. However, in practice, evaluation datasets are often limited in size and diversity because they rely on costly and time-consuming measurements. This paper explores how room-acoustic simulations can be used for evaluating ASP/AML algorithms. To this end, we evaluate three ASP/AML algorithms with room-acoustic measurements and data from different simulation engines, and assess the match between the evaluation results obtained from measurements and simulations. The presented investigation compares a numerical wave-based solver with two geometrical acoustics simulators. While numerical wave-based simulations yielded similar evaluation results as measurements for all three evaluated ASP/AML algorithms, geometrical acoustic simulations could not replicate the measured evaluation results as reliably.


【3】Lightweight DNN for Full-Band Speech Denoising on Mobile Devices: Exploiting Long and Short Temporal Patterns
标题:用于移动设备上全频段语音去噪的轻量级DNN:利用长时间和短时间模式
链接:https://arxiv.org/abs/2509.05079

作者:Konstantinos Drossos, Mikko Heikkinen, Paschalis Tsiaflakis
备注:Accepted for publication in Proceedings of the 2025 IEEE 27th International Workshop on Multimedia Signal Processing (MMSP)
摘要:语音去噪(SD)是许多(如果不是全部)用于设备和日常生活应用的现代信号处理链的重要任务。虽然有许多已发布的强大的基于深度神经网络(DNN)的SD方法,但很少有针对移动设备等资源受限平台进行优化。此外,用于SD的大多数基于DNN的方法并不关注全频带(FB)信号,即具有48 kHz采样率和/或低延迟情况。在本文中,我们提出了一种因果关系,低延迟,轻量级的基于DNN的方法,用于全频带SD,利用短期和长期的时间模式。该方法是基于一个修改后的UNet架构,采用回看帧,卷积核的时间跨度,和循环神经网络,利用短期和长期的信号和估计的去噪掩模的时间模式。DNN以STFT幅度作为输入,在因果逐帧的基础上进行操作,利用MobileNet启发的反向瓶颈,采用因果实例归一化进行信道归一化,并在部署在现代移动电话上时实现低于0.02的实时因子。使用已建立的语音去噪指标和公开可用的数据集对所提出的方法进行评估,证明其在实现优于现有FB和低延迟SD方法的(SI-)SDR值方面的有效性。
摘要:Speech denoising (SD) is an important task of many, if not all, modern signal processing chains used in devices and for everyday-life applications. While there are many published and powerful deep neural network (DNN)-based methods for SD, few are optimized for resource-constrained platforms such as mobile devices. Additionally, most DNN-based methods for SD are not focusing on full-band (FB) signals, i.e. having 48 kHz sampling rate, and/or low latency cases. In this paper we present a causal, low latency, and lightweight DNN-based method for full-band SD, leveraging both short and long temporal patterns. The method is based on a modified UNet architecture employing look-back frames, temporal spanning of convolutional kernels, and recurrent neural networks for exploiting short and long temporal patterns in the signal and estimated denoising mask. The DNN operates on a causal frame-by-frame basis taking as an input the STFT magnitude, utilizes inverted bottlenecks inspired by MobileNet, employs causal instance normalization for channel-wise normalization, and achieves a real-time factor below 0.02 when deployed on a modern mobile phone. The proposed method is evaluated using established speech denoising metrics and publicly available datasets, demonstrating its effectiveness in achieving an (SI-)SDR value that outperforms existing FB and low latency SD methods.


【4】Layer-wise Analysis for Quality of Multilingual Synthesized Speech
标题:多语言合成语音质量的分层分析
链接:https://arxiv.org/abs/2509.04830

作者:Erica Cooper, Takuma Okamoto, Yamato Ohtani, Tomoki Toda, Hisashi Kawai
备注:Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
摘要:虽然合成语音的监督质量预测器已经证明与人类评级具有很强的相关性,但它们对域内标记训练数据的要求阻碍了它们对新领域的泛化能力。基于预训练的自监督学习(SSL)模型和自动语音识别(ASR)模型的无监督方法是一种很有前途的选择;然而,人们对这些模型如何编码语音质量信息知之甚少。为了更好地理解语音质量的不同方面是如何在多语言环境中编码的,我们提出了一种基于参考建模的多语言预训练语音模型的分层分析。我们发现,从早期的SSL层提取的特征显示出与人类对合成语音的评级相关,而ASR模型的后期层可以预测非神经系统的质量以及可懂度。我们还证明了使用匹配良好的参考数据的重要性。
摘要:While supervised quality predictors for synthesized speech have demonstrated strong correlations with human ratings, their requirement for in-domain labeled training data hinders their generalization ability to new domains. Unsupervised approaches based on pretrained self-supervised learning (SSL) based models and automatic speech recognition (ASR) models are a promising alternative; however, little is known about how these models encode information about speech quality. Towards the goal of better understanding how different aspects of speech quality are encoded in a multilingual setting, we present a layer-wise analysis of multilingual pretrained speech models based on reference modeling. We find that features extracted from early SSL layers show correlations with human ratings of synthesized speech, and later layers of ASR models can predict quality of non-neural systems as well as intelligibility. We also demonstrate the importance of using well-matched reference data.


【5】Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
标题:少花钱多说:通过自适应集群和隐式持续时间编码的可变帧率语音令牌化
链接:https://arxiv.org/abs/2509.04685

作者:Rui-Chen Zheng, Wenrui Liu, Hui-Peng Du, Qinglin Zhang, Chong Deng, Qian Chen, Wen Wang, Yang Ai, Zhen-Hua Ling
摘要:现有的语音标记器通常每秒分配固定数量的标记,而不管语音信号中的变化的信息密度或时间波动。这种统一的令牌分配与语音的内在结构不匹配,其中信息随时间不均匀地分布。为了解决这个问题,我们提出了VARSTok,可变帧速率语音令牌,适应令牌分配的基础上,本地特征相似性。VARSTok引入了两项关键创新:(1)时间感知密度峰值聚类算法,可自适应地将语音分段为可变长度单元,以及(2)新颖的隐式持续时间编码方案,将内容和时间跨度嵌入单个令牌索引,无需辅助持续时间预测器。大量实验表明,VARSTok显著优于强固定速率基线。值得注意的是,它实现了卓越的重建自然度,同时使用的令牌比40 Hz固定帧速率基线少23%。VARSTok进一步降低了字错误率,并提高了zero-shot文本到语音合成的自然度。据我们所知,这是第一个证明完全动态,可变帧速率声学语音标记器可以无缝集成到下游语音语言模型的工作。语音样本可在https://zhengrachel.github.io/VARSTok上获得。
摘要:Existing speech tokenizers typically assign a fixed number of tokens per second, regardless of the varying information density or temporal fluctuations in the speech signal. This uniform token allocation mismatches the intrinsic structure of speech, where information is distributed unevenly over time. To address this, we propose VARSTok, a VAriable-frame-Rate Speech Tokenizer that adapts token allocation based on local feature similarity. VARSTok introduces two key innovations: (1) a temporal-aware density peak clustering algorithm that adaptively segments speech into variable-length units, and (2) a novel implicit duration coding scheme that embeds both content and temporal span into a single token index, eliminating the need for auxiliary duration predictors. Extensive experiments show that VARSTok significantly outperforms strong fixed-rate baselines. Notably, it achieves superior reconstruction naturalness while using up to 23% fewer tokens than a 40 Hz fixed-frame-rate baseline. VARSTok further yields lower word error rates and improved naturalness in zero-shot text-to-speech synthesis. To the best of our knowledge, this is the first work to demonstrate that a fully dynamic, variable-frame-rate acoustic speech tokenizer can be seamlessly integrated into downstream speech language models. Speech samples are available at https://zhengrachel.github.io/VARSTok.


【6】DarkStream: real-time speech anonymization with low latency
标题:DarkStream:低延迟的实时语音匿名化
链接:https://arxiv.org/abs/2509.04667

作者:Waris Quamer, Ricardo Gutierrez-Osuna
备注:Accepted for presentation at ASRU 2025
摘要:我们提出了DarkStream,一个用于实时说话人匿名化的流式语音合成模型。为了在严格的延迟约束下改进内容编码,DarkStream结合了因果波形编码器,短前瞻缓冲区和基于变换器的上下文层。为了进一步减少推理时间,该模型直接通过神经声码器生成波形,从而消除中间梅尔频谱图转换。最后,DarkStream通过将GAN生成的伪说话者嵌入到来自内容编码器的语言特征中来匿名化说话者身份。评估表明,我们的模型实现了强大的匿名化,在懒惰的通知攻击场景中产生接近50%的说话人验证EER(近机会性能),同时保持可接受的语言可理解性(WER在9%以内)。通过平衡低延迟,强大的隐私和最小的可理解性退化,DarkStream为保护隐私的实时语音通信提供了一个实用的解决方案。
摘要:We propose DarkStream, a streaming speech synthesis model for real-time speaker anonymization. To improve content encoding under strict latency constraints, DarkStream combines a causal waveform encoder, a short lookahead buffer, and transformer-based contextual layers. To further reduce inference time, the model generates waveforms directly via a neural vocoder, thus removing intermediate mel-spectrogram conversions. Finally, DarkStream anonymizes speaker identity by injecting a GAN-generated pseudo-speaker embedding into linguistic features from the content encoder. Evaluations show our model achieves strong anonymization, yielding close to 50% speaker verification EER (near-chance performance) on the lazy-informed attack scenario, while maintaining acceptable linguistic intelligibility (WER within 9%). By balancing low-latency, robust privacy, and minimal intelligibility degradation, DarkStream provides a practical solution for privacy-preserving real-time speech communication.


【7】On Time Delay Interpolation for Improved Acoustic Reflector Localization
标题:改进声反射器定位的延时插值
链接:https://arxiv.org/abs/2509.04629

作者:Hannes Rosseel, Toon van Waterschoot
备注:20 pages, 13 figures, 2 tables, submitted to J. Acoust. Soc. Am
摘要:声反射体的定位是室内声学分析、声源定位和声场景分析等应用中的一个基本组成部分。时间延迟估计(TDE)是确定反射器相对于传感器阵列的位置的关键。传统的TDE算法通常产生的时间延迟是操作采样周期的整数倍,可能缺乏足够的时间分辨率。为了实现子采样TDE精度,已经提出了各种插值方法,包括抛物线、高斯、频率和sinc插值。本文提出了一个全面的研究时间延迟插值,以实现混响条件下的声反射定位的子样本精度。我们推导出惠特克-香农插值公式,从先前提出的sinc插值的上下文中的短时间加窗时域估计声反射定位。仿真表明,sinc和Whittaker-Shannon插值在临界采样和带限反射的时延误差和位置误差方面优于现有方法。在MYRiAD数据集的真实测量结果上评估了性能,结果表明sinc和Whittaker-Shannon插值在不同的传感器-源对和扬声器位置上始终提供可靠的性能。这些结果可以提高声反射器定位系统的精度,对于诸如室内声学分析、声源定位和声学场景分析等应用至关重要。
摘要:The localization of acoustic reflectors is a fundamental component in various applications, including room acoustics analysis, sound source localization, and acoustic scene analysis. Time Delay Estimation (TDE) is essential for determining the position of reflectors relative to a sensor array. Traditional TDE algorithms generally yield time delays that are integer multiples of the operating sampling period, potentially lacking sufficient time resolution. To achieve subsample TDE accuracy, various interpolation methods, including parabolic, Gaussian, frequency, and sinc interpolation, have been proposed. This paper presents a comprehensive study on time delay interpolation to achieve subsample accuracy for acoustic reflector localization in reverberant conditions. We derive the Whittaker-Shannon interpolation formula from the previously proposed sinc interpolation in the context of short-time windowed TDE for acoustic reflector localization. Simulations show that sinc and Whittaker-Shannon interpolation outperform existing methods in terms of time delay error and positional error for critically sampled and band-limited reflections. Performance is evaluated on real-world measurements from the MYRiAD dataset, showing that sinc and Whittaker-Shannon interpolation consistently provide reliable performance across different sensor-source pairs and loudspeaker positions. These results can enhance the precision of acoustic reflector localization systems, vital for applications such as room acoustics analysis, sound source localization, and acoustic scene analysis.


【8】Recomposer: Event-roll-guided generative audio editing
标题:重新组合器:事件滚动引导生成音频编辑
链接:https://arxiv.org/abs/2509.05256

作者:Daniel P. W. Ellis, Eduardo Fonseca, Ron J. Weiss, Kevin Wilson, Scott Wisdom, Hakan Erdogan, John R. Hershey, Aren Jansen, R. Channing Moore, Manoj Plakal
备注:5 pages, 5 figures
摘要:编辑复杂的真实世界声音场景是困难的,因为各个声源在时间上重叠。生成模型可以根据其对数据域的强大先验理解来填充缺失或损坏的细节。我们提出了一种用于在复杂场景内编辑各个声音事件的系统,该系统能够基于文本编辑描述(例如,"增强门“)和从"事件卷”转录导出的事件定时的图形表示。我们提出了一个编码器-解码器Transformer,它工作在SoundStream表示上,在合成(输入,所需输出)音频示例对上进行训练,这些音频示例对是通过将孤立的声音事件添加到密集的真实世界背景中而形成的。评估揭示了编辑描述的每个部分的重要性--动作、类、时间。我们的工作表明“重组”是一个重要而实际的应用。
摘要:Editing complex real-world sound scenes is difficult because individual sound sources overlap in time. Generative models can fill-in missing or corrupted details based on their strong prior understanding of the data domain. We present a system for editing individual sound events within complex scenes able to delete, insert, and enhance individual sound events based on textual edit descriptions (e.g., ``enhance Door'') and a graphical representation of the event timing derived from an ``event roll'' transcription. We present an encoder-decoder transformer working on SoundStream representations, trained on synthetic (input, desired output) audio example pairs formed by adding isolated sound events to dense, real-world backgrounds. Evaluation reveals the importance of each part of the edit descriptions -- action, class, timing. Our work demonstrates ``recomposition'' is an important and practical application.


【9】Exploring Situated Stabilities of a Rhythm Generation System through Variational Cross-Examination
标题:通过变分盘问探索节奏生成系统的情境稳定性
链接:https://arxiv.org/abs/2509.05145

作者:Błazej Kotowski, Nicholas Evans, Behzad Haki, Frederic Font, Sergi Jordà
备注:AI Music Creativity 2025
摘要:本文通过变分交叉检验(VCE)的后现象学框架,研究了一个实时节奏生成系统Groupy Transformer。通过反思其部署在三个不同的艺术背景下,我们确定了三个稳定性:一个自主的鼓伴奏发生器,节奏控制电压音序器在Eurorack格式,和一个和谐的伴奏系统的节奏驱动程序。其应用的多功能性从项目一开始就不是一个明确的目标。因此,我们要问:这种多重稳定性是如何出现的?通过VCE,我们确定了其出现的三个关键因素:系统不变量的启示,跨学科的合作,以及其发展的定位性质。最后,我们反思VCE的可行性作为一种描述和分析方法,数字乐器(MIDI)设计,强调其价值,揭示技术如何调解,共同塑造,并共同塑造用户和上下文。
摘要:This paper investigates GrooveTransformer, a real-time rhythm generation system, through the postphenomenological framework of Variational Cross-Examination (VCE). By reflecting on its deployment across three distinct artistic contexts, we identify three stabilities: an autonomous drum accompaniment generator, a rhythmic control voltage sequencer in Eurorack format, and a rhythm driver for a harmonic accompaniment system. The versatility of its applications was not an explicit goal from the outset of the project. Thus, we ask: how did this multistability emerge? Through VCE, we identify three key contributors to its emergence: the affordances of system invariants, the interdisciplinary collaboration, and the situated nature of its development. We conclude by reflecting on the viability of VCE as a descriptive and analytical method for Digital Musical Instrument (DMI) design, emphasizing its value in uncovering how technologies mediate, co-shape, and are co-shaped by users and contexts.


【10】Training a Perceptual Model for Evaluating Auditory Similarity in Music Adversarial Attack
标题:训练用于评估音乐对抗性攻击中听觉相似性的感知模型
链接:https://arxiv.org/abs/2509.04985

作者:Yuxuan Liu, Rui Sang, Peihong Zhang, Zhixin Li, Shengchen Li
摘要:音乐信息检索(MIR)系统非常容易受到对抗性攻击,这些攻击通常是人类无法感知的,主要是由于模型特征空间和人类听觉感知之间的不一致。现有的防御措施和感知指标经常无法充分捕捉这些听觉细微差别,我们最初的听力测试支持这一限制,显示通用指标与人类判断之间的相关性较低。为了弥合这一差距,我们引入了感知对齐的MERT Transformer(PAMT),这是一种用于学习鲁棒的感知对齐音乐表示的新框架。我们的核心创新在于心理声学调节的顺序对比Transformer,这是一种建立在冻结MERT编码器上的轻量级投影头。PAMT与主观评分的斯皮尔曼相关系数为0.65,优于现有的感知指标。我们的方法在具有挑战性的MIR任务上的鲁棒准确性平均提高了9.15%,包括在各种感知对抗攻击下的封面歌曲识别和音乐流派分类。这项工作开创了架构集成的心理声学条件反射,产生的表示更符合人类的感知,并对音乐对抗性攻击具有鲁棒性。
摘要:Music Information Retrieval (MIR) systems are highly vulnerable to adversarial attacks that are often imperceptible to humans, primarily due to a misalignment between model feature spaces and human auditory perception. Existing defenses and perceptual metrics frequently fail to adequately capture these auditory nuances, a limitation supported by our initial listening tests showing low correlation between common metrics and human judgments. To bridge this gap, we introduce Perceptually-Aligned MERT Transformer (PAMT), a novel framework for learning robust, perceptually-aligned music representations. Our core innovation lies in the psychoacoustically-conditioned sequential contrastive transformer, a lightweight projection head built atop a frozen MERT encoder. PAMT achieves a Spearman correlation coefficient of 0.65 with subjective scores, outperforming existing perceptual metrics. Our approach also achieves an average of 9.15\% improvement in robust accuracy on challenging MIR tasks, including Cover Song Identification and Music Genre Classification, under diverse perceptual adversarial attacks. This work pioneers architecturally-integrated psychoacoustic conditioning, yielding representations significantly more aligned with human perception and robust against music adversarial attacks.


【11】MAIA: An Inpainting-Based Approach for Music Adversarial Attacks
标题:MAIA:一种基于修补的音乐对抗攻击方法
链接:https://arxiv.org/abs/2509.04980

作者:Yuxuan Liu, Peihong Zhang, Rui Sang, Zhixin Li, Shengchen Li
备注:Accepted at ISMIR2025
摘要:音乐对抗攻击在音乐信息检索(MIR)领域引起了极大的兴趣。在本文中,我们提出了音乐对抗性修复攻击(MAIA),一种新的对抗性攻击框架,支持白盒和黑盒攻击场景。MAIA从重要性分析开始,以识别关键的音频片段,然后将其作为修改的目标。利用生成修复模型,这些片段在受攻击模型输出的指导下进行重建,确保微妙而有效的对抗性扰动。我们在多个MIR任务上评估MAIA,在白盒和黑盒设置中展示了高攻击成功率,同时保持最小的感知失真。此外,主观听力测试证实了对抗样本的高音频保真度。我们的研究结果突出了当前MIR系统的漏洞,并强调需要更强大和安全的模型。
摘要:Music adversarial attacks have garnered significant interest in the field of Music Information Retrieval (MIR). In this paper, we present Music Adversarial Inpainting Attack (MAIA), a novel adversarial attack framework that supports both white-box and black-box attack scenarios. MAIA begins with an importance analysis to identify critical audio segments, which are then targeted for modification. Utilizing generative inpainting models, these segments are reconstructed with guidance from the output of the attacked model, ensuring subtle and effective adversarial perturbations. We evaluate MAIA on multiple MIR tasks, demonstrating high attack success rates in both white-box and black-box settings while maintaining minimal perceptual distortion. Additionally, subjective listening tests confirm the high audio fidelity of the adversarial samples. Our findings highlight vulnerabilities in current MIR systems and emphasize the need for more robust and secure models.


【12】Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper
标题:通过多个基础模型映射器高效生成视频到音频
链接:https://arxiv.org/abs/2509.04957

作者:Gehui Chen, Guan'an Wang, Xiaowen Huang, Jitao Sang
摘要:最近的视频到音频(V2 A)生成依赖于从视频中提取语义和时间特征到条件生成模型。从头开始训练这些模型是资源密集型的。因此,利用基础模型(FM)由于其跨模态知识转移和泛化能力而获得了吸引力。之前的一项工作探索了微调轻量级映射器网络,以将预训练的视觉编码器与V2 A的文本到音频生成模型连接起来。受此启发,我们引入了多基础模型映射器(MFM-Mapper)。与之前的映射器方法相比,MFM-Mapper通过融合来自双视觉编码器的特征,受益于更丰富的语义和时间信息。此外,通过用GPT-2替换线性映射器,MFM-Mapper改进了特征对齐,在跨模态特征映射和自回归翻译任务之间绘制了平行线。我们的MFM映射器具有显着的训练效率。它在语义和时间一致性方面取得了更好的性能,训练消耗更少,与以前的基于映射器的工作相比,只需要16%的训练规模,但在更大的规模上训练的模型却取得了有竞争力的性能。
摘要:Recent Video-to-Audio (V2A) generation relies on extracting semantic and temporal features from video to condition generative models. Training these models from scratch is resource intensive. Consequently, leveraging foundation models (FMs) has gained traction due to their cross-modal knowledge transfer and generalization capabilities. One prior work has explored fine-tuning a lightweight mapper network to connect a pre-trained visual encoder with a text-to-audio generation model for V2A. Inspired by this, we introduce the Multiple Foundation Model Mapper (MFM-Mapper). Compared to the previous mapper approach, MFM-Mapper benefits from richer semantic and temporal information by fusing features from dual visual encoders. Furthermore, by replacing a linear mapper with GPT-2, MFM-Mapper improves feature alignment, drawing parallels between cross-modal features mapping and autoregressive translation tasks. Our MFM-Mapper exhibits remarkable training efficiency. It achieves better performance in semantic and temporal consistency with fewer training consuming, requiring only 16\% of the training scale compared to previous mapper-based work, yet achieves competitive performance with models trained on a much larger scale.


【13】Learning and composing of classical music using restricted Boltzmann machines
标题:使用受限制的Boltzmann机器学习和创作古典音乐
链接:https://arxiv.org/abs/2509.04899

作者:Mutsumi Kobayashi, Hiroshi Watanabe
备注:19 pages, 10 figures
摘要:最近,已经开发出软件,利用机器学习来模仿特定作曲家的风格,例如J. S。巴赫.但是,由于这类软件往往采用结构复杂的机器学习模型,因此很难分析软件如何理解作曲家音乐的特征。在这项研究中,我们采用了J。巴赫的音乐用于限制玻尔兹曼机(RBM)的训练。由于RBM的结构简单,它允许我们研究学习后的内部状态。我们发现,学习RBM是能够创作音乐。
摘要:Recently, software has been developed that uses machine learning to mimic the style of a particular composer, such as J. S. Bach. However, since such software often adopts machine learning models with complex structures, it is difficult to analyze how the software understands the characteristics of the composer's music. In this study, we adopted J. S. Bach's music for training of a restricted Boltzmann machine (RBM). Since the structure of RBMs is simple, it allows us to investigate the internal states after learning. We found that the learned RBM is able to compose music.


【14】Quantum Fourier Transform Based Denoising: Unitary Filtering for Enhanced Speech Clarity
标题:基于量子傅里叶变换的去噪:用于增强语音清晰度的元化过滤
链接:https://arxiv.org/abs/2509.04851

作者:Rajeshwar Tripathi, Sahil Tomar, Sandeep Kumar, Monika Aggarwal
备注:8 pages
摘要:本文介绍了一种量子启发的去噪框架,该框架将量子傅立叶变换(QFT)集成到经典的音频增强管道中。与传统的基于快速傅立叶变换(FFT)的方法不同,QFT提供了具有全局相位相干性和能量保存的酉变换,使得能够改进语音和噪声之间的区分。所提出的方法用QFT算子取代Wiener和谱减法滤波器中的FFT,确保一致的超参数设置以进行公平比较。在不同信噪比(SNR)条件下对干净语音、合成音调和嘈杂混合物进行的实验表明,SNR在统计上有显著的提高,最高可达15 dB的改善,并减少了伪影的产生。结果证实,基于QFT的去噪在低SNR和非平稳噪声场景下提供了鲁棒性,而无需额外的计算开销,突出了其作为量子增强语音处理的可扩展途径的潜力。
摘要:This paper introduces a quantum-inspired denoising framework that integrates the Quantum Fourier Transform (QFT) into classical audio enhancement pipelines. Unlike conventional Fast Fourier Transform (FFT) based methods, QFT provides a unitary transformation with global phase coherence and energy preservation, enabling improved discrimination between speech and noise. The proposed approach replaces FFT in Wiener and spectral subtraction filters with a QFT operator, ensuring consistent hyperparameter settings for fair comparison. Experiments on clean speech, synthetic tones, and noisy mixtures across diverse signal to noise ratio (SNR) conditions, demonstrate statistically significant gains in SNR, with up to 15 dB improvement and reduced artifact generation. Results confirm that QFT based denoising offers robustness under low SNR and nonstationary noise scenarios without additional computational overhead, highlighting its potential as a scalable pathway toward quantum-enhanced speech processing.


【15】WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
标题:WildScore:狂野符号音乐推理中的MLLM基准
链接:https://arxiv.org/abs/2509.04744

作者:Gagan Mundada, Yash Vishe, Amit Namburi, Xin Xu, Zachary Novack, Julian McAuley, Junda Wu
摘要:多模态大型语言模型(MLLM)的最新进展已经在各种视觉语言任务中展示了令人印象深刻的能力。然而,他们在多模态符号音乐领域的推理能力仍然在很大程度上未被探索。我们介绍WildScore,第一个在野外多模态符号音乐推理和分析基准,旨在评估MLLM的能力,解释现实世界的乐谱和回答复杂的音乐学查询。WildScore中的每个实例都来自真实的音乐作品,并伴随着真实的用户生成的问题和讨论,捕捉实际音乐分析的复杂性。为了便于系统的评估,我们提出了一个系统的分类,包括高层次和细粒度的音乐本体。此外,我们框架复杂的音乐推理作为多项选择题的回答,使控制和可扩展的评估MLLM的符号音乐的理解。在WildScore上对最先进的MLLM进行的经验基准测试揭示了其视觉符号推理中的有趣模式,揭示了MLLM在符号音乐推理和分析中的有前途的方向和持续挑战。我们发布数据集和代码。
摘要:Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, their reasoning abilities in the multimodal symbolic music domain remain largely unexplored. We introduce WildScore, the first in-the-wild multimodal symbolic music reasoning and analysis benchmark, designed to evaluate MLLMs' capacity to interpret real-world music scores and answer complex musicological queries. Each instance in WildScore is sourced from genuine musical compositions and accompanied by authentic user-generated questions and discussions, capturing the intricacies of practical music analysis. To facilitate systematic evaluation, we propose a systematic taxonomy, comprising both high-level and fine-grained musicological ontologies. Furthermore, we frame complex music reasoning as multiple-choice question answering, enabling controlled and scalable assessment of MLLMs' symbolic music understanding. Empirical benchmarking of state-of-the-art MLLMs on WildScore reveals intriguing patterns in their visual-symbolic reasoning, uncovering both promising directions and persistent challenges for MLLMs in symbolic music reasoning and analysis. We release the dataset and code.


【16】A Multiclass Acoustic Dataset and Interactive Tool for Analyzing Drone Signatures in Real-World Environments
标题:用于分析现实环境中无人机签名的多类声学数据集和交互工具
链接:https://arxiv.org/abs/2509.04715

作者:Mia Y. Wang, Mackenzie Linn, Andrew P. Berg, Qian Zhang
备注:This article extends our previous work presented in the 2024 Artificial Intelligence x Humanities, Education, and Art (2024 AIxHeart) Conference
摘要:无人机在各个行业的迅速普及带来了与隐私、安全和噪音污染相关的重大挑战。目前的无人机探测系统主要基于视觉和雷达技术,在某些条件下面临局限性,这突出表明需要有效的基于声学的探测方法。本文提供了一个独特而全面的无人机声学特征数据集,包括按品牌和型号区分的32个不同类别。该数据集包括每架无人机的原始音频记录、声谱图和梅尔频率倒谱系数(MFCC)图。此外,我们还介绍了一个交互式Web应用程序,允许用户通过选择特定的无人机类别、收听相关音频以及查看相应的频谱图和MFCC图来探索此数据集。该工具旨在促进无人机检测,分类和声学分析的研究,支持技术进步和教育计划。本文详细介绍了数据集的创建过程,Web应用程序的设计和实现,并提供了实验结果和用户反馈。最后,我们讨论了潜在的应用和未来的工作,以扩大和加强该项目。
摘要:The rapid proliferation of drones across various industries has introduced significant challenges related to privacy, security, and noise pollution. Current drone detection systems, primarily based on visual and radar technologies, face limitations under certain conditions, highlighting the need for effective acoustic-based detection methods. This paper presents a unique and comprehensive dataset of drone acoustic signatures, encompassing 32 different categories differentiated by brand and model. The dataset includes raw audio recordings, spectrogram plots, and Mel-frequency cepstral coefficient (MFCC) plots for each drone. Additionally, we introduce an interactive web application that allows users to explore this dataset by selecting specific drone categories, listening to the associated audio, and viewing the corresponding spectrogram and MFCC plots. This tool aims to facilitate research in drone detection, classification, and acoustic analysis, supporting both technological advancements and educational initiatives. The paper details the dataset creation process, the design and implementation of the web application, and provides experimental results and user feedback. Finally, we discuss potential applications and future work to expand and enhance the project.


【17】Ecologically Valid Benchmarking and Adaptive Attention: Scalable Marine Bioacoustic Monitoring
标题:公共有效的基准和适应性注意力:可扩展的海洋生物声学监测
链接:https://arxiv.org/abs/2509.04682

作者:Nicholas R. Rasmussen, Rodrigue Rizk, Longwei Wang, KC Santosh
备注:Under review as an anonymous submission to IEEETAI - We are allowed an archive submission. Final formatting is yet to be determined
摘要:水下被动声监测(UPAM)为长期生态分析提供了丰富的时空数据,但固有的噪声和复杂的信号依赖性阻碍了模型的稳定性和推广。多层窗口改进了目标声音定位,但变化的环境噪声、不同的传播效应以及混合的生物和人为来源的可变性需要强大的架构和严格的评估。我们介绍GetNetUPAM,一个分层嵌套的交叉验证框架,旨在量化生态现实的变化下的模型稳定性。数据被划分为不同的站点年段,保留记录异质性,并确保每个验证折叠反映了一个独特的环境子集,减少局部噪声和传感器伪影的过拟合。站点年阻塞强制执行对真正的环境多样性的评估,而标准的交叉验证随机子集的措施,在UPAM的完整的信号分布,一个维度缺席目前的基准泛化。使用GetNetUPAM作为评估骨干,我们提出了自适应分辨率池和注意力网络(ARPA-N),这是一种用于不规则频谱图维度的神经架构。具有空间注意力的自适应池扩展了感受野,在没有过多参数的情况下捕获全局上下文。在GetNetUPAM下,ARPA-N在DenseNet基线上实现了14.4%的平均精度增益,并且在所有指标上的变异性都下降了log 2量级,从而实现了跨站点年折叠的一致检测,并推进了可扩展的准确生物声学监测。
摘要:Underwater Passive Acoustic Monitoring (UPAM) provides rich spatiotemporal data for long-term ecological analysis, but intrinsic noise and complex signal dependencies hinder model stability and generalization. Multilayered windowing has improved target sound localization, yet variability from shifting ambient noise, diverse propagation effects, and mixed biological and anthropogenic sources demands robust architectures and rigorous evaluation. We introduce GetNetUPAM, a hierarchical nested cross-validation framework designed to quantify model stability under ecologically realistic variability. Data are partitioned into distinct site-year segments, preserving recording heterogeneity and ensuring each validation fold reflects a unique environmental subset, reducing overfitting to localized noise and sensor artifacts. Site-year blocking enforces evaluation against genuine environmental diversity, while standard cross-validation on random subsets measures generalization across UPAM's full signal distribution, a dimension absent from current benchmarks. Using GetNetUPAM as the evaluation backbone, we propose the Adaptive Resolution Pooling and Attention Network (ARPA-N), a neural architecture for irregular spectrogram dimensions. Adaptive pooling with spatial attention extends the receptive field, capturing global context without excessive parameters. Under GetNetUPAM, ARPA-N achieves a 14.4% gain in average precision over DenseNet baselines and a log2-scale order-of-magnitude drop in variability across all metrics, enabling consistent detection across site-year folds and advancing scalable, accurate bioacoustic monitoring.


【18】Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
标题:基于大语言模型的多说话者语音识别的序列化输出预处理
链接:https://arxiv.org/abs/2509.04488

作者:Hao Shi, Yusuke Fujita, Tomoya Mizumoto, Lianbo Liu, Atsushi Kojima, Yui Sudo
摘要:摘要对于任务定义和提高基于大型语言模型(LLM)的系统的性能至关重要。然而,现有的基于LLM的多说话者(MT)自动语音识别(ASR)系统要么省略提示,要么依赖于简单的任务定义提示,没有先前的工作探索提示的设计,以提高性能。在本文中,我们提出了提取序列化的输出提示(SOP),并明确指导LLM使用结构化提示,以提高系统性能(SOP-MT-ASR)。在语音编码器之后插入分离器和序列化的连接主义时间分类(CTC)层,以按照先说先出的方式从混合语音编码中分离和提取MT内容。随后,SOP,这是作为一个提示LLM,通过使用贪婪搜索解码的串行CTC输出。为了有效地训练模型,我们设计了一个三阶段的训练策略,包括序列化输出训练(SOT)微调,序列化语音信息提取,和基于SOP的适应。LibriMix数据集上的实验结果表明,尽管基于LLM的SOT模型在两个谈话者场景中表现良好,但在更复杂的条件下,如三个谈话者场景,它无法充分利用LLM。所提出的SOP方法显着提高了性能,在两个和三个谈话者的条件下。
摘要:Prompts are crucial for task definition and for improving the performance of large language models (LLM)-based systems. However, existing LLM-based multi-talker (MT) automatic speech recognition (ASR) systems either omit prompts or rely on simple task-definition prompts, with no prior work exploring the design of prompts to enhance performance. In this paper, we propose extracting serialized output prompts (SOP) and explicitly guiding the LLM using structured prompts to improve system performance (SOP-MT-ASR). A Separator and serialized Connectionist Temporal Classification (CTC) layers are inserted after the speech encoder to separate and extract MT content from the mixed speech encoding in a first-speaking-first-out manner. Subsequently, the SOP, which serves as a prompt for LLMs, is obtained by decoding the serialized CTC outputs using greedy search. To train the model effectively, we design a three-stage training strategy, consisting of serialized output training (SOT) fine-tuning, serialized speech information extraction, and SOP-based adaptation. Experimental results on the LibriMix dataset show that, although the LLM-based SOT model performs well in the two-talker scenario, it fails to fully leverage LLMs under more complex conditions, such as the three-talker scenario. The proposed SOP approach significantly improved performance under both two- and three-talker conditions.


机器翻译由腾讯交互翻译提供,仅供参考