微信公众号:arXiv_Daily
cs.SD语音
标题: Stream-Omni:采用大范围-视觉-语音模型的同时多模式交互
链接:https://arxiv.org/abs/2506.13642
备注:Code: this https URL , Model: this https URL
摘要:GPT-40类大型多模态模型(LMRM)的出现引发了对整合文本、视觉和语音模态以支持更灵活的多模态交互的探索。现有的LLM通常沿着序列维度连接模态的表示,并将它们馈送到大型语言模型(LLM)主干中。虽然序列-维度连接对于模态集成是直接的,但它通常严重依赖于大规模数据来学习模态对齐。在本文中,我们的目标是更有目的地建模模态之间的关系,从而实现更有效和灵活的模态对齐。为此,我们提出了Stream-Omni,一个大型的语言-视觉-语音模型,具有高效的模态对齐,可以同时支持各种模态组合下的交互。Stream-Omni采用LLM作为骨干,并根据其关系将视觉和语音与文本对齐。对于与文本在语义上互补的视觉,Stream-Omni使用序列维度连接来实现视觉与文本对齐。对于与文本语义一致的语音,Stream-Omni引入了基于CTC的层-维度映射来实现语音-文本对齐。通过这种方式,Stream-Omni可以用更少的数据(特别是语音)实现模态对齐,从而将文本功能转移到其他模态。在各种基准测试上的实验表明,Stream-Omni在视觉理解、语音交互和基于视觉的语音交互任务上都取得了很好的性能。由于层维映射,Stream-Omni可以在语音交互过程中同时提供中间文本输出(如ASR transmittance和模型响应),为用户提供全面的多模态体验。
摘要:The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal interaction. Existing LMMs typically concatenate representation of modalities along the sequence dimension and feed them into a large language model (LLM) backbone. While sequence-dimension concatenation is straightforward for modality integration, it often relies heavily on large-scale data to learn modality alignments. In this paper, we aim to model the relationships between modalities more purposefully, thereby achieving more efficient and flexible modality alignments. To this end, we propose Stream-Omni, a large language-vision-speech model with efficient modality alignments, which can simultaneously support interactions under various modality combinations. Stream-Omni employs LLM as the backbone and aligns the vision and speech to the text based on their relationships. For vision that is semantically complementary to text, Stream-Omni uses sequence-dimension concatenation to achieve vision-text alignment. For speech that is semantically consistent with text, Stream-Omni introduces a CTC-based layer-dimension mapping to achieve speech-text alignment. In this way, Stream-Omni can achieve modality alignments with less data (especially speech), enabling the transfer of text capabilities to other modalities. Experiments on various benchmarks demonstrate that Stream-Omni achieves strong performance on visual understanding, speech interaction, and vision-grounded speech interaction tasks. Owing to the layer-dimensional mapping, Stream-Omni can simultaneously provide intermediate text outputs (such as ASR transcriptions and model responses) during speech interaction, offering users a comprehensive multimodal experience.
标题: Qwen与Gemma与Whisper的集成:多语言演讲LLM系统的比较研究
链接:https://arxiv.org/abs/2506.13596
备注:Technical report for Interspeech 2025 MLC-SLM Challenge
摘要:本文介绍了我们的MLC-SLM挑战2025系统,重点是多语言语音识别和大型语言模型(LLM)的语言建模。我们的方法结合了一个微调的耳语大v3编码器与高效的投影机架构和各种解码器配置。我们采用三阶段的培训方法,逐步优化编码器,投影仪和LLM组件。我们的系统实现了具有竞争力的性能,使用Gemma 3 - 12 B和18.6%使用Qwen2.5- 7 B作为仅解码器的语言模型的私人测试平均WER/CER结果为16.63%。
摘要:This paper presents our system for the MLC-SLM Challenge 2025, focusing on multilingual speech recognition and language modeling with large language models (LLMs). Our approach combines a fine-tuned Whisper-large-v3 encoder with efficient projector architectures and various decoder configurations. We employ a three-stage training methodology that progressively optimizes the encoder, projector, and LLM components. Our system achieves competitive performance with a private test average WER/CER result of 16.63% using the Gemma3-12B and 18.6% using the Qwen2.5-7B as decoder-only language model.
标题: 三种不同距离音乐网络的持久同源性
链接:https://arxiv.org/abs/2506.13595
摘要:持久同源性已被广泛用于发现跨各种应用的数据中隐藏的拓扑结构,包括音乐数据。要应用持久同源性,必须在点云中的点之间或图形网络中的节点之间定义距离或度量。这些定义不是唯一的,取决于给定问题的具体目标。换句话说,选择不同的度量定义允许多个拓扑推断。在这项工作中,我们专注于应用持久同源音乐图与预定义的权重。我们研究了三个不同的距离定义的基础上边明智的途径,并展示了这些定义如何影响持久性条形码,持久性图,出生/死亡边缘。我们发现这三种距离定义在一维持久同源性条码和持久同源性图上都存在包含关系。我们使用真实的音乐数据验证了这些发现。
摘要:Persistent homology has been widely used to discover hidden topological structures in data across various applications, including music data. To apply persistent homology, a distance or metric must be defined between points in a point cloud or between nodes in a graph network. These definitions are not unique and depend on the specific objectives of a given problem. In other words, selecting different metric definitions allows for multiple topological inferences. In this work, we focus on applying persistent homology to music graph with predefined weights. We examine three distinct distance definitions based on edge-wise pathways and demonstrate how these definitions affect persistent barcodes, persistence diagrams, and birth/death edges. We found that there exist inclusion relations in one-dimensional persistent homology reflected on persistence barcode and diagram among these three distance definitions. We verified these findings using real music data.
标题: Seewo提交给MLC-LAM:从语音推理语言模型中吸取的教训
链接:https://arxiv.org/abs/2506.13300
摘要:本文介绍了Seewo的系统,用于多语言对话语音语言模型挑战赛(MLC-SLM)的两个轨道,解决自动语音识别(ASR)和说话人日志化与ASR(SD-ASR)。我们引入了一个多阶段的训练管道,显式地增强了ASR语音语言模型中的推理和自我校正。我们的方法结合了课程学习以获得渐进的能力,思想链数据增强以促进中间反思,以及强化学习与可验证奖励(RLVR),以通过奖励驱动的优化进一步完善自我纠正。这种方法在官方挑战基线上实现了实质性改进。在评估集中,我们最好的系统在Track 1中获得了11.57%的WER/CER,在Track 2中获得了17.67%的tcpWER/tcpCER。全面的消融研究证明了每个组件在挑战限制条件下的有效性。
摘要:This paper presents Seewo's systems for both tracks of the Multilingual Conversational Speech Language Model Challenge (MLC-SLM), addressing automatic speech recognition (ASR) and speaker diarization with ASR (SD-ASR). We introduce a multi-stage training pipeline that explicitly enhances reasoning and self-correction in speech language models for ASR. Our approach combines curriculum learning for progressive capability acquisition, Chain-of-Thought data augmentation to foster intermediate reflection, and Reinforcement Learning with Verifiable Rewards (RLVR) to further refine self-correction through reward-driven optimization. This approach achieves substantial improvements over the official challenge baselines. On the evaluation set, our best system attains a WER/CER of 11.57% for Track 1 and a tcpWER/tcpCER of 17.67% for Track 2. Comprehensive ablation studies demonstrate the effectiveness of each component under challenge constraints.
标题: SONIC:针对人群噪音的声音优化
链接:https://arxiv.org/abs/2506.13272
摘要:介绍了一种基于ARM Cortex-M7的STM32 H753 ZI微控制器实现的嵌入式实时噪声抑制系统SONIC。使用自适应滤波(LMS),该系统提高了在嘈杂的环境中的语音清晰度。SONIC专注于音频信号中噪声抑制的新方法,特别是解决传统有源噪声消除(ANC)系统的局限性。本文探讨了各种信号处理算法在微控制器的角度来看,突出各种性能因素,并被认为是最佳的嵌入式系统。此外,我们还讨论了系统架构,解释了如何利用MCU的效率,以及如何在处理器内转换音频信号的深入概述。结果表明,提高了语音清晰度和实际的实时性能,显示低功耗DSP作为复杂AI去噪方法的替代方案。
摘要:This paper presents SONIC, an embedded real-time noise suppression system implemented on the ARM Cortex-M7-based STM32H753ZI microcontroller. Using adaptive filtering (LMS), the system improves speech intelligibility in noisy environments. SONIC focuses on a novel approach to noise suppression in audio signals, specifically addressing the limitations of traditional Active Noise Cancellation (ANC) systems. The paper explores various signal processing algorithms in a micro-controller point of view, highlighting various performance factors and which were considered optimal in our embedded system. Additionally we also discussed the system architecture, explaining how the MCU's efficiency was harnessed, along with an in-depth overview of how the audio signals were translated within the processor. The results demonstrate improved speech clarity and practical real-time performance, showing low-power DSP as an alternative to complex AI denoising methods.
标题: 音乐偏好反映文化价值观吗?使用音乐嵌入和世界价值观调查的跨国分析
链接:https://arxiv.org/abs/2506.13199
摘要:这项研究探讨了民族音乐偏好在多大程度上反映了潜在的文化价值观。我们从62个国家的YouTube音乐排行榜中收集了长期流行音乐数据,包括西方和非西方地区,并使用CLAP模型提取音频嵌入。为了补充这些定量表示,我们使用LP-MusicCaps和基于GPT的摘要为每个曲目生成语义字幕。各国根据突出偏离全球音乐规范的对比嵌入进行聚类。通过t-SNE将产生的集群投影到二维空间中进行可视化,并根据世界价值观调查(WVS)定义的文化区进行评估。统计分析,包括MANOVA和卡方检验,证实了基于音乐的集群表现出显着的一致性与既定的文化群体。此外,残差分析揭示了一致的模式,代表性过高,这表明特定集群和文化区之间的非随机关联。这些发现表明,国家层面的音乐偏好编码有意义的文化信号,可以作为理解全球文化边界的代理。
摘要:This study explores the extent to which national music preferences reflect underlying cultural values. We collected long-term popular music data from YouTube Music Charts across 62 countries, encompassing both Western and non-Western regions, and extracted audio embeddings using the CLAP model. To complement these quantitative representations, we generated semantic captions for each track using LP-MusicCaps and GPT-based summarization. Countries were clustered based on contrastive embeddings that highlight deviations from global musical norms. The resulting clusters were projected into a two-dimensional space via t-SNE for visualization and evaluated against cultural zones defined by the World Values Survey (WVS). Statistical analyses, including MANOVA and chi-squared tests, confirmed that music-based clusters exhibit significant alignment with established cultural groupings. Furthermore, residual analysis revealed consistent patterns of overrepresentation, suggesting non-random associations between specific clusters and cultural zones. These findings indicate that national-level music preferences encode meaningful cultural signals and can serve as a proxy for understanding global cultural boundaries.
标题: I$#2$S-TFCKD:具有语音增强时频校准的内-间集知识提取
链接:https://arxiv.org/abs/2506.13127
备注:submitted to IEEE Transactions on Neural Networks and Learning Systems
摘要:近年来,基于神经网络(NN)的语音增强(SE)模型的复杂度压缩问题逐渐引起研究者的关注,特别是在硬件资源有限或时延要求严格的场景下。主要的困难和挑战在于根据任务的特点在复杂性和性能之间取得平衡。在本文中,我们提出了一个内集知识提取(KD)框架与时频校准(I$^2$S-TFCKD)的SE。与以往的提取策略不同,该框架充分利用了语音的时频差分信息,同时促进了全局知识流动。首先,提出了一种基于双流时频交叉校准的多层交互式蒸馏方法,该方法分别在时域和频域计算师生相似度校准权值并进行交叉加权,从而实现了根据语音特征在不同层间精细分配蒸馏贡献。其次,我们构建了一个集内和集间相关性的协同蒸馏范式。在一个相关的集合中,多层教师-学生特征被成对匹配以进行校准蒸馏。随后,我们产生的代表性特征,从每个相关的集合,通过残差融合,形成融合的特征集,使集合间的知识交互。提出的蒸馏策略应用于双路径扩张卷积递归网络(DPDCRN),该网络在L3 DAS 23挑战赛的SE赛道中排名第一。客观的评价表明,所提出的KD策略一致,有效地提高了低复杂度的学生模型的性能,并优于其他蒸馏计划。
摘要:In recent years, complexity compression of neural network (NN)-based speech enhancement (SE) models has gradually attracted the attention of researchers, especially in scenarios with limited hardware resources or strict latency requirements. The main difficulties and challenges lie in achieving a balance between complexity and performance according to the characteristics of the task. In this paper, we propose an intra-inter set knowledge distillation (KD) framework with time-frequency calibration (I$^2$S-TFCKD) for SE. Different from previous distillation strategies for SE, the proposed framework fully utilizes the time-frequency differential information of speech while promoting global knowledge flow. Firstly, we propose a multi-layer interactive distillation based on dual-stream time-frequency cross-calibration, which calculates the teacher-student similarity calibration weights in the time and frequency domains respectively and performs cross-weighting, thus enabling refined allocation of distillation contributions across different layers according to speech characteristics. Secondly, we construct a collaborative distillation paradigm for intra-set and inter-set correlations. Within a correlated set, multi-layer teacher-student features are pairwise matched for calibrated distillation. Subsequently, we generate representative features from each correlated set through residual fusion to form the fused feature set that enables inter-set knowledge interaction. The proposed distillation strategy is applied to the dual-path dilated convolutional recurrent network (DPDCRN) that ranked first in the SE track of the L3DAS23 challenge. Objective evaluations demonstrate that the proposed KD strategy consistently and effectively improves the performance of the low-complexity student model and outperforms other distillation schemes.
标题: 用MIDI-RWKV填充可个性化的长上下文象征性音乐
链接:https://arxiv.org/abs/2506.13001
摘要:自动音乐生成的现有工作主要集中在产生完整作品或延续的端到端系统上。然而,由于音乐创作通常是一个迭代的过程,这样的系统使得很难参与人与机器之间的来回,而这对计算机辅助创造力至关重要。在这项研究中,我们解决的任务个性化,多轨道,长的上下文,可控的符号音乐填充,以提高计算机辅助作曲的过程。我们提出了MIDI-RWKV,一种基于RWKV-7线性架构的新型模型,以实现边缘设备上高效和连贯的音乐共同创作。我们还表明,MIDI-RWKV承认微调其初始状态的个性化在非常低的样本制度的有效方法。我们评估MIDI-RWKV及其状态调整的几个定量和定性指标,并发布模型的权重和代码在https://github.com/christianazinn/MIDI-RWKV。
摘要:Existing work in automatic music generation has primarily focused on end-to-end systems that produce complete compositions or continuations. However, because musical composition is typically an iterative process, such systems make it difficult to engage in the back-and-forth between human and machine that is essential to computer-assisted creativity. In this study, we address the task of personalizable, multi-track, long-context, and controllable symbolic music infilling to enhance the process of computer-assisted composition. We present MIDI-RWKV, a novel model based on the RWKV-7 linear architecture, to enable efficient and coherent musical cocreation on edge devices. We also demonstrate that MIDI-RWKV admits an effective method of finetuning its initial state for personalization in the very-low-sample regime. We evaluate MIDI-RWKV and its state tuning on several quantitative and qualitative metrics, and release model weights and code at https://github.com/christianazinn/MIDI-RWKV.
标题: SoundMind:音频语言模型的RL激励逻辑推理
链接:https://arxiv.org/abs/2506.12935
摘要:虽然大型语言模型已经显示出推理能力,但它们在音频模态中的应用,特别是在大型音频语言模型(ALM)中,仍然显着不发达。解决这一差距需要一种系统的方法,包括一个有能力的基础模型,高质量的面向推理的音频数据和有效的训练算法。在这项研究中,我们提出了一个全面的解决方案:我们引入了音频逻辑推理(ALR)数据集,由6,446个专门为复杂推理任务设计的文本音频注释样本组成。在此基础上,我们提出了SoundMind,一种基于规则的强化学习(RL)算法,旨在赋予ALM深度双峰推理能力。通过使用SoundMind在ALR数据集上训练Qwen2.5-Omni-7 B,我们的方法在音频逻辑推理方面实现了最先进的性能。这项工作突出了将高质量的以推理为中心的数据集与专业RL技术相结合的影响,推进了语言模型中听觉智能的前沿。我们的代码和建议的数据集可以在https://github.com/xid32/SoundMind上找到。
摘要:While large language models have shown reasoning capabilities, their application to the audio modality, particularly in large audio-language models (ALMs), remains significantly underdeveloped. Addressing this gap requires a systematic approach, involving a capable base model, high-quality reasoning-oriented audio data, and effective training algorithms. In this study, we present a comprehensive solution: we introduce the Audio Logical Reasoning (ALR) dataset, consisting of 6,446 text-audio annotated samples specifically designed for complex reasoning tasks. Building on this resource, we propose SoundMind, a rule-based reinforcement learning (RL) algorithm tailored to endow ALMs with deep bimodal reasoning abilities. By training Qwen2.5-Omni-7B on the ALR dataset using SoundMind, our approach achieves state-of-the-art performance in audio logical reasoning. This work highlights the impact of combining high-quality, reasoning-focused datasets with specialized RL techniques, advancing the frontier of auditory intelligence in language models. Our code and the proposed dataset are available at https://github.com/xid32/SoundMind.
标题: SC-SOT:根据数字化说话人信息调节解码器,以实现端到端重叠语音识别
链接:https://arxiv.org/abs/2506.12672
备注:Accepted by Interspeech 2025
摘要:我们提出了扬声器条件串行化输出训练(SC-SOT),这是一种增强的基于SOT的E2 E多说话者ASR训练。我们首先探讨SOT如何处理重叠语音,我们发现解码器执行隐式说话人分离。我们假设这种隐含的分离往往是不够的,由于模糊的声学线索重叠区域。为了解决这个问题,SC-SOT显式地根据说话者信息来调节解码器,提供关于“谁在什么时候说话”的详细信息。具体来说,我们通过合并以下内容来增强解码器:(1)扬声器嵌入,其允许模型关注目标扬声器的声学特性,以及(2)扬声器活动信息,其引导模型抑制非目标扬声器。说话人嵌入是从联合训练的E2 E说话人日记化模型导出的,从而减轻了对说话人登记的需要。实验结果表明,我们的条件化方法对重叠语音的有效性。
摘要:We propose Speaker-Conditioned Serialized Output Training (SC-SOT), an enhanced SOT-based training for E2E multi-talker ASR. We first probe how SOT handles overlapped speech, and we found the decoder performs implicit speaker separation. We hypothesize this implicit separation is often insufficient due to ambiguous acoustic cues in overlapping regions. To address this, SC-SOT explicitly conditions the decoder on speaker information, providing detailed information about "who spoke when". Specifically, we enhance the decoder by incorporating: (1) speaker embeddings, which allow the model to focus on the acoustic characteristics of the target speaker, and (2) speaker activity information, which guides the model to suppress non-target speakers. The speaker embeddings are derived from a jointly trained E2E speaker diarization model, mitigating the need for speaker enrollment. Experimental results demonstrate the effectiveness of our conditioning approach on overlapped speech.
标题: ANIRA:实时音频应用中神经网络推理的架构
链接:https://arxiv.org/abs/2506.12665
备注:8 pages, accepted to the Proceedings of the 5th IEEE International Symposium on the Internet of Sounds (2024) - repository: github.com/anira-project/anira
摘要:目前有许多用于神经网络推理的工具,但许多工具并不满足实时音频应用的要求。作为回应,我们引入了anira,一个高效的跨平台库。为了确保与各种神经网络架构和框架的兼容性,anira支持ONNX SDK、LibTorch和TensorFlow Lite作为后端。每个推理引擎都表现出实时违规,anira通过将音频回调的推理解耦到静态线程池来减轻这种违规。该库集成了内置的延迟管理和广泛的基准测试功能,这两者对于确保连续的信号流至关重要。然后,三种不同的音频效果仿真神经网络架构在各种配置中进行基准测试。采用统计建模来确定各种因素对性能的影响。研究结果表明,对于无状态模型,ONNX的运行时间最低。对于有状态模型,LibTorch表现出最快的性能。我们的研究结果还表明,对于某些模型发动机组合,初始推理需要更长的时间,特别是当这些推理表现出较高的实时违规发生率。
摘要:Numerous tools for neural network inference are currently available, yet many do not meet the requirements of real-time audio applications. In response, we introduce anira, an efficient cross-platform library. To ensure compatibility with a broad range of neural network architectures and frameworks, anira supports ONNX Runtime, LibTorch, and TensorFlow Lite as backends. Each inference engine exhibits real-time violations, which anira mitigates by decoupling the inference from the audio callback to a static thread pool. The library incorporates built-in latency management and extensive benchmarking capabilities, both crucial to ensure a continuous signal flow. Three different neural network architectures for audio effect emulation are then subjected to benchmarking across various configurations. Statistical modeling is employed to identify the influence of various factors on performance. The findings indicate that for stateless models, ONNX Runtime exhibits the lowest runtimes. For stateful models, LibTorch demonstrates the fastest performance. Our results also indicate that for certain model-engine combinations, the initial inferences take longer, particularly when these inferences exhibit a higher incidence of real-time violations.
标题: 使用公共领域电影集的视频引导文本到音乐生成
链接:https://arxiv.org/abs/2506.12573
备注:ISMIR 2025 regular paper. Dataset and code available at this https URL
摘要:尽管音乐生成系统最近取得了进展,但它们在电影制作中的应用仍然有限,因为它们难以捕捉现实世界电影制作的细微差别,电影制作人在为场景选择或创作音乐时会考虑多种因素,例如视觉内容,对话和情感基调。这种限制主要是由于缺乏综合这些要素的全面数据集。为了解决这一差距,我们引入了开放屏幕声音库(OSSL),这是一个由公共领域电影中的电影片段组成的数据集,总计约36.5小时,配有高质量的配乐和人类注释的情绪信息。为了证明我们的数据集在提高预训练模型在电影音乐生成任务中的性能方面的有效性,我们引入了一种新的视频适配器,该适配器通过添加基于视频的调节来增强基于自回归变换器的文本到音乐模型。我们的实验结果表明,我们提出的方法有效地提高了MusicGen媒体的分布和配对保真度的客观措施,和主观的兼容性的情绪和流派。数据集和代码可在https://havenpersona.github.io/ossl-v1上获得。
摘要:Despite recent advancements in music generation systems, their application in film production remains limited, as they struggle to capture the nuances of real-world filmmaking, where filmmakers consider multiple factors-such as visual content, dialogue, and emotional tone-when selecting or composing music for a scene. This limitation primarily stems from the absence of comprehensive datasets that integrate these elements. To address this gap, we introduce Open Screen Sound Library (OSSL), a dataset consisting of movie clips from public domain films, totaling approximately 36.5 hours, paired with high-quality soundtracks and human-annotated mood information. To demonstrate the effectiveness of our dataset in improving the performance of pre-trained models on film music generation tasks, we introduce a new video adapter that enhances an autoregressive transformer-based text-to-music model by adding video-based conditioning. Our experimental results demonstrate that our proposed approach effectively enhances MusicGen-Medium in terms of both objective measures of distributional and paired fidelity, and subjective compatibility in mood and genre. The dataset and code are available at https://havenpersona.github.io/ossl-v1.
标题: StreamMel:基于交织连续自回归模型的实时零触发文本到语音转换
链接:https://arxiv.org/abs/2506.12570
摘要:zero-shot文本到语音(TTS)合成的最新进展已经实现了高质量的语音生成看不见的扬声器,但大多数系统仍然不适合实时应用,因为他们的离线设计。当前的流TTS范例通常依赖于多级流水线和离散表示,导致计算成本增加和次优系统性能。在这项工作中,我们提出了StreamMel,一个开创性的单级流TTS框架,模型连续梅尔频谱。通过将文本标记与声学帧交错,StreamMel实现了低延迟、自回归合成,同时保持了较高的说话人相似性和自然度。LibriSpeech上的实验表明,StreamMel在质量和延迟方面都优于现有的流式TTS基线。它甚至可以达到与离线系统相当的性能,同时支持高效的实时生成,展示了与实时语音大语言模型集成的广阔前景。音频样本可在https://aka.ms/StreamMel上获得。
摘要:Recent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time applications because of their offline design. Current streaming TTS paradigms often rely on multi-stage pipelines and discrete representations, leading to increased computational cost and suboptimal system performance. In this work, we propose StreamMel, a pioneering single-stage streaming TTS framework that models continuous mel-spectrograms. By interleaving text tokens with acoustic frames, StreamMel enables low-latency, autoregressive synthesis while preserving high speaker similarity and naturalness. Experiments on LibriSpeech demonstrate that StreamMel outperforms existing streaming TTS baselines in both quality and latency. It even achieves performance comparable to offline systems while supporting efficient real-time generation, showcasing broad prospects for integration with real-time speech large language models. Audio samples are available at: https://aka.ms/StreamMel.
标题: 探索音频线索以增强测试时视频模型适应
链接:https://arxiv.org/abs/2506.12481
备注:14 pages, 7 figures
摘要:测试时自适应(TTA)旨在通过在测试阶段进行自/无监督学习来提高训练模型的泛化能力。虽然大多数现有的视频TTA方法主要利用视觉监控信号,但它们往往忽略了固有音频数据的潜在贡献。为了解决这个差距,我们提出了一种新的方法,将音频信息到视频TTA。我们的方法利用音频的丰富语义内容来生成音频辅助的伪标签,这是视频TTA背景下的一个新概念。具体来说,我们提出了一种音频到视频标签映射方法,首先采用预先训练的音频模型对从视频中提取的音频信号进行分类,然后通过大型语言模型将基于音频的预测映射到视频标签空间,从而建立音频类别和视频标签之间的联系。为了有效地利用生成的伪标签,我们提出了一个灵活的自适应周期,该周期根据不同视图之间的损失和一致性变化来确定每个样本的最佳自适应迭代次数。这使得能够为每个样本定制自适应过程。在两个广泛使用的数据集(UCF 101-C和Kinetics-Sounds-C)以及两个新构建的具有各种腐败类型的音频-视频TTA数据集(AVE-C和AVMIT-C)上的实验结果证明了我们方法的优越性。我们的方法始终提高了不同视频分类模型的适应性能,并代表了将音频信息集成到视频TTA中的重要一步。代码:https://github.com/keikeiqi/Audio-Assisted-TTA。
摘要:Test-time adaptation (TTA) aims to boost the generalization capability of a trained model by conducting self-/unsupervised learning during the testing phase. While most existing TTA methods for video primarily utilize visual supervisory signals, they often overlook the potential contribution of inherent audio data. To address this gap, we propose a novel approach that incorporates audio information into video TTA. Our method capitalizes on the rich semantic content of audio to generate audio-assisted pseudo-labels, a new concept in the context of video TTA. Specifically, we propose an audio-to-video label mapping method by first employing pre-trained audio models to classify audio signals extracted from videos and then mapping the audio-based predictions to video label spaces through large language models, thereby establishing a connection between the audio categories and video labels. To effectively leverage the generated pseudo-labels, we present a flexible adaptation cycle that determines the optimal number of adaptation iterations for each sample, based on changes in loss and consistency across different views. This enables a customized adaptation process for each sample. Experimental results on two widely used datasets (UCF101-C and Kinetics-Sounds-C), as well as on two newly constructed audio-video TTA datasets (AVE-C and AVMIT-C) with various corruption types, demonstrate the superiority of our approach. Our method consistently improves adaptation performance across different video classification models and represents a significant step forward in integrating audio information into video TTA. Code: https://github.com/keikeiqi/Audio-Assisted-TTA.
标题: 基于风格的作曲家识别与象征性音乐配乐归因:一项系统调查
链接:https://arxiv.org/abs/2506.12440
备注:Accepted at the TISMIR
摘要:本文首次对基于风格的作曲家识别和符号乐谱作者归属的文献进行了全面系统的综述。为了满足这一领域对提高可靠性和再现性的迫切需求,该综述严格分析了58篇在不同历史时期发表的同行评审论文,并根据不断发展的术语进行了检索。该分析批判性地评估了流行的剧目,计算方法和评价方法,突出了重大挑战。它揭示了现有研究的很大一部分受到不充分的验证协议和过度依赖简单的准确性指标的影响,这些指标通常是不平衡的数据集,这可能会破坏归因声明的可信度。强调了平衡准确性和严格交叉验证等强大指标在确保可靠结果方面的关键作用。该调查还详细介绍了不同的特征表示和所采用的机器学习模型的演变。值得注意的现实世界的作者归属的情况下,如涉及巴赫,Josquin Desprez和列侬-麦卡特尼的作品,具体讨论,说明应用计算技术来解决有争议的音乐出处的机会和陷阱。基于这些见解,提出了一套可操作的指导方针,为未来的研究。这些建议旨在显着提高作曲家识别和作者归属研究的可靠性、可重复性和音乐学有效性,促进更强大且可解释的计算风格分析。
摘要:This paper presents the first comprehensive systematic review of literature on style-based composer identification and authorship attribution in symbolic music scores. Addressing the critical need for improved reliability and reproducibility in this field, the review rigorously analyzes 58 peer-reviewed papers published across various historical periods, with the search adapted to evolving terminology. The analysis critically assesses prevailing repertoires, computational approaches, and evaluation methodologies, highlighting significant challenges. It reveals that a substantial portion of existing research suffers from inadequate validation protocols and an over-reliance on simple accuracy metrics for often imbalanced datasets, which can undermine the credibility of attribution claims. The crucial role of robust metrics like Balanced Accuracy and rigorous cross-validation in ensuring trustworthy results is emphasized. The survey also details diverse feature representations and the evolution of machine learning models employed. Notable real-world authorship attribution cases, such as those involving works attributed to Bach, Josquin Desprez, and Lennon-McCartney, are specifically discussed, illustrating the opportunities and pitfalls of applying computational techniques to resolve disputed musical provenance. Based on these insights, a set of actionable guidelines for future research are proposed. These recommendations are designed to significantly enhance the reliability, reproducibility, and musicological validity of composer identification and authorship attribution studies, fostering more robust and interpretable computational stylistic analysis.
标题: 当代流行音乐中音高分析的方法:维塔利克音乐中和声的多个音高
链接:https://arxiv.org/abs/2506.12405
备注:Pending review, Journal of the Audio Engineering Society
摘要:目标。这项研究表明,使用多个感知音高产生于一个单一的谐波复杂的音调是一个积极的和故意的当代流行音乐的特点。通过电子艺术家Vitalic等人的作品中的例子来说明这种现象。 方法.进行了两个听力测试:(1)评估从单个谐波音调感知的同时音高的数量,以及(2)谐波音调序列的手动音高转录。分析了信号特征与音高感知的关系。 结果在研究中的音乐序列中发现的合成谐波音调被观察到比它们的声学对应物传输更多的感知音高,并且在听众之间存在显着差异。多个模糊的音高与音调属性,如突出的上分音和特定的自相关曲线。 结论.在当代流行音乐的背景下,和声通常可以传达几个模糊的音高。感知音高的集合取决于收听者和收听条件两者。
摘要:Aims. This study suggests that the use of multiple perceived pitches arising from a single harmonic complex tone is an active and intentional feature of contemporary popular music. The phenomenon is illustrated through examples drawn from the work of electronic artist Vitalic and others. Methods. Two listening tests were conducted: (1) evaluation of the number of simultaneous pitches perceived from single harmonic tones, and (2) manual pitch transcription of sequences of harmonic tones. Relationships between signal characteristics and pitch perception were then analyzed. Results. The synthetic harmonic tones found in the musical sequences under study were observed to transmit more perceived pitches than their acoustic counterparts, with significant variation across listeners. Multiple ambiguous pitches were associated with tone properties such as prominent upper partials and particular autocorrelation profiles. Conclusions. Harmonic tones in a context of contemporary popular music can, in general, convey several ambiguous pitches. The set of perceived pitches depends on both the listener and the listening conditions.
标题: GSDNet:从图谱角度重新审视对话情绪识别的不完全多模式扩散
链接:https://arxiv.org/abs/2506.12325
摘要:会话中的多模态情感识别(MERC)旨在通过分析来自多个源的话语信息(即,视频、音频和文本)。与单模态相比,通过融合不同模态的互补语义信息可以获得更鲁棒的话语表示。然而,模态缺失问题严重限制了MERC在实际场景中的性能。最近的工作取得了令人印象深刻的性能模态完成使用图神经网络和扩散模型,分别。这启发我们通过图扩散模型将这两个维度结合起来,以获得更强大的模式恢复能力。然而,现有的图扩散模型直接在邻接矩阵中加入高斯噪声,可能会破坏图的连通性和局部结构,导致生成的图数据无法保留原图的语义和拓扑信息。为此,我们提出了一种新的图谱扩散网络(GSDNet),它将高斯噪声映射到丢失模态的图谱空间,并根据其原始分布恢复丢失数据。与以往的图扩散方法相比,GSDNet只影响邻接矩阵的特征值,而不直接破坏邻接矩阵,能够在扩散过程中保持全局拓扑信息和重要的谱特征。大量的实验表明,GSDNet在各种模态丢失场景中实现了最先进的情感识别性能。
摘要:Multimodal emotion recognition in conversations (MERC) aims to infer the speaker's emotional state by analyzing utterance information from multiple sources (i.e., video, audio, and text). Compared with unimodality, a more robust utterance representation can be obtained by fusing complementary semantic information from different modalities. However, the modality missing problem severely limits the performance of MERC in practical scenarios. Recent work has achieved impressive performance on modality completion using graph neural networks and diffusion models, respectively. This inspires us to combine these two dimensions through the graph diffusion model to obtain more powerful modal recovery capabilities. Unfortunately, existing graph diffusion models may destroy the connectivity and local structure of the graph by directly adding Gaussian noise to the adjacency matrix, resulting in the generated graph data being unable to retain the semantic and topological information of the original graph. To this end, we propose a novel Graph Spectral Diffusion Network (GSDNet), which maps Gaussian noise to the graph spectral space of missing modalities and recovers the missing data according to its original distribution. Compared with previous graph diffusion methods, GSDNet only affects the eigenvalues of the adjacency matrix instead of destroying the adjacency matrix directly, which can maintain the global topological information and important spectral features during the diffusion process. Extensive experiments have demonstrated that GSDNet achieves state-of-the-art emotion recognition performance in various modality loss scenarios.
标题: Phonikud:实时文本到语音的希伯来语字形到音素转换
链接:https://arxiv.org/abs/2506.12311
备注:Project page: this https URL
摘要:现代希伯来语的实时文本到语音(TTS)是具有挑战性的,由于语言的拼写复杂性。现有的解决方案忽略了关键的语音特征,如重音,即使添加了元音标记,重音也没有被充分指定。为了解决这些限制,我们引入Phonikud,一个轻量级的,开源的希伯来语字素到音素(G2 P)系统,输出完全指定的IPA音标。我们的方法适应现有的变音模型与轻量级的适配器,招致可忽略不计的额外延迟。我们还贡献了带有IPA注释的希伯来语转录语音的ILSpeech数据集,作为希伯来语G2 P的基准和TTS系统的训练数据。我们的研究结果表明,与以前的方法相比,Phonikud G2 P转换更准确地预测希伯来语文本中的音素,这使得有效的实时希伯来语TTS模型的训练具有优越的速度-准确性权衡。我们在https://phonikud.github.io上发布我们的代码、数据和模型。
摘要:Real-time text-to-speech (TTS) for Modern Hebrew is challenging due to the language's orthographic complexity. Existing solutions ignore crucial phonetic features such as stress that remain underspecified even when vowel marks are added. To address these limitations, we introduce Phonikud, a lightweight, open-source Hebrew grapheme-to-phoneme (G2P) system that outputs fully-specified IPA transcriptions. Our approach adapts an existing diacritization model with lightweight adaptors, incurring negligible additional latency. We also contribute the ILSpeech dataset of transcribed Hebrew speech with IPA annotations, serving as a benchmark for Hebrew G2P and as training data for TTS systems. Our results demonstrate that Phonikud G2P conversion more accurately predicts phonemes from Hebrew text compared to prior methods, and that this enables training of effective real-time Hebrew TTS models with superior speed-accuracy trade-offs. We release our code, data, and models at https://phonikud.github.io.
标题: 通过习得质量评估的多指标监督改善语音增强
链接:https://arxiv.org/abs/2506.12260
备注:Submitted to ASRU 2025
摘要:语音质量评估(SQA)旨在预测语音信号在大范围失真下的感知质量。它与语音增强(SE)有着内在的联系,后者试图通过去除不需要的信号分量来提高语音质量。虽然SQA模型被广泛用于评估SE绩效,但其指导SE培训的潜力仍有待开发。在这项工作中,我们研究了一个培训框架,该框架利用SQA模型,经过训练,可以从公共SE排行榜中预测多个评估指标,作为SE的监督信号。这种方法解决了传统SE目标的关键限制,例如SI-SNR,其通常无法与感知质量保持一致,并且在评估指标中概括性较差。此外,它可以在没有干净引用的情况下对真实世界的数据进行训练。在模拟和真实测试集上的实验表明,SQA引导的训练在一系列质量指标上持续提高性能。
摘要:Speech quality assessment (SQA) aims to predict the perceived quality of speech signals under a wide range of distortions. It is inherently connected to speech enhancement (SE), which seeks to improve speech quality by removing unwanted signal components. While SQA models are widely used to evaluate SE performance, their potential to guide SE training remains underexplored. In this work, we investigate a training framework that leverages a SQA model, trained to predict multiple evaluation metrics from a public SE leaderboard, as a supervisory signal for SE. This approach addresses a key limitation of conventional SE objectives, such as SI-SNR, which often fail to align with perceptual quality and generalize poorly across evaluation metrics. Moreover, it enables training on real-world data where clean references are unavailable. Experiments on both simulated and real-world test sets show that SQA-guided training consistently improves performance across a range of quality metrics.
标题: SSLAM:通过复音声景的音频混合增强自我监督模型
链接:https://arxiv.org/abs/2506.12222
备注:Accepted at ICLR 2025. Code and pre-trained models are available at \url{this https URL}
摘要:自我监督的预训练音频网络在现实世界的系统中得到了广泛的应用,特别是在多模态大型语言模型中。这些网络通常在冻结状态下使用,假设SSL预训练已经足够让它们处理真实世界的音频。然而,一个关键的问题仍然存在:这些模型在现实世界中的实际表现如何,其中音频通常是复调和复杂的,涉及多个重叠的声源?当前的音频SSL方法通常在主要以单声道音频为特征的数据集上进行基准测试,例如环境声音和语音。因此,SSL模型推广到复音音频(自然场景中的常见特征)的能力仍然没有得到充分探索。这种限制引起了人们对SSL模型在更真实的音频设置中的实际鲁棒性的担忧。为了解决这一差距,我们引入了音频混合自监督学习(SSLAM),这是音频SSL研究的一个新方向,旨在提高模型从复调数据中学习的能力,同时保持单声道数据的强大性能。我们在标准音频SSL基准数据集上彻底评估了SSLAM,这些数据集主要是单声道的,并使用一系列高质量,公开可用的复音数据集对SOTA方法进行了全面的比较分析。SSLAM不仅提高了复音音频的模型性能,而且在标准音频SSL基准测试中保持或超过了性能。值得注意的是,它比AudioSet-2 M(AS-2 M)提高了3.9%,平均精度(mAP)达到50.2。对于复调数据集,SSLAM在线性评估和微调机制中设置了新的SOTA,性能提高高达9.1\%(mAP)。
摘要:Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the SSL pre-training has sufficiently equipped them to handle real-world audio. However, a critical question remains: how well do these models actually perform in real-world conditions, where audio is typically polyphonic and complex, involving multiple overlapping sound sources? Current audio SSL methods are often benchmarked on datasets predominantly featuring monophonic audio, such as environmental sounds, and speech. As a result, the ability of SSL models to generalize to polyphonic audio, a common characteristic in natural scenarios, remains underexplored. This limitation raises concerns about the practical robustness of SSL models in more realistic audio settings. To address this gap, we introduce Self-Supervised Learning from Audio Mixtures (SSLAM), a novel direction in audio SSL research, designed to improve, designed to improve the model's ability to learn from polyphonic data while maintaining strong performance on monophonic data. We thoroughly evaluate SSLAM on standard audio SSL benchmark datasets which are predominantly monophonic and conduct a comprehensive comparative analysis against SOTA methods using a range of high-quality, publicly available polyphonic datasets. SSLAM not only improves model performance on polyphonic audio, but also maintains or exceeds performance on standard audio SSL benchmarks. Notably, it achieves up to a 3.9\% improvement on the AudioSet-2M (AS-2M), reaching a mean average precision (mAP) of 50.2. For polyphonic datasets, SSLAM sets new SOTA in both linear evaluation and fine-tuning regimes with performance improvements of up to 9.1\% (mAP).
标题: ViSAGee:视频到空间音频生成
链接:https://arxiv.org/abs/2506.12199
备注:ICLR 2025. Project page: this https URL
摘要:空间音频对于增强视听体验的沉浸感至关重要,但其制作通常需要复杂的录音系统和专业知识。在这项工作中,我们解决了直接从无声视频生成一阶立体混响(一种广泛使用的空间音频格式)的新问题。为了支持这一任务,我们引入了YT-Ambigen,这是一个包含102 K 5秒YouTube视频剪辑的数据集,与相应的一阶立体混响配对。我们还提出了新的评估指标,以评估基于音频能量图和显着性度量生成的音频的空间方面。此外,我们提出了视频到空间音频生成(ViSAGe),一个端到端的框架,通过利用CLIP视觉功能,自回归神经音频编解码器建模与方向和视觉指导,从无声视频帧生成一阶立体混响。实验结果表明,ViSAGe产生合理的和连贯的一阶立体混响,优于两个阶段的方法,包括视频到音频生成和音频空间化。定性示例进一步说明ViSAGe生成适应视点变化的时间对齐的高质量空间音频。
摘要:Spatial audio is essential for enhancing the immersiveness of audio-visual experiences, yet its production typically demands complex recording systems and specialized expertise. In this work, we address a novel problem of generating first-order ambisonics, a widely used spatial audio format, directly from silent videos. To support this task, we introduce YT-Ambigen, a dataset comprising 102K 5-second YouTube video clips paired with corresponding first-order ambisonics. We also propose new evaluation metrics to assess the spatial aspect of generated audio based on audio energy maps and saliency metrics. Furthermore, we present Video-to-Spatial Audio Generation (ViSAGe), an end-to-end framework that generates first-order ambisonics from silent video frames by leveraging CLIP visual features, autoregressive neural audio codec modeling with both directional and visual guidance. Experimental results demonstrate that ViSAGe produces plausible and coherent first-order ambisonics, outperforming two-stage approaches consisting of video-to-audio generation and audio spatialization. Qualitative examples further illustrate that ViSAGe generates temporally aligned high-quality spatial audio that adapts to viewpoint changes.
标题: 通过两遍解码将Whisper改编为流语音识别
链接:https://arxiv.org/abs/2506.12154
备注:Accepted to INTERSPEECH 2025
摘要:OpenAI Whisper是一系列强大的自动语音识别(ASR)模型,经过680,000小时的音频训练。然而,它的编码器-解码器架构,用序列到序列目标训练,缺乏对流式ASR的原生支持。在本文中,我们使用WeNet工具包通过采用统一的双通道(U2)结构来微调Whisper用于流式ASR。我们引入了一个额外的连接主义时间分类(CTC)解码器,用因果注意掩码训练来生成流部分转录本,而原始的Whisper解码器则对这些部分输出进行重新排序。我们在LibriSpeech和财报电话数据集上的实验表明,只要有足够的微调数据,Whisper就可以适应有能力的流ASR模型。我们还介绍了一种混合令牌化方法,它使用一个较小的令牌空间的CTC解码器,同时保留耳语的原始令牌空间的注意力解码器,从而提高数据效率和泛化。
摘要:OpenAI Whisper is a family of robust Automatic Speech Recognition (ASR) models trained on 680,000 hours of audio. However, its encoder-decoder architecture, trained with a sequence-to-sequence objective, lacks native support for streaming ASR. In this paper, we fine-tune Whisper for streaming ASR using the WeNet toolkit by adopting a Unified Two-pass (U2) structure. We introduce an additional Connectionist Temporal Classification (CTC) decoder trained with causal attention masks to generate streaming partial transcripts, while the original Whisper decoder reranks these partial outputs. Our experiments on LibriSpeech and an earnings call dataset demonstrate that, with adequate fine-tuning data, Whisper can be adapted into a capable streaming ASR model. We also introduce a hybrid tokenizer approach, which uses a smaller token space for the CTC decoder while retaining Whisper's original token space for the attention decoder, resulting in improved data efficiency and generalization.
标题: TuneGenie:基于推理的LLM代理,提供优先音乐生成
链接:https://arxiv.org/abs/2506.12083
备注:15 pages
摘要:最近,大型语言模型(LLM)在从生成图像到空间推理的各种任务中表现出了巨大的潜力。考虑到他们显着的(和不断增长的)文本推理能力,我们调查LLM在进行个人的音乐偏好分析(基于播放列表元数据,个人写作等)的潜力。并产生有效的提示(基于这些分析),以传递给Suno AI(用于音乐制作的生成AI工具)。我们提出了一种新的基于LLM的文本表示音乐模型(我们称之为TuneGenie),以及我们开发的各种方法来评估和基准类似的模型,增加了越来越多(越来越有争议)的关于使用AI生成艺术的研究语料库。
摘要:Recently, Large language models (LLMs) have shown great promise across a diversity of tasks, ranging from generating images to reasoning spatially. Considering their remarkable (and growing) textual reasoning capabilities, we investigate LLMs' potency in conducting analyses of an individual's preferences in music (based on playlist metadata, personal write-ups, etc.) and producing effective prompts (based on these analyses) to be passed to Suno AI (a generative AI tool for music production). Our proposition of a novel LLM-based textual representation to music model (which we call TuneGenie) and the various methods we develop to evaluate & benchmark similar models add to the increasing (and increasingly controversial) corpus of research on the use of AI in generating art.
标题: SpeechRefiner:前端算法的感知质量细化
链接:https://arxiv.org/abs/2506.13709
备注:Accepted by Interspeech 2025
摘要:诸如去噪、去混响和分离的语音预处理技术通常被用作各种下游语音处理任务的前端。然而,这些方法有时可能是不够的,导致残留噪声或引入新的伪影。这样的缺陷通常不被像SI-SNR这样的度量捕获,但是对于人类收听者是明显的。为了解决这个问题,我们引入了SpeechRefiner,这是一种后处理工具,它利用条件流匹配(CFM)来提高语音的感知质量。在这项研究中,我们将SpeechRefiner与最近的特定于任务的细化方法进行基准测试,并在我们的内部处理管道中评估其性能,该管道集成了多个前端算法。实验表明,SpeechRefiner在不同的损伤源中表现出很强的泛化能力,显著提高了语音感知质量。音频演示可以在https://speechrefiner.github.io/SpeechRefiner/上找到。
摘要:Speech pre-processing techniques such as denoising, de-reverberation, and separation, are commonly employed as front-ends for various downstream speech processing tasks. However, these methods can sometimes be inadequate, resulting in residual noise or the introduction of new artifacts. Such deficiencies are typically not captured by metrics like SI-SNR but are noticeable to human listeners. To address this, we introduce SpeechRefiner, a post-processing tool that utilizes Conditional Flow Matching (CFM) to improve the perceptual quality of speech. In this study, we benchmark SpeechRefiner against recent task-specific refinement methods and evaluate its performance within our internal processing pipeline, which integrates multiple front-end algorithms. Experiments show that SpeechRefiner exhibits strong generalization across diverse impairment sources, significantly enhancing speech perceptual quality. Audio demos can be found at https://speechrefiner.github.io/SpeechRefiner/.
标题: 基于PSELDnet预训练和BiMamba序列建模的立体声事件定位和检测
链接:https://arxiv.org/abs/2506.13455
备注:Technical report for DCASE 2025 Challenge Task 3
摘要:预训练方法在声音事件定位和检测(SELD)任务中取得了显着的性能改进,但现有的基于transformer的模型具有较高的计算复杂度。在这项工作中,我们提出了一个立体声声音事件定位和检测系统的基础上预先训练的PSELDnet和双向曼巴序列建模。我们用BiMamba模块替换Conformer模块,并引入非对称卷积来更有效地建模时间和频率维度之间的时空关系。在DCASE2025任务3开发数据集上的实验结果表明,该方法在降低计算复杂度的同时,显著优于基线和原始的PSELDnet与Conformer解码器架构。这些发现突出了BiMamba架构在应对SELD任务挑战方面的有效性。
摘要:Pre-training methods have achieved significant performance improvements in sound event localization and detection (SELD) tasks, but existing Transformer-based models suffer from high computational complexity. In this work, we propose a stereo sound event localization and detection system based on pre-trained PSELDnet and bidirectional Mamba sequence modeling. We replace the Conformer module with a BiMamba module and introduce asymmetric convolutions to more effectively model the spatiotemporal relationships between time and frequency dimensions. Experimental results demonstrate that the proposed method achieves significantly better performance than the baseline and the original PSELDnet with Conformer decoder architecture on the DCASE2025 Task 3 development dataset, while also reducing computational complexity. These findings highlight the effectiveness of the BiMamba architecture in addressing the challenges of the SELD task.
标题: 野外语音编辑的特定实例测试时训练
链接:https://arxiv.org/abs/2506.13295
备注:Submitted to IEEE Signal Processing Letters
摘要:语音编辑系统旨在自然地修改语音内容,同时保持声学一致性和说话者身份。然而,以前的研究往往难以适应看不见的和不同的声学条件,导致在现实世界中的编辑性能下降。为了解决这个问题,我们提出了一个特定于实例的测试时间训练方法,用于在野外进行语音编辑。我们的方法采用直接监督从地面实况声学特征在未经编辑的地区,和间接监督在编辑的地区通过辅助损失的基础上持续时间的限制和音素预测。该策略缓解了语音编辑中的带宽不连续问题,确保了未编辑区域和编辑区域之间的平滑声学过渡。此外,它通过在测试时间训练期间调整掩码长度使模型适应目标持续时间,从而实现对语音速率的精确控制。在野外基准数据集上的实验表明,我们的方法优于现有的语音编辑系统在客观和主观评价。
摘要:Speech editing systems aim to naturally modify speech content while preserving acoustic consistency and speaker identity. However, previous studies often struggle to adapt to unseen and diverse acoustic conditions, resulting in degraded editing performance in real-world scenarios. To address this, we propose an instance-specific test-time training method for speech editing in the wild. Our approach employs direct supervision from ground-truth acoustic features in unedited regions, and indirect supervision in edited regions via auxiliary losses based on duration constraints and phoneme prediction. This strategy mitigates the bandwidth discontinuity problem in speech editing, ensuring smooth acoustic transitions between unedited and edited regions. Additionally, it enables precise control over speech rate by adapting the model to target durations via mask length adjustment during test-time training. Experiments on in-the-wild benchmark datasets demonstrate that our method outperforms existing speech editing systems in both objective and subjective evaluations.
标题: ZipVoice:具有流匹配的快速高质量Zero-Shot文本到语音
链接:https://arxiv.org/abs/2506.13053
摘要:现有的大规模zero-shot文本到语音(TTS)模型提供高的语音质量,但遭受缓慢的推理速度,由于大量的参数。为了解决这个问题,本文介绍了ZipVoice,一个高质量的流匹配的zero-shot TTS模型,具有紧凑的模型大小和快速的推理速度。主要设计包括:1)基于Zipformer的流匹配解码器,以在受限大小下保持足够的建模能力; 2)基于平均上采样的初始语音-文本对齐和基于Zipformer的文本编码器,以提高语音可懂度; 3)流蒸馏方法,以减少采样步骤并消除与无分类器指导相关联的推理开销。在10万小时多语言数据集上的实验表明,ZipVoice在语音质量上与最先进的模型相匹配,同时比基于DiT的流匹配基线小3倍,快30倍。代码、模型检查点和演示样本都是公开的。
摘要:Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-quality flow-matching-based zero-shot TTS model with a compact model size and fast inference speed. Key designs include: 1) a Zipformer-based flow-matching decoder to maintain adequate modeling capabilities under constrained size; 2) Average upsampling-based initial speech-text alignment and Zipformer-based text encoder to improve speech intelligibility; 3) A flow distillation method to reduce sampling steps and eliminate the inference overhead associated with classifier-free guidance. Experiments on 100k hours multilingual datasets show that ZipVoice matches state-of-the-art models in speech quality, while being 3 times smaller and up to 30 times faster than a DiT-based flow-matching baseline. Codes, model checkpoints and demo samples are publicly available.
标题: 基于脑磁图(MEG)的无创中文语音解码
链接:https://arxiv.org/abs/2506.12817
摘要:语音脑机接口作为一种新兴的脑机接口模式,具有直接反映听觉感知和思维的潜力,为失语症患者提供了一种很有前途的交流方式。汉语是世界上使用最广泛的语言之一,而汉语语音脑机接口的研究却非常有限。本文报告了一个用于无创汉语语音脑机接口的文本脑磁图(MEG)数据集。它还提出了一种多模态辅助语音解码(MASD)算法,以捕获在语音活动过程中嵌入在大脑信号中的文本和声学信息。实验结果表明,我们的文本MEG数据集和我们提出的MASD算法的有效性。据我们所知,这是第一个研究模态辅助解码的非侵入性语音脑机接口。
摘要:As an emerging paradigm of brain-computer interfaces (BCIs), speech BCI has the potential to directly reflect auditory perception and thoughts, offering a promising communication alternative for patients with aphasia. Chinese is one of the most widely spoken languages in the world, whereas there is very limited research on speech BCIs for Chinese language. This paper reports a text-magnetoencephalography (MEG) dataset for non-invasive Chinese speech BCIs. It also proposes a multi-modality assisted speech decoding (MASD) algorithm to capture both text and acoustic information embedded in brain signals during speech activities. Experiment results demonstrated the effectiveness of both our text-MEG dataset and our proposed MASD algorithm. To our knowledge, this is the first study on modality-assisted decoding for non-invasive speech BCIs.
标题: 用于声音事件检测的频率动态卷积
链接:https://arxiv.org/abs/2506.12785
备注:Ph. D. Dissertation in English(KAIST)
摘要:最近基于深度学习的声音事件检测(SED)的研究主要集中在卷积递归神经网络(CRNN)和Transformer模型上。然而,传统的基于2D卷积的模型假设沿时间和频率轴的移位不变性,导致在处理声信号的频率相关特性时的不一致性。为了解决这个问题,本研究提出了频率动态卷积(FDY conv),它根据输入信号的频率组成动态调整卷积核,以提高SED性能。FDY conv通过基于频率特定注意力权重的自适应加权多个基核来构造最优频率响应。实验结果表明,将FDY conv应用于CRNN,与基线CRNN相比,DESED数据集的性能提高了7.56%。然而,FDY conv具有局限性,因为它在所有频率上组合了相同形状的基础内核,限制了其捕获不同频率特定特性的能力。此外,3\times3 $ basis内核大小不足以捕获更宽的频率范围。为了克服这些局限性,本研究介绍了一个扩展的家庭FDY转换模型。Dilated FDY conv(DFD conv)应用具有各种膨胀率的卷积核来沿着频率轴扩展感受野并增强频率特定的特征表示。实验结果表明,DFD conv比基线提高了9.27%的性能。部分FDY conv(PFD conv)解决了FDY conv的高计算成本,这是由于使用动态内核执行所有卷积运算而导致的。由于FDY conv可能会为准平稳声音事件引入不必要的自适应性,因此PFD conv将标准2D卷积与频率自适应内核集成在一起,以降低计算复杂度,同时保持性能。实验结果表明,PFD conv提高了7.80%的基线性能,同时减少了54.4%的FDY conv相比,参数的数量。多重扩张FDY conv(MDFD conv)通过解决其在所有频率上应用相同扩张的结构限制来扩展DFD conv。通过利用具有不同膨胀率的多个卷积核,MDFD conv有效地捕获了不同的频率相关模式。实验结果表明,MDFD conv实现了最高的性能,提高了10.98%的基准CRNN的性能。此外,标准FDY conv采用时间平均池化,其沿时间轴向所有帧分配相等权重,限制了其有效捕获瞬态事件的能力。为了克服这一点,本研究提出了TAP-FDY conv(TFD conv),它集成了专注于显著特征的时间注意力池(TA),强调瞬态特征的速度注意力池(VA)和捕获静态属性的平均池(AP)。TAP-FDY conv实现了与MDFD conv相同的性能,但减少了约30.01%的参数数量(12.703 M vs. 18.157 M),实现了相同的精度和更低的计算复杂度。类的性能分析表明,FDY conv提高了非平稳事件的检测,DFD conv是特别有效的事件具有广泛的频谱特征,和PFD conv增强了准平稳事件的检测。此外,TFD conv(TFD-CRNN)在检测瞬态事件方面表现出强大的性能。在案例研究中,PFD conv在坦克动力系统故障识别中有效地捕获稳定的信号模式,DFD conv在变速电机故障识别中识别宽谐波谱模式,而TFD conv在检测海上电弧检测中的瞬态信号方面优于其他模型。这些结果表明,频率自适应卷积及其扩展变体在基于深度学习的音频处理中提供了传统2D卷积的强大替代方案。
摘要:Recent research in deep learning-based Sound Event Detection (SED) has primarily focused on Convolutional Recurrent Neural Networks (CRNNs) and Transformer models. However, conventional 2D convolution-based models assume shift invariance along both the temporal and frequency axes, leadin to inconsistencies when dealing with frequency-dependent characteristics of acoustic signals. To address this issue, this study proposes Frequency Dynamic Convolution (FDY conv), which dynamically adjusts convolutional kernels based on the frequency composition of the input signal to enhance SED performance. FDY conv constructs an optimal frequency response by adaptively weighting multiple basis kernels based on frequency-specific attention weights. Experimental results show that applying FDY conv to CRNNs improves performance on the DESED dataset by 7.56% compared to the baseline CRNN. However, FDY conv has limitations in that it combines basis kernels of the same shape across all frequencies, restricting its ability to capture diverse frequency-specific characteristics. Additionally, the $3\times3$ basis kernel size is insufficient to capture a broader frequency range. To overcome these limitations, this study introduces an extended family of FDY conv models. Dilated FDY conv (DFD conv) applies convolutional kernels with various dilation rates to expand the receptive field along the frequency axis and enhance frequency-specific feature representation. Experimental results show that DFD conv improves performance by 9.27% over the baseline. Partial FDY conv (PFD conv) addresses the high computational cost of FDY conv, which results from performing all convolution operations with dynamic kernels. Since FDY conv may introduce unnecessary adaptivity for quasi-stationary sound events, PFD conv integrates standard 2D convolutions with frequency-adaptive kernels to reduce computational complexity while maintaining performance. Experimental results demonstrate that PFD conv improves performance by 7.80% over the baseline while reducing the number of parameters by 54.4% compared to FDY conv. Multi-Dilated FDY conv (MDFD conv) extends DFD conv by addressing its structural limitation of applying the same dilation across all frequencies. By utilizing multiple convolutional kernels with different dilation rates, MDFD conv effectively captures diverse frequency-dependent patterns. Experimental results indicate that MDFD conv achieves the highest performance, improving the baseline CRNN performance by 10.98%. Furthermore, standard FDY conv employs Temporal Average Pooling, which assigns equal weight to all frames along the time axis, limiting its ability to effectively capture transient events. To overcome this, this study proposes TAP-FDY conv (TFD conv), which integrates Temporal Attention Pooling (TA) that focuses on salient features, Velocity Attention Pooling (VA) that emphasizes transient characteristics, and Average Pooling (AP) that captures stationary properties. TAP-FDY conv achieves the same performance as MDFD conv but reduces the number of parameters by approximately 30.01% (12.703M vs. 18.157M), achieving equivalent accuracy with lower computational complexity. Class-wise performance analysis reveals that FDY conv improves detection of non-stationary events, DFD conv is particularly effective for events with broad spectral features, and PFD conv enhances the detection of quasi-stationary events. Additionally, TFD conv (TFD-CRNN) demonstrates strong performance in detecting transient events. In the case studies, PFD conv effectively captures stable signal patterns in tank powertrain fault recognition, DFD conv recognizes wide harmonic spectral patterns on speed-varying motor fault recognition, while TFD conv outperforms other models in detecting transient signals in offshore arc detection. These results suggest that frequency-adaptive convolutions and their extended variants provide a robust alternative to conventional 2D convolutions in deep learning-based audio processing.
标题: 使用神经图相似性指数测量(NSIM)建模听力损失和腰神经退行性变
链接:https://arxiv.org/abs/2506.12705
备注:Accepted for presentation at INTERSPEECH 2025
摘要:在嘈杂环境中的听力障碍仍然是听力损失和听力正常的人的共同投诉。这被假设是由于称为耳蜗神经变性(CND)的情况而引起的,CND也会导致助听器结果的显著变化。本文使用听觉周边的计算模型来模拟各种听觉任务。我们提出了一种客观的方法来量化听力损失和CND比较听觉神经纤维反应使用神经图相似性指数测量(NSIM)。具体来说,研究1表明,NSIM可以用于映射个人的听力损失的音素识别任务的性能与合理的准确性。在研究2中,我们表明NSIM是一种敏感的测量方法,也可以用于捕获CND导致的缺陷,并且可以作为听觉突触病的非侵入性生物标志物的候选者。
摘要:Trouble hearing in noisy situations remains a common complaint for both individuals with hearing loss and individuals with normal hearing. This is hypothesized to arise due to condition called: cochlear neural degeneration (CND) which can also result in significant variabilities in hearing aids outcomes. This paper uses computational models of auditory periphery to simulate various hearing tasks. We present an objective method to quantify hearing loss and CND by comparing auditory nerve fiber responses using a Neurogram Similarity Index Measure (NSIM). Specifically study 1, shows that NSIM can be used to map performance of individuals with hearing loss on phoneme recognition task with reasonable accuracy. In the study 2, we show that NSIM is a sensitive measure that can also be used to capture the deficits resulting from CND and can be a candidate for noninvasive biomarker of auditory synaptopathy.
标题: 走向神经音频编解码器源解析
链接:https://arxiv.org/abs/2506.12627
摘要:一类新的音频deepfakes-codecfakes(CF)-最近引起了人们的注意,它是由音频语言模型合成的,在后端利用神经音频编解码器(NAC)。作为回应,社区引入了专门的基准和量身定制的检测策略。随着该领域的发展,人们的努力已经超越了二元检测,转向了源归因,包括开集归因,其目的是识别负责生成的NAC,并在推理过程中标记新的、看不见的NAC。这种向来源归因的转变提高了法医的可解释性和问责制。然而,开集属性仍然存在根本性的局限性:虽然它可以检测到NAC是不熟悉的,但它无法表征或识别单个看不见的编解码器。它将这些输入视为一般的“未知数”,缺乏对其内部配置的洞察力。这导致了主要的缺点:对新NAC的推广有限,并且无法解决NAC家族内的细粒度变化。为了解决这些差距,我们提出了神经音频编解码器源解析(NACSP)-一种范式转变,将CF的源属性重新构建为生成NAC参数(如量化器,带宽和采样率)的结构化回归。我们将NACSP制定为用于预测这些NAC参数的多任务回归任务,并使用各种最先进的语音预训练模型(PTM)建立第一个全面的基准。为此,我们提出了HYDRONIC,一个新的框架,利用双曲几何解开复杂的潜在属性PTM表示。通过在多个曲率感知的双曲子空间上使用特定于任务的注意力,HYDRONIC实现了卓越的多任务泛化。我们广泛的实验表明,与在欧几里得空间中操作的基线相比,HYDRONIC在基准CFs数据集上获得了最佳结果。
摘要:A new class of audio deepfakes-codecfakes (CFs)-has recently caught attention, synthesized by Audio Language Models that leverage neural audio codecs (NACs) in the backend. In response, the community has introduced dedicated benchmarks and tailored detection strategies. As the field advances, efforts have moved beyond binary detection toward source attribution, including open-set attribution, which aims to identify the NAC responsible for generation and flag novel, unseen ones during inference. This shift toward source attribution improves forensic interpretability and accountability. However, open-set attribution remains fundamentally limited: while it can detect that a NAC is unfamiliar, it cannot characterize or identify individual unseen codecs. It treats such inputs as generic ``unknowns'', lacking insight into their internal configuration. This leads to major shortcomings: limited generalization to new NACs and inability to resolve fine-grained variations within NAC families. To address these gaps, we propose Neural Audio Codec Source Parsing (NACSP) - a paradigm shift that reframes source attribution for CFs as structured regression over generative NAC parameters such as quantizers, bandwidth, and sampling rate. We formulate NACSP as a multi-task regression task for predicting these NAC parameters and establish the first comprehensive benchmark using various state-of-the-art speech pre-trained models (PTMs). To this end, we propose HYDRA, a novel framework that leverages hyperbolic geometry to disentangle complex latent properties from PTM representations. By employing task-specific attention over multiple curvature-aware hyperbolic subspaces, HYDRA enables superior multi-task generalization. Our extensive experiments show HYDRA achieves top results on benchmark CFs datasets compared to baselines operating in Euclidean space.
标题: 缓解引导说话人嵌入中的非目标说话人偏见
链接:https://arxiv.org/abs/2506.12500
备注:Accepted to Interspeech 2025
摘要:在多说话人环境中获得高质量的说话人嵌入对于许多应用来说至关重要。最近提出的一种引导式说话人嵌入框架,利用目标和非目标说话人的语音活动作为线索,大大提高了严重重叠下的嵌入,在低重叠的情况下有小的退化。然而,由于极端的重叠在自然对话中是罕见的,这种退化不能被忽视。本文首先揭示了退化是由于广泛用于说话人嵌入提取器的基于全局统计的模块对仅包含非目标说话人的间隔过于敏感。作为一种对策,我们提出了这样的模块,利用目标扬声器活动线索的扩展,从目标是活跃的间隔计算统计。该方法提高了说话人确认在低重叠率和高重叠率下的性能,以及多数据集上的日志化性能。
摘要:Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target speakers as clues, drastically improved embeddings under severe overlap with small degradation in low-overlap cases. However, since extreme overlaps are rare in natural conversations, this degradation cannot be overlooked. This paper first reveals that the degradation is caused by the global-statistics-based modules, widely used in speaker embedding extractors, being overly sensitive to intervals containing only non-target speakers. As a countermeasure, we propose an extension of such modules that exploit the target speaker activity clues, to compute statistics from intervals where the target is active. The proposed method improves speaker verification performance in both low and high overlap ratios, and diarization performance on multiple datasets.
标题: CMI-Bench:评估音乐教学遵循的综合基准
链接:https://arxiv.org/abs/2506.12285
备注:Accepted by ISMIR 2025
摘要:音频文本大语言模型(LLM)的最新进展为音乐理解和生成开辟了新的可能性。然而,现有的基准在范围上是有限的,往往依赖于简化的任务或多选择的评估,无法反映现实世界的音乐分析的复杂性。我们重新解释了广泛的传统的MIR注释作为解释以下格式,并介绍了CMI-Bench,一个全面的音乐指令以下基准,旨在评估音频文本LLM的音乐信息检索(MIR)任务的不同集合。这些包括流派分类,情感回归,情感标记,乐器分类,音高估计,关键检测,歌词转录,旋律提取,声乐技术识别,乐器演奏技术检测,音乐标记,音乐字幕和(下)节拍跟踪:反映了MIR研究的核心挑战。与以前的基准不同,CMI-Bench采用了与以前最先进的MIR模型一致的标准化评估指标,确保与监督方法的直接可比性。我们提供了一个评估工具包,支持所有开源音频文本LLM,包括LTU,Qwen-audio,SALMONN,MusiLingo等实验结果显示LLM和监督模型之间存在显着的性能差距,以及它们的文化,时间顺序和性别偏见,突出了当前模型在解决MIR任务方面的潜力和局限性。CMI-Bench为评估音乐教学建立了统一的基础,推动了音乐感知LLM的进步。
摘要:Recent advances in audio-text large language models (LLMs) have opened new possibilities for music understanding and generation. However, existing benchmarks are limited in scope, often relying on simplified tasks or multi-choice evaluations that fail to reflect the complexity of real-world music analysis. We reinterpret a broad range of traditional MIR annotations as instruction-following formats and introduce CMI-Bench, a comprehensive music instruction following benchmark designed to evaluate audio-text LLMs on a diverse set of music information retrieval (MIR) tasks. These include genre classification, emotion regression, emotion tagging, instrument classification, pitch estimation, key detection, lyrics transcription, melody extraction, vocal technique recognition, instrument performance technique detection, music tagging, music captioning, and (down)beat tracking: reflecting core challenges in MIR research. Unlike previous benchmarks, CMI-Bench adopts standardized evaluation metrics consistent with previous state-of-the-art MIR models, ensuring direct comparability with supervised approaches. We provide an evaluation toolkit supporting all open-source audio-textual LLMs, including LTU, Qwen-audio, SALMONN, MusiLingo, etc. Experiment results reveal significant performance gaps between LLMs and supervised models, along with their culture, chronological and gender bias, highlighting the potential and limitations of current models in addressing MIR tasks. CMI-Bench establishes a unified foundation for evaluating music instruction following, driving progress in music-aware LLMs.
标题: 用于无序语音分析的无缝不流畅语音文本对齐
链接:https://arxiv.org/abs/2506.12073
备注:Accepted for Interspeech2025
摘要:不流利语音与预期文本的准确对齐对于神经退行性语音障碍的自动诊断至关重要。传统的方法往往不能有效地模拟音素相似性,限制了它们的性能。在这项工作中,我们提出了神经LCS,一种新的方法不流畅的文本文本和语音文本对齐。神经LCS通过利用强大的音素级建模来解决关键挑战,包括部分对齐和上下文感知相似性映射。我们评估我们的方法在一个大规模的模拟数据集,使用先进的数据模拟技术,和真正的PPA数据。神经LCS在对齐准确性和不流利语音分割方面均显着优于最先进的模型。我们的研究结果表明,神经LCS有潜力增强诊断和分析语音障碍的自动化系统,为不流利的语音对齐提供更准确和更有语言基础的解决方案。
摘要:Accurate alignment of dysfluent speech with intended text is crucial for automating the diagnosis of neurodegenerative speech disorders. Traditional methods often fail to model phoneme similarities effectively, limiting their performance. In this work, we propose Neural LCS, a novel approach for dysfluent text-text and speech-text alignment. Neural LCS addresses key challenges, including partial alignment and context-aware similarity mapping, by leveraging robust phoneme-level modeling. We evaluate our method on a large-scale simulated dataset, generated using advanced data simulation techniques, and real PPA data. Neural LCS significantly outperforms state-of-the-art models in both alignment accuracy and dysfluent speech segmentation. Our results demonstrate the potential of Neural LCS to enhance automated systems for diagnosing and analyzing speech disorders, offering a more accurate and linguistically grounded solution for dysfluent speech alignment.
标题: 评估基于Logit的GOP分数以检测发音错误
链接:https://arxiv.org/abs/2506.12067
备注:Accepted to Interspeech 2025. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)
摘要:发音评估依赖于发音良好度(GOP)分数,传统上来自基于softmax的后验概率。然而,后验概率可能会受到过度自信和音素分离不良的影响,从而限制了它们的有效性。本研究比较了基于logit的GOP分数与基于概率的GOP分数的发音错误检测。我们在荷兰人和普通话使用者的两个L2英语语音数据集上进行了实验,评估了分类性能和与人类评级的相关性。基于Logit的方法在分类方面优于基于概率的GOP,但它们的有效性取决于数据集的特征。最大logit GOP显示出与人类感知的最强对齐,而不同GOP分数的组合平衡了概率和logit特征。研究结果表明,混合GOP的方法,结合不确定性建模和音素特定的加权改善发音评估。
摘要:Pronunciation assessment relies on goodness of pronunciation (GOP) scores, traditionally derived from softmax-based posterior probabilities. However, posterior probabilities may suffer from overconfidence and poor phoneme separation, limiting their effectiveness. This study compares logit-based GOP scores with probability-based GOP scores for mispronunciation detection. We conducted our experiment on two L2 English speech datasets spoken by Dutch and Mandarin speakers, assessing classification performance and correlation with human ratings. Logit-based methods outperform probability-based GOP in classification, but their effectiveness depends on dataset characteristics. The maximum logit GOP shows the strongest alignment with human perception, while a combination of different GOP scores balances probability and logit features. The findings suggest that hybrid GOP methods incorporating uncertainty modeling and phoneme-specific weighting improve pronunciation assessment.
标题: CMT-LLM:利用大型语言模型的上下文多说话者ASB
链接:https://arxiv.org/abs/2506.12059
备注:Accepted by INTERSPEECH 2025
摘要:在实际应用中,自动语音识别(ASR)系统必须处理来自多个说话者的重叠语音,并识别技术术语等罕见单词。传统方法分别解决多说话者ASR和上下文偏置,限制了复杂场景中的性能。我们提出了一个统一的框架,结合多说话者重叠语音识别和上下文偏置到一个单一的任务。我们的ASR方法集成了预训练的语音编码器和大型语言模型(LLM),使用优化的微调策略。我们还引入了一个两阶段的过滤算法,以有效地识别相关的稀有词从大型偏置列表,并将它们纳入LLM的提示输入,提高稀有词识别。实验表明,我们的方法优于传统的上下文偏置方法,实现了7.9%的WER LibriMix和32.9%的AMI SDM偏置大小为1,000时,证明了其在复杂的语音场景的有效性。
摘要:In real-world applications, automatic speech recognition (ASR) systems must handle overlapping speech from multiple speakers and recognize rare words like technical terms. Traditional methods address multi-talker ASR and contextual biasing separately, limiting performance in complex scenarios. We propose a unified framework that combines multi-talker overlapping speech recognition and contextual biasing into a single task. Our ASR method integrates pretrained speech encoders and large language models (LLMs), using optimized finetuning strategies. We also introduce a two-stage filtering algorithm to efficiently identify relevant rare words from large biasing lists and incorporate them into the LLM's prompt input, enhancing rare word recognition. Experiments show that our approach outperforms traditional contextual biasing methods, achieving a WER of 7.9% on LibriMix and 32.9% on AMI SDM when the biasing size is 1,000, demonstrating its effectiveness in complex speech scenarios.
2.eess.AS音频处理:
标题: SpeechRefiner:前端算法的感知质量细化
链接:https://arxiv.org/abs/2506.13709
备注:Accepted by Interspeech 2025
摘要:诸如去噪、去混响和分离的语音预处理技术通常被用作各种下游语音处理任务的前端。然而,这些方法有时可能是不够的,导致残留噪声或引入新的伪影。这样的缺陷通常不会被像SI-SNR这样的度量捕获,但是对于人类收听者来说是明显的。为了解决这个问题,我们引入了SpeechRefiner,这是一种后处理工具,它利用条件流匹配(CFM)来提高语音的感知质量。在这项研究中,我们将SpeechRefiner与最近的特定于任务的细化方法进行基准测试,并在我们的内部处理管道中评估其性能,该管道集成了多个前端算法。实验表明,SpeechRefiner在不同的损伤源中表现出很强的泛化能力,显著提高了语音感知质量。音频演示可以在https://speechrefiner.github.io/SpeechRefiner/上找到。
摘要:Speech pre-processing techniques such as denoising, de-reverberation, and separation, are commonly employed as front-ends for various downstream speech processing tasks. However, these methods can sometimes be inadequate, resulting in residual noise or the introduction of new artifacts. Such deficiencies are typically not captured by metrics like SI-SNR but are noticeable to human listeners. To address this, we introduce SpeechRefiner, a post-processing tool that utilizes Conditional Flow Matching (CFM) to improve the perceptual quality of speech. In this study, we benchmark SpeechRefiner against recent task-specific refinement methods and evaluate its performance within our internal processing pipeline, which integrates multiple front-end algorithms. Experiments show that SpeechRefiner exhibits strong generalization across diverse impairment sources, significantly enhancing speech perceptual quality. Audio demos can be found at https://speechrefiner.github.io/SpeechRefiner/.
标题: 基于PSELDnet预训练和BiMamba序列建模的立体声事件定位和检测
链接:https://arxiv.org/abs/2506.13455
备注:Technical report for DCASE 2025 Challenge Task 3
摘要:预训练方法在声音事件定位和检测(SELD)任务中取得了显着的性能改进,但现有的基于transformer的模型具有较高的计算复杂度。在这项工作中,我们提出了一个立体声声音事件定位和检测系统的基础上预先训练的PSELDnet和双向曼巴序列建模。我们用BiMamba模块替换Conformer模块,并引入非对称卷积来更有效地建模时间和频率维度之间的时空关系。在DCASE2025任务3开发数据集上的实验结果表明,该方法在降低计算复杂度的同时,显著优于基线和原始的PSELDnet与Conformer解码器架构。这些发现突出了BiMamba架构在应对SELD任务挑战方面的有效性。
摘要:Pre-training methods have achieved significant performance improvements in sound event localization and detection (SELD) tasks, but existing Transformer-based models suffer from high computational complexity. In this work, we propose a stereo sound event localization and detection system based on pre-trained PSELDnet and bidirectional Mamba sequence modeling. We replace the Conformer module with a BiMamba module and introduce asymmetric convolutions to more effectively model the spatiotemporal relationships between time and frequency dimensions. Experimental results demonstrate that the proposed method achieves significantly better performance than the baseline and the original PSELDnet with Conformer decoder architecture on the DCASE2025 Task 3 development dataset, while also reducing computational complexity. These findings highlight the effectiveness of the BiMamba architecture in addressing the challenges of the SELD task.
标题: BUT针对MLC-LAM挑战的系统
链接:https://arxiv.org/abs/2506.13414
摘要:我们提出了一个双说话人自动语音识别(ASR)系统,该系统将DiCoW(Whisper的diarization-conditioned变体)与DiariZen(构建在Pyannote之上的diarization管道)相结合。我们首先评估这两个系统在域外(OOD)的多语言场景没有任何微调。在这种情况下,DiariZen始终优于基准Pyannote日志化模型,表现出强大的泛化能力。尽管针对目标说话者ASR的仅英语数据进行了微调,但DiCoW仍保持了坚实的多语言性能,这表明编码器修改保留了Whisper的多语言功能。然后,我们在MLC-SLM挑战数据上微调DiCoW和DiariZen。经过微调的DiariZen继续优于经过微调的Pyannote基线,而DiCoW则从域适应中看到了进一步的收益。我们的最终系统实现了16.75%的微平均tcpWER/CER,并在MLC-SLM挑战的任务2中排名第二。最后,我们确定了训练数据中的几个标签不一致性-例如缺失的语音片段和不正确的沉默注释-这可能会阻碍日记微调。我们提出了简单的缓解策略来解决这些问题,提高系统的鲁棒性。
摘要:We present a two-speaker automatic speech recognition (ASR) system that combines DiCoW -- a diarization-conditioned variant of Whisper -- with DiariZen, a diarization pipeline built on top of Pyannote. We first evaluate both systems in out-of-domain (OOD) multilingual scenarios without any fine-tuning. In this scenario, DiariZen consistently outperforms the baseline Pyannote diarization model, demonstrating strong generalization. Despite being fine-tuned on English-only data for target-speaker ASR, DiCoW retains solid multilingual performance, indicating that encoder modifications preserve Whisper's multilingual capabilities. We then fine-tune both DiCoW and DiariZen on the MLC-SLM challenge data. The fine-tuned DiariZen continues to outperform the fine-tuned Pyannote baseline, while DiCoW sees further gains from domain adaptation. Our final system achieves a micro-average tcpWER/CER of 16.75% and ranks second in Task 2 of the MLC-SLM challenge. Lastly, we identify several labeling inconsistencies in the training data -- such as missing speech segments and incorrect silence annotations -- which can hinder diarization fine-tuning. We propose simple mitigation strategies to address these issues and improve system robustness.
标题: 野外语音编辑的特定实例测试时训练
链接:https://arxiv.org/abs/2506.13295
备注:Submitted to IEEE Signal Processing Letters
摘要:语音编辑系统旨在自然修改语音内容,同时保留声学一致性和说话者身份。然而,以前的研究往往难以适应看不见的和不同的声学条件,导致在现实世界中的编辑性能下降。为了解决这个问题,我们提出了一个特定于实例的测试时间训练方法,用于在野外进行语音编辑。我们的方法采用直接监督从地面实况声学特征在未经编辑的地区,和间接监督在编辑的地区通过辅助损失的基础上持续时间的限制和音素预测。该策略缓解了语音编辑中的带宽不连续问题,确保了未编辑区域和编辑区域之间的平滑声学过渡。此外,它通过在测试时间训练期间调整掩码长度使模型适应目标持续时间,从而实现对语音速率的精确控制。在野外基准数据集上的实验表明,我们的方法优于现有的语音编辑系统在客观和主观评价。
摘要:Speech editing systems aim to naturally modify speech content while preserving acoustic consistency and speaker identity. However, previous studies often struggle to adapt to unseen and diverse acoustic conditions, resulting in degraded editing performance in real-world scenarios. To address this, we propose an instance-specific test-time training method for speech editing in the wild. Our approach employs direct supervision from ground-truth acoustic features in unedited regions, and indirect supervision in edited regions via auxiliary losses based on duration constraints and phoneme prediction. This strategy mitigates the bandwidth discontinuity problem in speech editing, ensuring smooth acoustic transitions between unedited and edited regions. Additionally, it enables precise control over speech rate by adapting the model to target durations via mask length adjustment during test-time training. Experiments on in-the-wild benchmark datasets demonstrate that our method outperforms existing speech editing systems in both objective and subjective evaluations.
标题: 边界知情的声学场重建
链接:https://arxiv.org/abs/2506.13279
备注:Accepted for publication at EUSIPCO 2025
摘要:我们考虑的问题,重建的声场在一个房间里使用先验信息的边界几何形状,表示为一个点云。通常,当没有边界信息可用时,在大空间区域上并且在高频下的精确声场重建需要大量麦克风测量。另一方面,如果边界的所有几何和声学方面都是已知的,那么理论上可以在没有任何测量的情况下模拟声场。在这项工作中,我们解决中间的情况下,只有部分或不确定的边界信息。此设置类似于虚拟现实应用中研究的设置,其目标是创建感知上令人信服的音频体验。在这项工作中,我们专注于空间声音控制应用程序,相比之下,需要一个准确的声场重建。因此,我们制定了一个线性贝叶斯框架内的问题,将来自阻抗边界条件的边界知情的先验。该公式允许未知超参数的联合优化,包括噪声和信号方差以及阻抗边界条件。使用数值实验,我们表明,将边界知情的先验显着增强重建,特别是即使只有几百个边界点可用,或当边界位置校准的不确定性高达1 dm。
摘要:We consider the problem of reconstructing the sound field in a room using prior information of the boundary geometry, represented as a point cloud. In general, when no boundary information is available, an accurate sound field reconstruction over a large spatial region and at high frequencies requires numerous microphone measurements. On the other hand, if all geometrical and acoustical aspects of the boundaries are known, the sound field could, in theory, be simulated without any measurements. In this work, we address the intermediate case, where only partial or uncertain boundary information is available. This setting is similar to one studied in virtual reality applications, where the goal is to create a perceptually convincing audio experience. In this work, we focus on spatial sound control applications, which in contrast require an accurate sound field reconstruction. Therefore, we formulate the problem within a linear Bayesian framework, incorporating a boundary-informed prior derived from impedance boundary conditions. The formulation allows for joint optimization of the unknown hyperparameters, including the noise and signal variances and the impedance boundary conditions. Using numerical experiments, we show that incorporating the boundary-informed prior significantly enhances the reconstruction, notably even when only a few hundreds of boundary points are available or when the boundary positions are calibrated with an uncertainty up to 1 dm.
标题: ZipVoice:具有流匹配的快速高质量Zero-Shot文本到语音
链接:https://arxiv.org/abs/2506.13053
摘要:现有的大规模zero-shot文本到语音(TTS)模型提供高的语音质量,但遭受缓慢的推理速度,由于大量的参数。为了解决这个问题,本文介绍了ZipVoice,一个高质量的流匹配的zero-shot TTS模型,具有紧凑的模型大小和快速的推理速度。主要设计包括:1)基于Zipformer的流匹配解码器,以在受限大小下保持足够的建模能力; 2)基于平均上采样的初始语音-文本对齐和基于Zipformer的文本编码器,以提高语音可懂度; 3)流蒸馏方法,以减少采样步骤并消除与无分类器指导相关联的推理开销。在10万小时多语言数据集上的实验表明,ZipVoice在语音质量上与最先进的模型相匹配,同时比基于DiT的流匹配基线小3倍,快30倍。代码、模型检查点和演示样本都是公开的。
摘要:Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-quality flow-matching-based zero-shot TTS model with a compact model size and fast inference speed. Key designs include: 1) a Zipformer-based flow-matching decoder to maintain adequate modeling capabilities under constrained size; 2) Average upsampling-based initial speech-text alignment and Zipformer-based text encoder to improve speech intelligibility; 3) A flow distillation method to reduce sampling steps and eliminate the inference overhead associated with classifier-free guidance. Experiments on 100k hours multilingual datasets show that ZipVoice matches state-of-the-art models in speech quality, while being 3 times smaller and up to 30 times faster than a DiT-based flow-matching baseline. Codes, model checkpoints and demo samples are publicly available.
标题: 基于脑磁图(MEG)的无创中文语音解码
链接:https://arxiv.org/abs/2506.12817
摘要:语音脑机接口作为一种新兴的脑机接口模式,具有直接反映听觉感知和思维的潜力,为失语症患者提供了一种很有前途的交流方式。汉语是世界上使用最广泛的语言之一,而汉语语音脑机接口的研究却非常有限。本文报告了一个用于无创汉语语音脑机接口的文本脑磁图(MEG)数据集。它还提出了一种多模态辅助语音解码(MASD)算法,以捕获在语音活动过程中嵌入在大脑信号中的文本和声学信息。实验结果表明,我们的文本MEG数据集和我们提出的MASD算法的有效性。据我们所知,这是第一个研究模态辅助解码的非侵入性语音脑机接口。
摘要:As an emerging paradigm of brain-computer interfaces (BCIs), speech BCI has the potential to directly reflect auditory perception and thoughts, offering a promising communication alternative for patients with aphasia. Chinese is one of the most widely spoken languages in the world, whereas there is very limited research on speech BCIs for Chinese language. This paper reports a text-magnetoencephalography (MEG) dataset for non-invasive Chinese speech BCIs. It also proposes a multi-modality assisted speech decoding (MASD) algorithm to capture both text and acoustic information embedded in brain signals during speech activities. Experiment results demonstrated the effectiveness of both our text-MEG dataset and our proposed MASD algorithm. To our knowledge, this is the first study on modality-assisted decoding for non-invasive speech BCIs.
标题: 用于声音事件检测的频率动态卷积
链接:https://arxiv.org/abs/2506.12785
备注:Ph. D. Dissertation in English(KAIST)
摘要:最近基于深度学习的声音事件检测(SED)的研究主要集中在卷积递归神经网络(CRNN)和Transformer模型上。然而,传统的基于2D卷积的模型假设沿时间和频率轴的移位不变性,导致在处理声信号的频率相关特性时的不一致性。为了解决这个问题,本研究提出了频率动态卷积(FDY conv),它根据输入信号的频率组成动态调整卷积核,以提高SED性能。FDY conv通过基于频率特定注意力权重的自适应加权多个基核来构造最优频率响应。实验结果表明,将FDY conv应用于CRNN,与基线CRNN相比,DESED数据集的性能提高了7.56%。然而,FDY conv具有局限性,因为它在所有频率上组合了相同形状的基础内核,限制了其捕获不同频率特定特性的能力。此外,3\times3 $ basis内核大小不足以捕获更宽的频率范围。为了克服这些局限性,本研究介绍了一个扩展的家庭FDY转换模型。Dilated FDY conv(DFD conv)应用具有各种膨胀率的卷积核来沿着频率轴扩展感受野并增强频率特定的特征表示。实验结果表明,DFD conv比基线提高了9.27%的性能。部分FDY conv(PFD conv)解决了FDY conv的高计算成本,这是由于使用动态内核执行所有卷积运算而导致的。由于FDY conv可能会为准平稳声音事件引入不必要的自适应性,因此PFD conv将标准2D卷积与频率自适应内核集成在一起,以降低计算复杂度,同时保持性能。实验结果表明,PFD conv提高了7.80%的基线性能,同时减少了54.4%的FDY conv相比,参数的数量。多重扩张FDY conv(MDFD conv)通过解决其在所有频率上应用相同扩张的结构限制来扩展DFD conv。通过利用具有不同膨胀率的多个卷积核,MDFD conv有效地捕获了不同的频率相关模式。实验结果表明,MDFD conv实现了最高的性能,提高了10.98%的基准CRNN的性能。此外,标准FDY conv采用时间平均池化,其沿时间轴向所有帧分配相等权重,限制了其有效捕获瞬态事件的能力。为了克服这一点,本研究提出了TAP-FDY conv(TFD conv),它集成了专注于显著特征的时间注意力池(TA),强调瞬态特征的速度注意力池(VA)和捕获静态属性的平均池(AP)。TAP-FDY conv实现了与MDFD conv相同的性能,但减少了约30.01%的参数数量(12.703 M vs. 18.157 M),实现了相同的精度和更低的计算复杂度。类的性能分析表明,FDY conv提高了非平稳事件的检测,DFD conv是特别有效的事件具有广泛的频谱特征,和PFD conv增强了准平稳事件的检测。此外,TFD conv(TFD-CRNN)在检测瞬态事件方面表现出强大的性能。在案例研究中,PFD conv在坦克动力系统故障识别中有效地捕获稳定的信号模式,DFD conv在变速电机故障识别中识别宽谐波谱模式,而TFD conv在检测海上电弧检测中的瞬态信号方面优于其他模型。这些结果表明,频率自适应卷积及其扩展变体为基于深度学习的音频处理中的传统2D卷积提供了一种强大的替代方案。
摘要:Recent research in deep learning-based Sound Event Detection (SED) has primarily focused on Convolutional Recurrent Neural Networks (CRNNs) and Transformer models. However, conventional 2D convolution-based models assume shift invariance along both the temporal and frequency axes, leadin to inconsistencies when dealing with frequency-dependent characteristics of acoustic signals. To address this issue, this study proposes Frequency Dynamic Convolution (FDY conv), which dynamically adjusts convolutional kernels based on the frequency composition of the input signal to enhance SED performance. FDY conv constructs an optimal frequency response by adaptively weighting multiple basis kernels based on frequency-specific attention weights. Experimental results show that applying FDY conv to CRNNs improves performance on the DESED dataset by 7.56% compared to the baseline CRNN. However, FDY conv has limitations in that it combines basis kernels of the same shape across all frequencies, restricting its ability to capture diverse frequency-specific characteristics. Additionally, the $3\times3$ basis kernel size is insufficient to capture a broader frequency range. To overcome these limitations, this study introduces an extended family of FDY conv models. Dilated FDY conv (DFD conv) applies convolutional kernels with various dilation rates to expand the receptive field along the frequency axis and enhance frequency-specific feature representation. Experimental results show that DFD conv improves performance by 9.27% over the baseline. Partial FDY conv (PFD conv) addresses the high computational cost of FDY conv, which results from performing all convolution operations with dynamic kernels. Since FDY conv may introduce unnecessary adaptivity for quasi-stationary sound events, PFD conv integrates standard 2D convolutions with frequency-adaptive kernels to reduce computational complexity while maintaining performance. Experimental results demonstrate that PFD conv improves performance by 7.80% over the baseline while reducing the number of parameters by 54.4% compared to FDY conv. Multi-Dilated FDY conv (MDFD conv) extends DFD conv by addressing its structural limitation of applying the same dilation across all frequencies. By utilizing multiple convolutional kernels with different dilation rates, MDFD conv effectively captures diverse frequency-dependent patterns. Experimental results indicate that MDFD conv achieves the highest performance, improving the baseline CRNN performance by 10.98%. Furthermore, standard FDY conv employs Temporal Average Pooling, which assigns equal weight to all frames along the time axis, limiting its ability to effectively capture transient events. To overcome this, this study proposes TAP-FDY conv (TFD conv), which integrates Temporal Attention Pooling (TA) that focuses on salient features, Velocity Attention Pooling (VA) that emphasizes transient characteristics, and Average Pooling (AP) that captures stationary properties. TAP-FDY conv achieves the same performance as MDFD conv but reduces the number of parameters by approximately 30.01% (12.703M vs. 18.157M), achieving equivalent accuracy with lower computational complexity. Class-wise performance analysis reveals that FDY conv improves detection of non-stationary events, DFD conv is particularly effective for events with broad spectral features, and PFD conv enhances the detection of quasi-stationary events. Additionally, TFD conv (TFD-CRNN) demonstrates strong performance in detecting transient events. In the case studies, PFD conv effectively captures stable signal patterns in tank powertrain fault recognition, DFD conv recognizes wide harmonic spectral patterns on speed-varying motor fault recognition, while TFD conv outperforms other models in detecting transient signals in offshore arc detection. These results suggest that frequency-adaptive convolutions and their extended variants provide a robust alternative to conventional 2D convolutions in deep learning-based audio processing.
标题: 使用神经图相似性指数测量(NSIM)建模听力损失和腰神经退行性变
链接:https://arxiv.org/abs/2506.12705
备注:Accepted for presentation at INTERSPEECH 2025
摘要:在嘈杂环境中的听力障碍仍然是听力损失和听力正常的人的共同投诉。这被假设是由于称为耳蜗神经变性(CND)的情况而引起的,CND也会导致助听器结果的显著变化。本文使用听觉周边的计算模型来模拟各种听觉任务。我们提出了一种客观的方法来量化听力损失和CND比较听觉神经纤维反应使用神经图相似性指数测量(NSIM)。具体来说,研究1表明,NSIM可以用于映射个人的听力损失的音素识别任务的性能与合理的准确性。在研究2中,我们表明NSIM是一种敏感的测量方法,也可以用于捕获CND导致的缺陷,并且可以作为听觉突触病的非侵入性生物标志物的候选者。
摘要:Trouble hearing in noisy situations remains a common complaint for both individuals with hearing loss and individuals with normal hearing. This is hypothesized to arise due to condition called: cochlear neural degeneration (CND) which can also result in significant variabilities in hearing aids outcomes. This paper uses computational models of auditory periphery to simulate various hearing tasks. We present an objective method to quantify hearing loss and CND by comparing auditory nerve fiber responses using a Neurogram Similarity Index Measure (NSIM). Specifically study 1, shows that NSIM can be used to map performance of individuals with hearing loss on phoneme recognition task with reasonable accuracy. In the study 2, we show that NSIM is a sensitive measure that can also be used to capture the deficits resulting from CND and can be a candidate for noninvasive biomarker of auditory synaptopathy.
标题: 走向神经音频编解码器源解析
链接:https://arxiv.org/abs/2506.12627
摘要:一类新的音频deepfakes-codecfakes(CF)-最近引起了人们的注意,它是由音频语言模型合成的,在后端利用神经音频编解码器(NAC)。作为回应,社区引入了专门的基准和量身定制的检测策略。随着该领域的发展,人们的努力已经超越了二元检测,转向了源归因,包括开集归因,其目的是识别负责生成的NAC,并在推理过程中标记新的、看不见的NAC。这种向来源归因的转变提高了法医的可解释性和问责制。然而,开集属性仍然存在根本性的局限性:虽然它可以检测到NAC是不熟悉的,但它无法表征或识别单个看不见的编解码器。它将这些输入视为一般的“未知数”,缺乏对其内部配置的洞察力。这导致了主要的缺点:对新NAC的推广有限,并且无法解决NAC家族内的细粒度变化。为了解决这些差距,我们提出了神经音频编解码器源解析(NACSP)-一种范式转变,将CF的源属性重新构建为生成NAC参数(如量化器,带宽和采样率)的结构化回归。我们将NACSP制定为用于预测这些NAC参数的多任务回归任务,并使用各种最先进的语音预训练模型(PTM)建立第一个全面的基准。为此,我们提出了HYDRONIC,一个新的框架,利用双曲几何解开复杂的潜在属性PTM表示。通过在多个曲率感知的双曲子空间上使用特定于任务的注意力,HYDRONIC实现了卓越的多任务泛化。我们广泛的实验表明,与在欧几里得空间中操作的基线相比,HYDRONIC在基准CFs数据集上获得了最佳结果。
摘要:A new class of audio deepfakes-codecfakes (CFs)-has recently caught attention, synthesized by Audio Language Models that leverage neural audio codecs (NACs) in the backend. In response, the community has introduced dedicated benchmarks and tailored detection strategies. As the field advances, efforts have moved beyond binary detection toward source attribution, including open-set attribution, which aims to identify the NAC responsible for generation and flag novel, unseen ones during inference. This shift toward source attribution improves forensic interpretability and accountability. However, open-set attribution remains fundamentally limited: while it can detect that a NAC is unfamiliar, it cannot characterize or identify individual unseen codecs. It treats such inputs as generic ``unknowns'', lacking insight into their internal configuration. This leads to major shortcomings: limited generalization to new NACs and inability to resolve fine-grained variations within NAC families. To address these gaps, we propose Neural Audio Codec Source Parsing (NACSP) - a paradigm shift that reframes source attribution for CFs as structured regression over generative NAC parameters such as quantizers, bandwidth, and sampling rate. We formulate NACSP as a multi-task regression task for predicting these NAC parameters and establish the first comprehensive benchmark using various state-of-the-art speech pre-trained models (PTMs). To this end, we propose HYDRA, a novel framework that leverages hyperbolic geometry to disentangle complex latent properties from PTM representations. By employing task-specific attention over multiple curvature-aware hyperbolic subspaces, HYDRA enables superior multi-task generalization. Our extensive experiments show HYDRA achieves top results on benchmark CFs datasets compared to baselines operating in Euclidean space.
标题: 缓解引导说话人嵌入中的非目标说话人偏见
链接:https://arxiv.org/abs/2506.12500
备注:Accepted to Interspeech 2025
摘要:在多说话人环境中获得高质量的说话人嵌入对于许多应用来说至关重要。最近提出的一种引导式说话人嵌入框架,利用目标和非目标说话人的语音活动作为线索,大大提高了严重重叠下的嵌入,在低重叠的情况下有小的退化。然而,由于极端的重叠在自然对话中是罕见的,这种退化不能被忽视。本文首先揭示了退化是由于广泛用于说话人嵌入提取器的基于全局统计的模块对仅包含非目标说话人的间隔过于敏感。作为一种对策,我们提出了这样的模块,利用目标扬声器活动线索的扩展,从目标是活跃的间隔计算统计。该方法提高了说话人确认在低重叠率和高重叠率下的性能,以及多数据集上的日志化性能。
摘要:Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target speakers as clues, drastically improved embeddings under severe overlap with small degradation in low-overlap cases. However, since extreme overlaps are rare in natural conversations, this degradation cannot be overlooked. This paper first reveals that the degradation is caused by the global-statistics-based modules, widely used in speaker embedding extractors, being overly sensitive to intervals containing only non-target speakers. As a countermeasure, we propose an extension of such modules that exploit the target speaker activity clues, to compute statistics from intervals where the target is active. The proposed method improves speaker verification performance in both low and high overlap ratios, and diarization performance on multiple datasets.
标题: CMI-Bench:评估音乐教学遵循的综合基准
链接:https://arxiv.org/abs/2506.12285
备注:Accepted by ISMIR 2025
摘要:音频文本大语言模型(LLM)的最新进展为音乐理解和生成开辟了新的可能性。然而,现有的基准在范围上是有限的,往往依赖于简化的任务或多选择的评估,无法反映现实世界的音乐分析的复杂性。我们重新解释了广泛的传统的MIR注释作为解释以下格式,并介绍了CMI-Bench,一个全面的音乐指令以下基准,旨在评估音频文本LLM的音乐信息检索(MIR)任务的不同集合。这些包括流派分类,情感回归,情感标记,乐器分类,音高估计,关键检测,歌词转录,旋律提取,声乐技术识别,乐器演奏技术检测,音乐标记,音乐字幕和(下)节拍跟踪:反映了MIR研究的核心挑战。与以前的基准不同,CMI-Bench采用了与以前最先进的MIR模型一致的标准化评估指标,确保与监督方法的直接可比性。我们提供了一个评估工具包,支持所有开源音频文本LLM,包括LTU,Qwen-audio,SALMONN,MusiLingo等实验结果显示LLM和监督模型之间存在显着的性能差距,以及它们的文化,时间顺序和性别偏见,突出了当前模型在解决MIR任务方面的潜力和局限性。CMI-Bench为评估音乐教学建立了统一的基础,推动了音乐感知LLM的进步。
摘要:Recent advances in audio-text large language models (LLMs) have opened new possibilities for music understanding and generation. However, existing benchmarks are limited in scope, often relying on simplified tasks or multi-choice evaluations that fail to reflect the complexity of real-world music analysis. We reinterpret a broad range of traditional MIR annotations as instruction-following formats and introduce CMI-Bench, a comprehensive music instruction following benchmark designed to evaluate audio-text LLMs on a diverse set of music information retrieval (MIR) tasks. These include genre classification, emotion regression, emotion tagging, instrument classification, pitch estimation, key detection, lyrics transcription, melody extraction, vocal technique recognition, instrument performance technique detection, music tagging, music captioning, and (down)beat tracking: reflecting core challenges in MIR research. Unlike previous benchmarks, CMI-Bench adopts standardized evaluation metrics consistent with previous state-of-the-art MIR models, ensuring direct comparability with supervised approaches. We provide an evaluation toolkit supporting all open-source audio-textual LLMs, including LTU, Qwen-audio, SALMONN, MusiLingo, etc. Experiment results reveal significant performance gaps between LLMs and supervised models, along with their culture, chronological and gender bias, highlighting the potential and limitations of current models in addressing MIR tasks. CMI-Bench establishes a unified foundation for evaluating music instruction following, driving progress in music-aware LLMs.
标题: 用于无序语音分析的无缝不流畅语音文本对齐
链接:https://arxiv.org/abs/2506.12073
备注:Accepted for Interspeech2025
摘要:不流利语音与预期文本的准确对齐对于神经退行性语音障碍的自动诊断至关重要。传统的方法往往不能有效地模拟音素相似性,限制了它们的性能。在这项工作中,我们提出了神经LCS,一种新的方法不流畅的文本文本和语音文本对齐。神经LCS通过利用强大的音素级建模解决了关键挑战,包括部分对齐和上下文感知相似性映射。我们评估我们的方法在一个大规模的模拟数据集,使用先进的数据模拟技术,和真正的PPA数据。神经LCS在对齐准确性和不流利语音分割方面均显着优于最先进的模型。我们的研究结果表明,神经LCS有潜力增强诊断和分析语音障碍的自动化系统,为不流利的语音对齐提供更准确和更有语言基础的解决方案。
摘要:Accurate alignment of dysfluent speech with intended text is crucial for automating the diagnosis of neurodegenerative speech disorders. Traditional methods often fail to model phoneme similarities effectively, limiting their performance. In this work, we propose Neural LCS, a novel approach for dysfluent text-text and speech-text alignment. Neural LCS addresses key challenges, including partial alignment and context-aware similarity mapping, by leveraging robust phoneme-level modeling. We evaluate our method on a large-scale simulated dataset, generated using advanced data simulation techniques, and real PPA data. Neural LCS significantly outperforms state-of-the-art models in both alignment accuracy and dysfluent speech segmentation. Our results demonstrate the potential of Neural LCS to enhance automated systems for diagnosing and analyzing speech disorders, offering a more accurate and linguistically grounded solution for dysfluent speech alignment.
标题: 评估基于Logit的GOP分数以检测发音错误
链接:https://arxiv.org/abs/2506.12067
备注:Accepted to Interspeech 2025. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)
摘要:发音评估依赖于发音良好度(GOP)分数,传统上来自基于softmax的后验概率。然而,后验概率可能会受到过度自信和音素分离不良的影响,从而限制了它们的有效性。本研究比较了基于logit的GOP分数和基于概率的GOP分数来检测发音错误。我们在荷兰人和普通话使用者的两个L2英语语音数据集上进行了实验,评估了分类性能和与人类评级的相关性。基于Logit的方法在分类方面优于基于概率的GOP,但它们的有效性取决于数据集的特征。最大logit GOP显示出与人类感知的最强一致性,而不同GOP分数的组合平衡了概率和logit特征。研究结果表明,混合GOP的方法,结合不确定性建模和音素特定的加权改善发音评估。
摘要:Pronunciation assessment relies on goodness of pronunciation (GOP) scores, traditionally derived from softmax-based posterior probabilities. However, posterior probabilities may suffer from overconfidence and poor phoneme separation, limiting their effectiveness. This study compares logit-based GOP scores with probability-based GOP scores for mispronunciation detection. We conducted our experiment on two L2 English speech datasets spoken by Dutch and Mandarin speakers, assessing classification performance and correlation with human ratings. Logit-based methods outperform probability-based GOP in classification, but their effectiveness depends on dataset characteristics. The maximum logit GOP shows the strongest alignment with human perception, while a combination of different GOP scores balances probability and logit features. The findings suggest that hybrid GOP methods incorporating uncertainty modeling and phoneme-specific weighting improve pronunciation assessment.
标题: CMT-LLM:利用大型语言模型的上下文多说话者ASB
链接:https://arxiv.org/abs/2506.12059
备注:Accepted by INTERSPEECH 2025
摘要:在实际应用中,自动语音识别(ASR)系统必须处理来自多个说话者的重叠语音,并识别技术术语等罕见单词。传统方法分别解决多说话者ASR和上下文偏置,限制了复杂场景中的性能。我们提出了一个统一的框架,结合多说话者重叠语音识别和上下文偏置到一个单一的任务。我们的ASR方法集成了预训练的语音编码器和大型语言模型(LLM),使用优化的微调策略。我们还引入了一个两阶段的过滤算法,以有效地识别相关的稀有词从大型偏置列表,并将它们纳入LLM的提示输入,提高稀有词识别。实验表明,我们的方法优于传统的上下文偏置方法,实现了7.9%的WER LibriMix和32.9%的AMI SDM偏置大小为1,000时,证明了其在复杂的语音场景的有效性。
摘要:In real-world applications, automatic speech recognition (ASR) systems must handle overlapping speech from multiple speakers and recognize rare words like technical terms. Traditional methods address multi-talker ASR and contextual biasing separately, limiting performance in complex scenarios. We propose a unified framework that combines multi-talker overlapping speech recognition and contextual biasing into a single task. Our ASR method integrates pretrained speech encoders and large language models (LLMs), using optimized finetuning strategies. We also introduce a two-stage filtering algorithm to efficiently identify relevant rare words from large biasing lists and incorporate them into the LLM's prompt input, enhancing rare word recognition. Experiments show that our approach outperforms traditional contextual biasing methods, achieving a WER of 7.9% on LibriMix and 32.9% on AMI SDM when the biasing size is 1,000, demonstrating its effectiveness in complex speech scenarios.
标题: Stream-Omni:采用大范围-视觉-语音模型的同时多模式交互
链接:https://arxiv.org/abs/2506.13642
备注:Code: this https URL , Model: this https URL
摘要:GPT-40类大型多模态模型(LMRM)的出现引发了对整合文本、视觉和语音模态以支持更灵活的多模态交互的探索。现有的LLM通常沿着序列维度连接模态的表示,并将它们馈送到大型语言模型(LLM)主干中。虽然序列-维度连接对于模态集成是直接的,但它通常严重依赖于大规模数据来学习模态对齐。在本文中,我们的目标是更有目的地建模模态之间的关系,从而实现更有效和灵活的模态对齐。为此,我们提出了Stream-Omni,一个大型的语言-视觉-语音模型,具有高效的模态对齐,可以同时支持各种模态组合下的交互。Stream-Omni采用LLM作为骨干,并根据其关系将视觉和语音与文本对齐。对于在语义上与文本互补的视觉,Stream-Omni使用序列维度连接来实现视觉-文本对齐。对于与文本语义一致的语音,Stream-Omni引入了基于CTC的层-维度映射来实现语音-文本对齐。通过这种方式,Stream-Omni可以用更少的数据(特别是语音)实现模态对齐,从而将文本功能转移到其他模态。在各种基准测试上的实验表明,Stream-Omni在视觉理解、语音交互和基于视觉的语音交互任务上都取得了很好的性能。由于层维映射,Stream-Omni可以在语音交互过程中同时提供中间文本输出(如ASR transmittance和模型响应),为用户提供全面的多模态体验。
摘要:The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal interaction. Existing LMMs typically concatenate representation of modalities along the sequence dimension and feed them into a large language model (LLM) backbone. While sequence-dimension concatenation is straightforward for modality integration, it often relies heavily on large-scale data to learn modality alignments. In this paper, we aim to model the relationships between modalities more purposefully, thereby achieving more efficient and flexible modality alignments. To this end, we propose Stream-Omni, a large language-vision-speech model with efficient modality alignments, which can simultaneously support interactions under various modality combinations. Stream-Omni employs LLM as the backbone and aligns the vision and speech to the text based on their relationships. For vision that is semantically complementary to text, Stream-Omni uses sequence-dimension concatenation to achieve vision-text alignment. For speech that is semantically consistent with text, Stream-Omni introduces a CTC-based layer-dimension mapping to achieve speech-text alignment. In this way, Stream-Omni can achieve modality alignments with less data (especially speech), enabling the transfer of text capabilities to other modalities. Experiments on various benchmarks demonstrate that Stream-Omni achieves strong performance on visual understanding, speech interaction, and vision-grounded speech interaction tasks. Owing to the layer-dimensional mapping, Stream-Omni can simultaneously provide intermediate text outputs (such as ASR transcriptions and model responses) during speech interaction, offering users a comprehensive multimodal experience.
标题: Qwen与Gemma与Whisper的集成:多语言演讲LLM系统的比较研究
链接:https://arxiv.org/abs/2506.13596
备注:Technical report for Interspeech 2025 MLC-SLM Challenge
摘要:本文介绍了我们的MLC-SLM挑战2025系统,重点是多语言语音识别和大型语言模型(LLM)的语言建模。我们的方法结合了一个微调的耳语大v3编码器与高效的投影机架构和各种解码器配置。我们采用三阶段的培训方法,逐步优化编码器,投影仪和LLM组件。我们的系统实现了具有竞争力的性能,使用Gemma 3 - 12 B和18.6%使用Qwen2.5- 7 B作为仅解码器的语言模型的私人测试平均WER/CER结果为16.63%。
摘要:This paper presents our system for the MLC-SLM Challenge 2025, focusing on multilingual speech recognition and language modeling with large language models (LLMs). Our approach combines a fine-tuned Whisper-large-v3 encoder with efficient projector architectures and various decoder configurations. We employ a three-stage training methodology that progressively optimizes the encoder, projector, and LLM components. Our system achieves competitive performance with a private test average WER/CER result of 16.63% using the Gemma3-12B and 18.6% using the Qwen2.5-7B as decoder-only language model.
标题: 三种不同距离音乐网络的持久同源性
链接:https://arxiv.org/abs/2506.13595
摘要:持久同源性已被广泛用于发现跨各种应用的数据中隐藏的拓扑结构,包括音乐数据。要应用持久同源性,必须在点云中的点之间或图形网络中的节点之间定义距离或度量。这些定义不是唯一的,取决于给定问题的具体目标。换句话说,选择不同的度量定义允许多个拓扑推断。在这项工作中,我们专注于应用持久同源音乐图与预定义的权重。我们研究了三个不同的距离定义的基础上边明智的途径,并展示了这些定义如何影响持久性条形码,持久性图,出生/死亡边缘。我们发现这三种距离定义在一维持久同源性条码和持久同源性图上都存在包含关系。我们使用真实的音乐数据验证了这些发现。
摘要:Persistent homology has been widely used to discover hidden topological structures in data across various applications, including music data. To apply persistent homology, a distance or metric must be defined between points in a point cloud or between nodes in a graph network. These definitions are not unique and depend on the specific objectives of a given problem. In other words, selecting different metric definitions allows for multiple topological inferences. In this work, we focus on applying persistent homology to music graph with predefined weights. We examine three distinct distance definitions based on edge-wise pathways and demonstrate how these definitions affect persistent barcodes, persistence diagrams, and birth/death edges. We found that there exist inclusion relations in one-dimensional persistent homology reflected on persistence barcode and diagram among these three distance definitions. We verified these findings using real music data.
标题: 面向多语言会话ASR的双向上下文增强语音大语言模型
链接:https://arxiv.org/abs/2506.13396
备注:Submitted to Interspeech 2025 MLC-SLM workshop as a Research Paper
摘要:本文介绍了语言特定的双向上下文集成到语音大语言模型(SLLM),以提高多语种连续会话自动语音识别(ASR)。我们提出了一个字符级的上下文掩蔽策略在训练过程中,随机删除部分的上下文,以提高鲁棒性和更好地模拟有缺陷的transmittance可能发生在推理过程中。对于解码,利用两阶段流水线:初始隔离段解码,然后是使用相邻假设的上下文感知重新解码。在涵盖11种语言的1500小时多语言会话语音和语言模型(MLC-SLM)语料库上进行评估,与强大的基线相比,我们的方法实现了18%的相对改进,甚至超过了在MLC-SLM竞赛中使用6000小时数据训练的模型。这些结果强调了将上下文信息纳入多语言连续会话ASR的显着好处。
摘要:This paper introduces the integration of language-specific bi-directional context into a speech large language model (SLLM) to improve multilingual continuous conversational automatic speech recognition (ASR). We propose a character-level contextual masking strategy during training, which randomly removes portions of the context to enhance robustness and better emulate the flawed transcriptions that may occur during inference. For decoding, a two-stage pipeline is utilized: initial isolated segment decoding followed by context-aware re-decoding using neighboring hypotheses. Evaluated on the 1500-hour Multilingual Conversational Speech and Language Model (MLC-SLM) corpus covering eleven languages, our method achieves an 18% relative improvement compared to a strong baseline, outperforming even the model trained on 6000 hours of data for the MLC-SLM competition. These results underscore the significant benefit of incorporating contextual information in multilingual continuous conversational ASR.
标题: NTU Speechlab基于LLM的多语言ASB系统Interspeech MLC-LAM挑战赛2025
链接:https://arxiv.org/abs/2506.13339
备注:Submitted to Interspeech 2025 MLC-SLM challenge (5th place). System report
摘要:本报告详细介绍了为Interspeech 2025多语言对话语音和语言模型(MLC-SLM)挑战赛(任务I)开发的NTU Speechlab系统,我们在该挑战赛中获得了第五名。我们对我们的多语言自动语音识别系统进行了全面的分析,突出了模型架构,数据选择和训练策略的关键进展。特别是,语言特定的提示和模型平均技术有助于提高不同语言的系统性能。与初始基线系统相比,我们的最终模型将平均混合错误率从20.2%降低到10.6%,表示评估集的绝对改进为9.6%(相对改进为48%)。我们的结果证明了我们方法的有效性,并为未来的语音大语言模型提供了实用的见解。
摘要:This report details the NTU Speechlab system developed for the Interspeech 2025 Multilingual Conversational Speech and Language Model (MLC-SLM) Challenge (Task I), where we achieved 5th place. We present comprehensive analyses of our multilingual automatic speech recognition system, highlighting key advancements in model architecture, data selection, and training strategies. In particular, language-specific prompts and model averaging techniques were instrumental in boosting system performance across diverse languages. Compared to the initial baseline system, our final model reduced the average Mix Error Rate from 20.2% to 10.6%, representing an absolute improvement of 9.6% (a relative improvement of 48%) on the evaluation set. Our results demonstrate the effectiveness of our approach and offer practical insights for future Speech Large Language Models.
标题: I$#2$S-TFCKD:具有语音增强时频校准的内-间集知识提取
链接:https://arxiv.org/abs/2506.13127
备注:submitted to IEEE Transactions on Neural Networks and Learning Systems
摘要:近年来,基于神经网络(NN)的语音增强(SE)模型的复杂度压缩问题逐渐引起研究者的关注,特别是在硬件资源有限或时延要求严格的场景下。主要的困难和挑战在于根据任务的特点在复杂性和性能之间取得平衡。在本文中,我们提出了一个内集知识提取(KD)框架与时频校准(I$^2$S-TFCKD)的SE。与以往的提取策略不同,该框架充分利用了语音的时频差分信息,同时促进了全局知识流动。首先,提出了一种基于双流时频交叉校准的多层交互式蒸馏方法,该方法分别在时域和频域计算师生相似度校准权值并进行交叉加权,从而实现了根据语音特征在不同层间精细分配蒸馏贡献。其次,我们构建了一个集内和集间相关性的协同蒸馏范式。在一个相关的集合中,多层教师-学生特征被成对匹配以进行校准蒸馏。随后,我们产生的代表性特征,从每个相关的集合,通过残差融合,形成融合的特征集,使集合间的知识交互。提出的蒸馏策略应用于双路径扩张卷积递归网络(DPDCRN),该网络在L3 DAS 23挑战赛的SE赛道中排名第一。客观的评价表明,所提出的KD策略一致,有效地提高了低复杂度的学生模型的性能,并优于其他蒸馏计划。
摘要:In recent years, complexity compression of neural network (NN)-based speech enhancement (SE) models has gradually attracted the attention of researchers, especially in scenarios with limited hardware resources or strict latency requirements. The main difficulties and challenges lie in achieving a balance between complexity and performance according to the characteristics of the task. In this paper, we propose an intra-inter set knowledge distillation (KD) framework with time-frequency calibration (I$^2$S-TFCKD) for SE. Different from previous distillation strategies for SE, the proposed framework fully utilizes the time-frequency differential information of speech while promoting global knowledge flow. Firstly, we propose a multi-layer interactive distillation based on dual-stream time-frequency cross-calibration, which calculates the teacher-student similarity calibration weights in the time and frequency domains respectively and performs cross-weighting, thus enabling refined allocation of distillation contributions across different layers according to speech characteristics. Secondly, we construct a collaborative distillation paradigm for intra-set and inter-set correlations. Within a correlated set, multi-layer teacher-student features are pairwise matched for calibrated distillation. Subsequently, we generate representative features from each correlated set through residual fusion to form the fused feature set that enables inter-set knowledge interaction. The proposed distillation strategy is applied to the dual-path dilated convolutional recurrent network (DPDCRN) that ranked first in the SE track of the L3DAS23 challenge. Objective evaluations demonstrate that the proposed KD strategy consistently and effectively improves the performance of the low-complexity student model and outperforms other distillation schemes.
标题: 用MIDI-RWKV填充可个性化的长上下文象征性音乐
链接:https://arxiv.org/abs/2506.13001
摘要:自动音乐生成的现有工作主要集中在产生完整作品或延续的端到端系统上。然而,由于音乐创作通常是一个迭代的过程,这样的系统使得很难参与人与机器之间的来回,而这对计算机辅助创造力至关重要。在这项研究中,我们解决的任务个性化,多轨道,长的上下文,可控的符号音乐填充,以提高计算机辅助作曲的过程。我们提出了MIDI-RWKV,一种基于RWKV-7线性架构的新型模型,以实现边缘设备上高效和连贯的音乐共同创作。我们还表明,MIDI-RWKV承认微调其初始状态的个性化在非常低的样本制度的有效方法。我们评估MIDI-RWKV及其状态调整的几个定量和定性指标,并发布模型的权重和代码在https://github.com/christianazinn/MIDI-RWKV。
摘要:Existing work in automatic music generation has primarily focused on end-to-end systems that produce complete compositions or continuations. However, because musical composition is typically an iterative process, such systems make it difficult to engage in the back-and-forth between human and machine that is essential to computer-assisted creativity. In this study, we address the task of personalizable, multi-track, long-context, and controllable symbolic music infilling to enhance the process of computer-assisted composition. We present MIDI-RWKV, a novel model based on the RWKV-7 linear architecture, to enable efficient and coherent musical cocreation on edge devices. We also demonstrate that MIDI-RWKV admits an effective method of finetuning its initial state for personalization in the very-low-sample regime. We evaluate MIDI-RWKV and its state tuning on several quantitative and qualitative metrics, and release model weights and code at https://github.com/christianazinn/MIDI-RWKV.
标题: SoundMind:音频语言模型的RL激励逻辑推理
链接:https://arxiv.org/abs/2506.12935
摘要:虽然大型语言模型已经显示出推理能力,但它们在音频模态中的应用,特别是在大型音频语言模型(ALM)中,仍然显着不发达。解决这一差距需要一种系统的方法,包括一个有能力的基础模型,高质量的面向推理的音频数据和有效的训练算法。在这项研究中,我们提出了一个全面的解决方案:我们引入了音频逻辑推理(ALR)数据集,由6,446个专门为复杂推理任务设计的文本音频注释样本组成。在此基础上,我们提出了SoundMind,一种基于规则的强化学习(RL)算法,旨在赋予ALM深度双峰推理能力。通过使用SoundMind在ALR数据集上训练Qwen2.5-Omni-7 B,我们的方法在音频逻辑推理方面实现了最先进的性能。这项工作突出了将高质量的以推理为中心的数据集与专业RL技术相结合的影响,推进了语言模型中听觉智能的前沿。我们的代码和建议的数据集可以在https://github.com/xid32/SoundMind上找到。
摘要:While large language models have shown reasoning capabilities, their application to the audio modality, particularly in large audio-language models (ALMs), remains significantly underdeveloped. Addressing this gap requires a systematic approach, involving a capable base model, high-quality reasoning-oriented audio data, and effective training algorithms. In this study, we present a comprehensive solution: we introduce the Audio Logical Reasoning (ALR) dataset, consisting of 6,446 text-audio annotated samples specifically designed for complex reasoning tasks. Building on this resource, we propose SoundMind, a rule-based reinforcement learning (RL) algorithm tailored to endow ALMs with deep bimodal reasoning abilities. By training Qwen2.5-Omni-7B on the ALR dataset using SoundMind, our approach achieves state-of-the-art performance in audio logical reasoning. This work highlights the impact of combining high-quality, reasoning-focused datasets with specialized RL techniques, advancing the frontier of auditory intelligence in language models. Our code and the proposed dataset are available at https://github.com/xid32/SoundMind.
标题: SC-SOT:根据数字化说话人信息调节解码器,以实现端到端重叠语音识别
链接:https://arxiv.org/abs/2506.12672
备注:Accepted by Interspeech 2025
摘要:我们提出了扬声器条件串行输出训练(SC-SOT),这是一种针对E2 E多说话者ASR的增强型基于SOT的训练。我们首先探讨SOT如何处理重叠语音,我们发现解码器执行隐式说话人分离。我们假设这种隐含的分离往往是不够的,由于模糊的声学线索重叠区域。为了解决这个问题,SC-SOT显式地根据说话者信息来调节解码器,提供关于“谁在什么时候说话”的详细信息。具体来说,我们通过合并以下内容来增强解码器:(1)扬声器嵌入,其允许模型关注目标扬声器的声学特性,以及(2)扬声器活动信息,其引导模型抑制非目标扬声器。说话人嵌入是从联合训练的E2 E说话人日记化模型导出的,从而减轻了对说话人登记的需要。实验结果表明,我们的条件化方法对重叠语音的有效性。
摘要:We propose Speaker-Conditioned Serialized Output Training (SC-SOT), an enhanced SOT-based training for E2E multi-talker ASR. We first probe how SOT handles overlapped speech, and we found the decoder performs implicit speaker separation. We hypothesize this implicit separation is often insufficient due to ambiguous acoustic cues in overlapping regions. To address this, SC-SOT explicitly conditions the decoder on speaker information, providing detailed information about "who spoke when". Specifically, we enhance the decoder by incorporating: (1) speaker embeddings, which allow the model to focus on the acoustic characteristics of the target speaker, and (2) speaker activity information, which guides the model to suppress non-target speakers. The speaker embeddings are derived from a jointly trained E2E speaker diarization model, mitigating the need for speaker enrollment. Experimental results demonstrate the effectiveness of our conditioning approach on overlapped speech.
标题: ANIRA:实时音频应用中神经网络推理的架构
链接:https://arxiv.org/abs/2506.12665
备注:8 pages, accepted to the Proceedings of the 5th IEEE International Symposium on the Internet of Sounds (2024) - repository: github.com/anira-project/anira
摘要:目前有许多用于神经网络推理的工具,但许多工具并不满足实时音频应用的要求。作为回应,我们引入了anira,一个高效的跨平台库。为了确保与各种神经网络架构和框架的兼容性,anira支持ONNX SDK、LibTorch和TensorFlow Lite作为后端。每个推理引擎都表现出实时违规,anira通过将音频回调的推理解耦到静态线程池来减轻这种违规。该库集成了内置的延迟管理和广泛的基准测试功能,这两者对于确保连续的信号流至关重要。三种不同的神经网络架构的音频效果仿真,然后在各种配置进行基准测试。采用统计建模来确定各种因素对性能的影响。研究结果表明,对于无状态模型,ONNX的运行时间最低。对于有状态模型,LibTorch表现出最快的性能。我们的研究结果还表明,对于某些模型发动机组合,初始推理需要更长的时间,特别是当这些推理表现出较高的实时违规发生率。
摘要:Numerous tools for neural network inference are currently available, yet many do not meet the requirements of real-time audio applications. In response, we introduce anira, an efficient cross-platform library. To ensure compatibility with a broad range of neural network architectures and frameworks, anira supports ONNX Runtime, LibTorch, and TensorFlow Lite as backends. Each inference engine exhibits real-time violations, which anira mitigates by decoupling the inference from the audio callback to a static thread pool. The library incorporates built-in latency management and extensive benchmarking capabilities, both crucial to ensure a continuous signal flow. Three different neural network architectures for audio effect emulation are then subjected to benchmarking across various configurations. Statistical modeling is employed to identify the influence of various factors on performance. The findings indicate that for stateless models, ONNX Runtime exhibits the lowest runtimes. For stateful models, LibTorch demonstrates the fastest performance. Our results also indicate that for certain model-engine combinations, the initial inferences take longer, particularly when these inferences exhibit a higher incidence of real-time violations.
标题: 使用公共领域电影集的视频引导文本到音乐生成
链接:https://arxiv.org/abs/2506.12573
备注:ISMIR 2025 regular paper. Dataset and code available at this https URL
摘要:尽管音乐生成系统最近取得了进展,但它们在电影制作中的应用仍然有限,因为它们难以捕捉现实世界电影制作的细微差别,电影制作人在为场景选择或创作音乐时会考虑多种因素,例如视觉内容,对话和情感基调。这种限制主要是由于缺乏综合这些要素的全面数据集。为了解决这一差距,我们引入了开放屏幕声音库(OSSL),这是一个由公共领域电影中的电影片段组成的数据集,总计约36.5小时,配有高质量的配乐和人类注释的情绪信息。为了证明我们的数据集在提高预训练模型在电影音乐生成任务中的性能方面的有效性,我们引入了一种新的视频适配器,该适配器通过添加基于视频的调节来增强基于自回归变换器的文本到音乐模型。我们的实验结果表明,我们提出的方法有效地提高了MusicGen媒体的分布和配对保真度的客观措施,和主观的兼容性的情绪和流派。数据集和代码可在https://havenpersona.github.io/ossl-v1上获得。
摘要:Despite recent advancements in music generation systems, their application in film production remains limited, as they struggle to capture the nuances of real-world filmmaking, where filmmakers consider multiple factors-such as visual content, dialogue, and emotional tone-when selecting or composing music for a scene. This limitation primarily stems from the absence of comprehensive datasets that integrate these elements. To address this gap, we introduce Open Screen Sound Library (OSSL), a dataset consisting of movie clips from public domain films, totaling approximately 36.5 hours, paired with high-quality soundtracks and human-annotated mood information. To demonstrate the effectiveness of our dataset in improving the performance of pre-trained models on film music generation tasks, we introduce a new video adapter that enhances an autoregressive transformer-based text-to-music model by adding video-based conditioning. Our experimental results demonstrate that our proposed approach effectively enhances MusicGen-Medium in terms of both objective measures of distributional and paired fidelity, and subjective compatibility in mood and genre. The dataset and code are available at https://havenpersona.github.io/ossl-v1.
标题: StreamMel:基于交织连续自回归模型的实时零触发文本到语音转换
链接:https://arxiv.org/abs/2506.12570
摘要:zero-shot文本到语音(TTS)合成的最新进展已经实现了高质量的语音生成看不见的扬声器,但大多数系统仍然不适合实时应用,因为他们的离线设计。目前的流TTS模式往往依赖于多级流水线和离散表示,导致计算成本增加和次优的系统性能。在这项工作中,我们提出了StreamMel,一个开创性的单级流TTS框架,模型连续梅尔频谱。通过将文本标记与声学帧交错,StreamMel实现了低延迟、自回归合成,同时保持了较高的说话人相似性和自然度。LibriSpeech上的实验表明,StreamMel在质量和延迟方面都优于现有的流式TTS基线。它甚至可以达到与离线系统相当的性能,同时支持高效的实时生成,展示了与实时语音大语言模型集成的广阔前景。音频样本可在https://aka.ms/StreamMel上获得。
摘要:Recent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time applications because of their offline design. Current streaming TTS paradigms often rely on multi-stage pipelines and discrete representations, leading to increased computational cost and suboptimal system performance. In this work, we propose StreamMel, a pioneering single-stage streaming TTS framework that models continuous mel-spectrograms. By interleaving text tokens with acoustic frames, StreamMel enables low-latency, autoregressive synthesis while preserving high speaker similarity and naturalness. Experiments on LibriSpeech demonstrate that StreamMel outperforms existing streaming TTS baselines in both quality and latency. It even achieves performance comparable to offline systems while supporting efficient real-time generation, showcasing broad prospects for integration with real-time speech large language models. Audio samples are available at: https://aka.ms/StreamMel.
标题: 具有去耦合令牌器和多令牌预测的语音语言模型
链接:https://arxiv.org/abs/2506.12537
摘要:语音语言模型(SLMs)为统一语音和文本理解和生成提供了一条有前途的道路。然而,在实现有效的跨模态对齐和高质量的语音生成方面仍然存在挑战。在这项工作中,我们系统地研究了关键组件(即,语音标记器、语音头和说话者建模)对以LLM为中心的SLM的性能的影响。我们比较耦合,半解耦,完全解耦的语音标记器下一个公平的SLM框架,发现解耦标记显着提高对齐和合成质量。为了解决语音和文本之间的信息密度不匹配,我们引入多令牌预测(MTP)到SLM,使每个隐藏状态解码多个语音令牌。这使得解码速度提高了12倍,并且字错误率大幅下降(从6.07降至3.01)。此外,我们提出了一个扬声器感知生成范式,并介绍RoleTriviaQA,一个大规模的角色扮演知识QA基准与不同的扬声器身份。实验表明,我们的方法提高了知识的理解和说话人的一致性。
摘要:Speech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we systematically investigate the impact of key components (i.e., speech tokenizers, speech heads, and speaker modeling) on the performance of LLM-centric SLMs. We compare coupled, semi-decoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This leads to up to 12$\times$ faster decoding and a substantial drop in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce RoleTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency.
标题: 探索音频线索以增强测试时视频模型适应
链接:https://arxiv.org/abs/2506.12481
备注:14 pages, 7 figures
摘要:测试时自适应(TTA)旨在通过在测试阶段进行自/无监督学习来提高训练模型的泛化能力。虽然大多数现有的视频TTA方法主要利用视觉监控信号,但它们往往忽略了固有音频数据的潜在贡献。为了解决这个差距,我们提出了一种新的方法,将音频信息到视频TTA。我们的方法利用音频的丰富语义内容来生成音频辅助的伪标签,这是视频TTA背景下的一个新概念。具体来说,我们提出了一种音频到视频标签映射方法,首先采用预先训练的音频模型对从视频中提取的音频信号进行分类,然后通过大型语言模型将基于音频的预测映射到视频标签空间,从而建立音频类别和视频标签之间的联系。为了有效地利用生成的伪标签,我们提出了一个灵活的自适应周期,该周期根据不同视图之间的损失和一致性变化来确定每个样本的最佳自适应迭代次数。这使得能够为每个样本定制自适应过程。在两个广泛使用的数据集(UCF 101-C和Kinetics-Sounds-C)以及两个新构建的具有各种腐败类型的音频-视频TTA数据集(AVE-C和AVMIT-C)上的实验结果证明了我们方法的优越性。我们的方法始终提高了不同视频分类模型的适应性能,并代表了将音频信息集成到视频TTA中的重要一步。代码:https://github.com/keikeiqi/Audio-Assisted-TTA。
摘要:Test-time adaptation (TTA) aims to boost the generalization capability of a trained model by conducting self-/unsupervised learning during the testing phase. While most existing TTA methods for video primarily utilize visual supervisory signals, they often overlook the potential contribution of inherent audio data. To address this gap, we propose a novel approach that incorporates audio information into video TTA. Our method capitalizes on the rich semantic content of audio to generate audio-assisted pseudo-labels, a new concept in the context of video TTA. Specifically, we propose an audio-to-video label mapping method by first employing pre-trained audio models to classify audio signals extracted from videos and then mapping the audio-based predictions to video label spaces through large language models, thereby establishing a connection between the audio categories and video labels. To effectively leverage the generated pseudo-labels, we present a flexible adaptation cycle that determines the optimal number of adaptation iterations for each sample, based on changes in loss and consistency across different views. This enables a customized adaptation process for each sample. Experimental results on two widely used datasets (UCF101-C and Kinetics-Sounds-C), as well as on two newly constructed audio-video TTA datasets (AVE-C and AVMIT-C) with various corruption types, demonstrate the superiority of our approach. Our method consistently improves adaptation performance across different video classification models and represents a significant step forward in integrating audio information into video TTA. Code: https://github.com/keikeiqi/Audio-Assisted-TTA.
标题: 基于风格的作曲家识别与象征性音乐配乐归因:一项系统调查
链接:https://arxiv.org/abs/2506.12440
备注:Accepted at the TISMIR
摘要:本文首次对基于风格的作曲家识别和符号乐谱作者归属的文献进行了全面系统的综述。为了满足这一领域对提高可靠性和再现性的迫切需求,该综述严格分析了58篇在不同历史时期发表的同行评审论文,并根据不断发展的术语进行了检索。该分析批判性地评估了流行的剧目,计算方法和评价方法,突出了重大挑战。它揭示了现有研究的很大一部分受到不充分的验证协议和过度依赖简单的准确性指标的影响,这些指标通常是不平衡的数据集,这可能会破坏归因声明的可信度。强调了平衡准确性和严格交叉验证等强大指标在确保可靠结果方面的关键作用。该调查还详细介绍了不同的特征表示和所采用的机器学习模型的演变。值得注意的现实世界的作者归属的情况下,如涉及巴赫,Josquin Desprez和列侬-麦卡特尼的作品,具体讨论,说明应用计算技术来解决有争议的音乐出处的机会和陷阱。基于这些见解,提出了一套可操作的指导方针,为未来的研究。这些建议的目的是显着提高作曲家识别和作者归属研究的可靠性,再现性和音乐学的有效性,促进更强大和可解释的计算风格分析。
摘要:This paper presents the first comprehensive systematic review of literature on style-based composer identification and authorship attribution in symbolic music scores. Addressing the critical need for improved reliability and reproducibility in this field, the review rigorously analyzes 58 peer-reviewed papers published across various historical periods, with the search adapted to evolving terminology. The analysis critically assesses prevailing repertoires, computational approaches, and evaluation methodologies, highlighting significant challenges. It reveals that a substantial portion of existing research suffers from inadequate validation protocols and an over-reliance on simple accuracy metrics for often imbalanced datasets, which can undermine the credibility of attribution claims. The crucial role of robust metrics like Balanced Accuracy and rigorous cross-validation in ensuring trustworthy results is emphasized. The survey also details diverse feature representations and the evolution of machine learning models employed. Notable real-world authorship attribution cases, such as those involving works attributed to Bach, Josquin Desprez, and Lennon-McCartney, are specifically discussed, illustrating the opportunities and pitfalls of applying computational techniques to resolve disputed musical provenance. Based on these insights, a set of actionable guidelines for future research are proposed. These recommendations are designed to significantly enhance the reliability, reproducibility, and musicological validity of composer identification and authorship attribution studies, fostering more robust and interpretable computational stylistic analysis.
标题: GSDNet:从图谱角度重新审视对话情绪识别的不完全多模式扩散
链接:https://arxiv.org/abs/2506.12325
摘要:会话中的多模态情感识别(MERC)旨在通过分析来自多个源的话语信息(即,视频、音频和文本)。与单模态相比,通过融合不同模态的互补语义信息可以获得更鲁棒的话语表示。然而,模态缺失问题严重限制了MERC在实际场景中的性能。最近的工作取得了令人印象深刻的性能模态完成使用图神经网络和扩散模型,分别。这启发我们通过图扩散模型将这两个维度结合起来,以获得更强大的模式恢复能力。然而,现有的图扩散模型直接在邻接矩阵中加入高斯噪声,可能会破坏图的连通性和局部结构,导致生成的图数据无法保留原图的语义和拓扑信息。为此,我们提出了一种新的图谱扩散网络(GSDNet),它将高斯噪声映射到丢失模态的图谱空间,并根据其原始分布恢复丢失数据。与以往的图扩散方法相比,GSDNet只影响邻接矩阵的特征值,而不直接破坏邻接矩阵,能够在扩散过程中保持全局拓扑信息和重要的谱特征。大量的实验表明,GSDNet在各种模态丢失场景中实现了最先进的情感识别性能。
摘要:Multimodal emotion recognition in conversations (MERC) aims to infer the speaker's emotional state by analyzing utterance information from multiple sources (i.e., video, audio, and text). Compared with unimodality, a more robust utterance representation can be obtained by fusing complementary semantic information from different modalities. However, the modality missing problem severely limits the performance of MERC in practical scenarios. Recent work has achieved impressive performance on modality completion using graph neural networks and diffusion models, respectively. This inspires us to combine these two dimensions through the graph diffusion model to obtain more powerful modal recovery capabilities. Unfortunately, existing graph diffusion models may destroy the connectivity and local structure of the graph by directly adding Gaussian noise to the adjacency matrix, resulting in the generated graph data being unable to retain the semantic and topological information of the original graph. To this end, we propose a novel Graph Spectral Diffusion Network (GSDNet), which maps Gaussian noise to the graph spectral space of missing modalities and recovers the missing data according to its original distribution. Compared with previous graph diffusion methods, GSDNet only affects the eigenvalues of the adjacency matrix instead of destroying the adjacency matrix directly, which can maintain the global topological information and important spectral features during the diffusion process. Extensive experiments have demonstrated that GSDNet achieves state-of-the-art emotion recognition performance in various modality loss scenarios.
标题: Phonikud:实时文本到语音的希伯来语字形到音素转换
链接:https://arxiv.org/abs/2506.12311
备注:Project page: this https URL
摘要:现代希伯来语的实时文本到语音(TTS)是具有挑战性的,由于语言的拼写复杂性。现有的解决方案忽略了关键的语音特征,如重音,即使添加了元音标记,重音也没有被充分指定。为了解决这些限制,我们引入Phonikud,一个轻量级的,开源的希伯来语字素到音素(G2 P)系统,输出完全指定的IPA音标。我们的方法适应现有的变音模型与轻量级的适配器,招致可忽略不计的额外延迟。我们还贡献了带有IPA注释的希伯来语转录语音的ILSpeech数据集,作为希伯来语G2 P的基准和TTS系统的训练数据。我们的研究结果表明,与以前的方法相比,Phonikud G2 P转换更准确地预测希伯来语文本中的音素,这使得有效的实时希伯来语TTS模型的训练具有优越的速度-准确性权衡。我们在https://phonikud.github.io上发布我们的代码、数据和模型。
摘要:Real-time text-to-speech (TTS) for Modern Hebrew is challenging due to the language's orthographic complexity. Existing solutions ignore crucial phonetic features such as stress that remain underspecified even when vowel marks are added. To address these limitations, we introduce Phonikud, a lightweight, open-source Hebrew grapheme-to-phoneme (G2P) system that outputs fully-specified IPA transcriptions. Our approach adapts an existing diacritization model with lightweight adaptors, incurring negligible additional latency. We also contribute the ILSpeech dataset of transcribed Hebrew speech with IPA annotations, serving as a benchmark for Hebrew G2P and as training data for TTS systems. Our results demonstrate that Phonikud G2P conversion more accurately predicts phonemes from Hebrew text compared to prior methods, and that this enables training of effective real-time Hebrew TTS models with superior speed-accuracy trade-offs. We release our code, data, and models at https://phonikud.github.io.
标题: 通过习得质量评估的多指标监督改善语音增强
链接:https://arxiv.org/abs/2506.12260
备注:Submitted to ASRU 2025
摘要:语音质量评估(SQA)旨在预测语音信号在大范围失真下的感知质量。它与语音增强(SE)有着内在的联系,后者试图通过去除不需要的信号分量来提高语音质量。虽然SQA模型被广泛用于评估SE绩效,但其指导SE培训的潜力仍有待开发。在这项工作中,我们研究了一个培训框架,该框架利用SQA模型,经过训练,可以从公共SE排行榜中预测多个评估指标,作为SE的监督信号。这种方法解决了传统SE目标(例如SI-SNR)的关键限制,这些目标通常无法与感知质量保持一致,并且在评估指标中概括性较差。此外,它可以在没有干净引用的情况下对真实世界的数据进行训练。在模拟和真实测试集上的实验表明,SQA引导的训练在一系列质量指标上持续提高性能。
摘要:Speech quality assessment (SQA) aims to predict the perceived quality of speech signals under a wide range of distortions. It is inherently connected to speech enhancement (SE), which seeks to improve speech quality by removing unwanted signal components. While SQA models are widely used to evaluate SE performance, their potential to guide SE training remains underexplored. In this work, we investigate a training framework that leverages a SQA model, trained to predict multiple evaluation metrics from a public SE leaderboard, as a supervisory signal for SE. This approach addresses a key limitation of conventional SE objectives, such as SI-SNR, which often fail to align with perceptual quality and generalize poorly across evaluation metrics. Moreover, it enables training on real-world data where clean references are unavailable. Experiments on both simulated and real-world test sets show that SQA-guided training consistently improves performance across a range of quality metrics.
标题: SSLAM:通过复音声景的音频混合增强自我监督模型
链接:https://arxiv.org/abs/2506.12222
备注:Accepted at ICLR 2025. Code and pre-trained models are available at \url{this https URL}
摘要:自我监督的预训练音频网络在现实世界的系统中得到了广泛的应用,特别是在多模态大型语言模型中。这些网络通常在冻结状态下使用,假设SSL预训练已经足够让它们处理真实世界的音频。然而,一个关键的问题仍然存在:这些模型在现实世界中的实际表现如何,其中音频通常是复调和复杂的,涉及多个重叠的声源?当前的音频SSL方法通常在主要以单声道音频为特征的数据集上进行基准测试,例如环境声音和语音。因此,SSL模型推广到复音音频(自然场景中的常见特征)的能力仍然没有得到充分探索。这种限制引起了人们对SSL模型在更真实的音频设置中的实际鲁棒性的担忧。为了解决这一差距,我们引入了音频混合自监督学习(SSLAM),这是音频SSL研究的一个新方向,旨在提高模型从复调数据中学习的能力,同时保持单声道数据的强大性能。我们在标准音频SSL基准数据集上彻底评估了SSLAM,这些数据集主要是单声道的,并使用一系列高质量,公开可用的复音数据集对SOTA方法进行了全面的比较分析。SSLAM不仅提高了复音音频的模型性能,而且在标准音频SSL基准测试中保持或超过了性能。值得注意的是,它比AudioSet-2M(AS-2M)提高了3.9%,平均精度(mAP)达到50.2。对于复调数据集,SSLAM在线性评估和微调机制中设置了新的SOTA,性能提高高达9.1\ %(mAP)。
摘要:Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the SSL pre-training has sufficiently equipped them to handle real-world audio. However, a critical question remains: how well do these models actually perform in real-world conditions, where audio is typically polyphonic and complex, involving multiple overlapping sound sources? Current audio SSL methods are often benchmarked on datasets predominantly featuring monophonic audio, such as environmental sounds, and speech. As a result, the ability of SSL models to generalize to polyphonic audio, a common characteristic in natural scenarios, remains underexplored. This limitation raises concerns about the practical robustness of SSL models in more realistic audio settings. To address this gap, we introduce Self-Supervised Learning from Audio Mixtures (SSLAM), a novel direction in audio SSL research, designed to improve, designed to improve the model's ability to learn from polyphonic data while maintaining strong performance on monophonic data. We thoroughly evaluate SSLAM on standard audio SSL benchmark datasets which are predominantly monophonic and conduct a comprehensive comparative analysis against SOTA methods using a range of high-quality, publicly available polyphonic datasets. SSLAM not only improves model performance on polyphonic audio, but also maintains or exceeds performance on standard audio SSL benchmarks. Notably, it achieves up to a 3.9\% improvement on the AudioSet-2M (AS-2M), reaching a mean average precision (mAP) of 50.2. For polyphonic datasets, SSLAM sets new SOTA in both linear evaluation and fine-tuning regimes with performance improvements of up to 9.1\% (mAP).
标题: ViSAGee:视频到空间音频生成
链接:https://arxiv.org/abs/2506.12199
备注:ICLR 2025. Project page: this https URL
摘要:空间音频对于增强视听体验的沉浸感至关重要,但其制作通常需要复杂的录音系统和专业知识。在这项工作中,我们解决了一个新的问题,产生一阶立体混响,广泛使用的空间音频格式,直接从无声视频。为了支持这一任务,我们引入了YT-Ambigen,这是一个包含102 K 5秒YouTube视频剪辑的数据集,与相应的一阶立体混响配对。我们还提出了新的评估指标,以评估基于音频能量图和显着性度量生成的音频的空间方面。此外,我们提出了视频到空间音频生成(ViSAGe),一个端到端的框架,通过利用CLIP视觉功能,自回归神经音频编解码器建模与方向和视觉指导,从无声视频帧生成一阶立体混响。实验结果表明,ViSAGe产生合理的和连贯的一阶立体混响,优于两个阶段的方法,包括视频到音频生成和音频空间化。定性示例进一步说明ViSAGe生成适应视点变化的时间对齐的高质量空间音频。
摘要:Spatial audio is essential for enhancing the immersiveness of audio-visual experiences, yet its production typically demands complex recording systems and specialized expertise. In this work, we address a novel problem of generating first-order ambisonics, a widely used spatial audio format, directly from silent videos. To support this task, we introduce YT-Ambigen, a dataset comprising 102K 5-second YouTube video clips paired with corresponding first-order ambisonics. We also propose new evaluation metrics to assess the spatial aspect of generated audio based on audio energy maps and saliency metrics. Furthermore, we present Video-to-Spatial Audio Generation (ViSAGe), an end-to-end framework that generates first-order ambisonics from silent video frames by leveraging CLIP visual features, autoregressive neural audio codec modeling with both directional and visual guidance. Experimental results demonstrate that ViSAGe produces plausible and coherent first-order ambisonics, outperforming two-stage approaches consisting of video-to-audio generation and audio spatialization. Qualitative examples further illustrate that ViSAGe generates temporally aligned high-quality spatial audio that adapts to viewpoint changes.
标题: 通过两遍解码将Whisper改编为流语音识别
链接:https://arxiv.org/abs/2506.12154
备注:Accepted to INTERSPEECH 2025
摘要:OpenAI Whisper是一系列强大的自动语音识别(ASR)模型,经过680,000小时的音频训练。然而,它的编码器-解码器架构,用序列到序列目标训练,缺乏对流式ASR的原生支持。在本文中,我们使用WeNet工具包通过采用统一的双通道(U2)结构来微调Whisper用于流式ASR。我们引入了一个额外的连接主义时间分类(CTC)解码器,用因果注意掩码训练来生成流部分转录本,而原始的Whisper解码器则对这些部分输出进行重新排序。我们在LibriSpeech和财报电话数据集上的实验表明,只要有足够的微调数据,Whisper就可以适应有能力的流ASR模型。我们还介绍了一种混合令牌化方法,它使用一个较小的令牌空间的CTC解码器,同时保留耳语的原始令牌空间的注意力解码器,从而提高数据效率和泛化。
摘要:OpenAI Whisper is a family of robust Automatic Speech Recognition (ASR) models trained on 680,000 hours of audio. However, its encoder-decoder architecture, trained with a sequence-to-sequence objective, lacks native support for streaming ASR. In this paper, we fine-tune Whisper for streaming ASR using the WeNet toolkit by adopting a Unified Two-pass (U2) structure. We introduce an additional Connectionist Temporal Classification (CTC) decoder trained with causal attention masks to generate streaming partial transcripts, while the original Whisper decoder reranks these partial outputs. Our experiments on LibriSpeech and an earnings call dataset demonstrate that, with adequate fine-tuning data, Whisper can be adapted into a capable streaming ASR model. We also introduce a hybrid tokenizer approach, which uses a smaller token space for the CTC decoder while retaining Whisper's original token space for the attention decoder, resulting in improved data efficiency and generalization.
标题: TuneGenie:基于推理的LLM代理,提供优先音乐生成
链接:https://arxiv.org/abs/2506.12083
备注:15 pages
摘要:最近,大型语言模型(LLM)在从生成图像到空间推理的各种任务中表现出了巨大的潜力。考虑到他们显着的(和不断增长的)文本推理能力,我们调查LLM在进行个人的音乐偏好分析(基于播放列表元数据,个人写作等)的潜力。并产生有效的提示(基于这些分析),以传递给Suno AI(用于音乐制作的生成AI工具)。我们提出了一种新的基于LLM的文本表示音乐模型(我们称之为TuneGenie),以及我们开发的各种方法来评估和基准类似的模型,增加了越来越多(越来越有争议)的关于使用AI生成艺术的研究语料库。
摘要:Recently, Large language models (LLMs) have shown great promise across a diversity of tasks, ranging from generating images to reasoning spatially. Considering their remarkable (and growing) textual reasoning capabilities, we investigate LLMs' potency in conducting analyses of an individual's preferences in music (based on playlist metadata, personal write-ups, etc.) and producing effective prompts (based on these analyses) to be passed to Suno AI (a generative AI tool for music production). Our proposition of a novel LLM-based textual representation to music model (which we call TuneGenie) and the various methods we develop to evaluate & benchmark similar models add to the increasing (and increasingly controversial) corpus of research on the use of AI in generating art.
机器翻译由腾讯交互翻译提供,仅供参考
