今日论文合集:cs.SD语音10篇,eess.AS音频处理13篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound  Source Localization
标题: 通过CLIP听和看:自我监督的声音源定位框架
链接:https://arxiv.org/abs/2505.05343
作者: Sooyoung Park,  Arda Senocak,  Joon Son Chung 
备注:Journal Extension of WACV 2024 paper (arXiv:2311.04066). Code is available at this https URL
摘要:大型视觉语言模型在不同的任务中表现出强大的多模态对齐和泛化能力。其中,CLIP是最成功的方法之一。在这项工作中,我们将CLIP的应用扩展到声源定位,提出了一种无需显式文本输入的自监督方法。我们引入了一个框架,将音频映射到与CLIP的文本编码器兼容的令牌,产生音频驱动的嵌入。这些嵌入用于生成发声区域掩码,从该发声区域掩码中提取视觉特征,并通过对比性视听对应目标与音频嵌入对齐。我们的研究结果表明,预训练的多模态基础模型的对齐知识使我们的方法能够为探测对象生成更完整和紧凑的定位。我们进一步提出了一个LLM指导的扩展,在训练过程中将对象感知的视听场景理解提取到模型中,以增强对齐。五个不同任务的广泛实验表明,我们的方法,在所有的变种,优于国家的最先进的方法,并实现了强大的泛化在zero-shot设置。
摘要:Large-scale vision-language models demonstrate strong multimodal alignment and generalization across diverse tasks. Among them, CLIP stands out as one of the most successful approaches. In this work, we extend the application of CLIP to sound source localization, proposing a self-supervised method operates without explicit text input. We introduce a framework that maps audios into tokens compatible with CLIP's text encoder, producing audio-driven embeddings. These embeddings are used to generate sounding region masks, from which visual features are extracted and aligned with the audio embeddings through a contrastive audio-visual correspondence objective. Our findings show that alignment knowledge of pre-trained multimodal foundation model enables our method to generate more complete and compact localization for sounding objects. We further propose an LLM-guided extension that distills object-aware audio-visual scene understanding into the model during training to enhance alignment. Extensive experiments across five diverse tasks demonstrate that our method, in all variants, outperforms state-of-the-art approaches and achieves strong generalization in zero-shot settings.


【2】 FLAM: Frame-Wise Language-Audio Modeling

标题: FLAM:帧级格式音频建模
链接:https://arxiv.org/abs/2505.05335
作者: Yusong Wu,  Christos Tsirigotis,  Ke Chen,  Cheng-Zhi Anna Huang,  Aaron Courville,  Oriol Nieto,  Prem Seetharaman,  Justin Salamon 
备注:Accepted at ICML 2025
摘要:最近的多模态音频语言模型(ALM)擅长文本音频检索,但与帧明智的音频理解的斗争。以前的工作使用时间感知标签或无监督训练来提高帧的能力,但它们仍然缺乏细粒度的标签能力来确定事件发生的时间。虽然传统的声音事件检测模型可以精确地定位事件,但它们仅限于预定义的类别,这使得它们对于具有分布外事件的真实场景无效。在这项工作中,我们介绍FLAM,一个开放的词汇对比音频语言模型能够本地化特定的声音事件。FLAM采用了一个高效的内存和校准的逐帧目标与logit调整,以解决虚假的相关性,如事件依赖性和标签在训练过程中的不平衡。为了实现逐帧监督,我们利用了具有不同音频事件、LLM生成的字幕和模拟的大规模数据集。实验结果和案例研究表明,FLAM显着提高了开放词汇的本地化能力,同时保持了强大的性能,在全球检索和下游任务。
摘要:Recent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervised training to improve frame-wise capabilities, but they still lack fine-grained labeling capability to pinpoint when an event occurs. While traditional sound event detection models can precisely localize events, they are limited to pre-defined categories, making them ineffective for real-world scenarios with out-of-distribution events. In this work, we introduce FLAM, an open-vocabulary contrastive audio-language model capable of localizing specific sound events. FLAM employs a memory-efficient and calibrated frame-wise objective with logit adjustment to address spurious correlations, such as event dependencies and label imbalances during training. To enable frame-wise supervision, we leverage a large-scale dataset with diverse audio events, LLM-generated captions and simulation. Experimental results and case studies demonstrate that FLAM significantly improves the open-vocabulary localization capability while maintaining strong performance in global retrieval and downstream tasks.


【3】 Pairing Real-Time Piano Transcription with Symbol-level Tracking for  Precise and Robust Score Following

标题: 将实时钢琴抄写与符号级跟踪相结合,实现精确且稳健的乐谱跟踪
链接:https://arxiv.org/abs/2505.05078
作者: Silvan Peter,  Patricia Hu,  Gerhard Widmer 
备注:5 pages, 3 tables, 2 pseudocodes, to be published at the Sound and Music Computing Conference 2025
摘要:实时音乐跟踪系统跟踪音乐表演,并随时报告相应乐谱中的当前位置。大多数现有方法专门在音频域中解决这个问题,通常对传入音频和分数的音频表示使用在线时间规整(OLTW)技术。音频OLTW技术在过去十年中在功能和模型性能方面都有了不断的改进,达到了性能平台。我们认为,转换和代表的表现在象征性的域-从而将音乐跟踪到一个象征性的任务-可以是一个更有效的方法,即使域转换是不完美的。我们的音乐跟踪系统结合了两个实时组件:一个处理音频到音符的转录和其他转录输入和得分之间的一种新的符号级跟踪器。我们将这种混合音频符号方法的性能与其等效的仅音频对应方法进行了比较,并证明我们的方法在精度,即,绝对跟踪误差和鲁棒性,即,追踪成功
摘要:Real-time music tracking systems follow a musical performance and at any time report the current position in a corresponding score. Most existing methods approach this problem exclusively in the audio domain, typically using online time warping (OLTW) techniques on incoming audio and an audio representation of the score. Audio OLTW techniques have seen incremental improvements both in features and model heuristics which reached a performance plateau in the past ten years. We argue that converting and representing the performance in the symbolic domain -- thereby transforming music tracking into a symbolic task -- can be a more effective approach, even when the domain transformation is imperfect. Our music tracking system combines two real-time components: one handling audio-to-note transcription and the other a novel symbol-level tracker between transcribed input and score. We compare the performance of this mixed audio-symbolic approach with its equivalent audio-only counterpart, and demonstrate that our method outperforms the latter in terms of both precision, i.e., absolute tracking error, and robustness, i.e., tracking success.


【4】 ReverbMiipher: Generative Speech Restoration meets Reverberation  Characteristics Controllability

标题: ReverbMiipher:生成性语音恢复满足回响特征可控制性
链接:https://arxiv.org/abs/2505.05077
作者: Wataru Nakata,  Yuma Koizumi,  Shigeki Karita,  Robin Scheibler,  Haruko Ishikawa,  Adriana Guevara-Rukoz,  Heiga Zen,  Michiel Bacchiani 
备注:5 pages, 5 figures
摘要:混响编码的声源环境的空间信息,而传统的语音恢复(SR)通常完全消除混响。我们提出了ReverbMiipher,一个扩展参数再合成框架的SR模型,旨在对语音进行降噪,同时保留并控制混响。ReverbMiipher集成了一个专用的ReverbEncoder,用于从嘈杂的输入中提取混响特征向量。该特征调节声码器以重构语音信号,在保留原始混响特性的同时去除噪声。训练过程中的随机零向量替换策略确保该特征专门编码混响,将其与其他语音属性分离。这种学习的表示通过诸如特征之间的插值、用来自其他话语的特征替换或从潜在空间采样的技术来促进混响控制。客观和主观评估证实ReverbMiipher有效地保留混响,去除其他伪影,并优于传统的两阶段SR和卷积模拟房间脉冲响应方法。我们进一步证明了它的能力,通过功能操作产生新的混响效果。
摘要:Reverberation encodes spatial information regarding the acoustic source environment, yet traditional Speech Restoration (SR) usually completely removes reverberation. We propose ReverbMiipher, an SR model extending parametric resynthesis framework, designed to denoise speech while preserving and enabling control over reverberation. ReverbMiipher incorporates a dedicated ReverbEncoder to extract a reverb feature vector from noisy input. This feature conditions a vocoder to reconstruct the speech signal, removing noise while retaining the original reverberation characteristics. A stochastic zero-vector replacement strategy during training ensures the feature specifically encodes reverberation, disentangling it from other speech attributes. This learned representation facilitates reverberation control via techniques such as interpolation between features, replacement with features from other utterances, or sampling from a latent space. Objective and subjective evaluations confirm ReverbMiipher effectively preserves reverberation, removes other artifacts, and outperforms the conventional two-stage SR and convolving simulated room impulse response approach. We further demonstrate its ability to generate novel reverberation effects through feature manipulation.


【5】 How to Infer Repeat Structures in MIDI Performances

标题: 如何推断工作组性能中的重复结构
链接:https://arxiv.org/abs/2505.05055
作者: Silvan Peter,  Patricia Hu,  Gerhard Widmer 
备注:3 pages, 1 figure, 1 table, to be published in the Music Encoding Conference 2025
摘要:在演奏研究和音乐信息检索中,演奏者的演奏通常是有利的,如果他们能与乐谱联系起来,那就更有利了。这种联系通常是通过对齐的方式建立的,将乐谱和演奏之间的音符或时间点联系起来。当试图建立这样的对齐时,第一个障碍是表演实现了乐谱的一个(许多)结构版本,该结构版本可以由诸如重复、变化和导航标记(如“dal segno/da capo al coda”)之类的指令来产生。在可以应用比对算法之前,需要展开分数,即,需要明确写出其重复和导航标记,以创建没有与性能匹配的跳跃的单个时间轴。在大型表演语料库的管理中,这个过程是手动进行的,因为没有工具可以推断表演的重复结构。为了简化这一过程,我们开发了一种方法来自动推断重复结构的一个重复的性能,给出了一个象征性的编码得分,包括重复和导航标记。指导我们设计的直觉是:1)乐谱的每个连续部分与包含相同材料的表演部分的局部对准应该接收高对准增益,而与任何其他表演部分的局部对准应该产生低增益或零增益。以及2)如果结构版本对应于性能,则根据得分的有效结构版本将局部比对拼接在一起应导致近似的完全比对和相应的高全局累积增益,并且对于所有其他不合适的结构版本,导致低增益。
摘要:MIDI performances are generally expedient in performance research and music information retrieval, and even more so if they can be connected to a score. This connection is usually established by means of alignment, linking either notes or time points between the score and the performance. The first obstacle when trying to establish such an alignment is that a performance realizes one (out of many) structural versions of the score that can plausibly result from instructions such as repeats, variations, and navigation markers like 'dal segno/da capo al coda'. A score needs to be unfolded, that is, its repeats and navigation markers need to be explicitly written out to create a single timeline without jumps matching the performance, before alignment algorithms can be applied. In the curation of large performance corpora this process is carried out manually, as no tools are available to infer the repeat structure of the performance. To ease this process, we develop a method to automatically infer the repeat structure of a MIDI performance, given a symbolically encoded score including repeat and navigation markers. The intuition guiding our design is: 1) local alignment of every contiguous section of the score with a section of a performance containing the same material should receive high alignment gain, whereas local alignment with any other performance section should accrue a low or zero gain. And 2) stitching local alignments together according to a valid structural version of the score should result in an approximate full alignment and correspondingly high global accumulated gain if the structural version corresponds to the performance, and low gain for all other, ill-fitting structural versions.


【6】 Inter-Diffusion Generation Model of Speakers and Listeners for Effective  Communication

标题: 有效沟通的说话者和听众的相互扩散生成模型
链接:https://arxiv.org/abs/2505.04996
作者: Jinhe Huang,  Yongkang Cheng,  Yuming Hang,  Gaoge Han,  Jinewei Li,  Jing Zhang,  Xingjian Gu 
备注:accepted by ICMR 2025
摘要:全身手势在自然交互中起着关键作用,对于实现有效沟通至关重要。然而,现有的研究大多集中在说话人的手势生成上,忽视了听话人在互动过程中的重要作用,未能充分探讨他们之间的动态互动。本文创新性地提出了一种有效交际的说者和听者互扩散生成模型。我们第一次将听众的全身手势整合到生成框架中。通过设计一种新的交互扩散机制,该模型可以准确地捕捉说话人和听话人在交流过程中的复杂交互模式。在模型构建过程中,基于先进的扩散模型架构,创新性地引入交互条件和GAN模型,增加去噪步长。因此,在生成手势序列时,模型不仅可以基于说话者的语音信息动态生成,而且还可以实时响应收听者的反馈,从而实现两者之间的协同交互。大量的实验结果表明,与现有的手势生成方法相比,本文提出的模型在手势的自然度、连贯性和语音-手势同步性等方面都有了显著的提高。在主观评价实验中,用户对生成的交互场景给予了高度评价,认为它们更接近于现实生活中的人际交往场景。客观的指标评价也表明,我们的模型在多个关键指标上优于基线方法,为有效沟通提供了更有力的支持。
摘要:Full-body gestures play a pivotal role in natural interactions and are crucial for achieving effective communication. Nevertheless, most existing studies primarily focus on the gesture generation of speakers, overlooking the vital role of listeners in the interaction process and failing to fully explore the dynamic interaction between them. This paper innovatively proposes an Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication. For the first time, we integrate the full-body gestures of listeners into the generation framework. By devising a novel inter-diffusion mechanism, this model can accurately capture the complex interaction patterns between speakers and listeners during communication. In the model construction process, based on the advanced diffusion model architecture, we innovatively introduce interaction conditions and the GAN model to increase the denoising step size. As a result, when generating gesture sequences, the model can not only dynamically generate based on the speaker's speech information but also respond in realtime to the listener's feedback, enabling synergistic interaction between the two. Abundant experimental results demonstrate that compared with the current state-of-the-art gesture generation methods, the model we proposed has achieved remarkable improvements in the naturalness, coherence, and speech-gesture synchronization of the generated gestures. In the subjective evaluation experiments, users highly praised the generated interaction scenarios, believing that they are closer to real life human communication situations. Objective index evaluations also show that our model outperforms the baseline methods in multiple key indicators, providing more powerful support for effective communication.


【7】 A Multi-Agent AI Framework for Immersive Audiobook Production through  Spatial Audio and Neural Narration

标题: 基于空间音频和神经叙事的沉浸式有声读物制作的多智能体AI框架
链接:https://arxiv.org/abs/2505.04885
作者: Shaja Arul Selvamani,  Nia D'Souza Ganapathy 
摘要:这项研究介绍了一种创新的AI驱动的多代理框架,专门用于创建沉浸式有声读物。该框架利用FastSpeech 2和VALL-E的神经文本到语音合成来进行富有表现力的叙述和特定于角色的声音,采用先进的语言模型来自动解释文本叙述并生成逼真的空间音频效果。这些声音效果通过复杂的时间整合方法与故事情节动态同步,包括动态时间弯曲(DTW)和递归神经网络(RNN)。基于扩散的生成模型与高阶立体混响(HOA)和散射延迟网络(SDN)相结合,实现了高度逼真的3D音景,大大增强了听众的沉浸感和叙事真实感。这项技术大大推进了有声读物应用,为视障观众提供了更丰富的教育内容体验、讲故事平台和无障碍解决方案。未来的工作将涉及个性化,合成语音的道德管理以及与多感官平台的整合。
摘要:This research introduces an innovative AI-driven multi-agent framework specifically designed for creating immersive audiobooks. Leveraging neural text-to-speech synthesis with FastSpeech 2 and VALL-E for expressive narration and character-specific voices, the framework employs advanced language models to automatically interpret textual narratives and generate realistic spatial audio effects. These sound effects are dynamically synchronized with the storyline through sophisticated temporal integration methods, including Dynamic Time Warping (DTW) and recurrent neural networks (RNNs). Diffusion-based generative models combined with higher-order ambisonics (HOA) and scattering delay networks (SDN) enable highly realistic 3D soundscapes, substantially enhancing listener immersion and narrative realism. This technology significantly advances audiobook applications, providing richer experiences for educational content, storytelling platforms, and accessibility solutions for visually impaired audiences. Future work will address personalization, ethical management of synthesized voices, and integration with multi-sensory platforms.


【8】 Data Standards in Audiology: A Mixed-Methods Exploration of Community  Perspectives and Implementation Considerations

标题: 听力学数据标准:社区观点和实施考虑因素的混合方法探索
链接:https://arxiv.org/abs/2505.04728
作者: Charlotte Vercammen,  Antje Heinrich,  Christophe Lesimple,  Alessia Paglialonga,  Jan-Willem A. Wasmann,  Mareike Buhl 
摘要:目的:本研究的目的是探索听力学数据标准化的选择,并记录全球听力学界对数据标准的当前知识和观点,探索他们的需求和偏好,并因此制定数据标准化建议。   设计图:在2024年计算听力学虚拟会议的“听力学大数据和数据标准”特别会议期间,采用混合方法,将结构化调查与专家对主题的深入探索相结合。   研究样本:调查样本由全球听力学社区的82名成员组成; 5名专家参加了小组讨论。   结果:调查结果强调了听力学数据标准化的必要性,旨在促进研究和改善患者护理。对现有举措的了解程度较低:38%的人了解这些举措。然而,90%的人希望为他们的发展做出贡献。小组讨论探讨了听力学中新兴的标准化倡议(OMOP,openEHR,HIMSA的Noah标准),挑战(例如,数据质量和隐私),以及机会(例如,方法之间的转换以及与其他医学领域的协同作用)。   结论:本研究中确定的社区支持可以用来进一步制定听力学标准化计划,确保计划之间以及与其他医学领域的一致性。
摘要:Objective: The purpose of this study was to explore options for data standardisation in audiology and document the global audiology community's current knowledge and views of data standards, explore their needs and preferences, and develop recommendations for data standardisation as a result.   Design: A mixed-methods approach, combining a structured survey with an in-depth exploration of themes by experts during a special session on "Big Data and Data Standards in Audiology" at the 2024 Virtual Conference of Computational Audiology.   Study Sample: The survey sample consisted of 82 members of the global audiology community; five experts joined the panel discussion.   Results: Survey results emphasized the need for data standardisation in audiology aimed at facilitating research and improving patient care. Knowledge of existing initiatives was low: 38% were aware of initiatives. Yet, 90% envisioned contributing to them moving forward. The panel discussion explored emerging standardisation initiatives in audiology (OMOP, openEHR, HIMSA's Noah standard), challenges (e.g., data quality and privacy), and opportunities (e.g., conversion between approaches and synergies with other medical fields).   Conclusions: The community support identified in this study could be leveraged to further develop standardisation initiatives for audiology, ensuring alignment between initiatives and with other medical fields.


【9】 Normalize Everything: A Preconditioned Magnitude-Preserving Architecture  for Diffusion-Based Speech Enhancement

标题: Normalize Everything:一种基于扩散的预处理幅度保持语音增强结构
链接:https://arxiv.org/abs/2505.05216
作者: Julius Richter,  Danilo de Oliveira,  Timo Gerkmann 
备注:Submitted to WASPAA 2025
摘要:本文提出了一种新的基于扩散的语音增强框架。我们的方法采用了一个薛定谔桥转换到干净的语音分布的嘈杂的语音分布。为了稳定和改进训练,我们对网络的输入和输出进行了时间相关的缩放,称为预处理。我们考虑两个跳过连接配置,其中包括或省略当前处理状态的降噪器的输出,使网络预测环境噪声或干净的语音。每种方法都可以提高不同语音增强指标的性能。为了在训练过程中保持稳定的幅度水平和平衡,我们使用了一个幅度保持网络架构,将所有激活和网络权重归一化为单位长度。此外,我们建议学习每个网络块内的噪声输入的贡献,以实现有效的输入调节。在训练之后,我们应用一种方法来近似不同的指数移动平均(EMA)轮廓,并研究它们对语音增强性能的影响。与图像生成任务相比,较长的EMA长度通常会增强模式覆盖范围,我们观察到较短的EMA长度始终会导致标准语音增强指标的更好性能。代码、音频示例和检查点可在线获取。
摘要:This paper presents a new framework for diffusion-based speech enhancement. Our method employs a Schroedinger bridge to transform the noisy speech distribution into the clean speech distribution. To stabilize and improve training, we employ time-dependent scalings of the inputs and outputs of the network, known as preconditioning. We consider two skip connection configurations, which either include or omit the current process state in the denoiser's output, enabling the network to predict either environmental noise or clean speech. Each approach leads to improved performance on different speech enhancement metrics. To maintain stable magnitude levels and balance during training, we use a magnitude-preserving network architecture that normalizes all activations and network weights to unit length. Additionally, we propose learning the contribution of the noisy input within each network block for effective input conditioning. After training, we apply a method to approximate different exponential moving average (EMA) profiles and investigate their effects on the speech enhancement performance. In contrast to image generation tasks, where longer EMA lengths often enhance mode coverage, we observe that shorter EMA lengths consistently lead to better performance on standard speech enhancement metrics. Code, audio examples, and checkpoints are available online.


【10】 Listen to Extract: Onset-Prompted Target Speaker Extraction

标题: 收听摘录:起始点确定的目标说话人提取
链接:https://arxiv.org/abs/2505.05114
作者: Pengjie Shen,  Kangrui Chen,  Shulin He,  Pengru Chen,  Shuqi Yuan,  He Kong,  Xueliang Zhang,  Zhong-Qiu Wang 
备注:in submission
摘要:我们提出了$\textit{listen to extract}$(LExt),一种高效而又非常简单的单声道目标说话人提取(TSE)算法。给定一个目标说话人的注册话语,LExt的目的是从说话人与其他说话人的混合语音中提取目标说话人。对于每个混合,LExt将目标说话者的登记话语连接到波形级别的混合信号,并训练深度神经网络(DNN)以基于连接的混合信号提取目标语音。其基本原理是,通过这种方式,为目标说话者创建人工语音起始,并且它可以提示DNN(a)提取哪个说话者是目标;以及(b)可以帮助提取的目标说话者的频谱-时间模式。这种简单的方法在多个公共TSE数据集上产生了强大的TSE性能,包括WSJ 0 - 2 mix,WHAM!然后“轰”!
摘要:We propose $\textit{listen to extract}$ (LExt), a highly-effective while extremely-simple algorithm for monaural target speaker extraction (TSE). Given an enrollment utterance of a target speaker, LExt aims at extracting the target speaker from the speaker's mixed speech with other speakers. For each mixture, LExt concatenates an enrollment utterance of the target speaker to the mixture signal at the waveform level, and trains deep neural networks (DNN) to extract the target speech based on the concatenated mixture signal. The rationale is that, this way, an artificial speech onset is created for the target speaker and it could prompt the DNN (a) which speaker is the target to extract; and (b) spectral-temporal patterns of the target speaker that could help extraction. This simple approach produces strong TSE performance on multiple public TSE datasets including WSJ0-2mix, WHAM! and WHAMR!.


eess.AS音频处理


【1】 Normalize Everything: A Preconditioned Magnitude-Preserving Architecture  for Diffusion-Based Speech Enhancement

标题: Normalize Everything:一种基于扩散的预处理幅度保持语音增强结构
链接:https://arxiv.org/abs/2505.05216
作者: Julius Richter,  Danilo de Oliveira,  Timo Gerkmann 
备注:Submitted to WASPAA 2025
摘要:本文提出了一种新的基于扩散的语音增强框架。我们的方法采用了一个薛定谔桥转换到干净的语音分布的嘈杂的语音分布。为了稳定和改进训练,我们对网络的输入和输出进行了时间相关的缩放,称为预处理。我们考虑两个跳过连接配置,其中包括或省略当前处理状态的降噪器的输出,使网络预测环境噪声或干净的语音。每种方法都可以提高不同语音增强指标的性能。为了在训练过程中保持稳定的幅度水平和平衡,我们使用了一个幅度保持网络架构,将所有激活和网络权重归一化为单位长度。此外,我们建议学习每个网络块内的噪声输入的贡献,以实现有效的输入调节。在训练之后,我们应用一种方法来近似不同的指数移动平均(EMA)轮廓,并研究它们对语音增强性能的影响。与图像生成任务相比,较长的EMA长度通常会增强模式覆盖范围,我们观察到较短的EMA长度始终会导致标准语音增强指标的更好性能。代码、音频示例和检查点可在线获取。
摘要:This paper presents a new framework for diffusion-based speech enhancement. Our method employs a Schroedinger bridge to transform the noisy speech distribution into the clean speech distribution. To stabilize and improve training, we employ time-dependent scalings of the inputs and outputs of the network, known as preconditioning. We consider two skip connection configurations, which either include or omit the current process state in the denoiser's output, enabling the network to predict either environmental noise or clean speech. Each approach leads to improved performance on different speech enhancement metrics. To maintain stable magnitude levels and balance during training, we use a magnitude-preserving network architecture that normalizes all activations and network weights to unit length. Additionally, we propose learning the contribution of the noisy input within each network block for effective input conditioning. After training, we apply a method to approximate different exponential moving average (EMA) profiles and investigate their effects on the speech enhancement performance. In contrast to image generation tasks, where longer EMA lengths often enhance mode coverage, we observe that shorter EMA lengths consistently lead to better performance on standard speech enhancement metrics. Code, audio examples, and checkpoints are available online.


【2】 FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech

标题: FlexSpeech:迈向稳定、可控和富有表达力的文本到语音
链接:https://arxiv.org/abs/2505.05159
作者: Linhan Ma,  Dake Guo,  He Wang,  Jin Xu,  Lei Xie 
备注:10 pages, 5 figures
摘要:目前的语音生成研究可以分为两大类:非自回归和自回归。这些方法之间的根本区别在于用于可预测长度序列的持续时间预测策略。NAR方法通过明确地和独立地对每个语音单元的持续时间建模来确保语音生成的稳定性。相反,AR方法采用自回归范式来预测压缩语音令牌隐式建模持续时间与马尔可夫特性。虽然这种方法改善了韵律,但它不能提供稳定性所需的结构保证。为了同时解决语音生成中的稳定性和自然性问题,我们提出了FlexSpeech,一个稳定,可控和表达的TTS模型。FlexSpeech背后的动机是将马尔可夫依赖关系和偏好优化直接纳入持续时间预测器,以提高其自然性,同时保持语音单元的显式建模,以确保稳定性。具体来说,我们将语音生成任务分解为两个部分:AR持续时间预测器和NAR声学模型。声学模型在大量数据上进行训练,以学习在给定参考音频韵律和音素持续时间的情况下更稳定地呈现音频。持续时间预测器针对不同的风格变化以轻量级方式进行优化,从而实现快速风格转移,同时保持与指定扬声器音色的解耦关系。实验结果表明,该方法在zero-shot TTS中实现了SOTA的稳定性和自然度。更重要的是,当转换到特定的风格领域时,我们可以仅用大约100个数据样本来完成持续时间模块的轻量级优化,而无需调整声学模型,从而实现快速稳定的风格转换。
摘要:Current speech generation research can be categorized into two primary classes: non-autoregressive and autoregressive. The fundamental distinction between these approaches lies in the duration prediction strategy employed for predictable-length sequences. The NAR methods ensure stability in speech generation by explicitly and independently modeling the duration of each phonetic unit. Conversely, AR methods employ an autoregressive paradigm to predict the compressed speech token by implicitly modeling duration with Markov properties. Although this approach improves prosody, it does not provide the structural guarantees necessary for stability. To simultaneously address the issues of stability and naturalness in speech generation, we propose FlexSpeech, a stable, controllable, and expressive TTS model. The motivation behind FlexSpeech is to incorporate Markov dependencies and preference optimization directly on the duration predictor to boost its naturalness while maintaining explicit modeling of the phonetic units to ensure stability. Specifically, we decompose the speech generation task into two components: an AR duration predictor and a NAR acoustic model. The acoustic model is trained on a substantial amount of data to learn to render audio more stably, given reference audio prosody and phone durations. The duration predictor is optimized in a lightweight manner for different stylistic variations, thereby enabling rapid style transfer while maintaining a decoupled relationship with the specified speaker timbre. Experimental results demonstrate that our approach achieves SOTA stability and naturalness in zero-shot TTS. More importantly, when transferring to a specific stylistic domain, we can accomplish lightweight optimization of the duration module solely with about 100 data samples, without the need to adjust the acoustic model, thereby enabling rapid and stable style transfer.


【3】 Regression-based Melody Estimation with Uncertainty Quantification

标题: 具有不确定性量化的基于回归的旋律估计
链接:https://arxiv.org/abs/2505.05156
作者: Kavya Ranjan Saxena,  Vipul Arora 
摘要:现有的机器学习模型将从复调音频中估计旋律的任务作为分类问题通过离散化音高值来处理,这导致旋律中存在的更精细的频率变化的损失。为了更好地捕捉这些变化,我们建议将此任务视为回归问题。除了仅预测音频中特定区域的音高外,我们还预测其不确定性,以增强模型的可信度。为了进行基于回归的旋律估计,我们提出了三种不同的方法,使用直方图表示模型的音高值。这种表示要求直方图的支持范围是连续的。前两种方法通过将清音和浊音频率范围映射到连续范围来解决清音和浊音频率范围之间的突然不连续。第三种方法将旋律估计重新表述为完全贝叶斯任务,将浊音检测建模为分类问题,将浊音音高估计建模为回归问题。此外,我们引入了一种新的方法来估计直方图表示的不确定性,该直方图表示与预测分布的平均值与地面实况的偏差相关。实验结果表明,重新制定旋律估计为回归问题显着提高了基于分类的方法的性能。比较所提出的方法与国家的最先进的回归模型,据观察,贝叶斯方法在估计旋律及其相关的不确定性方面表现最好。
摘要:Existing machine learning models approach the task of melody estimation from polyphonic audio as a classification problem by discretizing the pitch values, which results in the loss of finer frequency variations present in the melody. To better capture these variations, we propose to approach this task as a regression problem. Apart from predicting only the pitch for a particular region in the audio, we also predict its uncertainty to enhance the trustworthiness of the model. To perform regression-based melody estimation, we propose three different methods that use histogram representation to model the pitch values. Such a representation requires the support range of the histogram to be continuous. The first two methods address the abrupt discontinuity between unvoiced and voiced frequency ranges by mapping them to a continuous range. The third method reformulates melody estimation as a fully Bayesian task, modeling voicing detection as a classification problem, and voiced pitch estimation as a regression problem. Additionally, we introduce a novel method to estimate the uncertainty from the histogram representation that correlates well with the deviation of the mean of the predicted distribution from the ground truth. Experimental results demonstrate that reformulating melody estimation as a regression problem significantly improves the performance over classification-based approaches. Comparing the proposed methods with a state-of-the-art regression model, it is observed that the Bayesian method performs the best at estimating both the melody and its associated uncertainty.


【4】 Listen to Extract: Onset-Prompted Target Speaker Extraction

标题: 收听摘录:起始点确定的目标说话人提取
链接:https://arxiv.org/abs/2505.05114
作者: Pengjie Shen,  Kangrui Chen,  Shulin He,  Pengru Chen,  Shuqi Yuan,  He Kong,  Xueliang Zhang,  Zhong-Qiu Wang 
备注:in submission
摘要:我们提出了$\textit{listen to extract}$(LExt),一种高效而又非常简单的单声道目标说话人提取(TSE)算法。给定一个目标说话人的注册话语,LExt的目的是从说话人与其他说话人的混合语音中提取目标说话人。对于每个混合,LExt将目标说话者的登记话语连接到波形级别的混合信号,并训练深度神经网络(DNN)以基于连接的混合信号提取目标语音。其基本原理是,通过这种方式,为目标说话者创建人工语音开始,并且它可以提示DNN(a)提取哪个说话者是目标;以及(b)可以帮助提取的目标说话者的频谱时间模式。这种简单的方法在多个公共TSE数据集上产生了强大的TSE性能,包括WSJ 0 - 2 mix,WHAM!哇!
摘要:We propose $\textit{listen to extract}$ (LExt), a highly-effective while extremely-simple algorithm for monaural target speaker extraction (TSE). Given an enrollment utterance of a target speaker, LExt aims at extracting the target speaker from the speaker's mixed speech with other speakers. For each mixture, LExt concatenates an enrollment utterance of the target speaker to the mixture signal at the waveform level, and trains deep neural networks (DNN) to extract the target speech based on the concatenated mixture signal. The rationale is that, this way, an artificial speech onset is created for the target speaker and it could prompt the DNN (a) which speaker is the target to extract; and (b) spectral-temporal patterns of the target speaker that could help extraction. This simple approach produces strong TSE performance on multiple public TSE datasets including WSJ0-2mix, WHAM! and WHAMR!.


【5】 From Dialect Gaps to Identity Maps: Tackling Variability in Speaker  Verification

标题: 从方言差距到身份地图:解决说话人验证中的可变性
链接:https://arxiv.org/abs/2505.04629
作者: Abdulhady Abas Abdullah,  Soran Badawi,  Dana A. Abdullah,  Dana Rasul Hamad,  Hanan Abdulrahman Taher,  Sabat Salih Muhamad,  Aram Mahmood Ahmed,  Bryar A. Hassan,  Sirwan Abdolwahed Aula,  Tarik A. Rashid 
摘要:研究了库尔德语说话人检测的复杂性和困难性。由于库尔德语在语音和词汇上的巨大差异,库尔德语的几种方言(包括Kurmanji、Sorani和Hawrami)给说话人识别系统带来了特殊的挑战。在这项工作中的主要困难,建立一个强大的说话人识别系统,能够精确地识别说话人跨越几种方言。为了提高这些系统的准确性和可靠性,它还提出了一些解决方案,如复杂的机器学习方法,数据增强策略,以及建立全面的方言语料库。结果表明,为每种方言定制的策略以及跨方言训练大大提高了识别性能。
摘要:The complexity and difficulties of Kurdish speaker detection among its several dialects are investigated in this work. Because of its great phonetic and lexical differences, Kurdish with several dialects including Kurmanji, Sorani, and Hawrami offers special challenges for speaker recognition systems. The main difficulties in building a strong speaker identification system capable of precisely identifying speakers across several dialects are investigated in this work. To raise the accuracy and dependability of these systems, it also suggests solutions like sophisticated machine learning approaches, data augmentation tactics, and the building of thorough dialect-specific corpus. The results show that customized strategies for every dialect together with cross-dialect training greatly enhance recognition performance.


【6】 Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound  Source Localization

标题: 通过CLIP听和看:自我监督的声音源定位框架
链接:https://arxiv.org/abs/2505.05343
作者: Sooyoung Park,  Arda Senocak,  Joon Son Chung 
备注:Journal Extension of WACV 2024 paper (arXiv:2311.04066). Code is available at this https URL
摘要:大型视觉语言模型在不同的任务中表现出强大的多模态对齐和泛化能力。其中,CLIP是最成功的方法之一。在这项工作中,我们将CLIP的应用扩展到声源定位,提出了一种无需显式文本输入的自监督方法。我们引入了一个框架,将音频映射到与CLIP的文本编码器兼容的令牌,产生音频驱动的嵌入。这些嵌入用于生成发声区域掩码,从该发声区域掩码中提取视觉特征,并通过对比性视听对应目标与音频嵌入对齐。我们的研究结果表明,预训练的多模态基础模型的对齐知识使我们的方法能够为探测对象生成更完整和紧凑的定位。我们进一步提出了一个LLM指导的扩展,在训练过程中将对象感知的视听场景理解提取到模型中,以增强对齐。五个不同任务的广泛实验表明,我们的方法,在所有的变种,优于国家的最先进的方法,并实现了强大的泛化在zero-shot设置。
摘要:Large-scale vision-language models demonstrate strong multimodal alignment and generalization across diverse tasks. Among them, CLIP stands out as one of the most successful approaches. In this work, we extend the application of CLIP to sound source localization, proposing a self-supervised method operates without explicit text input. We introduce a framework that maps audios into tokens compatible with CLIP's text encoder, producing audio-driven embeddings. These embeddings are used to generate sounding region masks, from which visual features are extracted and aligned with the audio embeddings through a contrastive audio-visual correspondence objective. Our findings show that alignment knowledge of pre-trained multimodal foundation model enables our method to generate more complete and compact localization for sounding objects. We further propose an LLM-guided extension that distills object-aware audio-visual scene understanding into the model during training to enhance alignment. Extensive experiments across five diverse tasks demonstrate that our method, in all variants, outperforms state-of-the-art approaches and achieves strong generalization in zero-shot settings.


【7】 FLAM: Frame-Wise Language-Audio Modeling

标题: FLAM:帧级格式音频建模
链接:https://arxiv.org/abs/2505.05335
作者: Yusong Wu,  Christos Tsirigotis,  Ke Chen,  Cheng-Zhi Anna Huang,  Aaron Courville,  Oriol Nieto,  Prem Seetharaman,  Justin Salamon 
备注:Accepted at ICML 2025
摘要:最近的多模态音频语言模型(ALM)擅长文本音频检索,但与帧明智的音频理解的斗争。以前的工作使用时间感知标签或无监督训练来提高帧的能力,但它们仍然缺乏细粒度的标签能力来确定事件发生的时间。虽然传统的声音事件检测模型可以精确地定位事件,但它们仅限于预定义的类别,这使得它们对于具有分布外事件的真实场景无效。在这项工作中,我们介绍FLAM,一个开放的词汇对比音频语言模型能够本地化特定的声音事件。FLAM采用了一个高效的内存和校准的逐帧目标与logit调整,以解决虚假的相关性,如事件依赖性和标签在训练过程中的不平衡。为了实现逐帧监督,我们利用了具有不同音频事件、LLM生成的字幕和模拟的大规模数据集。实验结果和案例研究表明,FLAM显着提高了开放词汇的本地化能力,同时保持了强大的性能,在全球检索和下游任务。
摘要:Recent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervised training to improve frame-wise capabilities, but they still lack fine-grained labeling capability to pinpoint when an event occurs. While traditional sound event detection models can precisely localize events, they are limited to pre-defined categories, making them ineffective for real-world scenarios with out-of-distribution events. In this work, we introduce FLAM, an open-vocabulary contrastive audio-language model capable of localizing specific sound events. FLAM employs a memory-efficient and calibrated frame-wise objective with logit adjustment to address spurious correlations, such as event dependencies and label imbalances during training. To enable frame-wise supervision, we leverage a large-scale dataset with diverse audio events, LLM-generated captions and simulation. Experimental results and case studies demonstrate that FLAM significantly improves the open-vocabulary localization capability while maintaining strong performance in global retrieval and downstream tasks.


【8】 Pairing Real-Time Piano Transcription with Symbol-level Tracking for  Precise and Robust Score Following

标题: 将实时钢琴抄写与符号级跟踪相结合,实现精确且稳健的乐谱跟踪
链接:https://arxiv.org/abs/2505.05078
作者: Silvan Peter,  Patricia Hu,  Gerhard Widmer 
备注:5 pages, 3 tables, 2 pseudocodes, to be published at the Sound and Music Computing Conference 2025
摘要:实时音乐跟踪系统跟踪音乐表演,并随时报告相应乐谱中的当前位置。大多数现有方法专门在音频域中解决这个问题,通常对传入音频和分数的音频表示使用在线时间规整(OLTW)技术。音频OLTW技术在过去十年中在功能和模型性能方面都有了不断的改进,达到了性能平台。我们认为,转换和代表的表现在象征性的域-从而将音乐跟踪到一个象征性的任务-可以是一个更有效的方法,即使域转换是不完美的。我们的音乐跟踪系统结合了两个实时组件:一个处理音频到音符的转录和其他转录输入和得分之间的一种新的符号级跟踪器。我们将这种混合音频符号方法的性能与其等效的仅音频对应方法进行了比较,并证明我们的方法在精度,即,绝对跟踪误差和鲁棒性,即,追踪成功
摘要:Real-time music tracking systems follow a musical performance and at any time report the current position in a corresponding score. Most existing methods approach this problem exclusively in the audio domain, typically using online time warping (OLTW) techniques on incoming audio and an audio representation of the score. Audio OLTW techniques have seen incremental improvements both in features and model heuristics which reached a performance plateau in the past ten years. We argue that converting and representing the performance in the symbolic domain -- thereby transforming music tracking into a symbolic task -- can be a more effective approach, even when the domain transformation is imperfect. Our music tracking system combines two real-time components: one handling audio-to-note transcription and the other a novel symbol-level tracker between transcribed input and score. We compare the performance of this mixed audio-symbolic approach with its equivalent audio-only counterpart, and demonstrate that our method outperforms the latter in terms of both precision, i.e., absolute tracking error, and robustness, i.e., tracking success.


【9】 ReverbMiipher: Generative Speech Restoration meets Reverberation  Characteristics Controllability

标题: ReverbMiipher:生成性语音恢复满足回响特征可控制性
链接:https://arxiv.org/abs/2505.05077
作者: Wataru Nakata,  Yuma Koizumi,  Shigeki Karita,  Robin Scheibler,  Haruko Ishikawa,  Adriana Guevara-Rukoz,  Heiga Zen,  Michiel Bacchiani 
备注:5 pages, 5 figures
摘要:混响编码的声源环境的空间信息,而传统的语音恢复(SR)通常完全消除混响。我们提出了ReverbMiipher,一个扩展参数再合成框架的SR模型,旨在对语音进行降噪,同时保留并控制混响。ReverbMiipher集成了一个专用的ReverbEncoder,用于从嘈杂的输入中提取混响特征向量。该功能可以调节声码器重建语音信号,去除噪音,同时保留原始混响特征。训练过程中的随机零向量替换策略确保该特征专门编码混响,将其与其他语音属性分离。这种学习的表示通过诸如特征之间的插值、用来自其他话语的特征替换或从潜在空间采样的技术来促进混响控制。客观和主观评估证实ReverbMiipher有效地保留混响,去除其他伪影,并优于传统的两阶段SR和卷积模拟房间脉冲响应方法。我们进一步证明了它的能力,通过功能操作产生新的混响效果。
摘要:Reverberation encodes spatial information regarding the acoustic source environment, yet traditional Speech Restoration (SR) usually completely removes reverberation. We propose ReverbMiipher, an SR model extending parametric resynthesis framework, designed to denoise speech while preserving and enabling control over reverberation. ReverbMiipher incorporates a dedicated ReverbEncoder to extract a reverb feature vector from noisy input. This feature conditions a vocoder to reconstruct the speech signal, removing noise while retaining the original reverberation characteristics. A stochastic zero-vector replacement strategy during training ensures the feature specifically encodes reverberation, disentangling it from other speech attributes. This learned representation facilitates reverberation control via techniques such as interpolation between features, replacement with features from other utterances, or sampling from a latent space. Objective and subjective evaluations confirm ReverbMiipher effectively preserves reverberation, removes other artifacts, and outperforms the conventional two-stage SR and convolving simulated room impulse response approach. We further demonstrate its ability to generate novel reverberation effects through feature manipulation.


【10】 How to Infer Repeat Structures in MIDI Performances

标题: 如何推断工作组性能中的重复结构
链接:https://arxiv.org/abs/2505.05055
作者: Silvan Peter,  Patricia Hu,  Gerhard Widmer 
备注:3 pages, 1 figure, 1 table, to be published in the Music Encoding Conference 2025
摘要:在演奏研究和音乐信息检索中,演奏者的演奏通常是有利的,如果他们能与乐谱联系起来,那就更有利了。这种联系通常是通过对齐的方式建立的,将乐谱和演奏之间的音符或时间点联系起来。当试图建立这样的对齐时,第一个障碍是表演实现了乐谱的一个(许多)结构版本,该结构版本可以由诸如重复、变化和导航标记(如“dal segno/da capo al coda”)之类的指令来产生。在可以应用比对算法之前,需要展开分数,即,需要明确写出其重复和导航标记,以创建没有与性能匹配的跳跃的单个时间轴。在大型表演语料库的管理中,这个过程是手动进行的,因为没有工具可以推断表演的重复结构。为了简化这一过程,我们开发了一种方法来自动推断重复结构的一个重复的性能,给出了一个象征性的编码得分,包括重复和导航标记。指导我们设计的直觉是:1)乐谱的每个连续部分与包含相同材料的表演部分的局部对准应该接收高对准增益,而与任何其他表演部分的局部对准应该产生低增益或零增益。以及2)如果结构版本对应于性能,则根据得分的有效结构版本将局部比对拼接在一起应导致近似的完全比对和相应的高全局累积增益,并且对于所有其他不合适的结构版本,导致低增益。
摘要:MIDI performances are generally expedient in performance research and music information retrieval, and even more so if they can be connected to a score. This connection is usually established by means of alignment, linking either notes or time points between the score and the performance. The first obstacle when trying to establish such an alignment is that a performance realizes one (out of many) structural versions of the score that can plausibly result from instructions such as repeats, variations, and navigation markers like 'dal segno/da capo al coda'. A score needs to be unfolded, that is, its repeats and navigation markers need to be explicitly written out to create a single timeline without jumps matching the performance, before alignment algorithms can be applied. In the curation of large performance corpora this process is carried out manually, as no tools are available to infer the repeat structure of the performance. To ease this process, we develop a method to automatically infer the repeat structure of a MIDI performance, given a symbolically encoded score including repeat and navigation markers. The intuition guiding our design is: 1) local alignment of every contiguous section of the score with a section of a performance containing the same material should receive high alignment gain, whereas local alignment with any other performance section should accrue a low or zero gain. And 2) stitching local alignments together according to a valid structural version of the score should result in an approximate full alignment and correspondingly high global accumulated gain if the structural version corresponds to the performance, and low gain for all other, ill-fitting structural versions.


【11】 Inter-Diffusion Generation Model of Speakers and Listeners for Effective  Communication

标题: 有效沟通的说话者和听众的相互扩散生成模型
链接:https://arxiv.org/abs/2505.04996
作者: Jinhe Huang,  Yongkang Cheng,  Yuming Hang,  Gaoge Han,  Jinewei Li,  Jing Zhang,  Xingjian Gu 
备注:accepted by ICMR 2025
摘要:全身手势在自然交互中起着关键作用,对于实现有效沟通至关重要。然而,现有的研究大多集中在说话人的手势生成上,忽视了听话人在互动过程中的重要作用,未能充分探讨他们之间的动态互动。本文创新性地提出了一种有效交际的说者和听者互扩散生成模型。我们第一次将听众的全身手势整合到生成框架中。通过设计一种新的交互扩散机制,该模型可以准确地捕捉说话人和听话人在交流过程中的复杂交互模式。在模型构建过程中,基于先进的扩散模型架构,创新性地引入交互条件和GAN模型,增加去噪步长。因此,在生成手势序列时,模型不仅可以基于说话者的语音信息动态生成,而且还可以实时响应收听者的反馈,从而实现两者之间的协同交互。大量的实验结果表明,与现有的手势生成方法相比,本文提出的模型在手势的自然度、连贯性和语音-手势同步性等方面都有了显著的提高。在主观评价实验中,用户对生成的交互场景给予了高度评价,认为它们更接近于现实生活中的人际交往场景。客观的指标评价也表明,我们的模型在多个关键指标上优于基线方法,为有效沟通提供了更有力的支持。
摘要:Full-body gestures play a pivotal role in natural interactions and are crucial for achieving effective communication. Nevertheless, most existing studies primarily focus on the gesture generation of speakers, overlooking the vital role of listeners in the interaction process and failing to fully explore the dynamic interaction between them. This paper innovatively proposes an Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication. For the first time, we integrate the full-body gestures of listeners into the generation framework. By devising a novel inter-diffusion mechanism, this model can accurately capture the complex interaction patterns between speakers and listeners during communication. In the model construction process, based on the advanced diffusion model architecture, we innovatively introduce interaction conditions and the GAN model to increase the denoising step size. As a result, when generating gesture sequences, the model can not only dynamically generate based on the speaker's speech information but also respond in realtime to the listener's feedback, enabling synergistic interaction between the two. Abundant experimental results demonstrate that compared with the current state-of-the-art gesture generation methods, the model we proposed has achieved remarkable improvements in the naturalness, coherence, and speech-gesture synchronization of the generated gestures. In the subjective evaluation experiments, users highly praised the generated interaction scenarios, believing that they are closer to real life human communication situations. Objective index evaluations also show that our model outperforms the baseline methods in multiple key indicators, providing more powerful support for effective communication.


【12】 A Multi-Agent AI Framework for Immersive Audiobook Production through  Spatial Audio and Neural Narration

标题: 基于空间音频和神经叙事的沉浸式有声读物制作的多智能体AI框架
链接:https://arxiv.org/abs/2505.04885
作者: Shaja Arul Selvamani,  Nia D'Souza Ganapathy 
摘要:这项研究介绍了一种创新的AI驱动的多代理框架,专门用于创建沉浸式有声读物。该框架利用FastSpeech 2和VALL-E的神经文本到语音合成来进行富有表现力的叙述和特定于角色的声音,采用先进的语言模型来自动解释文本叙述并生成逼真的空间音频效果。这些声音效果通过复杂的时间整合方法与故事情节动态同步,包括动态时间弯曲(DTW)和递归神经网络(RNN)。基于扩散的生成模型与高阶立体混响(HOA)和散射延迟网络(SDN)相结合,实现了高度逼真的3D音景,大大增强了听众的沉浸感和叙事真实感。这项技术大大推进了有声读物应用,为视障观众提供了更丰富的教育内容体验、讲故事平台和无障碍解决方案。未来的工作将解决个性化、合成声音的道德管理以及与多感官平台的集成问题。
摘要:This research introduces an innovative AI-driven multi-agent framework specifically designed for creating immersive audiobooks. Leveraging neural text-to-speech synthesis with FastSpeech 2 and VALL-E for expressive narration and character-specific voices, the framework employs advanced language models to automatically interpret textual narratives and generate realistic spatial audio effects. These sound effects are dynamically synchronized with the storyline through sophisticated temporal integration methods, including Dynamic Time Warping (DTW) and recurrent neural networks (RNNs). Diffusion-based generative models combined with higher-order ambisonics (HOA) and scattering delay networks (SDN) enable highly realistic 3D soundscapes, substantially enhancing listener immersion and narrative realism. This technology significantly advances audiobook applications, providing richer experiences for educational content, storytelling platforms, and accessibility solutions for visually impaired audiences. Future work will address personalization, ethical management of synthesized voices, and integration with multi-sensory platforms.


【13】 Data Standards in Audiology: A Mixed-Methods Exploration of Community  Perspectives and Implementation Considerations

标题: 听力学数据标准:社区观点和实施考虑因素的混合方法探索
链接:https://arxiv.org/abs/2505.04728
作者: Charlotte Vercammen,  Antje Heinrich,  Christophe Lesimple,  Alessia Paglialonga,  Jan-Willem A. Wasmann,  Mareike Buhl 
摘要:目的:本研究的目的是探索听力学数据标准化的选择,并记录全球听力学界对数据标准的当前知识和观点,探索他们的需求和偏好,并因此制定数据标准化建议。   设计图:在2024年计算听力学虚拟会议的“听力学大数据和数据标准”特别会议期间,采用混合方法,将结构化调查与专家对主题的深入探索相结合。   研究样本:调查样本由全球听力学社区的82名成员组成; 5名专家参加了小组讨论。   结果:调查结果强调了听力学数据标准化的必要性,旨在促进研究和改善患者护理。对现有举措的了解程度较低:38%的人了解这些举措。然而,90%的人希望为他们的发展做出贡献。小组讨论探讨了听力学中新兴的标准化倡议(OMOP,openEHR,HIMSA的Noah标准),挑战(例如,数据质量和隐私),以及机会(例如,方法之间的转换以及与其他医学领域的协同作用)。   结论:本研究中确定的社区支持可以用来进一步制定听力学标准化计划,确保计划之间以及与其他医学领域的一致性。
摘要:Objective: The purpose of this study was to explore options for data standardisation in audiology and document the global audiology community's current knowledge and views of data standards, explore their needs and preferences, and develop recommendations for data standardisation as a result.   Design: A mixed-methods approach, combining a structured survey with an in-depth exploration of themes by experts during a special session on "Big Data and Data Standards in Audiology" at the 2024 Virtual Conference of Computational Audiology.   Study Sample: The survey sample consisted of 82 members of the global audiology community; five experts joined the panel discussion.   Results: Survey results emphasized the need for data standardisation in audiology aimed at facilitating research and improving patient care. Knowledge of existing initiatives was low: 38% were aware of initiatives. Yet, 90% envisioned contributing to them moving forward. The panel discussion explored emerging standardisation initiatives in audiology (OMOP, openEHR, HIMSA's Noah standard), challenges (e.g., data quality and privacy), and opportunities (e.g., conversion between approaches and synergies with other medical fields).   Conclusions: The community support identified in this study could be leveraged to further develop standardisation initiatives for audiology, ensuring alignment between initiatives and with other medical fields.


机器翻译由腾讯交互翻译提供,仅供参考