今日论文合集:cs.SD语音20篇,eess.AS音频处理17篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音



【1】 A Variational Framework for Improving Naturalness in Generative Spoken  Language Models
标题: 提高生成性口语模型自然性的变分框架
链接:https://arxiv.org/abs/2506.14767
作者: Li-Wei Chen,  Takuya Higuchi,  Zakaria Aldeneh,  Ahmed Hussen Abdelaziz,  Alexander Rudnicky 
备注:International Conference on Machine Learning (ICML) 2025
摘要:大型语言模型在文本处理中的成功激发了它们对语音建模的适应。然而,由于语音是连续和复杂的,它往往是离散化的自回归建模。从自监督模型中得到的语音标记(称为语义标记)通常关注语音的语言学方面,但忽略了韵律信息。因此,在这些令牌上训练的模型可以生成自然度降低的语音。现有的方法试图通过向语义标记添加音高特征来解决这个问题。然而,音高本身并不能完全代表语言属性的范围,选择正确的特征需要仔细的手工设计。为了克服这一点,我们提出了一种端到端的变分方法,自动学习编码这些连续的语音属性,以增强语义令牌。我们的方法消除了手动提取和选择的非语言特征的需要。此外,它根据人类评分员产生优选的语音延续。代码、示例和模型可在https://github.com/b04901014/vae-gslm上获得。
摘要:The success of large language models in text processing has inspired their adaptation to speech modeling. However, since speech is continuous and complex, it is often discretized for autoregressive modeling. Speech tokens derived from self-supervised models (known as semantic tokens) typically focus on the linguistic aspects of speech but neglect prosodic information. As a result, models trained on these tokens can generate speech with reduced naturalness. Existing approaches try to fix this by adding pitch features to the semantic tokens. However, pitch alone cannot fully represent the range of paralinguistic attributes, and selecting the right features requires careful hand-engineering. To overcome this, we propose an end-to-end variational approach that automatically learns to encode these continuous speech attributes to enhance the semantic tokens. Our approach eliminates the need for manual extraction and selection of paralinguistic features. Moreover, it produces preferred speech continuations according to human raters. Code, samples and models are available at https://github.com/b04901014/vae-gslm.


【2】 Exploring Speaker Diarization with Mixture of Experts

标题: 与混合专家一起探索发言者日记化
链接:https://arxiv.org/abs/2506.14750
作者: Gaobin Yang,  Maokui He,  Shutong Niu,  Ruoyu Wang,  Hang Chen,  Jun Du 
摘要:在本文中,我们提出了一种新的神经说话人日记系统使用记忆感知的多说话人嵌入与序列到序列架构(NSD-MS 2S),它集成了一个记忆感知的多说话人嵌入模块与序列到序列架构。该系统利用存储器模块来增强扬声器嵌入,并采用Seq 2Seq框架来有效地将声学特征映射到扬声器标签。此外,我们探讨了混合专家在说话人日志中的应用,并引入了共享和软混合专家(SS-MoE)模块,以进一步减轻模型偏差,提高性能。对SS-MoE模型进行了简化,得到了扩展模型NSD-MS 2S-SSMoE。在包括CHiME-6、DiPCo、Mixer 6和DIHARD-III评估集在内的多个复杂声学数据集上进行的实验表明,在鲁棒性和泛化方面都有了有意义的改进。所提出的方法实现了最先进的结果,展示了它们在具有挑战性的现实世界场景中的有效性。
摘要:In this paper, we propose a novel neural speaker diarization system using memory-aware multi-speaker embedding with sequence-to-sequence architecture (NSD-MS2S), which integrates a memory-aware multi-speaker embedding module with a sequence-to-sequence architecture. The system leverages a memory module to enhance speaker embeddings and employs a Seq2Seq framework to efficiently map acoustic features to speaker labels. Additionally, we explore the application of mixture of experts in speaker diarization, and introduce a Shared and Soft Mixture of Experts (SS-MoE) module to further mitigate model bias and enhance performance. Incorporating SS-MoE leads to the extended model NSD-MS2S-SSMoE. Experiments on multiple complex acoustic datasets, including CHiME-6, DiPCo, Mixer 6 and DIHARD-III evaluation sets, demonstrate meaningful improvements in robustness and generalization. The proposed methods achieve state-of-the-art results, showcasing their effectiveness in challenging real-world scenarios.


【3】 Adaptive Accompaniment with ReaLchords

标题: 带有ReaLchords的自适应伴奏
链接:https://arxiv.org/abs/2506.14723
作者: Yusong Wu,  Tim Cooijmans,  Kyle Kastner,  Adam Roberts,  Ian Simon,  Alexander Scarlatos,  Chris Donahue,  Cassie Tarakajian,  Shayegan Omidshafiei,  Aaron Courville,  Pablo Samuel Castro,  Natasha Jaques,  Cheng-Zhi Anna Huang 
备注:Accepted by ICML 2024
摘要:干扰需要音乐家之间的协调,预期和合作创造力。当前的音乐生成模型产生表达性输出,但不能以在线方式生成,这意味着与其他音乐家(人类或其他)同时生成。我们提出了ReaLchords,一个在线生成模型即兴和弦伴奏用户旋律。我们从最大似然预训练的在线模型开始,并使用强化学习来微调模型以供在线使用。微调目标利用了一个新的奖励模型,该模型提供了对旋律和和弦之间的谐波和时间相干性的反馈,以及一个发散项,该发散项实现了一种新的类型的从教师模型中提取的方法,该教师模型可以看到未来的旋律。通过定量实验和听力测试,我们证明了所得到的模型能够很好地适应不熟悉的输入,并产生合适的伴奏。ReaLchords为现场干扰以及其他形式的同步共同创作打开了大门。
摘要:Jamming requires coordination, anticipation, and collaborative creativity between musicians. Current generative models of music produce expressive output but are not able to generate in an \emph{online} manner, meaning simultaneously with other musicians (human or otherwise). We propose ReaLchords, an online generative model for improvising chord accompaniment to user melody. We start with an online model pretrained by maximum likelihood, and use reinforcement learning to finetune the model for online use. The finetuning objective leverages both a novel reward model that provides feedback on both harmonic and temporal coherency between melody and chord, and a divergence term that implements a novel type of distillation from a teacher model that can see the future melody. Through quantitative experiments and listening tests, we demonstrate that the resulting model adapts well to unfamiliar input and produce fitting accompaniment. ReaLchords opens the door to live jamming, as well as simultaneous co-creation in other modalities.


【4】 Refining music sample identification with a self-supervised graph neural  network

标题: 使用自我监督图神经网络完善音乐样本识别
链接:https://arxiv.org/abs/2506.14684
作者: Aditya Bhattacharjee,  Ivan Meresman Higgs,  Mark Sandler,  Emmanouil Benetos 
备注:Accepted at International Conference for Music Information Retrieval (ISMIR) 2025
摘要:自动样本识别(ASID),检测和识别已被重新使用在新的音乐作品的音频记录的部分,是一个基本的,但具有挑战性的任务,在音频查询为基础的检索领域。虽然相关的任务,音频指纹,在“真实世界”(嘈杂,混响)条件下准确检索音乐内容取得了重大进展,ASID系统难以识别经历了音乐修改的样本。因此,一个对常见的音乐制作转换(如时间拉伸、音高转换、效果处理以及底层或覆盖音乐)具有鲁棒性的系统是一个重要的开放性挑战。   在这项工作中,我们提出了一个轻量级的和可扩展的编码架构,采用一个对比学习框架内的图神经网络。与当前最先进的系统相比,我们的模型仅使用了9%的可训练参数,同时实现了相当的性能,达到了44.2%的平均精度(mAP)。   为了提高检索质量,我们引入了一个两阶段的方法,包括一个初步的粗相似性搜索的候选人选择,然后由一个交叉注意力分类器,拒绝不相关的匹配和细化检索到的候选人的排名-一个基本的能力,在以前的模型中没有。此外,由于现实应用中的查询通常持续时间很短,因此我们使用Sample 100数据集的新细粒度注释对我们的系统进行了短查询基准测试,我们将其作为这项工作的一部分发布。
摘要:Automatic sample identification (ASID), the detection and identification of portions of audio recordings that have been reused in new musical works, is an essential but challenging task in the field of audio query-based retrieval. While a related task, audio fingerprinting, has made significant progress in accurately retrieving musical content under "real world" (noisy, reverberant) conditions, ASID systems struggle to identify samples that have undergone musical modifications. Thus, a system robust to common music production transformations such as time-stretching, pitch-shifting, effects processing, and underlying or overlaying music is an important open challenge.   In this work, we propose a lightweight and scalable encoding architecture employing a Graph Neural Network within a contrastive learning framework. Our model uses only 9% of the trainable parameters compared to the current state-of-the-art system while achieving comparable performance, reaching a mean average precision (mAP) of 44.2%.   To enhance retrieval quality, we introduce a two-stage approach consisting of an initial coarse similarity search for candidate selection, followed by a cross-attention classifier that rejects irrelevant matches and refines the ranking of retrieved candidates - an essential capability absent in prior models. In addition, because queries in real-world applications are often short in duration, we benchmark our system for short queries using new fine-grained annotations for the Sample100 dataset, which we publish as part of this work.


【5】 Evolving music theory for emerging musical languages

标题: 为新兴音乐语言发展音乐理论
链接:https://arxiv.org/abs/2506.14504
作者: Emmanuel Deruty 
备注:In Music 2025, Innovation in Music Conference. 20-22 June, 2025, Bath Spa University, Bath, UK
摘要:本章重新考虑了当代流行音乐(CPM)中的音高概念,特别是在传统假设可能失败的电子背景下。运用现象学和归纳法,本文认为音高不是一个本体论上的客观属性,而是一个由听者和环境塑造的知觉建构。对准调和调的分析表明,一个声调可以传达多个音高,从而产生声调分裂。音调的感知也可能是多稳态的,随着时间的推移,同一听众的感知也会发生变化。在这个框架中,调律系统可以从音调的内部结构中出现。与海岸线悖论的平行支持了基于感知变化的音高模型,挑战了继承的理论规范。
摘要:This chapter reconsiders the concept of pitch in contemporary popular music (CPM), particularly in electronic contexts where traditional assumptions may fail. Drawing on phenomenological and inductive methods, it argues that pitch is not an ontologically objective property but a perceptual construct shaped by listeners and conditions. Analyses of quasi-harmonic tones reveal that a single tone can convey multiple pitches, giving rise to tonal fission. The perception of pitch may also be multistable, varying for the same listener over time. In this framework, the tuning system may emerge from a tone's internal structure. A parallel with the coastline paradox supports a model of pitch grounded in perceptual variability, challenging inherited theoretical norms.


【6】 An Open Research Dataset of the 1932 Cairo Congress of Arab Music

标题: 1932年开罗阿拉伯音乐大会的开放研究数据集
链接:https://arxiv.org/abs/2506.14503
作者: Baris Bozkurt (College of Interdisciplinary Studies, Zayed University, Dubai, United Arab Emirates) 
备注:14 pages, 4 figures, 4 tables
摘要:本文介绍了ord-cc 32,一个开放的研究数据集来自1932年开罗大会的阿拉伯音乐录音,一个具有历史意义的集合,代表了不同的阿拉伯音乐传统。该数据集包括结构化元数据、旋律和节奏模式标签(maqam和iqa)、手动标记的主音信息以及使用最先进的音高检测方法提取的声学特征。这些资源支持阿拉伯音乐的调音、音律和区域变化的计算研究。使用音高直方图的案例研究表明,跨区域的微色调差异的数据驱动分析的潜力。通过公开提供该数据集,我们的目标是实现计算民族音乐学,音乐信息检索(MIR),文化研究和数字遗产保护的跨学科研究。Order-CC 32在Zenodo上与用于特征提取和元数据检索的工具共享。
摘要:This paper introduces ORD-CC32 , an open research dataset derived from the 1932 Cairo Congress of Arab Music recordings, a historically significant collection representing diverse Arab musical traditions. The dataset includes structured metadata, melodic and rhythmic mode tags (maqam and iqa), manually labeled tonic information, and acoustic features extracted using state-of-the-art pitch detection methods. These resources support computational studies of tuning, temperament, and regional variations in Arab music. A case study using pitch histograms demonstrates the potential for data-driven analysis of microtonal differences across regions. By making this dataset openly available, we aim to enable interdisciplinary research in computational ethnomusicology, music information retrieval (MIR), cultural studies, and digital heritage preservation. ORD-CC32 is shared on Zenodo with tools for feature extraction and metadata retrieval.


【7】 Unifying Streaming and Non-streaming Zipformer-based ASR

标题: 统一流媒体和非流媒体基于Zipformer的ASB
链接:https://arxiv.org/abs/2506.14434
作者: Bidisha Sharma,  Karthik Pandia Durai,  Shankar Venkatesan,  Jeena J Prakash,  Shashi Kumar,  Malolan Chetlur,  Andreas Stolcke 
备注:Accepted in ACL2025 Industry track
摘要:人们对统一流和非流自动语音识别(ASR)模型以降低开发、培训和部署成本的兴趣越来越大。我们提出了一个统一的框架,训练一个单一的端到端的ASR模型流和非流应用程序,利用未来的上下文信息。我们建议在基于zipformer的ASR模型的训练中通过分块注意掩蔽来使用动态右上下文。我们证明,使用右上下文是更有效的zipformer模型相比,其他构象模型,由于其多尺度的性质。我们分析了右上下文帧的数量变化对流式ASR模型的准确性和延迟的影响。我们使用Librispeech和大型内部会话数据集来训练不同版本的流和非流模型,并在不同领域的各种测试集上的生产级服务器-客户端设置中对其进行评估。所提出的策略减少了相对7.9%的字错误,在用户感知的延迟一个小的退化。通过添加更多的右上下文帧,我们能够实现接近非流模型的流性能。我们的方法还允许根据客户的要求灵活控制延迟-准确性权衡。
摘要:There has been increasing interest in unifying streaming and non-streaming automatic speech recognition (ASR) models to reduce development, training, and deployment costs. We present a unified framework that trains a single end-to-end ASR model for both streaming and non-streaming applications, leveraging future context information. We propose to use dynamic right-context through the chunked attention masking in the training of zipformer-based ASR models. We demonstrate that using right-context is more effective in zipformer models compared to other conformer models due to its multi-scale nature. We analyze the effect of varying the number of right-context frames on accuracy and latency of the streaming ASR models. We use Librispeech and large in-house conversational datasets to train different versions of streaming and non-streaming models and evaluate them in a production grade server-client setup across diverse testsets of different domains. The proposed strategy reduces word error by relative 7.9\% with a small degradation in user-perceived latency. By adding more right-context frames, we are able to achieve streaming performance close to that of non-streaming models. Our approach also allows flexible control of the latency-accuracy tradeoff according to customers requirements.


【8】 A Comparative Study on Proactive and Passive Detection of Deepfake  Speech

标题: Deepfake语音主动和被动检测的比较研究
链接:https://arxiv.org/abs/2506.14398
作者: Chia-Hua Wu,  Wanying Ge,  Xin Wang,  Junichi Yamagishi,  Yu Tsao,  Hsin-Min Wang 
摘要:防御deepfake语音的解决方案分为两类:主动水印模型和被动传统deepfake检测器。虽然两者都解决了共同的威胁,但它们在训练,优化和评估方面的差异阻碍了联合评估和为不同情况选择最佳解决方案的统一协议。这项工作提出了一个框架来评估deepfake语音检测中的两种模型类型。为了确保公平比较并最大限度地减少差异,所有模型都在通用数据集上进行了训练和测试,并使用共享指标进行了性能评估。我们还分析了它们对各种对抗性攻击的鲁棒性,表明不同的模型对不同的语音属性失真表现出不同的脆弱性。我们的培训和评估代码可以在Github上找到。
摘要:Solutions for defending against deepfake speech fall into two categories: proactive watermarking models and passive conventional deepfake detectors. While both address common threats, their differences in training, optimization, and evaluation prevent a unified protocol for joint evaluation and selecting the best solutions for different cases. This work proposes a framework to evaluate both model types in deepfake speech detection. To ensure fair comparison and minimize discrepancies, all models were trained and tested on common datasets, with performance evaluated using a shared metric. We also analyze their robustness against various adversarial attacks, showing that different models exhibit distinct vulnerabilities to different speech attribute distortions. Our training and evaluation code is available at Github.


【9】 Manipulated Regions Localization For Partially Deepfake Audio: A Survey

标题: 部分Deepfake音频的操纵区域定位:一项调查
链接:https://arxiv.org/abs/2506.14396
作者: Jiayi He,  Jiangyan Yi,  Jianhua Tao,  Siding Zeng,  Hao Gu 
摘要:随着音频deepfake技术的发展,使用部分deepfake音频的攻击开始增加。与完全deepfake相比,由于部分加密操作,检测器更难识别,导致更高的安全风险。虽然已经开展了一些研究,但没有全面的审查,系统地介绍解决这一问题的现状和发展趋势。因此,在本次调查中,我们首次对部分deepfake音频操作区域定位任务进行了系统的介绍,包括现有方法的基本原理、分支、当前限制和潜在趋势,为这一领域提供了一个有启发性的见解。
摘要:With the development of audio deepfake techniques, attacks with partially deepfake audio are beginning to rise. Compared to fully deepfake, it is much harder to be identified by the detector due to the partially cryptic manipulation, resulting in higher security risks. Although some studies have been launched, there is no comprehensive review to systematically introduce the current situations and development trends for addressing this issue. Thus, in this survey, we are the first to outline a systematic introduction for partially deepfake audio manipulated region localization tasks, including the fundamentals, branches of existing methods, current limitations and potential trends, providing a revealing insight into this scope.


【10】 SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative  music modeling

标题: SLEEPING-DISCO 9 M:用于生成式音乐建模的大规模预训练数据集
链接:https://arxiv.org/abs/2506.14293
作者: Tawsif Ahmed,  Andrej Radonjic,  Gollam Rabby 
摘要:我们介绍了Sleeping-DISCO 9 M,这是一个用于音乐和歌曲的大规模预训练数据集。据我们所知,没有开源的高质量数据集代表流行和知名的歌曲,用于生成音乐建模任务,如文本音乐,音乐字幕,歌唱语音合成,旋律重建和跨模型检索。过去的贡献集中在孤立和受约束的因素,其核心观点是创建合成或重新录制的音乐语料库(例如GTSinger,M4 Singer)和任意大规模的音频数据集(例如DISCO-10 M和LAIONDISCO-12 M)一直是社区的另一个焦点。不幸的是,这些数据集在生成音乐社区中的采用率一直低于实质性水平,因为这些数据集未能反映真实世界的音乐及其风味。我们的数据集改变了这种叙述,并提供了一个使用实际流行音乐和世界知名艺术家构建的数据集。
摘要:We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song. To the best of our knowledge, there are no open-source high-quality dataset representing popular and well-known songs for generative music modeling tasks such as text-music, music-captioning, singing-voice synthesis, melody reconstruction and cross-model retrieval. Past contributions focused on isolated and constrained factors whose core perspective was to create synthetic or re-recorded music corpus (e.g. GTSinger, M4Singer) and arbitrarily large-scale audio datasets (e.g. DISCO-10M and LAIONDISCO-12M) had been another focus for the community. Unfortunately, adoption of these datasets has been below substantial in the generative music community as these datasets fail to reflect real-world music and its flavour. Our dataset changes this narrative and provides a dataset that is constructed using actual popular music and world-renowned artists.


【11】 Investigation of Zero-shot Text-to-Speech Models for Enhancing  Short-Utterance Speaker Verification

标题: 用于增强短言语说话人验证的Zero-Shot文本到语音模型研究
链接:https://arxiv.org/abs/2506.14226
作者: Yiyang Zhao,  Shuai Wang,  Guangzhi Sun,  Zehua Chen,  Chao Zhang,  Mingxing Xu,  Thomas Fang Zheng 
摘要:短话语说话人确认是一个很大的挑战,因为短话语段中的信息有限,这会影响准确性和可靠性。近年来,zero-shot文本到语音(TTS)系统在保持说话人身份方面取得了相当大的进展。在这项研究中,我们探索,第一次,使用语音合成TTS系统的测试时间数据增强说话人验证。我们在VoxCeleb 1数据集上评估了三个最先进的预训练的语音合成TTS系统,NatureSpeech 3,CosyVoice和MaskGCT。我们的实验结果表明,结合真实和合成语音样本导致10%-16%的相对相等错误率(EER)降低在所有持续时间,特别是显着的改善短话语,所有没有重新训练任何现有的系统。然而,我们的分析表明,较长的合成语音在降低EER方面并不像较长的真实语音那样具有相同的好处。这些研究结果突出了使用语音合成TTS进行测试时说话人确认的潜力和挑战,为未来的研究提供了见解。
摘要:Short-utterance speaker verification presents significant challenges due to the limited information in brief speech segments, which can undermine accuracy and reliability. Recently, zero-shot text-to-speech (ZS-TTS) systems have made considerable progress in preserving speaker identity. In this study, we explore, for the first time, the use of ZS-TTS systems for test-time data augmentation for speaker verification. We evaluate three state-of-the-art pre-trained ZS-TTS systems, NatureSpeech 3, CosyVoice, and MaskGCT, on the VoxCeleb 1 dataset. Our experimental results show that combining real and synthetic speech samples leads to 10%-16% relative equal error rate (EER) reductions across all durations, with particularly notable improvements for short utterances, all without retraining any existing systems. However, our analysis reveals that longer synthetic speech does not yield the same benefits as longer real speech in reducing EERs. These findings highlight the potential and challenges of using ZS-TTS for test-time speaker verification, offering insights for future research.


【12】 Fretting-Transformer: Encoder-Decoder Model for MIDI to Tablature  Transcription

标题: Fretting-Transformer:字形到表谱转录的编码器-解码器模型
链接:https://arxiv.org/abs/2506.14223
作者: Anna Hamberger,  Sebastian Murgul,  Jochen Schmidt,  Michael Heizmann 
备注:Accepted to the 50th International Computer Music Conference (ICMC), 2025
摘要:音乐转录在音乐信息检索(MIR)中起着关键作用,特别是对于像吉他这样的弦乐器,其中符号化的音乐符号(如吉他)缺乏关键的可演奏性信息。这篇文章介绍了微动Transformer,一个encoderdecoder模型,利用T5 Transformer架构,自动转录成吉他指法序列。通过将任务框定为符号翻译问题,该模型解决了关键挑战,包括弦-品歧义和物理可玩性。该系统利用了不同的数据集,包括DadaGP、GuitarToday和Leduc,并采用了新颖的数据预处理和标记化策略。我们已经制定了指法准确性和可玩性的指标,以定量评估的性能。实验结果表明,Fretting-Transformer超越了A* 等基线方法和Guitar Pro等商业应用。上下文敏感处理和调音/Capo调节的集成进一步增强了模型的性能,为自动吉他转录的未来发展奠定了坚实的基础。
摘要:Music transcription plays a pivotal role in Music Information Retrieval (MIR), particularly for stringed instruments like the guitar, where symbolic music notations such as MIDI lack crucial playability information. This contribution introduces the Fretting-Transformer, an encoderdecoder model that utilizes a T5 transformer architecture to automate the transcription of MIDI sequences into guitar tablature. By framing the task as a symbolic translation problem, the model addresses key challenges, including string-fret ambiguity and physical playability. The proposed system leverages diverse datasets, including DadaGP, GuitarToday, and Leduc, with novel data pre-processing and tokenization strategies. We have developed metrics for tablature accuracy and playability to quantitatively evaluate the performance. The experimental results demonstrate that the Fretting-Transformer surpasses baseline methods like A* and commercial applications like Guitar Pro. The integration of context-sensitive processing and tuning/capo conditioning further enhances the model's performance, laying a robust foundation for future developments in automated guitar transcription.


【13】 AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR

标题: AsyncSwitch:代码交换ASB的同步文本语音自适应
链接:https://arxiv.org/abs/2506.14190
作者: Tuan Nguyen,  Huy-Dat Tran 
备注:This work has been submitted to the IEEE for possible publication. This paper is a preprint version submitted to the 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2025)
摘要:开发代码转换ASR系统是具有挑战性的,由于语言的模糊性和有限的接触多语言,代码转换数据,而收集这样的语音是昂贵的。先前的工作从文本生成合成音频,但这些方法计算密集且难以扩展。我们介绍AsyncSwitch,一种新型的异步自适应框架,利用大规模的,文本丰富的Web数据预暴露ASR模型到不同的代码切换域,然后对配对的语音文本语料库进行微调。我们的三阶段过程(1)在代码切换文本上训练解码器自注意和前馈层,(2)使用有限的语音文本数据通过交叉注意对齐解码器和编码器,以及(3)完全微调整个模型。Whisper对马来语-英语码转换的实验表明,相对WER减少了9.02%,同时提高了新加坡式英语,马来语和其他英语变体的单语性能。
摘要:Developing code-switched ASR systems is challenging due to language ambiguity and limited exposure to multilingual, code-switched data, while collecting such speech is costly. Prior work generates synthetic audio from text, but these methods are computationally intensive and hard to scale. We introduce AsyncSwitch, a novel asynchronous adaptation framework that leverages large-scale, text-rich web data to pre-expose ASR models to diverse code-switched domains before fine-tuning on paired speech-text corpora. Our three-stage process (1) trains decoder self-attention and feedforward layers on code-switched text, (2) aligns decoder and encoder via cross-attention using limited speech-text data, and (3) fully fine-tunes the entire model. Experiments with Whisper on Malay-English code-switching demonstrate a 9.02% relative WER reduction, while improving monolingual performance in Singlish, Malay, and other English variants.


【14】 Can we train ASR systems on Code-switch without real code-switch data?  Case study for Singapore's languages

标题: 我们能否在没有真正代码交换数据的情况下在代码交换上训练ASB系统?  新加坡语言案例研究
链接:https://arxiv.org/abs/2506.14177
作者: Tuan Nguyen,  Huy-Dat Tran 
备注:Accepted by Interspeech 2025
摘要:语码转换是多语言环境中常见的现象,由于语言的复杂性导致转录数据的稀缺和昂贵,这给自动语音识别带来了挑战。本研究探讨使用合成CS数据构建CS-ASR。我们提出了一种短语级混合方法来生成模拟自然模式的合成CS数据。利用合成短语混合CS数据增强的单语来微调大型预训练ASR模型(Whisper,MMS,M4 T)。本文重点研究了三种资源不足的东南亚语言对:马来语-英语(BM-EN),普通话-马来语(ZH-BM)和泰米尔语-英语(TA-EN),为CS-ASR建立了一个新的综合基准,以评估领先的ASR模型的性能。实验结果表明,该训练策略提高了单语和CS测试中的ASR性能,其中BM-EN表现出最高的增益,然后是TA-EN和ZH-BM。这一发现为CS-ASR的开发提供了一种具有成本效益的方法,使研究和工业受益。
摘要:Code-switching (CS), common in multilingual settings, presents challenges for ASR due to scarce and costly transcribed data caused by linguistic complexity. This study investigates building CS-ASR using synthetic CS data. We propose a phrase-level mixing method to generate synthetic CS data that mimics natural patterns. Utilizing monolingual augmented with synthetic phrase-mixed CS data to fine-tune large pretrained ASR models (Whisper, MMS, SeamlessM4T). This paper focuses on three under-resourced Southeast Asian language pairs: Malay-English (BM-EN), Mandarin-Malay (ZH-BM), and Tamil-English (TA-EN), establishing a new comprehensive benchmark for CS-ASR to evaluate the performance of leading ASR models. Experimental results show that the proposed training strategy enhances ASR performance on monolingual and CS tests, with BM-EN showing highest gains, then TA-EN and ZH-BM. This finding offers a cost-effective approach for CS-ASR development, benefiting research and industry.


【15】 Pushing the Performance of Synthetic Speech Detection with  Kolmogorov-Arnold Networks and Self-Supervised Learning Models

标题: 利用Kolmogorov-Arnold网络和自我监督学习模型提高合成语音检测的性能
链接:https://arxiv.org/abs/2506.14153
作者: Tuan Dat Phuong,  Long-Vu Hoang,  Huy Dat Tran 
备注:Accepted to Interspeech 2025
摘要:语音合成技术的最新进展导致了越来越先进的欺骗攻击,给自动说话人验证系统带来了重大挑战。虽然基于自监督学习(SSL)模型的系统,特别是XLSR-Conformer模型,在合成语音检测方面表现出了卓越的性能,但仍有架构改进的空间。在本文中,我们提出了一种新的方法,取代传统的多层感知器的XLSR-构象模型与Kolmogorov-Arnold网络(KAN),一种新的架构的基础上Kolmogorov-Arnold表示定理。我们在ASVspoof 2021上的结果表明,将KAN集成到基于SSL的模型中可以在LA和DF集上相对提高60.55%的性能,并在21 LA集上进一步实现0.70%的EER。这些研究结果表明,将KAN到基于SSL的模型是一个很有前途的方向,在合成语音检测的进步。
摘要:Recent advancements in speech synthesis technologies have led to increasingly advanced spoofing attacks, posing significant challenges for automatic speaker verification systems. While systems based on self-supervised learning (SSL) models, particularly the XLSR-Conformer model, have demonstrated remarkable performance in synthetic speech detection, there remains room for architectural improvements. In this paper, we propose a novel approach that replaces the traditional Multi-Layer Perceptron in the XLSR-Conformer model with a Kolmogorov-Arnold Network (KAN), a novel architecture based on the Kolmogorov-Arnold representation theorem. Our results on ASVspoof2021 demonstrate that integrating KAN into the SSL-based models can improve the performance by 60.55% relatively on LA and DF sets, further achieving 0.70% EER on the 21LA set. These findings suggest that incorporating KAN into SSL-based models is a promising direction for advances in synthetic speech detection.


【16】 Acoustic scattering AI for non-invasive object classifications: A case  study on hair assessment

标题: 用于非侵入性物体分类的声散射人工智能:头发评估的案例研究
链接:https://arxiv.org/abs/2506.14148
作者: Long-Vu Hoang,  Tuan Nguyen,  Tran Huy Dat 
备注:Accepted to Interspeech 2025
摘要:本文提出了一种新的非侵入性的物体分类方法,使用声散射,通过头发评估的案例研究证明。当入射波与物体相互作用时,它会产生编码结构和材料特性的散射声场。通过发出声学刺激并捕获来自头部和头发样本对象的散射信号,我们使用AI驱动的基于深度学习的声音分类对头发类型和水分进行分类。我们对综合方法进行了基准测试,包括(i)完全监督的深度学习,(ii)基于嵌入的分类,(iii)监督的基础模型微调,以及(iv)自监督模型微调。我们的最佳策略通过微调自监督模型的所有参数实现了近90%的分类准确率。这些结果突出了声散射作为隐私保护,非接触替代视觉分类,在各个行业的应用打开了巨大的潜力。
摘要:This paper presents a novel non-invasive object classification approach using acoustic scattering, demonstrated through a case study on hair assessment. When an incident wave interacts with an object, it generates a scattered acoustic field encoding structural and material properties. By emitting acoustic stimuli and capturing the scattered signals from head-with-hair-sample objects, we classify hair type and moisture using AI-driven, deep-learning-based sound classification. We benchmark comprehensive methods, including (i) fully supervised deep learning, (ii) embedding-based classification, (iii) supervised foundation model fine-tuning, and (iv) self-supervised model fine-tuning. Our best strategy achieves nearly 90% classification accuracy by fine-tuning all parameters of a self-supervised model. These results highlight acoustic scattering as a privacy-preserving, non-contact alternative to visual classification, opening huge potential for applications in various industries.


【17】 Making deep neural networks work for medical audio: representation,  compression and domain adaptation

标题: 让深度神经网络用于医疗音频:表示、压缩和领域适应
链接:https://arxiv.org/abs/2506.13970
作者: Charles C Onu 
备注:PhD Thesis
摘要:本论文解决了应用机器学习来理解和解释医疗音频信号的技术挑战。我们的肺,心脏和声音的声音传达有关我们健康的重要信息。然而,在当代医学中,这些声音主要是通过专家使用听诊器等设备进行听觉解释来分析的。自动化分析提供了标准化医疗声音处理的潜力,能够在医生稀缺的低资源环境中进行筛查,并检测可能无法感知的细微模式,从而促进早期诊断和治疗。   本论文的重点是分析婴儿哭声来预测医疗状况,主要有四个方面的贡献。首先,在低数据环境中,我们证明了可以通过神经迁移学习来利用成人语音的大型数据库,以开发更准确和更强大的婴儿哭声分析模型。其次,在成本有效的建模中,我们引入了一种使用张量分解的递归网络的端到端模型压缩方法。我们的方法不需要事后处理,实现了几百倍的压缩率,并提供了准确的,便携式的模型,适合资源受限的设备。第三,我们提出了为音频模型量身定制的新型域自适应技术,并从计算机视觉中调整现有方法。这些方法解决了数据集偏差,增强了跨领域的泛化能力,同时保持了对原始数据的强大性能。最后,为了推进这一领域的研究,我们发布了一个独特的、开源的婴儿哭声数据集,该数据集是与世界各地的临床医生合作开发的。   这项工作为将婴儿哭声识别为生命体征奠定了基础,并突出了人工智能驱动的音频监测在塑造可获得和负担得起的医疗保健未来方面的变革潜力。
摘要:This thesis addresses the technical challenges of applying machine learning to understand and interpret medical audio signals. The sounds of our lungs, heart, and voice convey vital information about our health. Yet, in contemporary medicine, these sounds are primarily analyzed through auditory interpretation by experts using devices like stethoscopes. Automated analysis offers the potential to standardize the processing of medical sounds, enable screening in low-resource settings where physicians are scarce, and detect subtle patterns that may elude human perception, thereby facilitating early diagnosis and treatment.   Focusing on the analysis of infant cry sounds to predict medical conditions, this thesis contributes on four key fronts. First, in low-data settings, we demonstrate that large databases of adult speech can be harnessed through neural transfer learning to develop more accurate and robust models for infant cry analysis. Second, in cost-effective modeling, we introduce an end-to-end model compression approach for recurrent networks using tensor decomposition. Our method requires no post-hoc processing, achieves compression rates of several hundred-fold, and delivers accurate, portable models suitable for resource-constrained devices. Third, we propose novel domain adaptation techniques tailored for audio models and adapt existing methods from computer vision. These approaches address dataset bias and enhance generalization across domains while maintaining strong performance on the original data. Finally, to advance research in this domain, we release a unique, open-source dataset of infant cry sounds, developed in collaboration with clinicians worldwide.   This work lays the foundation for recognizing the infant cry as a vital sign and highlights the transformative potential of AI-driven audio monitoring in shaping the future of accessible and affordable healthcare.


【18】 Set theoretic solution for the tuning problem

标题: 为调整问题设定理论解决方案
链接:https://arxiv.org/abs/2506.13969
作者: Vsevolod Vladimirovich Deriushkin 
摘要:在这篇论文中,我想提出一个新的解决音乐调音问题的方法。一方面,我把它看作是对不和谐木材的公正语调(JI)的概括,另一方面,在一个单一的框架内,作为频谱干扰和和谐性对协和音的贡献的统一。主要成就的工作是能够数学量化的现象,音乐和谐使用集理论。这种量化是通过定义和谐的两种度量来完成的:亲和力和和谐性。这些措施自然会生成可用作动态调谐系统的间隔集。该文件是针对广大观众的人谁可能不擅长音乐和调音理论或数学。因此,我试图提供尽可能多的细节和解释,同时尽可能减少页数。
摘要:In this paper I want to suggest a new solution to the problem of musical tuning. On one hand, I see it as a generalization of Just Intonation (JI) to inharmonic timbers, on another, as a unification of spectral interference and harmonicity contributions to consonance within a single framework. The main achievement of the work is the ability to mathematically quantify the phenomenon of musical consonance using set theory. That quantification is done by defining two measures of consonance: affinity and harmonicity. These measures naturally generate sets of intervals that can be used as dynamic tuning systems. The paper is aimed at a broad audience of people who may not be skilled in music and tuning theory or mathematics. Thus, I attempt to give as much details and explanations as I can, while keeping the number of pages as low as possible.


【19】 A Survey on World Models Grounded in Acoustic Physical Information

标题: 基于声学物理信息的世界模型概览
链接:https://arxiv.org/abs/2506.13833
作者: Xiaoliang Chen,  Le Chang,  Xin Yu,  Yunhe Huang,  Xianling Tu 
备注:28 pages,11 equations
摘要:本调查提供了一个全面的概述的新兴领域的世界模型的基础上的声学物理信息。它探讨了理论基础,基本的方法框架,以及利用声学信号进行高保真环境感知,因果物理推理和动态事件预测模拟的最新技术进步。该调查解释了声信号作为物理事件机械波能量的直接载体,如何编码有关材料特性、内部几何结构和复杂相互作用动力学的丰富潜在信息。具体来说,这项调查建立了理论基础,解释基本物理定律如何管理声信号内的物理信息的编码。然后,它回顾了核心方法支柱,包括物理信息神经网络(PINN),生成模型和自监督多模态学习框架。此外,该调查详细介绍了声学世界模型在机器人、自动驾驶、医疗保健和金融领域的重要应用。最后,它系统地概述了重要的技术和伦理挑战,同时为未来的研究方向提出了一个具体的路线图,以实现强大的,因果的,不确定性感知的和负责任的声学智能。这些元素共同指向了一条通往体现主动声学智能的研究途径,使人工智能系统能够通过声音构建内部“直观物理”引擎。
摘要:This survey provides a comprehensive overview of the emerging field of world models grounded in the foundation of acoustic physical information. It examines the theoretical underpinnings, essential methodological frameworks, and recent technological advancements in leveraging acoustic signals for high-fidelity environmental perception, causal physical reasoning, and predictive simulation of dynamic events. The survey explains how acoustic signals, as direct carriers of mechanical wave energy from physical events, encode rich, latent information about material properties, internal geometric structures, and complex interaction dynamics. Specifically, this survey establishes the theoretical foundation by explaining how fundamental physical laws govern the encoding of physical information within acoustic signals. It then reviews the core methodological pillars, including Physics-Informed Neural Networks (PINNs), generative models, and self-supervised multimodal learning frameworks. Furthermore, the survey details the significant applications of acoustic world models in robotics, autonomous driving, healthcare, and finance. Finally, it systematically outlines the important technical and ethical challenges while proposing a concrete roadmap for future research directions toward robust, causal, uncertainty-aware, and responsible acoustic intelligence. These elements collectively point to a research pathway towards embodied active acoustic intelligence, empowering AI systems to construct an internal "intuitive physics" engine through sound.


【20】 Improving Practical Aspects of End-to-End Multi-Talker Speech  Recognition for Online and Offline Scenarios

标题: 改进在线和线下场景的端到端多说话者语音识别的实用方面
链接:https://arxiv.org/abs/2506.14204
作者: Aswin Shanmugam Subramanian,  Amit Das,  Naoyuki Kanda,  Jinyu Li,  Xiaofei Wang,  Yifan Gong 
备注:Accepted to Interspeech 2025
摘要:我们扩展了串行输出训练(SOT)的框架,以满足流媒体和离线自动语音识别(ASR)应用程序的实际需求。我们的方法侧重于平衡延迟和准确性,满足实时字幕和摘要要求。我们提出了几个关键的改进:(1)利用连续语音分离(CSS)单通道前端与端到端(E2 E)系统的高度重叠的情况下,挑战传统的智慧E2 E与级联设置。CSS框架通过分离来自多个说话者的重叠语音来提高ASR系统的准确性。(2)实现双模型--用于流式传输的Conformer Transducer和用于离线的Sequence-to-Sequence--或者基于级联编码器的双通道模型。(3)探索基于段的SOT(segSOT),它更适合离线场景,同时还增强了多人传输的可读性。
摘要:We extend the frameworks of Serialized Output Training (SOT) to address practical needs of both streaming and offline automatic speech recognition (ASR) applications. Our approach focuses on balancing latency and accuracy, catering to real-time captioning and summarization requirements. We propose several key improvements: (1) Leveraging Continuous Speech Separation (CSS) single-channel front-end with end-to-end (E2E) systems for highly overlapping scenarios, challenging the conventional wisdom of E2E versus cascaded setups. The CSS framework improves the accuracy of the ASR system by separating overlapped speech from multiple speakers. (2) Implementing dual models -- Conformer Transducer for streaming and Sequence-to-Sequence for offline -- or alternatively, a two-pass model based on cascaded encoders. (3) Exploring segment-based SOT (segSOT) which is better suited for offline scenarios while also enhancing readability of multi-talker transcriptions.


eess.AS音频处理


【1】 ASAP-FE: Energy-Efficient Feature Extraction Enabling Multi-Channel  Keyword Spotting on Edge Processors
标题: ASAP-FE:节能特征提取,支持边缘处理器上的多通道关键字发现
链接:https://arxiv.org/abs/2506.14657
作者: Jongin Choi,  Jina Park,  Woojoo Lee,  Jae-Jin Lee,  Massoud Pedram 
备注:7 pages, 11 figures, ISLPED 2025
摘要:多通道关键字识别(KWS)对于边缘环境中基于语音的应用程序至关重要。然而,其大量的计算和能量需求构成了重大挑战。我们介绍ASAP-FE(敏捷稀疏感知的特征提取器),一个面向硬件的前端,旨在解决这些挑战。我们的框架结合了三个关键的创新:(1)半重叠无限脉冲响应(IIR)框架:这减少了大约25%的冗余数据,同时保持基本的音素过渡线索。(2)稀疏感知数据减少:我们利用帧级稀疏性,通过将跳帧与基于步幅的过滤相结合,实现额外50%的数据减少。(3)动态并行处理:我们引入了一个参数化的过滤器集群和基于优先级的调度算法,允许并行执行IIR过滤任务,减少延迟和优化能源效率。ASAP-FE在边缘处理器上使用各种滤波器集群大小实现,功能在FPGA原型上验证,设计在45 nm处合成。使用TC-ResNet 8,DS-CNN和KWT-1的实验结果表明,ASAP-FE将平均工作负载减少了62.73%,同时支持多达32个通道的实时处理。与传统的完全重叠基线相比,ASAP-FE实现了小于1%的精度下降(例如,96.22%与DS-CNN的97.13%),这完全在边缘AI的可接受范围内。通过调整滤波器模块的数量,我们的设计优化了性能和能量之间的平衡,15个并行滤波器为多达25个通道提供最佳性能。总的来说,ASAP-FE为能源受限的边缘设备上的多通道KWS提供了一个实用而高效的解决方案。
摘要:Multi-channel keyword spotting (KWS) has become crucial for voice-based applications in edge environments. However, its substantial computational and energy requirements pose significant challenges. We introduce ASAP-FE (Agile Sparsity-Aware Parallelized-Feature Extractor), a hardware-oriented front-end designed to address these challenges. Our framework incorporates three key innovations: (1) Half-overlapped Infinite Impulse Response (IIR) Framing: This reduces redundant data by approximately 25% while maintaining essential phoneme transition cues. (2) Sparsity-aware Data Reduction: We exploit frame-level sparsity to achieve an additional 50% data reduction by combining frame skipping with stride-based filtering. (3) Dynamic Parallel Processing: We introduce a parameterizable filter cluster and a priority-based scheduling algorithm that allows parallel execution of IIR filtering tasks, reducing latency and optimizing energy efficiency. ASAP-FE is implemented with various filter cluster sizes on edge processors, with functionality verified on FPGA prototypes and designs synthesized at 45 nm. Experimental results using TC-ResNet8, DS-CNN, and KWT-1 demonstrate that ASAP-FE reduces the average workload by 62.73% while supporting real-time processing for up to 32 channels. Compared to a conventional fully overlapped baseline, ASAP-FE achieves less than a 1% accuracy drop (e.g., 96.22% vs. 97.13% for DS-CNN), which is well within acceptable limits for edge AI. By adjusting the number of filter modules, our design optimizes the trade-off between performance and energy, with 15 parallel filters providing optimal performance for up to 25 channels. Overall, ASAP-FE offers a practical and efficient solution for multi-channel KWS on energy-constrained edge devices.


【2】 The Perception of Phase Intercept Distortion and its Application in Data  Augmentation

标题: 相截失真的感知及其在数据增强中的应用
链接:https://arxiv.org/abs/2506.14571
作者: Venkatakrishnan Vaidyanathapuram Krishnan,  Nathaniel Condit-Schultz 
备注:Submitted to the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2025
摘要:相位失真是指信号中频率之间的相位关系的改变,这是可以感知的。在本文中,我们讨论了一种特殊情况下的相位失真称为相位截距失真,这是由一个频率无关的相移。我们假设,虽然这种形式的失真会显著改变信号的波形,但失真是不可感知的。人类受试者的实验结果与这一假设是一致的。此外,我们还讨论了相位截距失真的不可感知性如何对机器学习有用,特别是对数据增强。我们使用相位截距失真作为一种新的数据增强方法进行了多次实验,并获得了音频机器学习任务的改进结果。
摘要:Phase distortion refers to the alteration of the phase relationships between frequencies in a signal, which can be perceptible. In this paper, we discuss a special case of phase distortion known as phase-intercept distortion, which is created by a frequency-independent phase shift. We hypothesize that, though this form of distortion changes a signal's waveform significantly, the distortion is imperceptible. Human-subject experiment results are reported which are consistent with this hypothesis. Furthermore, we discuss how the imperceptibility of phase-intercept distortion can be useful for machine learning, specifically for data augmentation. We conducted multiple experiments using phase-intercept distortion as a novel approach to data augmentation, and obtained improved results for audio machine learning tasks.


【3】 M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization  Dataset

标题: M3 SD:多模式、多场景和多语言说话者拨号数据集
链接:https://arxiv.org/abs/2506.14427
作者: Shilong Wu,  Hang Chen,  Jun Du 
备注:11 pages, 5 figures
摘要:在说话人日志化领域,技术的发展受到两个问题的制约:数据资源不足和深度学习模型泛化能力差。为了解决这两个问题,首先,我们提出了一种自动构建说话人日志数据集的方法,该方法通过音频和视频的组合为海量数据生成更准确的伪标签。基于这种方法,我们发布了多模态、多场景和多语言的说话人日记(M3SD)数据集。该数据集来源于真实的网络视频,具有高度多样性。此外,我们进一步提出了一个与网络相关的模型微调策略。基于使用上述数据集预训练的通用模型,我们结合目标场景的特定数据(例如,会议),并通过使用Adapter和LoRA联合微调实现针对性优化,从而实现模型的领域自适应。我们的数据集和代码已在https://huggingface.co/spaces/OldDragon/m3sd上开源。
摘要:In the field of speaker diarization, the development of technology is constrained by two problems: insufficient data resources and poor generalization ability of deep learning models. To address these two problems, firstly, we propose an automated method for constructing speaker diarization datasets, which generates more accurate pseudo-labels for massive data through the combination of audio and video. Relying on this method, we have released Multi-modal, Multi-scenario and Multi-language Speaker Diarization (M3SD) datasets. This dataset is derived from real network videos and is highly diverse. In addition, we further propose a scenario-related model fine-tuning strategy. Based on the general model pre-trained using the above dataset, we combine the specific data of the target scenario (e.g., meetings) and achieve targeted optimization by using Adapter and LoRA joint fine-tuning, thus achieving the model's domain adaptation. Our dataset and code have been open-sourced at https://huggingface.co/spaces/OldDragon/m3sd.


【4】 Improving Practical Aspects of End-to-End Multi-Talker Speech  Recognition for Online and Offline Scenarios

标题: 改进在线和线下场景的端到端多说话者语音识别的实用方面
链接:https://arxiv.org/abs/2506.14204
作者: Aswin Shanmugam Subramanian,  Amit Das,  Naoyuki Kanda,  Jinyu Li,  Xiaofei Wang,  Yifan Gong 
备注:Accepted to Interspeech 2025
摘要:我们扩展了串行输出训练(SOT)的框架,以满足流媒体和离线自动语音识别(ASR)应用程序的实际需求。我们的方法侧重于平衡延迟和准确性,满足实时字幕和摘要要求。我们提出了几个关键的改进:(1)利用连续语音分离(CSS)单通道前端与端到端(E2 E)系统的高度重叠的情况下,挑战传统的智慧E2 E与级联设置。CSS框架通过分离来自多个说话者的重叠语音来提高ASR系统的准确性。(2)实现双模型--用于流式传输的Conformer Transducer和用于离线的Sequence-to-Sequence--或者基于级联编码器的双通道模型。(3)探索基于段的SOT(segSOT),它更适合离线场景,同时还增强了多人传输的可读性。
摘要:We extend the frameworks of Serialized Output Training (SOT) to address practical needs of both streaming and offline automatic speech recognition (ASR) applications. Our approach focuses on balancing latency and accuracy, catering to real-time captioning and summarization requirements. We propose several key improvements: (1) Leveraging Continuous Speech Separation (CSS) single-channel front-end with end-to-end (E2E) systems for highly overlapping scenarios, challenging the conventional wisdom of E2E versus cascaded setups. The CSS framework improves the accuracy of the ASR system by separating overlapped speech from multiple speakers. (2) Implementing dual models -- Conformer Transducer for streaming and Sequence-to-Sequence for offline -- or alternatively, a two-pass model based on cascaded encoders. (3) Exploring segment-based SOT (segSOT) which is better suited for offline scenarios while also enhancing readability of multi-talker transcriptions.


【5】 Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation  Quantity for Modeling Videoconference Conversation Experience

标题: 采用半监督学习的多模式融合最大限度地减少了视频会议对话体验建模的注释数量
链接:https://arxiv.org/abs/2506.13971
作者: Andrew Chang,  Chenkai Hu,  Ji Qi,  Zhuojian Wei,  Kexin Zhang,  Viswadruth Akkaraju,  David Poeppel,  Dustin Freeman 
备注:Interspeech 2025
摘要:通过视频会议进行的群组对话是一种复杂的社会行为。然而,消极体验的主观时刻,即谈话失去流动性或乐趣的时刻,仍然没有得到充分的研究。这些时刻在自然数据中并不常见,因此训练监督学习(SL)模型需要昂贵的手动数据注释。我们应用半监督学习(SSL)来利用有针对性的标记和未标记的剪辑来训练多模态(音频,面部,文本)深度特征,以预测视频会议会话中的非流体或不愉快的时刻。模态融合的协同训练SSL实现了0.9的ROC-AUC和0.6的F1得分,在相同数量的标记数据下,比SL模型的性能高出4%。值得注意的是,只有8%标记数据的最佳SSL模型与SL模型的全数据性能匹配96%。这显示了一个用于对视频会议体验建模的注释高效框架。
摘要:Group conversations over videoconferencing are a complex social behavior. However, the subjective moments of negative experience, where the conversation loses fluidity or enjoyment remain understudied. These moments are infrequent in naturalistic data, and thus training a supervised learning (SL) model requires costly manual data annotation. We applied semi-supervised learning (SSL) to leverage targeted labeled and unlabeled clips for training multimodal (audio, facial, text) deep features to predict non-fluid or unenjoyable moments in holdout videoconference sessions. The modality-fused co-training SSL achieved an ROC-AUC of 0.9 and an F1 score of 0.6, outperforming SL models by up to 4% with the same amount of labeled data. Remarkably, the best SSL model with just 8% labeled data matched 96% of the SL model's full-data performance. This shows an annotation-efficient framework for modeling videoconference experience.


【6】 A Variational Framework for Improving Naturalness in Generative Spoken  Language Models

标题: 提高生成性口语模型自然性的变分框架
链接:https://arxiv.org/abs/2506.14767
作者: Li-Wei Chen,  Takuya Higuchi,  Zakaria Aldeneh,  Ahmed Hussen Abdelaziz,  Alexander Rudnicky 
备注:International Conference on Machine Learning (ICML) 2025
摘要:大型语言模型在文本处理中的成功激发了它们对语音建模的适应。然而,由于语音是连续和复杂的,它往往是离散化的自回归建模。从自监督模型中得到的语音标记(称为语义标记)通常关注语音的语言学方面,但忽略了韵律信息。因此,在这些令牌上训练的模型可以生成自然度降低的语音。现有的方法试图通过向语义标记添加音高特征来解决这个问题。然而,音高本身并不能完全代表语言属性的范围,选择正确的特征需要仔细的手工设计。为了克服这一点,我们提出了一种端到端的变分方法,自动学习编码这些连续的语音属性,以增强语义令牌。我们的方法消除了手动提取和选择的非语言特征的需要。此外,它根据人类评分员产生优选的语音延续。代码、示例和模型可在https://github.com/b04901014/vae-gslm上获得。
摘要:The success of large language models in text processing has inspired their adaptation to speech modeling. However, since speech is continuous and complex, it is often discretized for autoregressive modeling. Speech tokens derived from self-supervised models (known as semantic tokens) typically focus on the linguistic aspects of speech but neglect prosodic information. As a result, models trained on these tokens can generate speech with reduced naturalness. Existing approaches try to fix this by adding pitch features to the semantic tokens. However, pitch alone cannot fully represent the range of paralinguistic attributes, and selecting the right features requires careful hand-engineering. To overcome this, we propose an end-to-end variational approach that automatically learns to encode these continuous speech attributes to enhance the semantic tokens. Our approach eliminates the need for manual extraction and selection of paralinguistic features. Moreover, it produces preferred speech continuations according to human raters. Code, samples and models are available at https://github.com/b04901014/vae-gslm.


【7】 An Open Research Dataset of the 1932 Cairo Congress of Arab Music

标题: 1932年开罗阿拉伯音乐大会的开放研究数据集
链接:https://arxiv.org/abs/2506.14503
作者: Baris Bozkurt (College of Interdisciplinary Studies, Zayed University, Dubai, United Arab Emirates) 
备注:14 pages, 4 figures, 4 tables
摘要:本文介绍了ord-cc 32,一个开放的研究数据集来自1932年开罗大会的阿拉伯音乐录音,一个具有历史意义的集合,代表了不同的阿拉伯音乐传统。该数据集包括结构化元数据、旋律和节奏模式标签(maqam和iqa)、手动标记的主音信息以及使用最先进的音高检测方法提取的声学特征。这些资源支持阿拉伯音乐的调音、音律和区域变化的计算研究。使用音高直方图的案例研究表明,跨区域的微色调差异的数据驱动分析的潜力。通过公开提供该数据集,我们的目标是实现计算民族音乐学,音乐信息检索(MIR),文化研究和数字遗产保护的跨学科研究。Order-CC 32在Zenodo上与用于特征提取和元数据检索的工具共享。
摘要:This paper introduces ORD-CC32 , an open research dataset derived from the 1932 Cairo Congress of Arab Music recordings, a historically significant collection representing diverse Arab musical traditions. The dataset includes structured metadata, melodic and rhythmic mode tags (maqam and iqa), manually labeled tonic information, and acoustic features extracted using state-of-the-art pitch detection methods. These resources support computational studies of tuning, temperament, and regional variations in Arab music. A case study using pitch histograms demonstrates the potential for data-driven analysis of microtonal differences across regions. By making this dataset openly available, we aim to enable interdisciplinary research in computational ethnomusicology, music information retrieval (MIR), cultural studies, and digital heritage preservation. ORD-CC32 is shared on Zenodo with tools for feature extraction and metadata retrieval.


【8】 Unifying Streaming and Non-streaming Zipformer-based ASR

标题: 统一流媒体和非流媒体基于Zipformer的ASB
链接:https://arxiv.org/abs/2506.14434
作者: Bidisha Sharma,  Karthik Pandia Durai,  Shankar Venkatesan,  Jeena J Prakash,  Shashi Kumar,  Malolan Chetlur,  Andreas Stolcke 
备注:Accepted in ACL2025 Industry track
摘要:人们对统一流和非流自动语音识别(ASR)模型以降低开发、培训和部署成本的兴趣越来越大。我们提出了一个统一的框架,训练一个单一的端到端的ASR模型流和非流应用程序,利用未来的上下文信息。我们建议在基于zipformer的ASR模型的训练中通过分块注意掩蔽来使用动态右上下文。我们证明,使用右上下文是更有效的zipformer模型相比,其他构象模型,由于其多尺度的性质。我们分析了右上下文帧的数量变化对流式ASR模型的准确性和延迟的影响。我们使用Librisepeech和大型内部对话数据集来训练不同版本的流媒体和非流媒体模型,并在生产级服务器-客户端设置中跨不同域的不同测试集对其进行评估。所提出的策略减少了相对7.9%的字错误,在用户感知的延迟一个小的退化。通过添加更多的右上下文帧,我们能够实现接近非流模型的流性能。我们的方法还允许根据客户的要求灵活控制延迟-准确性权衡。
摘要:There has been increasing interest in unifying streaming and non-streaming automatic speech recognition (ASR) models to reduce development, training, and deployment costs. We present a unified framework that trains a single end-to-end ASR model for both streaming and non-streaming applications, leveraging future context information. We propose to use dynamic right-context through the chunked attention masking in the training of zipformer-based ASR models. We demonstrate that using right-context is more effective in zipformer models compared to other conformer models due to its multi-scale nature. We analyze the effect of varying the number of right-context frames on accuracy and latency of the streaming ASR models. We use Librispeech and large in-house conversational datasets to train different versions of streaming and non-streaming models and evaluate them in a production grade server-client setup across diverse testsets of different domains. The proposed strategy reduces word error by relative 7.9\% with a small degradation in user-perceived latency. By adding more right-context frames, we are able to achieve streaming performance close to that of non-streaming models. Our approach also allows flexible control of the latency-accuracy tradeoff according to customers requirements.


【9】 SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative  music modeling

标题: SLEEPING-DISCO 9 M:用于生成式音乐建模的大规模预训练数据集
链接:https://arxiv.org/abs/2506.14293
作者: Tawsif Ahmed,  Andrej Radonjic,  Gollam Rabby 
摘要:我们介绍了Sleeping-DISCO 9 M,这是一个用于音乐和歌曲的大规模预训练数据集。据我们所知,没有开源的高质量数据集代表流行和知名的歌曲,用于生成音乐建模任务,如文本音乐,音乐字幕,歌唱语音合成,旋律重建和跨模型检索。过去的贡献集中在孤立和受约束的因素,其核心观点是创建合成或重新录制的音乐语料库(例如GTSinger,M4 Singer)和任意大规模的音频数据集(例如DISCO-10 M和LAIONDISCO-12 M)一直是社区的另一个焦点。不幸的是,这些数据集在生成音乐社区中的采用率一直低于实质性水平,因为这些数据集未能反映真实世界的音乐及其风味。我们的数据集改变了这种叙述,并提供了一个使用实际流行音乐和世界知名艺术家构建的数据集。
摘要:We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song. To the best of our knowledge, there are no open-source high-quality dataset representing popular and well-known songs for generative music modeling tasks such as text-music, music-captioning, singing-voice synthesis, melody reconstruction and cross-model retrieval. Past contributions focused on isolated and constrained factors whose core perspective was to create synthetic or re-recorded music corpus (e.g. GTSinger, M4Singer) and arbitrarily large-scale audio datasets (e.g. DISCO-10M and LAIONDISCO-12M) had been another focus for the community. Unfortunately, adoption of these datasets has been below substantial in the generative music community as these datasets fail to reflect real-world music and its flavour. Our dataset changes this narrative and provides a dataset that is constructed using actual popular music and world-renowned artists.


【10】 Fretting-Transformer: Encoder-Decoder Model for MIDI to Tablature  Transcription

标题: Fretting-Transformer:字形到表谱转录的编码器-解码器模型
链接:https://arxiv.org/abs/2506.14223
作者: Anna Hamberger,  Sebastian Murgul,  Jochen Schmidt,  Michael Heizmann 
备注:Accepted to the 50th International Computer Music Conference (ICMC), 2025
摘要:音乐转录在音乐信息检索(MIR)中起着关键作用,特别是对于像吉他这样的弦乐器,其中符号化的音乐符号(如吉他)缺乏关键的可演奏性信息。这篇文章介绍了微动Transformer,一个encoderdecoder模型,利用T5 Transformer架构,自动转录成吉他指法序列。通过将任务框定为符号翻译问题,该模型解决了关键挑战,包括弦-品歧义和物理可玩性。该系统利用了不同的数据集,包括DadaGP、GuitarToday和Leduc,并采用了新颖的数据预处理和标记化策略。我们已经制定了指法准确性和可玩性的指标,以定量评估的性能。实验结果表明,Fretting-Transformer超越了A* 等基线方法和Guitar Pro等商业应用。上下文敏感处理和调音/Capo调节的集成进一步增强了模型的性能,为自动吉他转录的未来发展奠定了坚实的基础。
摘要:Music transcription plays a pivotal role in Music Information Retrieval (MIR), particularly for stringed instruments like the guitar, where symbolic music notations such as MIDI lack crucial playability information. This contribution introduces the Fretting-Transformer, an encoderdecoder model that utilizes a T5 transformer architecture to automate the transcription of MIDI sequences into guitar tablature. By framing the task as a symbolic translation problem, the model addresses key challenges, including string-fret ambiguity and physical playability. The proposed system leverages diverse datasets, including DadaGP, GuitarToday, and Leduc, with novel data pre-processing and tokenization strategies. We have developed metrics for tablature accuracy and playability to quantitatively evaluate the performance. The experimental results demonstrate that the Fretting-Transformer surpasses baseline methods like A* and commercial applications like Guitar Pro. The integration of context-sensitive processing and tuning/capo conditioning further enhances the model's performance, laying a robust foundation for future developments in automated guitar transcription.


【11】 AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR

标题: AsyncSwitch:代码交换ASB的同步文本语音自适应
链接:https://arxiv.org/abs/2506.14190
作者: Tuan Nguyen,  Huy-Dat Tran 
备注:This work has been submitted to the IEEE for possible publication. This paper is a preprint version submitted to the 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2025)
摘要:开发代码转换ASR系统是具有挑战性的,因为语言模糊性和有限的接触多语言,代码转换数据,而收集这样的语音是昂贵的。先前的工作从文本生成合成音频,但这些方法计算密集且难以扩展。我们介绍AsyncSwitch,一种新型的异步自适应框架,利用大规模的,文本丰富的Web数据预暴露ASR模型到不同的代码切换域,然后对配对的语音文本语料库进行微调。我们的三阶段过程(1)在代码切换文本上训练解码器自注意和前馈层,(2)使用有限的语音文本数据通过交叉注意对齐解码器和编码器,以及(3)完全微调整个模型。Whisper对马来语-英语码转换的实验表明,相对WER减少了9.02%,同时提高了新加坡式英语,马来语和其他英语变体的单语性能。
摘要:Developing code-switched ASR systems is challenging due to language ambiguity and limited exposure to multilingual, code-switched data, while collecting such speech is costly. Prior work generates synthetic audio from text, but these methods are computationally intensive and hard to scale. We introduce AsyncSwitch, a novel asynchronous adaptation framework that leverages large-scale, text-rich web data to pre-expose ASR models to diverse code-switched domains before fine-tuning on paired speech-text corpora. Our three-stage process (1) trains decoder self-attention and feedforward layers on code-switched text, (2) aligns decoder and encoder via cross-attention using limited speech-text data, and (3) fully fine-tunes the entire model. Experiments with Whisper on Malay-English code-switching demonstrate a 9.02% relative WER reduction, while improving monolingual performance in Singlish, Malay, and other English variants.


【12】 Can we train ASR systems on Code-switch without real code-switch data?  Case study for Singapore's languages

标题: 我们能否在没有真正代码交换数据的情况下在代码交换上训练ASB系统?  新加坡语言案例研究
链接:https://arxiv.org/abs/2506.14177
作者: Tuan Nguyen,  Huy-Dat Tran 
备注:Accepted by Interspeech 2025
摘要:语码转换是多语言环境中常见的现象,由于语言的复杂性导致转录数据的稀缺和昂贵,这给自动语音识别带来了挑战。本研究探讨使用合成CS数据构建CS-ASR。我们提出了一种短语级混合方法来生成模拟自然模式的合成CS数据。利用合成短语混合CS数据增强的单语来微调大型预训练ASR模型(Whisper,MMS,M4 T)。本文重点研究了三种资源不足的东南亚语言对:马来语-英语(BM-EN),普通话-马来语(ZH-BM)和泰米尔语-英语(TA-EN),为CS-ASR建立了一个新的综合基准,以评估领先的ASR模型的性能。实验结果表明,该训练策略提高了单语和CS测试中的ASR性能,其中BM-EN表现出最高的增益,然后是TA-EN和ZH-BM。这一发现为CS-ASR的开发提供了一种具有成本效益的方法,使研究和工业受益。
摘要:Code-switching (CS), common in multilingual settings, presents challenges for ASR due to scarce and costly transcribed data caused by linguistic complexity. This study investigates building CS-ASR using synthetic CS data. We propose a phrase-level mixing method to generate synthetic CS data that mimics natural patterns. Utilizing monolingual augmented with synthetic phrase-mixed CS data to fine-tune large pretrained ASR models (Whisper, MMS, SeamlessM4T). This paper focuses on three under-resourced Southeast Asian language pairs: Malay-English (BM-EN), Mandarin-Malay (ZH-BM), and Tamil-English (TA-EN), establishing a new comprehensive benchmark for CS-ASR to evaluate the performance of leading ASR models. Experimental results show that the proposed training strategy enhances ASR performance on monolingual and CS tests, with BM-EN showing highest gains, then TA-EN and ZH-BM. This finding offers a cost-effective approach for CS-ASR development, benefiting research and industry.


【13】 Pushing the Performance of Synthetic Speech Detection with  Kolmogorov-Arnold Networks and Self-Supervised Learning Models

标题: 利用Kolmogorov-Arnold网络和自我监督学习模型提高合成语音检测的性能
链接:https://arxiv.org/abs/2506.14153
作者: Tuan Dat Phuong,  Long-Vu Hoang,  Huy Dat Tran 
备注:Accepted to Interspeech 2025
摘要:语音合成技术的最新进展导致了越来越先进的欺骗攻击,对自动说话人验证系统提出了重大挑战。虽然基于自监督学习(SSL)模型的系统,特别是XLSR-Conformer模型,在合成语音检测方面表现出了卓越的性能,但仍有架构改进的空间。在本文中,我们提出了一种新的方法,取代传统的多层感知器的XLSR-构象模型与Kolmogorov-Arnold网络(KAN),一种新的架构的基础上Kolmogorov-Arnold表示定理。我们在ASVspoof 2021上的结果表明,将KAN集成到基于SSL的模型中可以在LA和DF集上相对提高60.55%的性能,并在21 LA集上进一步实现0.70%的EER。这些研究结果表明,将KAN到基于SSL的模型是一个很有前途的方向,在合成语音检测的进步。
摘要:Recent advancements in speech synthesis technologies have led to increasingly advanced spoofing attacks, posing significant challenges for automatic speaker verification systems. While systems based on self-supervised learning (SSL) models, particularly the XLSR-Conformer model, have demonstrated remarkable performance in synthetic speech detection, there remains room for architectural improvements. In this paper, we propose a novel approach that replaces the traditional Multi-Layer Perceptron in the XLSR-Conformer model with a Kolmogorov-Arnold Network (KAN), a novel architecture based on the Kolmogorov-Arnold representation theorem. Our results on ASVspoof2021 demonstrate that integrating KAN into the SSL-based models can improve the performance by 60.55% relatively on LA and DF sets, further achieving 0.70% EER on the 21LA set. These findings suggest that incorporating KAN into SSL-based models is a promising direction for advances in synthetic speech detection.


【14】 Acoustic scattering AI for non-invasive object classifications: A case  study on hair assessment

标题: 用于非侵入性物体分类的声散射人工智能:头发评估的案例研究
链接:https://arxiv.org/abs/2506.14148
作者: Long-Vu Hoang,  Tuan Nguyen,  Tran Huy Dat 
备注:Accepted to Interspeech 2025
摘要:本文提出了一种新的非侵入性的物体分类方法,使用声散射,通过头发评估的案例研究证明。当入射波与物体相互作用时,它会产生编码结构和材料特性的散射声场。通过发出声学刺激并捕获来自头部和头发样本对象的散射信号,我们使用AI驱动的基于深度学习的声音分类对头发类型和水分进行分类。我们对综合方法进行了基准测试,包括(i)完全监督的深度学习,(ii)基于嵌入的分类,(iii)监督的基础模型微调,以及(iv)自监督模型微调。我们的最佳策略通过微调自监督模型的所有参数实现了近90%的分类准确率。这些结果突出了声散射作为隐私保护,非接触替代视觉分类,在各个行业的应用打开了巨大的潜力。
摘要:This paper presents a novel non-invasive object classification approach using acoustic scattering, demonstrated through a case study on hair assessment. When an incident wave interacts with an object, it generates a scattered acoustic field encoding structural and material properties. By emitting acoustic stimuli and capturing the scattered signals from head-with-hair-sample objects, we classify hair type and moisture using AI-driven, deep-learning-based sound classification. We benchmark comprehensive methods, including (i) fully supervised deep learning, (ii) embedding-based classification, (iii) supervised foundation model fine-tuning, and (iv) self-supervised model fine-tuning. Our best strategy achieves nearly 90% classification accuracy by fine-tuning all parameters of a self-supervised model. These results highlight acoustic scattering as a privacy-preserving, non-contact alternative to visual classification, opening huge potential for applications in various industries.


【15】 Making deep neural networks work for medical audio: representation,  compression and domain adaptation

标题: 让深度神经网络用于医疗音频:表示、压缩和领域适应
链接:https://arxiv.org/abs/2506.13970
作者: Charles C Onu 
备注:PhD Thesis
摘要:本论文解决了应用机器学习来理解和解释医疗音频信号的技术挑战。我们的肺,心脏和声音的声音传达有关我们健康的重要信息。然而,在当代医学中,这些声音主要是通过专家使用听诊器等设备进行听觉解释来分析的。自动化分析提供了标准化医疗声音处理的潜力,能够在医生稀缺的低资源环境中进行筛查,并检测可能无法感知的细微模式,从而促进早期诊断和治疗。   本论文的重点是分析婴儿哭声来预测医疗状况,主要有四个方面的贡献。首先,在低数据环境中,我们证明了可以通过神经迁移学习来利用成人语音的大型数据库,以开发更准确和更强大的婴儿哭声分析模型。其次,在成本有效的建模中,我们引入了一种使用张量分解的递归网络的端到端模型压缩方法。我们的方法不需要事后处理,实现了几百倍的压缩率,并提供了准确的,便携式的模型,适合资源受限的设备。第三,我们提出了针对音频模型的新的域自适应技术,并从计算机视觉中调整现有的方法。这些方法解决了数据集偏差,增强了跨领域的泛化能力,同时保持了对原始数据的强大性能。最后,为了推进这一领域的研究,我们发布了一个独特的、开源的婴儿哭声数据集,该数据集是与世界各地的临床医生合作开发的。   这项工作为将婴儿哭声识别为生命体征奠定了基础,并突出了人工智能驱动的音频监测在塑造可获得和负担得起的医疗保健未来方面的变革潜力。
摘要:This thesis addresses the technical challenges of applying machine learning to understand and interpret medical audio signals. The sounds of our lungs, heart, and voice convey vital information about our health. Yet, in contemporary medicine, these sounds are primarily analyzed through auditory interpretation by experts using devices like stethoscopes. Automated analysis offers the potential to standardize the processing of medical sounds, enable screening in low-resource settings where physicians are scarce, and detect subtle patterns that may elude human perception, thereby facilitating early diagnosis and treatment.   Focusing on the analysis of infant cry sounds to predict medical conditions, this thesis contributes on four key fronts. First, in low-data settings, we demonstrate that large databases of adult speech can be harnessed through neural transfer learning to develop more accurate and robust models for infant cry analysis. Second, in cost-effective modeling, we introduce an end-to-end model compression approach for recurrent networks using tensor decomposition. Our method requires no post-hoc processing, achieves compression rates of several hundred-fold, and delivers accurate, portable models suitable for resource-constrained devices. Third, we propose novel domain adaptation techniques tailored for audio models and adapt existing methods from computer vision. These approaches address dataset bias and enhance generalization across domains while maintaining strong performance on the original data. Finally, to advance research in this domain, we release a unique, open-source dataset of infant cry sounds, developed in collaboration with clinicians worldwide.   This work lays the foundation for recognizing the infant cry as a vital sign and highlights the transformative potential of AI-driven audio monitoring in shaping the future of accessible and affordable healthcare.


【16】 Set theoretic solution for the tuning problem

标题: 为调整问题设定理论解决方案
链接:https://arxiv.org/abs/2506.13969
作者: Vsevolod Vladimirovich Deriushkin 
摘要:在这篇论文中,我想提出一个新的解决音乐调音问题的方法。一方面,我把它看作是对不和谐木材的公正语调(JI)的概括,另一方面,在一个单一的框架内,作为频谱干扰和和谐性对协和音的贡献的统一。主要成就的工作是能够数学量化的现象,音乐和谐使用集理论。这种量化是通过定义和谐的两种度量来完成的:亲和力和和谐性。这些措施自然产生的间隔,可用作动态调整系统。该文件是针对广大观众的人谁可能不擅长音乐和调音理论或数学。因此,我试图提供尽可能多的细节和解释,同时尽可能减少页数。
摘要:In this paper I want to suggest a new solution to the problem of musical tuning. On one hand, I see it as a generalization of Just Intonation (JI) to inharmonic timbers, on another, as a unification of spectral interference and harmonicity contributions to consonance within a single framework. The main achievement of the work is the ability to mathematically quantify the phenomenon of musical consonance using set theory. That quantification is done by defining two measures of consonance: affinity and harmonicity. These measures naturally generate sets of intervals that can be used as dynamic tuning systems. The paper is aimed at a broad audience of people who may not be skilled in music and tuning theory or mathematics. Thus, I attempt to give as much details and explanations as I can, while keeping the number of pages as low as possible.


【17】 A Survey on World Models Grounded in Acoustic Physical Information

标题: 基于声学物理信息的世界模型概览
链接:https://arxiv.org/abs/2506.13833
作者: Xiaoliang Chen,  Le Chang,  Xin Yu,  Yunhe Huang,  Xianling Tu 
备注:28 pages,11 equations
摘要:本调查提供了一个全面的概述的新兴领域的世界模型的基础上的声学物理信息。它探讨了理论基础,基本的方法框架,以及利用声学信号进行高保真环境感知,因果物理推理和动态事件预测模拟的最新技术进步。该调查解释了声信号作为物理事件中机械波能量的直接载体,如何编码有关材料特性,内部几何结构和复杂相互作用动力学的丰富潜在信息。具体来说,这项调查建立了理论基础,解释基本物理定律如何管理声信号内的物理信息的编码。然后,它回顾了核心方法支柱,包括物理信息神经网络(PINN),生成模型和自监督多模态学习框架。此外,该调查详细介绍了声学世界模型在机器人、自动驾驶、医疗保健和金融领域的重要应用。最后,它系统地概述了重要的技术和伦理挑战,同时为未来的研究方向提出了一个具体的路线图,以实现强大的,因果的,不确定性感知的和负责任的声学智能。这些元素共同指向了一条通往体现主动声学智能的研究途径,使人工智能系统能够通过声音构建内部“直观物理”引擎。
摘要:This survey provides a comprehensive overview of the emerging field of world models grounded in the foundation of acoustic physical information. It examines the theoretical underpinnings, essential methodological frameworks, and recent technological advancements in leveraging acoustic signals for high-fidelity environmental perception, causal physical reasoning, and predictive simulation of dynamic events. The survey explains how acoustic signals, as direct carriers of mechanical wave energy from physical events, encode rich, latent information about material properties, internal geometric structures, and complex interaction dynamics. Specifically, this survey establishes the theoretical foundation by explaining how fundamental physical laws govern the encoding of physical information within acoustic signals. It then reviews the core methodological pillars, including Physics-Informed Neural Networks (PINNs), generative models, and self-supervised multimodal learning frameworks. Furthermore, the survey details the significant applications of acoustic world models in robotics, autonomous driving, healthcare, and finance. Finally, it systematically outlines the important technical and ethical challenges while proposing a concrete roadmap for future research directions toward robust, causal, uncertainty-aware, and responsible acoustic intelligence. These elements collectively point to a research pathway towards embodied active acoustic intelligence, empowering AI systems to construct an internal "intuitive physics" engine through sound.


机器翻译由腾讯交互翻译提供,仅供参考