微信公众号:arXiv_Daily
cs.SD语音
标题:StreamFlow:使用逐块引导注意力屏蔽进行流媒体流匹配,用于语音令牌解码
链接:https://arxiv.org/abs/2506.23986
摘要:基于离散令牌的语音生成的最新进展突出了令牌到波形生成对于音频质量的重要性,特别是在实时交互中。将语义令牌与流匹配(FM)集成的传统框架由于依赖于全局接受域而难以实现流功能。此外,直接实现逐令牌流式语音生成通常导致音频质量下降。为了解决这些挑战,我们提出了StreamFlow,这是一种新型的神经架构,可以促进与扩散Transformers(DiT)的流匹配。为了减轻长期的历史依赖性所产生的长序列外推问题,我们设计了一个本地块明智的感受野策略。具体来说,序列首先被分割成块,我们引入了块级注意掩码,使当前块能够接收来自前一个或后一个块的信息。这些注意力掩模在不同的DiT块中分层组合,以调节DiT的感受野。主观和客观的实验结果表明,我们的方法实现的性能与非流方法相比,同时超过其他流方法的语音质量,同时有效地管理推理时间在长序列生成。此外,我们的方法实现了仅180 ms的显着的第一个数据包延迟。footnote{语音示例:https://dukguo.github.io/StreamFlow/}
摘要:Recent advancements in discrete token-based speech generation have highlighted the importance of token-to-waveform generation for audio quality, particularly in real-time interactions. Traditional frameworks integrating semantic tokens with flow matching (FM) struggle with streaming capabilities due to their reliance on a global receptive field. Additionally, directly implementing token-by-token streaming speech generation often results in degraded audio quality. To address these challenges, we propose StreamFlow, a novel neural architecture that facilitates streaming flow matching with diffusion transformers (DiT). To mitigate the long-sequence extrapolation issues arising from lengthy historical dependencies, we design a local block-wise receptive field strategy. Specifically, the sequence is first segmented into blocks, and we introduce block-wise attention masks that enable the current block to receive information from the previous or subsequent block. These attention masks are combined hierarchically across different DiT-blocks to regulate the receptive field of DiTs. Both subjective and objective experimental results demonstrate that our approach achieves performance comparable to non-streaming methods while surpassing other streaming methods in terms of speech quality, all the while effectively managing inference time during long-sequence generation. Furthermore, our method achieves a notable first-packet latency of only 180 ms.\footnote{Speech samples: https://dukguo.github.io/StreamFlow/}
【2】Emergent musical properties of a transformer under contrastive self-supervised learning
链接:https://arxiv.org/abs/2506.23873
备注:Accepted at ISMIR 2025
摘要:在音乐信息检索(MIR)中,通用表示模型的对比自监督学习对于自动标注等全局任务是有效的。然而,对于和弦估计等本地任务,人们普遍认为对比训练的通用自监督模型是不够的,需要更复杂的SSL;例如,蒙面建模。我们的论文挑战这一假设,揭示了潜在的对比SSL与一个Transformer在本地MIR任务配对。我们考虑一个轻量级的Vision Transformer与一维补丁在时间-频率域(ViT-1D)和训练它与简单的对比SSL通过归一化的温度缩放的交叉熵损失(NT-Xent)。虽然NT-Xent只在类令牌上运行,但我们观察到,可能由于权重共享,信息丰富的音乐属性出现在ViT-1D的序列令牌中。在全局任务中,类和序列标记的时间平均值与单独的类标记相比提供了性能提高,显示了序列标记中的有用属性。在本地任务中,序列令牌的表现出乎意料地好,尽管没有专门训练。此外,高层次的音乐特征,如开始出现从逐层注意力地图和自相似矩阵显示不同的层捕捉不同的音乐维度。我们的论文并不专注于提高性能,但推进Transformers的音乐解释,并揭示了一些被忽视的对比SSL与Transformers配对的MIR序列建模的能力。
摘要:In music information retrieval (MIR), contrastive self-supervised learning for general-purpose representation models is effective for global tasks such as automatic tagging. However, for local tasks such as chord estimation, it is widely assumed that contrastively trained general-purpose self-supervised models are inadequate and that more sophisticated SSL is necessary; e.g., masked modeling. Our paper challenges this assumption by revealing the potential of contrastive SSL paired with a transformer in local MIR tasks. We consider a lightweight vision transformer with one-dimensional patches in the time--frequency domain (ViT-1D) and train it with simple contrastive SSL through normalized temperature-scaled cross-entropy loss (NT-Xent). Although NT-Xent operates only over the class token, we observe that, potentially thanks to weight sharing, informative musical properties emerge in ViT-1D's sequence tokens. On global tasks, the temporal average of class and sequence tokens offers a performance increase compared to the class token alone, showing useful properties in the sequence tokens. On local tasks, sequence tokens perform unexpectedly well, despite not being specifically trained for. Furthermore, high-level musical features such as onsets emerge from layer-wise attention maps and self-similarity matrices show different layers capture different musical dimensions. Our paper does not focus on improving performance but advances the musical interpretation of transformers and sheds light on some overlooked abilities of contrastive SSL paired with transformers for sequence modeling in MIR.
【3】Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
链接:https://arxiv.org/abs/2506.23869
备注:ISMIR (2025)
摘要:我们研究了生成自回归Transformer模型在大量符号化钢琴独奏transertion上训练的能力。在对大约60,000小时的音乐进行第一次预训练后,我们使用相对较小的高质量子集来微调模型,以产生音乐延续,执行符号分类任务,并通过将Simplified框架适应符号音乐来产生通用对比嵌入。在评估钢琴连续连贯性时,我们的生成模型优于领先的符号生成技术,并与专有音频生成模型保持竞争力。在MIR分类基准测试中,来自我们对比模型的冻结表示在线性探测实验中获得了最先进的结果,而直接微调则展示了预训练表示的可推广性,通常只需要几百个带标签的示例就可以专门用于下游任务。
摘要:We study the capabilities of generative autoregressive transformer models trained on large amounts of symbolic solo-piano transcriptions. After first pretraining on approximately 60,000 hours of music, we use a comparatively smaller, high-quality subset, to finetune models to produce musical continuations, perform symbolic classification tasks, and produce general-purpose contrastive MIDI embeddings by adapting the SimCLR framework to symbolic music. When evaluating piano continuation coherence, our generative model outperforms leading symbolic generation techniques and remains competitive with proprietary audio generation models. On MIR classification benchmarks, frozen representations from our contrastive model achieve state-of-the-art results in linear probe experiments, while direct finetuning demonstrates the generalizability of pretrained representations, often requiring only a few hundred labeled examples to specialize to downstream tasks.
【4】Efficient Interleaved Speech Modeling through Knowledge Distillation
链接:https://arxiv.org/abs/2506.23670
摘要:当前的语音语言模型超过了许多部署环境的大小和延迟限制。我们通过层对齐蒸馏,匹配隐藏状态,注意力地图和软化logits来构建紧凑,富有表现力的语音生成模型,以最小的性能损失将大型多模态Transformers压缩3倍。我们介绍TinyWave,这是一系列用于语音到语音和交错语音文本生成的2B参数模型,在50,000小时的公共音频上进行了训练。TinyWave支持(i)使用语音或表达令牌的纯语音生成和(ii)混合语音文本延续。Libri-Light上的评估显示TinyWave在其教师的1.4标准化困惑点内。口语StoryCloze和SALMon的准确率达到教师表现的93-97%,优于大小匹配的基线。这些模型针对商用硬件上的部署进行了优化,支持实时会话代理、辅助技术和低资源环境中的应用程序。我们发布模型,训练代码和评估脚本,以支持对紧凑,富有表现力的语音生成的可重复研究。
摘要:Current speech language models exceed the size and latency constraints of many deployment environments. We build compact, expressive speech generation models through layer-aligned distillation, matching hidden states, attention maps, and softened logits to compress large multimodal transformers by 3x with minimal loss in performance. We introduce TinyWave, a family of 2B-parameter models for speech-to-speech and interleaved speech-text generation, trained on 50,000 hours of public audio. TinyWave supports (i) speech-only generation using phonetic or expressive tokens and (ii) mixed speech-text continuations. Evaluation on Libri-Light shows TinyWave within 1.4 normalized perplexity points of its teacher. Accuracy on spoken StoryCloze and SALMon reaches 93-97% of the teacher's performance, outperforming size-matched baselines. These models are optimized for deployment on commodity hardware, enabling applications in real-time conversational agents, assistive technologies, and low-resource environments. We release models, training code, and evaluation scripts to support reproducible research on compact, expressive speech generation.
【5】RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
链接:https://arxiv.org/abs/2506.23582
备注:Accepted to INTERSPEECH2025
摘要:在文本到音频(TTA)的研究中,输入文本和输出音频之间的相关性是一个重要的评价方面。传统的评价方法是从主观和客观两个方面进行的。然而,主观评价在金钱和时间方面是昂贵的,并且客观评价与主观评价分数的相关性不清楚。在这项研究中,我们构建了一个开源的数据集,它主观地评估了相关性。此外,我们还对一个用于从合成音频自动预测主观评价分数的模型进行了基准测试。我们的模型优于传统的CLAPScore模型,这种趋势扩展到许多声音类别。
摘要:In text-to-audio (TTA) research, the relevance between input text and output audio is an important evaluation aspect. Traditionally, it has been evaluated from both subjective and objective perspectives. However, subjective evaluation is costly in terms of money and time, and objective evaluation is unclear regarding the correlation to subjective evaluation scores. In this study, we construct RELATE, an open-sourced dataset that subjectively evaluates the relevance. Also, we benchmark a model for automatically predicting the subjective evaluation score from synthesized audio. Our model outperforms a conventional CLAPScore model, and that trend extends to many sound categories.
【6】JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
链接:https://arxiv.org/abs/2506.23552
备注:project page: this https URL Under review. Preprint published on arXiv
摘要:在生成建模中,面部运动和语音之间的内在联系经常被忽视,其中说话的头部合成和文本到语音(TTS)通常被视为单独的任务。本文介绍了JAM-Flow,一个统一的框架,同时合成和条件的面部运动和语音。我们的方法利用流匹配和一种新的多模扩散Transformer(MM-DiT)架构,集成了专门的运动DiT和音频DiT模块。这些通过选择性联合注意层耦合,并结合关键的架构选择,如时间对齐的位置嵌入和局部联合注意掩蔽,以实现有效的跨模态交互,同时保持模态特定的优势。JAM-Flow采用inpainting风格的目标进行训练,支持各种条件输入(包括文本、参考音频和参考运动),可在单个连贯模型中执行促进任务,例如从文本同步生成说话的头部、音频驱动的动画等。JAM-Flow通过为整体视听合成提供实用的解决方案,大大推进了多模态生成建模。项目页面:https://joonghyuk.com/jamflow-web
摘要:The intrinsic link between facial motion and speech is often overlooked in generative modeling, where talking head synthesis and text-to-speech (TTS) are typically addressed as separate tasks. This paper introduces JAM-Flow, a unified framework to simultaneously synthesize and condition on both facial motion and speech. Our approach leverages flow matching and a novel Multi-Modal Diffusion Transformer (MM-DiT) architecture, integrating specialized Motion-DiT and Audio-DiT modules. These are coupled via selective joint attention layers and incorporate key architectural choices, such as temporally aligned positional embeddings and localized joint attention masking, to enable effective cross-modal interaction while preserving modality-specific strengths. Trained with an inpainting-style objective, JAM-Flow supports a wide array of conditioning inputs-including text, reference audio, and reference motion-facilitating tasks such as synchronized talking head generation from text, audio-driven animation, and much more, within a single, coherent model. JAM-Flow significantly advances multi-modal generative modeling by providing a practical solution for holistic audio-visual synthesis. project page: https://joonghyuk.com/jamflow-web
【7】From Large-scale Audio Tagging to Real-Time Explainable Emergency Vehicle Sirens Detection
链接:https://arxiv.org/abs/2506.23437
备注:pre-print (submitted to the IEEE/ACM Transactions on Audio, Speech, and Language Processing)
摘要:紧急车辆(EV)警报器的准确识别对于智能交通系统,智能城市监控系统和自动驾驶技术的集成至关重要。现代的自动解决方案受到缺乏大规模的策划数据集以及最先进的声音事件检测模型的计算需求的限制。这项工作引入了E2 PANN(高效紧急预训练音频神经网络),这是一种源自PANN框架的轻量级卷积神经网络架构,专门针对二进制电动汽车警报器检测进行了优化。利用我们专用的AudioSet子集(AudioSet EV),我们在多个参考数据集上微调和评估E2PANN,并在嵌入式硬件上测试其可行性。实验活动包括消融研究,跨域基准测试和边缘设备上的实时推理部署。利用引导反向传播和ScoreCAM算法的可解释性分析提供了对模型内部表示的见解,并验证了其捕获与不同类型EV警报相关的不同光谱时间模式的能力。实时性能通过逐帧和基于事件的检测指标以及对假阳性激活的详细分析进行评估。结果表明,E2PANNs建立了一个新的国家的艺术在这个研究领域,具有高计算效率,并适用于基于边缘的音频监测和安全关键应用。
摘要:Accurate recognition of Emergency Vehicle (EV) sirens is critical for the integration of intelligent transportation systems, smart city monitoring systems, and autonomous driving technologies. Modern automatic solutions are limited by the lack of large scale, curated datasets and by the computational demands of state of the art sound event detection models. This work introduces E2PANNs (Efficient Emergency Pre trained Audio Neural Networks), a lightweight Convolutional Neural Network architecture derived from the PANNs framework, specifically optimized for binary EV siren detection. Leveraging our dedicated subset of AudioSet (AudioSet EV) we fine-tune and evaluate E2PANNs across multiple reference datasets and test its viability on embedded hardware. The experimental campaign includes ablation studies, cross-domain benchmarking, and real-time inference deployment on edge device. Interpretability analyses exploiting Guided Backpropagation and ScoreCAM algorithms provide insights into the model internal representations and validate its ability to capture distinct spectrotemporal patterns associated with different types of EV sirens. Real time performance is assessed through frame wise and event based detection metrics, as well as a detailed analysis of false positive activations. Results demonstrate that E2PANNs establish a new state of the art in this research domain, with high computational efficiency, and suitability for edge-based audio monitoring and safety-critical applications.
【8】You Sound a Little Tense: L2 Tailored Clear TTS Using Durational Vowel Properties
链接:https://arxiv.org/abs/2506.23367
备注:Accepted to ISCA Speech Synthesis Workshop, 2025
摘要:我们提出了第一个为第二语言(L2)使用者量身定制的文本到语音(TTS)系统。我们使用美国英语紧张(长)和宽松(短)元音之间的持续时间差异,以创建一个“清晰度模式”的Matcha-TTS。我们的感知研究表明,当使用我们的清晰度模式时,法语第一语言、英语第二语言听众的转录错误较少(至少9.15%),并且发现它比整体放慢的语音更令人鼓舞和尊重。值得注意的是,听众并没有意识到这些影响:尽管在清晰度模式下单词错误率降低,但听众仍然认为放慢所有目标单词是最容易理解的,这表明实际的可理解性与感知的可理解性无关。此外,我们发现,耳语ASR没有使用相同的线索,L2扬声器区分困难的元音,是不足以评估这些人的TTS系统的可懂度。
摘要:We present the first text-to-speech (TTS) system tailored to second language (L2) speakers. We use duration differences between American English tense (longer) and lax (shorter) vowels to create a "clarity mode" for Matcha-TTS. Our perception studies showed that French-L1, English-L2 listeners had fewer (at least 9.15%) transcription errors when using our clarity mode, and found it more encouraging and respectful than overall slowed down speech. Remarkably, listeners were not aware of these effects: despite the decreased word error rate in clarity mode, listeners still believed that slowing all target words was the most intelligible, suggesting that actual intelligibility does not correlate with perceived intelligibility. Additionally, we found that Whisper-ASR did not use the same cues as L2 speakers to differentiate difficult vowels and is not sufficient to assess the intelligibility of TTS systems for these individuals.
【9】XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
链接:https://arxiv.org/abs/2506.23325
摘要:语音编解码器是语音信号和大型语言模型之间的桥梁。一个理想的语音语言模型编解码器不仅要保留声学信息,而且要捕捉丰富的语义信息。然而,现有的语音编解码器难以平衡高质量的音频重建与易于通过语言模型建模。在这项研究中,我们分析了以前的编解码器在平衡语义丰富性和声学保真度方面的局限性。我们提出了XY-Tokenizer,一种新的编解码器,通过多阶段,多任务学习减轻语义和声学能力之间的冲突。实验结果表明,XY-Tokenizer在语义和声学任务中的性能与以类似比特率操作的最先进的编解码器相当,即使现有的编解码器通常只在一个方面表现出色。具体来说,XY-Tokenizer实现了强大的文本对齐,超过了基于蒸馏的语义建模方法,如SpeechTokenizer和Mimi,同时在重建和原始音频之间保持了0.83的说话人相似性得分。XY-Tokenizer的重建性能与BigCodec相当,BigCodec是目前最先进的纯声学编解码器,在类似的比特率下实现了0.84的扬声器相似性得分。代码和型号可从https://github.com/gyt1145028706/XY-Tokenizer获得。
摘要:Speech codecs serve as bridges between speech signals and large language models. An ideal codec for speech language models should not only preserve acoustic information but also capture rich semantic information. However, existing speech codecs struggle to balance high-quality audio reconstruction with ease of modeling by language models. In this study, we analyze the limitations of previous codecs in balancing semantic richness and acoustic fidelity. We propose XY-Tokenizer, a novel codec that mitigates the conflict between semantic and acoustic capabilities through multi-stage, multi-task learning. Experimental results demonstrate that XY-Tokenizer achieves performance in both semantic and acoustic tasks comparable to that of state-of-the-art codecs operating at similar bitrates, even though those existing codecs typically excel in only one aspect. Specifically, XY-Tokenizer achieves strong text alignment, surpassing distillation-based semantic modeling methods such as SpeechTokenizer and Mimi, while maintaining a speaker similarity score of 0.83 between reconstructed and original audio. The reconstruction performance of XY-Tokenizer is comparable to that of BigCodec, the current state-of-the-art among acoustic-only codecs, which achieves a speaker similarity score of 0.84 at a similar bitrate. Code and models are available at https://github.com/gyt1145028706/XY-Tokenizer.
【10】The Florence Price Art Song Dataset and Piano Accompaniment Generator
链接:https://arxiv.org/abs/2506.23130
备注:8 pages, 4 figures. To appear in the proceedings of ISMIR 2025
摘要:佛罗伦萨湾普赖斯是20世纪初的一位作曲家,她的音乐反映了她在美国南部的成长经历、她的非洲遗产和她的西方古典训练。她被认为是第一位由大型管弦乐队演奏交响乐的非洲裔美国女性。在她去世几十年后,她的音乐最近再次受到公众和研究界的关注。除了其他流派,普莱斯还是一位多产的独唱和钢琴作曲家。音乐历史学家记录了普赖斯创作的134首艺术歌曲和钢琴/人声编曲的灵歌和民歌。我们以MuseScore、MusicXML、XML和PDF格式发布了其中112件作品的数字目录。我们还使用这个数据集微调符号音乐生成模型,以生成旋律的替代品,我们进行了盲听实验,表明我们的模型生成的替代品被认为是反映佛罗伦萨价格的风格比基线模型生成的替代品更频繁。我们将我们的模型作为Florence Price钢琴伴奏生成器与我们的数据集一起发布。
摘要:Florence B. Price was a composer in the early 20th century whose music reflects her upbringing in the American South, her African heritage, and her Western classical training. She is noted as the first African-American woman to have a symphony performed by a major orchestra. Her music has recently received renewed attention from both the public and the research community, decades after her death. In addition to other genres, Price was a prolific composer for solo voice and piano. Music historians have documented the existence of 134 art songs and piano/voice arrangements for spirituals and folk songs written by Price. We release a digital catalog of 112 of these works in MuseScore, MusicXML, MIDI, and PDF format. We also use this dataset to fine-tune a symbolic music generation model to generate accompaniments to melodies, and we conduct a blind listening experiment that shows that accompaniments generated by our model are perceived as being reflective of Florence Price's style more frequently than accompaniments generated by a baseline model. We release our model as the Florence Price Piano Accompaniment Generator alongside our dataset.
【11】TOMI: Transforming and Organizing Music Ideas for Multi-Track Compositions with Full-Song Structure
链接:https://arxiv.org/abs/2506.23094
备注:9 pages, 4 figures, 2 tables. To be published in ISMIR 2025
摘要:分层规划是一种从结构上对长序列建模的有效方法。除了考虑音乐的时间结构的层次,本文探讨了一个更重要的方面:概念层次,这涉及到产生音乐的想法,转换它们,并最终组织它们-跨越音乐的时间和空间-成为一个完整的组成。为此,我们引入了TOMI(转换和组织音乐思想)作为深度音乐生成的一种新方法,并通过调整基础LLM开发了一个基于TOMI的模型。形式上,我们通过一个稀疏的四维空间来表示多轨作曲过程,该四维空间的特征在于剪辑(短音频或音频片段),部分(时间位置),轨道(乐器层)和转换(精心制作方法)。我们的模型能够生成具有完整歌曲结构的多音轨电子音乐,并且我们进一步将基于TOMI的模型与REAPER数字音频工作站集成,实现交互式的人-AI共同创作。实验结果表明,我们的方法产生更高质量的电子音乐与更强的结构一致性相比,基线。
摘要:Hierarchical planning is a powerful approach to model long sequences structurally. Aside from considering hierarchies in the temporal structure of music, this paper explores an even more important aspect: concept hierarchy, which involves generating music ideas, transforming them, and ultimately organizing them--across musical time and space--into a complete composition. To this end, we introduce TOMI (Transforming and Organizing Music Ideas) as a novel approach in deep music generation and develop a TOMI-based model via instruction-tuned foundation LLM. Formally, we represent a multi-track composition process via a sparse, four-dimensional space characterized by clips (short audio or MIDI segments), sections (temporal positions), tracks (instrument layers), and transformations (elaboration methods). Our model is capable of generating multi-track electronic music with full-song structure, and we further integrate the TOMI-based model with the REAPER digital audio workstation, enabling interactive human-AI co-creation. Experimental results demonstrate that our approach produces higher-quality electronic music with stronger structural coherence compared to baselines.
【12】AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
链接:https://arxiv.org/abs/2506.23049
摘要:尽管语言和语音技术取得了进步,但没有开源系统能够实现完整的语音到语音,多轮对话,集成工具使用和代理推理。AURA(Agent for Understanding,Reasoning,and Automated Tool Use)是第一个开源的语音助手,能够通过动态工具调用和多轮对话完成复杂的目标驱动任务。AURA将开放权重的ASR、TTS和LLM组合在一个级联管道中,并支持日历预订、联系人查找、Web搜索和电子邮件等工具。其模块化设计允许使用自然语言提示和操作类轻松集成新工具。在VoiceBench上,AURA在OpenBookQA上的得分为92.75%,优于所有开放重量系统,接近GPT-40,在AlpacaEval上为4.39,与其他开放重量系统竞争。人工评估显示,在复杂的多回合语音任务中,任务成功率为90%。
摘要:Despite advances in language and speech technologies, no open-source system enables full speech-to-speech, multi-turn dialogue with integrated tool use and agentic reasoning. We introduce AURA (Agent for Understanding, Reasoning, and Automated Tool Use), the first open-source, speech-native assistant capable of completing complex, goal-driven tasks through dynamic tool invocation and multi-turn conversation. AURA combines open-weight ASR, TTS, and LLMs in a cascaded pipeline and supports tools such as calendar booking, contact lookup, web search, and email. Its modular design allows easy integration of new tools using natural language prompts and action classes. On VoiceBench, AURA scores 92.75% on OpenBookQA-outperforming all open-weight systems and nearing GPT-4o-and 4.39 on AlpacaEval, competitive with other open-weight systems. Human evaluation shows 90% task success on complex, multi-turn speech tasks.
【13】VisionScores -- A system-segmented image score dataset for deep learning tasks
链接:https://arxiv.org/abs/2506.23030
备注:Comments: 5 pages, 3 figures. Accepted for presentation at the 2025 IEEE International Conference on Image Processing (ICIP). \c{opyright} 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for any other use
摘要:VisionScores提出了一个新的建议,作为第一个系统分割的图像分数数据集,旨在为机器和深度学习任务提供结构丰富的高信息密度图像。限定为双手钢琴作品,它不仅要考虑某些图形相似性,还要考虑构图模式,因为这个创作过程高度依赖于乐器。它提供了两个与作曲家和作曲类型有关的场景。第一个由14 k个样本组成,考虑来自不同作者但相同作曲类型的作品,特别是奏鸣曲。后者由10.8K样本组成,呈现相反的情况,来自同一作者的各种作曲类型,是弗朗茨·李斯特的一个。所有24.8k样本都被格式化为$128 \times 512$像素的灰度jpg图像。VisionScores不仅为用户提供格式化的样本,还提供系统的顺序和片段的元数据。此外,未分割的整版分数和预先格式化的图像包括进一步分析。
摘要:VisionScores presents a novel proposal being the first system-segmented image score dataset, aiming to offer structure-rich, high information-density images for machine and deep learning tasks. Delimited to two-handed piano pieces, it was built to consider not only certain graphic similarity but also composition patterns, as this creative process is highly instrument-dependent. It provides two scenarios in relation to composer and composition type. The first, formed by 14k samples, considers works from different authors but the same composition type, specifically, Sonatinas. The latter, consisting of 10.8K samples, presents the opposite case, various composition types from the same author, being the one selected Franz Liszt. All of the 24.8k samples are formatted as grayscale jpg images of $128 \times 512$ pixels. VisionScores supplies the users not only the formatted samples but the systems' order and pieces' metadata. Moreover, unsegmented full-page scores and the pre-formatted images are included for further analysis.
【14】Feasibility of spectral-element modeling of wave propagation through the anatomy of marine mammals
链接:https://arxiv.org/abs/2506.22944
摘要:本文首次采用三维谱元法(SEM)模拟了超声波在短吻海豚(Tursiops truncatus)头部的传播。与传统的有限元方法(FEM)不同,由于昂贵的线性系统反演和较慢的收敛速度,SEM提供了指数收敛和高效的并行计算。使用计算机断层扫描(CT)扫描数据,我们开发了一个详细的六面体网格捕捉复杂的解剖特征,如声学脂肪和颌骨。我们的平面波和球面波的模拟证实了SEM的有效性超声时域建模。这种方法为海洋生物学开辟了新的途径,有助于回声定位、人为海洋噪音污染的影响以及海洋哺乳动物听觉和咔哒声产生的生物物理学研究。通过克服FEM的局限性,SEM提供了一个强大的可扩展的工具来测试海豚生物声学的假设,具有重要意义的保护和理解海洋哺乳动物的听觉系统下日益增加的环境挑战。
摘要:This study introduces the first 3D spectral-element method (SEM) simulation of ultrasonic wave propagation in a bottlenose dolphin (Tursiops truncatus) head. Unlike traditional finite-element methods (FEM), which struggle with high-frequency simulations due to costly linear-system inversions and slower convergence, SEM offers exponential convergence and efficient parallel computation. Using Computed Tomography (CT) scan data, we developed a detailed hexahedral mesh capturing complex anatomical features, such as acoustic fats and jaws. Our simulations of plane and spherical waves confirm SEM's effectiveness for ultrasonic time-domain modeling. This approach opens new avenues for marine biology, contributing to research in echolocation, the impacts of anthropogenic marine noise pollution and the biophysics of hearing and click generation in marine mammals. By overcoming FEM's limitations, SEM provides a powerful scalable tool to test hypotheses about dolphin bioacoustics, with significant implications for conservation and understanding marine mammal auditory systems under increasing environmental challenges.
【15】Mind the Gap: Entity-Preserved Context-Aware ASR Structured Transcriptions
链接:https://arxiv.org/abs/2506.22858
备注:This is the accepted version of an article accepted to the TSD 2025 conference, published in Springer Lecture Notes in Artificial Intelligence (LNAI). The final authenticated version is available online at SpringerLink
摘要:自动语音识别(ASR)系统(例如Whisper)可以实现高转录准确性,但在命名实体和数字数据方面遇到了困难,尤其是在需要正确格式化时。这些问题增加了单词错误率(WER),并损害了法律,金融和医疗应用等关键领域的语义理解。我们提出了一种新的训练方法,通过在训练过程中添加重叠的上下文窗口来扩展ASR模型的语义上下文。通过在30秒块的两侧滑动5秒重叠,我们创建了一个40秒的“有效语义窗口”,提高实体识别和格式,同时将预测集中在中间30秒。为了解决跨越块边界的实体,我们将这些实体完全重新分配给右侧块,以确保正确的格式。此外,具有嵌入式实体标签的丰富训练数据使模型能够学习识别和特定类型的格式。在Spoken Wikipedia数据集上进行评估,我们的方法提高了语义任务的性能,包括命名实体识别(NER)和实体格式化。这些结果突出了上下文感知训练在解决长形式转录和复杂实体识别任务的ASR限制方面的有效性。
摘要:Automatic Speech Recognition (ASR) systems, such as Whisper, achieve high transcription accuracy but struggle with named entities and numerical data, especially when proper formatting is required. These issues increase word error rate (WER) and impair semantic understanding in critical domains like legal, financial, and medical applications. We propose a novel training approach that extends the semantic context of ASR models by adding overlapping context windows during training. By sliding 5-second overlaps on both sides of 30-second chunks, we create a 40-second "effective semantic window," improving entity recognition and formatting while focusing predictions on the central 30 seconds. To address entities spanning chunk boundaries, we reassign such entities entirely to the right-hand chunk, ensuring proper formatting. Additionally, enriched training data with embedded entity labels enables the model to learn both recognition and type-specific formatting. Evaluated on the Spoken Wikipedia dataset, our method improves performance across semantic tasks, including named entity recognition (NER) and entity formatting. These results highlight the effectiveness of context-aware training in addressing ASR limitations for long-form transcription and complex entity recognition tasks.
【16】Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization
链接:https://arxiv.org/abs/2506.22846
备注:This is the accepted version of an article accepted to the TSD 2025 conference, published in Springer Lecture Notes in Artificial Intelligence (LNAI). The final authenticated version is available online at SpringerLink
摘要:端到端(E2 E)自动语音识别(ASR)系统通过将所有组件集成到单个神经网络中,使该领域发生了革命性变化,基于注意力的编码器-解码器模型实现了最先进的性能。然而,它们的自回归解码过程限制了推理速度,使其不适合实时应用。相比之下,基于CTC的模型提供更快的非自回归解码,但难以有效地对语言依赖性进行建模。为了应对这一挑战,我们提出了一种新的辅助损失框架,称为存储感知中间损失(LAIL),以增强基于CTC的ASR使用大型语言模型(LLM)的语言知识。通过将连接器层连接到中间编码器层,LAIL将输出映射到LLM的嵌入空间,并在训练期间计算因果语言建模损失。这种方法增强了语言建模,同时保持CTC解码的计算效率。使用Conformer架构和各种LLaMA模型,我们在LibriSpeech,TEDLIUM 2和WSJ语料库上展示了字错误率(WER)的显着改善,以最小的计算开销实现了基于CTC的ASR的最先进的性能。
摘要:End-to-end (E2E) automatic speech recognition (ASR) systems have revolutionized the field by integrating all components into a single neural network, with attention-based encoder-decoder models achieving state-of-the-art performance. However, their autoregressive decoding process limits inference speed, making them unsuitable for real-time applications. In contrast, CTC-based models offer faster, non-autoregressive decoding but struggle to model linguistic dependencies effectively. Addressing this challenge, we propose a novel auxiliary loss framework called Language-Aware Intermediate Loss (LAIL) to enhance CTC-based ASR using the linguistic knowledge of large language models (LLMs). By attaching connector layers to intermediate encoder layers, LAIL maps outputs to the embedding space of an LLM and computes a causal language modeling loss during training. This approach enhances linguistic modeling while preserving the computational efficiency of CTC decoding. Using the Conformer architecture and various LLaMA models, we demonstrate significant improvements in Word Error Rate (WER) on the LibriSpeech, TEDLIUM2, and WSJ corpora, achieving state-of-the-art performance for CTC-based ASR with minimal computational overhead.
【17】A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
链接:https://arxiv.org/abs/2506.22810
备注:accepted by Interspeech 2025
摘要:构音障碍语音识别(DSR)增强了移动性受限的构音障碍说话者对智能设备的可访问性。以前,DSR研究受到现有数据集通常由孤立的单词,命令短语和少数人所说的有限数量的句子组成的事实的限制。这限制了对命令交互系统和说话人自适应的研究。Speech Accessibility Project(SAP)通过发布一个大型且多样化的英语构音障碍数据集改变了这一点,导致SAP挑战赛构建说话者和文本独立的DSR系统。我们通过一种新的自我训练方法增强了Whisper模型在长时间构音障碍语音上的性能。这种方法增加了训练数据,并调整了模型,以处理在推理过程中遇到的可能不完整的语音片段。我们的系统在SAP Challenge中的单词错误率和语义得分均获得第二名。
摘要:Dysarthric speech recognition (DSR) enhances the accessibility of smart devices for dysarthric speakers with limited mobility. Previously, DSR research was constrained by the fact that existing datasets typically consisted of isolated words, command phrases, and a limited number of sentences spoken by a few individuals. This constrained research to command-interaction systems and speaker adaptation. The Speech Accessibility Project (SAP) changed this by releasing a large and diverse English dysarthric dataset, leading to the SAP Challenge to build speaker- and text-independent DSR systems. We enhanced the Whisper model's performance on long dysarthric speech via a novel self-training method. This method increased training data and adapted the model to handle potentially incomplete speech segments encountered during inference. Our system achieved second place in both Word Error Rate and Semantic Score in the SAP Challenge.
【18】WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
链接:https://arxiv.org/abs/2506.22789
备注:5 pages, 4 figures, Published at The Proceedings of Interspeech 2025, code is available at http://www.github.com/UTAustin-SwarmLab/WavShape
摘要:语音嵌入通常会保留敏感属性,如说话者身份、口音或人口统计信息,从而带来有偏见的模型训练和隐私泄露的风险。我们提出了WavShape,这是一个信息论语音表示学习框架,它优化了嵌入的公平性和隐私性,同时保留了任务相关的信息。我们利用互信息(MI)估计使用Donsker-Varadhan公式来指导基于MI的编码器,该编码器系统地过滤敏感属性,同时保持下游任务所必需的语音内容。在三个已知数据集上的实验结果表明,WavShape将嵌入和敏感属性之间的MI降低了81%,同时保留了97%的任务相关信息。通过将信息理论与自监督语音模型相结合,这项工作推进了公平、隐私感知和资源高效的语音系统的开发。
摘要:Speech embeddings often retain sensitive attributes such as speaker identity, accent, or demographic information, posing risks in biased model training and privacy leakage. We propose WavShape, an information-theoretic speech representation learning framework that optimizes embeddings for fairness and privacy while preserving task-relevant information. We leverage mutual information (MI) estimation using the Donsker-Varadhan formulation to guide an MI-based encoder that systematically filters sensitive attributes while maintaining speech content essential for downstream tasks. Experimental results on three known datasets show that WavShape reduces MI between embeddings and sensitive attributes by up to 81% while retaining 97% of task-relevant information. By integrating information theory with self-supervised speech models, this work advances the development of fair, privacy-aware, and resource-efficient speech systems.
【19】Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
链接:https://arxiv.org/abs/2506.22661
备注:Accepted to ISMIR2025
摘要:音频指纹识别(AFP)允许通过提取被称为音频指纹的紧凑表示来识别未知音频内容,所述音频指纹被设计成针对常见音频降级保持鲁棒性。神经AFP方法通常采用度量学习,其中表示质量受到监督的性质和所使用的损失函数的影响。然而,最近的工作不切实际地模拟训练期间的真实音频降级,导致次优监督。此外,尽管已经提出了几种现代度量学习方法,但当前的神经AFP方法仍然依赖于NT-Xent损失,而没有探索最新的进展或经典的替代方案。在这项工作中,我们提出了一系列的最佳实践,以加强自我监督,利用音乐信号的属性和现实的房间声学。然后,我们提出了第一个系统的评估各种度量学习方法的背景下,AFP,证明了自我监督适应的三重损失产生卓越的性能。我们的研究结果还表明,每个锚点使用多个阳性样本进行训练,在损失函数中具有截然不同的效果。我们的方法建立在这些见解的基础上,并在一个大型的综合退化数据集和一个在不同音乐场所使用麦克风记录的真实数据集上实现了最先进的性能。
摘要:Audio fingerprinting (AFP) allows the identification of unknown audio content by extracting compact representations, termed audio fingerprints, that are designed to remain robust against common audio degradations. Neural AFP methods often employ metric learning, where representation quality is influenced by the nature of the supervision and the utilized loss function. However, recent work unrealistically simulates real-life audio degradation during training, resulting in sub-optimal supervision. Additionally, although several modern metric learning approaches have been proposed, current neural AFP methods continue to rely on the NT-Xent loss without exploring the recent advances or classical alternatives. In this work, we propose a series of best practices to enhance the self-supervision by leveraging musical signal properties and realistic room acoustics. We then present the first systematic evaluation of various metric learning approaches in the context of AFP, demonstrating that a self-supervised adaptation of the triplet loss yields superior performance. Our results also reveal that training with multiple positive samples per anchor has critically different effects across loss functions. Our approach is built upon these insights and achieves state-of-the-art performance on both a large, synthetically degraded dataset and a real-world dataset recorded using microphones in diverse music venues.
【20】Evaluating Sound Similarity Metrics for Differentiable, Iterative Sound-Matching
链接:https://arxiv.org/abs/2506.22628
摘要:使用合成器的手动声音设计本质上是迭代的:艺术家将合成的输出与心理目标进行比较,调整参数,重复直到满意。迭代声音匹配通过在损失函数(或相似性度量)的指导下不断地对合成器进行编程来自动化该工作流程。先前对损失函数的比较通常倾向于一个度量而不是另一个度量,但仅限于狭窄的设置:有限的合成方法,很少的损失类型,通常没有盲听测试。这留下了一个开放的问题,是否存在普遍的最佳损失,或损失的选择仍然是一个创造性的决定条件的合成方法和声音设计师的偏好。我们提出可微迭代声音匹配作为现有文献的自然延伸,因为它结合了声音设计的手动方法和机器学习的现代进步。为了分析不同合成器之间损失函数性能的变化,我们实现了四种新颖的和已建立的可微分损失函数的混合,并将它们与可微分减法,加法和AM合成器配对。对于16种合成器损失组合中的每一种,我们进行了300次随机声音匹配试验。性能进行了测量,使用参数差异,频谱距离指标,并手动分配的听力分数。我们观察到的三个性能指标之间的一致性中等水平。我们的事后分析表明,损失函数的性能是高度依赖于合成器。这些发现强调了扩大声音匹配实验范围和开发针对特定合成技术的新相似性度量的价值,而不是追求一刀切的解决方案。
摘要:Manual sound design with a synthesizer is inherently iterative: an artist compares the synthesized output to a mental target, adjusts parameters, and repeats until satisfied. Iterative sound-matching automates this workflow by continually programming a synthesizer under the guidance of a loss function (or similarity measure) toward a target sound. Prior comparisons of loss functions have typically favored one metric over another, but only within narrow settings: limited synthesis methods, few loss types, often without blind listening tests. This leaves open the question of whether a universally optimal loss exists, or the choice of loss remains a creative decision conditioned on the synthesis method and the sound designer's preference. We propose differentiable iterative sound-matching as the natural extension of the available literature, since it combines the manual approach to sound design with modern advances in machine learning. To analyze the variability of loss function performance across synthesizers, we implemented a mix of four novel and established differentiable loss functions, and paired them with differentiable subtractive, additive, and AM synthesizers. For each of the sixteen synthesizer--loss combinations, we ran 300 randomized sound-matching trials. Performance was measured using parameter differences, spectrogram-distance metrics, and manually assigned listening scores. We observed a moderate level of consistency among the three performance measures. Our post-hoc analysis shows that the loss function performance is highly dependent on the synthesizer. These findings underscore the value of expanding the scope of sound-matching experiments and developing new similarity metrics tailored to specific synthesis techniques rather than pursuing one-size-fits-all solutions.
【21】Aria-MIDI: A Dataset of Piano MIDI Files for Symbolic Music Modeling
链接:https://arxiv.org/abs/2504.15071
备注:None
摘要:我们介绍了一个广泛的新的数据集的录音文件,通过转录钢琴演奏的录音到其组成的音符。我们使用的数据管道是多阶段的,采用语言模型,根据元数据从互联网上自动抓取和评分音频记录,然后使用音频分类器进行修剪和分割。由此产生的数据集包含超过100万个不同的音频文件,包括大约10万小时的转录音频。我们对我们的技术进行深入分析,提供统计见解,并通过提取我们还提供的元数据标签来调查内容。数据集可在https://github.com/loubbrad/aria-midi获得。
摘要:We introduce an extensive new dataset of MIDI files, created by transcribing audio recordings of piano performances into their constituent notes. The data pipeline we use is multi-stage, employing a language model to autonomously crawl and score audio recordings from the internet based on their metadata, followed by a stage of pruning and segmentation using an audio classifier. The resulting dataset contains over one million distinct MIDI files, comprising roughly 100,000 hours of transcribed audio. We provide an in-depth analysis of our techniques, offering statistical insights, and investigate the content by extracting metadata tags, which we also provide. Dataset available at https://github.com/loubbrad/aria-midi.
【22】URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition
链接:https://arxiv.org/abs/2506.23874
备注:Submitted to ASRU2025
摘要:平均意见得分(MOS)是语音质量评估的基础。然而,它的获取需要大量的人工注释。虽然已经开发了深度神经网络方法(如DNSMOS和UTMOS)来预测MOS以避免这个问题,但它们通常会受到训练数据不足的影响。认识到语音增强(SE)系统的比较优先考虑一个可靠的系统比较绝对分数,我们提出了URGENT-PK,一种新的排名方法,利用成对比较。URGENT-PK以同源增强语音对作为输入来预测相对质量排名。这种成对范式有效地利用了有限的训练数据,因为多个系统的所有成对排列构成了一个训练实例。在多个开放测试集上的实验表明,尽管URGENT-PK的网络架构简单,训练数据有限,但其系统级排名性能优于最先进的基线。
摘要:The Mean Opinion Score (MOS) is fundamental to speech quality assessment. However, its acquisition requires significant human annotation. Although deep neural network approaches, such as DNSMOS and UTMOS, have been developed to predict MOS to avoid this issue, they often suffer from insufficient training data. Recognizing that the comparison of speech enhancement (SE) systems prioritizes a reliable system comparison over absolute scores, we propose URGENT-PK, a novel ranking approach leveraging pairwise comparisons. URGENT-PK takes homologous enhanced speech pairs as input to predict relative quality rankings. This pairwise paradigm efficiently utilizes limited training data, as all pairwise permutations of multiple systems constitute a training instance. Experiments across multiple open test sets demonstrate URGENT-PK's superior system-level ranking performance over state-of-the-art baselines, despite its simple network architecture and limited training data.
【23】Less is More: Data Curation Matters in Scaling Speech Enhancement
链接:https://arxiv.org/abs/2506.23859
备注:Submitted to ASRU2025
摘要:绝大多数现代语音增强系统依赖于数据驱动的神经网络模型。传统上,较大的数据集被认为会产生更好的模型性能,这一观察结果在其他领域的许多任务中得到了经验验证。然而,最近的研究表明,当缩放语音增强数据时,收益递减。我们专注于一个关键因素:大规模数据集中“干净”训练标签中普遍存在的质量问题。这项工作重新审视了这一现象,并证明了在大规模训练集中,优先考虑高质量的训练数据比仅仅扩大数据量更重要。实验结果表明,在精心策划的700小时子集上训练的模型可以优于在2,500小时完整数据集上训练的模型。这一结果突出了数据管理在有效扩展语音增强系统中的关键作用。
摘要:The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling speech enhancement data. We focus on a critical factor: prevalent quality issues in ``clean'' training labels within large-scale datasets. This work re-examines this phenomenon and demonstrates that, within large-scale training sets, prioritizing high-quality training data is more important than merely expanding the data volume. Experimental findings suggest that models trained on a carefully curated subset of 700 hours can outperform models trained on the 2,500-hour full dataset. This outcome highlights the crucial role of data curation in scaling speech enhancement systems effectively.
【24】Human-CLAP: Human-perception-based contrastive language-audio pretraining
链接:https://arxiv.org/abs/2506.23553
摘要:对比语言-音频预训练(CLAP)被广泛用于音频生成和识别任务。例如,CLAPScore利用CLAP嵌入的相似性,已经成为文本到音频中评估音频和文本之间相关性的主要度量。然而,CLAPScore和人类主观评价分数之间的关系仍然不清楚。我们发现,CLAPScore与人类主观评价分数的相关性较低。此外,我们提出了一个基于人类感知的CLAP称为人类CLAP通过使用主观评价分数训练对比语言音频模型。在我们的实验中,结果表明,与传统CLAP相比,我们的Human-CLAP将CLAPScore与主观评价分数之间的斯皮尔曼等级相关系数(SRCC)提高了0.25以上。
摘要:Contrastive language-audio pretraining (CLAP) is widely used for audio generation and recognition tasks. For example, CLAPScore, which utilizes the similarity of CLAP embeddings, has been a major metric for the evaluation of the relevance between audio and text in text-to-audio. However, the relationship between CLAPScore and human subjective evaluation scores is still unclarified. We show that CLAPScore has a low correlation with human subjective evaluation scores. Additionally, we propose a human-perception-based CLAP called Human-CLAP by training a contrastive language-audio model using the subjective evaluation score. In our experiments, the results indicate that our Human-CLAP improved the Spearman's rank correlation coefficient (SRCC) between the CLAPScore and the subjective evaluation scores by more than 0.25 compared with the conventional CLAP.
【25】Investigating an Overfitting and Degeneration Phenomenon in Self-Supervised Multi-Pitch Estimation
链接:https://arxiv.org/abs/2506.23371
备注:Accepted to ISMIR 2025
摘要:多音高估计(MPE)仍然是音乐信息检索(MIR)系统的一种受欢迎的能力,并且对于涉及音高的许多应用和下游任务(包括音乐转录)是至关重要的。然而,现有的方法在很大程度上是基于监督学习的,并且在为任务收集注释数据方面存在重大挑战。最近,利用音调和谐波信号的固有特性的自监督技术已经显示出对单声道和多声道音调估计的希望,但是这些仍然劣于监督方法。在这项工作中,我们通过结合几个基于音调不变和音调等变属性的自监督目标来扩展经典的监督MPE范式。这种联合训练在封闭的训练条件下取得了很大的改进,这自然表明,将同样的目标应用于更广泛的数据收集将产生进一步的改进。然而,在这样做的过程中,我们发现了一个现象,即我们的模型同时过拟合监督数据,而退化的数据仅用于自我监督。我们证明和调查这一点,并提供我们的见解的根本问题。
摘要:Multi-Pitch Estimation (MPE) continues to be a sought after capability of Music Information Retrieval (MIR) systems, and is critical for many applications and downstream tasks involving pitch, including music transcription. However, existing methods are largely based on supervised learning, and there are significant challenges in collecting annotated data for the task. Recently, self-supervised techniques exploiting intrinsic properties of pitch and harmonic signals have shown promise for both monophonic and polyphonic pitch estimation, but these still remain inferior to supervised methods. In this work, we extend the classic supervised MPE paradigm by incorporating several self-supervised objectives based on pitch-invariant and pitch-equivariant properties. This joint training results in a substantial improvement under closed training conditions, which naturally suggests that applying the same objectives to a broader collection of data will yield further improvements. However, in doing so we uncover a phenomenon whereby our model simultaneously overfits to the supervised data while degenerating on data used for self-supervision only. We demonstrate and investigate this and offer our insights on the underlying problem.
【26】Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
链接:https://arxiv.org/abs/2506.22646
备注:Accepted by INTERSPEECH 2025
摘要:我们提出了一种自说话人自适应方法流多说话人自动语音识别(ASR),消除了显式扬声器查询的需要。与需要目标说话者嵌入或注册音频的传统方法不同,我们的技术通过说话者语音活动预测动态地适应单个ASR实例。关键的创新包括将通过扬声器监督激活生成的特定于扬声器的内核注入选定的ASR编码器层。这使得即使在流场景中也能够在处理完全重叠的语音的同时实现对目标扬声器的瞬时扬声器自适应。实验表明,在离线和流媒体场景中的最先进的性能,表明我们的自适应方法有效地解决了严重的语音重叠,通过精简的扬声器为重点的识别。实验结果验证了所提出的自说话人自适应方法在严重语音重叠条件下是一种鲁棒的多说话人ASR解决方案。
摘要:We propose a self-speaker adaptation method for streaming multi-talker automatic speech recognition (ASR) that eliminates the need for explicit speaker queries. Unlike conventional approaches requiring target speaker embeddings or enrollment audio, our technique dynamically adapts individual ASR instances through speaker-wise speech activity prediction. The key innovation involves injecting speaker-specific kernels generated via speaker supervision activations into selected ASR encoder layers. This enables instantaneous speaker adaptation to target speakers while handling fully overlapped speech even in a streaming scenario. Experiments show state-of-the-art performance in both offline and streaming scenarios, demonstrating that our self-adaptive method effectively addresses severe speech overlap through streamlined speaker-focused recognition. The results validate the proposed self-speaker adaptation approach as a robust solution for multi-talker ASR under severe overlapping speech conditions.
【1】URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition
链接:https://arxiv.org/abs/2506.23874
备注:Submitted to ASRU2025
摘要:平均意见得分(MOS)是语音质量评估的基础。然而,它的获取需要大量的人工注释。虽然已经开发了深度神经网络方法,如DNSMOS和UTMOS,以预测MOS来避免这个问题,但它们通常会受到训练数据不足的影响。认识到语音增强(SE)系统的比较优先考虑一个可靠的系统比较绝对分数,我们提出了URGENT-PK,一种新的排名方法,利用成对比较。URGENT-PK以同源增强语音对作为输入来预测相对质量排名。这种成对模式有效地利用了有限的训练数据,因为多个系统的所有成对排列构成了一个训练实例。在多个开放测试集上的实验表明,尽管URGENT-PK的网络架构简单,训练数据有限,但其系统级排名性能优于最先进的基线。
摘要:The Mean Opinion Score (MOS) is fundamental to speech quality assessment. However, its acquisition requires significant human annotation. Although deep neural network approaches, such as DNSMOS and UTMOS, have been developed to predict MOS to avoid this issue, they often suffer from insufficient training data. Recognizing that the comparison of speech enhancement (SE) systems prioritizes a reliable system comparison over absolute scores, we propose URGENT-PK, a novel ranking approach leveraging pairwise comparisons. URGENT-PK takes homologous enhanced speech pairs as input to predict relative quality rankings. This pairwise paradigm efficiently utilizes limited training data, as all pairwise permutations of multiple systems constitute a training instance. Experiments across multiple open test sets demonstrate URGENT-PK's superior system-level ranking performance over state-of-the-art baselines, despite its simple network architecture and limited training data.
【2】Less is More: Data Curation Matters in Scaling Speech Enhancement
链接:https://arxiv.org/abs/2506.23859
备注:Submitted to ASRU2025
摘要:绝大多数现代语音增强系统依赖于数据驱动的神经网络模型。传统上,较大的数据集被认为会产生更好的模型性能,这一观察结果在其他领域的许多任务中得到了经验验证。然而,最近的研究表明,当缩放语音增强数据时,收益递减。我们专注于一个关键因素:大规模数据集中“干净”训练标签中普遍存在的质量问题。这项工作重新审视了这一现象,并证明了在大规模训练集中,优先考虑高质量的训练数据比仅仅扩大数据量更重要。实验结果表明,在精心策划的700小时子集上训练的模型可以优于在2,500小时完整数据集上训练的模型。这一结果突出了数据管理在有效扩展语音增强系统中的关键作用。
摘要:The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling speech enhancement data. We focus on a critical factor: prevalent quality issues in ``clean'' training labels within large-scale datasets. This work re-examines this phenomenon and demonstrates that, within large-scale training sets, prioritizing high-quality training data is more important than merely expanding the data volume. Experimental findings suggest that models trained on a carefully curated subset of 700 hours can outperform models trained on the 2,500-hour full dataset. This outcome highlights the crucial role of data curation in scaling speech enhancement systems effectively.
【3】Human-CLAP: Human-perception-based contrastive language-audio pretraining
链接:https://arxiv.org/abs/2506.23553
摘要:对比语言-音频预训练(CLAP)被广泛用于音频生成和识别任务。例如,CLAPScore利用CLAP嵌入的相似性,已经成为文本到音频中评估音频和文本之间相关性的主要度量。然而,CLAPScore和人类主观评价分数之间的关系仍然不清楚。我们表明CLAPScore与人类主观评价分数的相关性较低。此外,我们提出了一个基于人类感知的CLAP称为人类CLAP通过使用主观评价分数训练对比语言音频模型。在我们的实验中,结果表明,我们的人类CLAP提高了斯皮尔曼的等级相关系数(SRCC)的CLAPSCore和主观评价分数超过0.25相比,传统的CLAP。
摘要:Contrastive language-audio pretraining (CLAP) is widely used for audio generation and recognition tasks. For example, CLAPScore, which utilizes the similarity of CLAP embeddings, has been a major metric for the evaluation of the relevance between audio and text in text-to-audio. However, the relationship between CLAPScore and human subjective evaluation scores is still unclarified. We show that CLAPScore has a low correlation with human subjective evaluation scores. Additionally, we propose a human-perception-based CLAP called Human-CLAP by training a contrastive language-audio model using the subjective evaluation score. In our experiments, the results indicate that our Human-CLAP improved the Spearman's rank correlation coefficient (SRCC) between the CLAPScore and the subjective evaluation scores by more than 0.25 compared with the conventional CLAP.
【4】Investigating an Overfitting and Degeneration Phenomenon in Self-Supervised Multi-Pitch Estimation
链接:https://arxiv.org/abs/2506.23371
备注:Accepted to ISMIR 2025
摘要:多音高估计(MPE)仍然是音乐信息检索(MIR)系统的一种受欢迎的能力,并且对于涉及音高的许多应用和下游任务(包括音乐转录)是至关重要的。然而,现有的方法在很大程度上是基于监督学习的,并且在为任务收集注释数据方面存在重大挑战。最近,利用音调和谐波信号的固有特性的自监督技术已经显示出对单声道和多声道音调估计的希望,但是这些仍然劣于监督方法。在这项工作中,我们扩展了经典的监督MPE范式,将几个自我监督的目标的基础上音高不变和音高等变属性。这种联合训练在封闭的训练条件下取得了很大的改进,这自然表明,将同样的目标应用于更广泛的数据收集将产生进一步的改进。然而,在这样做的过程中,我们发现了一个现象,即我们的模型同时过拟合监督数据,而退化的数据仅用于自我监督。我们证明和调查这一点,并提供我们的见解的根本问题。
摘要:Multi-Pitch Estimation (MPE) continues to be a sought after capability of Music Information Retrieval (MIR) systems, and is critical for many applications and downstream tasks involving pitch, including music transcription. However, existing methods are largely based on supervised learning, and there are significant challenges in collecting annotated data for the task. Recently, self-supervised techniques exploiting intrinsic properties of pitch and harmonic signals have shown promise for both monophonic and polyphonic pitch estimation, but these still remain inferior to supervised methods. In this work, we extend the classic supervised MPE paradigm by incorporating several self-supervised objectives based on pitch-invariant and pitch-equivariant properties. This joint training results in a substantial improvement under closed training conditions, which naturally suggests that applying the same objectives to a broader collection of data will yield further improvements. However, in doing so we uncover a phenomenon whereby our model simultaneously overfits to the supervised data while degenerating on data used for self-supervision only. We demonstrate and investigate this and offer our insights on the underlying problem.
【5】Adaptable Non-parametric Approach for Speech-based Symptom Assessment: Isolating Private Medical Data in a Retrieval Datastore
链接:https://arxiv.org/abs/2506.22972
备注:IEEE MLSP 2025
摘要:与健康相关的声学线索的自动评估有可能提高医疗保健的可及性和可负担性。虽然参数模型很有前途,但它们面临着隐私和适应性方面的挑战。为了解决这些问题,我们提出了一个基于语音的症状评估(NoNPSA)的NoN参数框架。通过在检索表中隔离医疗数据,NoNPSA避免了在模型参数中编码私有信息,并实现了有效的数据更新。在通用数据集上预训练的自监督学习(SSL)模型提取用于基于相似性的检索的特征。元数据感知的细化过滤检索到的数据,相关的标签用于计算评估分数。实验结果表明,与微调基于SSL的方法相比,NoNPSA实现了具有竞争力的性能,同时实现了更高的隐私性,更新效率和适应性-展示了非参数方法在医疗保健中的潜力。
摘要:The automatic assessment of health-related acoustic cues has the potential to improve healthcare accessibility and affordability. Although parametric models are promising, they face challenges in privacy and adaptability. To address these, we propose a NoN-Parametric framework for Speech-based symptom Assessment (NoNPSA). By isolating medical data in a retrieval datastore, NoNPSA avoids encoding private information in model parameters and enables efficient data updates. A self-supervised learning (SSL) model pre-trained on general-purpose datasets extracts features, which are used for similarity-based retrieval. Metadata-aware refinement filters the retrieved data, and associated labels are used to compute an assessment score. Experimental results show that NoNPSA achieves competitive performance compared to fine-tuning SSL-based methods, while enabling greater privacy, update efficiency, and adaptability--showcasing the potential of non-parametric approaches in healthcare.
【6】Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
链接:https://arxiv.org/abs/2506.22646
备注:Accepted by INTERSPEECH 2025
摘要:我们提出了一种自说话人自适应方法流多说话人自动语音识别(ASR),消除了显式扬声器查询的需要。与需要目标说话者嵌入或注册音频的传统方法不同,我们的技术通过说话者语音活动预测动态地适应单个ASR实例。关键的创新包括将通过扬声器监督激活生成的特定于扬声器的内核注入选定的ASR编码器层。这使得即使在流场景中也能够在处理完全重叠的语音的同时实现对目标扬声器的瞬时扬声器自适应。实验表明,在离线和流媒体场景中的最先进的性能,表明我们的自适应方法有效地解决了严重的语音重叠,通过精简的扬声器为重点的识别。实验结果验证了所提出的自说话人自适应方法在严重语音重叠条件下是一种鲁棒的多说话人ASR解决方案。
摘要:We propose a self-speaker adaptation method for streaming multi-talker automatic speech recognition (ASR) that eliminates the need for explicit speaker queries. Unlike conventional approaches requiring target speaker embeddings or enrollment audio, our technique dynamically adapts individual ASR instances through speaker-wise speech activity prediction. The key innovation involves injecting speaker-specific kernels generated via speaker supervision activations into selected ASR encoder layers. This enables instantaneous speaker adaptation to target speakers while handling fully overlapped speech even in a streaming scenario. Experiments show state-of-the-art performance in both offline and streaming scenarios, demonstrating that our self-adaptive method effectively addresses severe speech overlap through streamlined speaker-focused recognition. The results validate the proposed self-speaker adaptation approach as a robust solution for multi-talker ASR under severe overlapping speech conditions.
【7】StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
链接:https://arxiv.org/abs/2506.23986
摘要:基于离散令牌的语音生成的最新进展突出了令牌到波形生成对于音频质量的重要性,特别是在实时交互中。将语义令牌与流匹配(FM)集成的传统框架由于依赖于全局接受域而难以实现流功能。此外,直接实现逐令牌流式语音生成通常导致音频质量下降。为了解决这些挑战,我们提出了StreamFlow,这是一种新型的神经架构,可以促进与扩散Transformers(DiT)的流匹配。为了减轻长期的历史依赖性所产生的长序列外推问题,我们设计了一个本地块明智的感受野策略。具体来说,序列首先被分割成块,我们引入了块级注意掩码,使当前块能够接收来自前一个或后一个块的信息。这些注意力面具在不同的DiT块中分层组合,以调节DiT的感受野。主观和客观的实验结果表明,我们的方法实现的性能与非流方法相比,同时超过其他流方法的语音质量,同时有效地管理推理时间在长序列生成。此外,我们的方法实现了仅180 ms的显着的第一个数据包延迟。footnote{语音示例:https://dukguo.github.io/StreamFlow/}
摘要:Recent advancements in discrete token-based speech generation have highlighted the importance of token-to-waveform generation for audio quality, particularly in real-time interactions. Traditional frameworks integrating semantic tokens with flow matching (FM) struggle with streaming capabilities due to their reliance on a global receptive field. Additionally, directly implementing token-by-token streaming speech generation often results in degraded audio quality. To address these challenges, we propose StreamFlow, a novel neural architecture that facilitates streaming flow matching with diffusion transformers (DiT). To mitigate the long-sequence extrapolation issues arising from lengthy historical dependencies, we design a local block-wise receptive field strategy. Specifically, the sequence is first segmented into blocks, and we introduce block-wise attention masks that enable the current block to receive information from the previous or subsequent block. These attention masks are combined hierarchically across different DiT-blocks to regulate the receptive field of DiTs. Both subjective and objective experimental results demonstrate that our approach achieves performance comparable to non-streaming methods while surpassing other streaming methods in terms of speech quality, all the while effectively managing inference time during long-sequence generation. Furthermore, our method achieves a notable first-packet latency of only 180 ms.\footnote{Speech samples: https://dukguo.github.io/StreamFlow/}
【8】Emergent musical properties of a transformer under contrastive self-supervised learning
链接:https://arxiv.org/abs/2506.23873
备注:Accepted at ISMIR 2025
摘要:在音乐信息检索(MIR)中,通用表示模型的对比自监督学习对于自动标注等全局任务是有效的。然而,对于和弦估计等本地任务,人们普遍认为对比训练的通用自监督模型是不够的,需要更复杂的SSL;例如,蒙面建模。我们的论文挑战这一假设,揭示了潜在的对比SSL与一个Transformer在本地MIR任务配对。我们考虑一个轻量级的Vision Transformer与一维补丁在时间-频率域(ViT-1D)和训练它与简单的对比SSL通过归一化的温度缩放的交叉熵损失(NT-Xent)。虽然NT-Xent只在类令牌上运行,但我们观察到,可能由于权重共享,信息丰富的音乐属性出现在ViT-1D的序列令牌中。在全局任务中,类和序列标记的时间平均值与单独的类标记相比提供了性能提高,显示了序列标记中的有用属性。在本地任务中,序列令牌的表现出乎意料地好,尽管没有专门训练。此外,高层次的音乐特征,如开始出现从逐层注意力地图和自相似矩阵显示不同的层捕捉不同的音乐维度。我们的论文并不专注于提高性能,但推进Transformers的音乐解释,并揭示了一些被忽视的对比SSL与Transformers配对的MIR序列建模的能力。
摘要:In music information retrieval (MIR), contrastive self-supervised learning for general-purpose representation models is effective for global tasks such as automatic tagging. However, for local tasks such as chord estimation, it is widely assumed that contrastively trained general-purpose self-supervised models are inadequate and that more sophisticated SSL is necessary; e.g., masked modeling. Our paper challenges this assumption by revealing the potential of contrastive SSL paired with a transformer in local MIR tasks. We consider a lightweight vision transformer with one-dimensional patches in the time--frequency domain (ViT-1D) and train it with simple contrastive SSL through normalized temperature-scaled cross-entropy loss (NT-Xent). Although NT-Xent operates only over the class token, we observe that, potentially thanks to weight sharing, informative musical properties emerge in ViT-1D's sequence tokens. On global tasks, the temporal average of class and sequence tokens offers a performance increase compared to the class token alone, showing useful properties in the sequence tokens. On local tasks, sequence tokens perform unexpectedly well, despite not being specifically trained for. Furthermore, high-level musical features such as onsets emerge from layer-wise attention maps and self-similarity matrices show different layers capture different musical dimensions. Our paper does not focus on improving performance but advances the musical interpretation of transformers and sheds light on some overlooked abilities of contrastive SSL paired with transformers for sequence modeling in MIR.
【9】Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
链接:https://arxiv.org/abs/2506.23869
备注:ISMIR (2025)
摘要:我们研究了生成自回归Transformer模型在大量符号化钢琴独奏transertion上训练的能力。在对大约60,000小时的音乐进行第一次预训练后,我们使用相对较小的高质量子集来微调模型,以产生音乐延续,执行符号分类任务,并通过将Simplified框架适应符号音乐来产生通用对比嵌入。在评估钢琴连续连贯性时,我们的生成模型优于领先的符号生成技术,并与专有音频生成模型保持竞争力。在MIR分类基准测试中,我们的对比模型中的冻结表示在线性探测实验中获得了最先进的结果,而直接微调则展示了预训练表示的可推广性,通常只需要几百个标记的示例就可以专门用于下游任务。
摘要:We study the capabilities of generative autoregressive transformer models trained on large amounts of symbolic solo-piano transcriptions. After first pretraining on approximately 60,000 hours of music, we use a comparatively smaller, high-quality subset, to finetune models to produce musical continuations, perform symbolic classification tasks, and produce general-purpose contrastive MIDI embeddings by adapting the SimCLR framework to symbolic music. When evaluating piano continuation coherence, our generative model outperforms leading symbolic generation techniques and remains competitive with proprietary audio generation models. On MIR classification benchmarks, frozen representations from our contrastive model achieve state-of-the-art results in linear probe experiments, while direct finetuning demonstrates the generalizability of pretrained representations, often requiring only a few hundred labeled examples to specialize to downstream tasks.
【10】Efficient Interleaved Speech Modeling through Knowledge Distillation
链接:https://arxiv.org/abs/2506.23670
摘要:当前的语音语言模型超过了许多部署环境的大小和延迟限制。我们通过层对齐蒸馏,匹配隐藏状态,注意力地图和软化logits来构建紧凑,富有表现力的语音生成模型,以最小的性能损失将大型多模态Transformers压缩3倍。我们介绍TinyWave,这是一系列用于语音到语音和交错语音文本生成的2B参数模型,在50,000小时的公共音频上进行了训练。TinyWave支持(i)使用语音或表达令牌生成纯语音,以及(ii)混合语音文本延续。Libri-Light上的评估显示TinyWave在其教师的1.4标准化困惑点内。口语StoryCloze和SALMon的准确率达到教师表现的93-97%,优于大小匹配的基线。这些模型针对商用硬件上的部署进行了优化,支持实时会话代理、辅助技术和低资源环境中的应用程序。我们发布模型,训练代码和评估脚本,以支持对紧凑,富有表现力的语音生成的可重复研究。
摘要:Current speech language models exceed the size and latency constraints of many deployment environments. We build compact, expressive speech generation models through layer-aligned distillation, matching hidden states, attention maps, and softened logits to compress large multimodal transformers by 3x with minimal loss in performance. We introduce TinyWave, a family of 2B-parameter models for speech-to-speech and interleaved speech-text generation, trained on 50,000 hours of public audio. TinyWave supports (i) speech-only generation using phonetic or expressive tokens and (ii) mixed speech-text continuations. Evaluation on Libri-Light shows TinyWave within 1.4 normalized perplexity points of its teacher. Accuracy on spoken StoryCloze and SALMon reaches 93-97% of the teacher's performance, outperforming size-matched baselines. These models are optimized for deployment on commodity hardware, enabling applications in real-time conversational agents, assistive technologies, and low-resource environments. We release models, training code, and evaluation scripts to support reproducible research on compact, expressive speech generation.
【11】RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
链接:https://arxiv.org/abs/2506.23582
备注:Accepted to INTERSPEECH2025
摘要:在文本到音频(TTA)的研究中,输入文本和输出音频之间的相关性是一个重要的评价方面。传统的评价方法是从主观和客观两个方面进行的。然而,主观评价在金钱和时间方面是昂贵的,并且客观评价与主观评价分数的相关性不清楚。在这项研究中,我们构建了一个开源的数据集,它主观地评估了相关性。此外,我们基准的模型,自动预测的主观评价得分从合成的音频。我们的模型优于传统的CLAPScore模型,这种趋势扩展到许多声音类别。
摘要:In text-to-audio (TTA) research, the relevance between input text and output audio is an important evaluation aspect. Traditionally, it has been evaluated from both subjective and objective perspectives. However, subjective evaluation is costly in terms of money and time, and objective evaluation is unclear regarding the correlation to subjective evaluation scores. In this study, we construct RELATE, an open-sourced dataset that subjectively evaluates the relevance. Also, we benchmark a model for automatically predicting the subjective evaluation score from synthesized audio. Our model outperforms a conventional CLAPScore model, and that trend extends to many sound categories.
【12】JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
链接:https://arxiv.org/abs/2506.23552
备注:project page: this https URL Under review. Preprint published on arXiv
摘要:在生成建模中,面部运动和语音之间的内在联系经常被忽视,其中说话的头部合成和文本到语音(TTS)通常被视为单独的任务。本文介绍了JAM-Flow,一个统一的框架,同时合成和条件的面部运动和语音。我们的方法利用流匹配和一种新的多模扩散Transformer(MM-DiT)架构,集成了专门的运动DiT和音频DiT模块。这些通过选择性联合注意层耦合,并结合关键的架构选择,如时间对齐的位置嵌入和局部联合注意掩蔽,以实现有效的跨模态交互,同时保持模态特定的优势。JAM-Flow采用inpainting风格的目标进行训练,支持各种条件输入(包括文本、参考音频和参考运动),可在单个连贯模型中执行促进任务,例如从文本同步生成说话的头部、音频驱动的动画等。JAM-Flow通过为整体视听合成提供实用的解决方案,大大推进了多模态生成建模。项目页面:https://joonghyuk.com/jamflow-web
摘要:The intrinsic link between facial motion and speech is often overlooked in generative modeling, where talking head synthesis and text-to-speech (TTS) are typically addressed as separate tasks. This paper introduces JAM-Flow, a unified framework to simultaneously synthesize and condition on both facial motion and speech. Our approach leverages flow matching and a novel Multi-Modal Diffusion Transformer (MM-DiT) architecture, integrating specialized Motion-DiT and Audio-DiT modules. These are coupled via selective joint attention layers and incorporate key architectural choices, such as temporally aligned positional embeddings and localized joint attention masking, to enable effective cross-modal interaction while preserving modality-specific strengths. Trained with an inpainting-style objective, JAM-Flow supports a wide array of conditioning inputs-including text, reference audio, and reference motion-facilitating tasks such as synchronized talking head generation from text, audio-driven animation, and much more, within a single, coherent model. JAM-Flow significantly advances multi-modal generative modeling by providing a practical solution for holistic audio-visual synthesis. project page: https://joonghyuk.com/jamflow-web
【13】From Large-scale Audio Tagging to Real-Time Explainable Emergency Vehicle Sirens Detection
链接:https://arxiv.org/abs/2506.23437
备注:pre-print (submitted to the IEEE/ACM Transactions on Audio, Speech, and Language Processing)
摘要:紧急车辆(EV)警报器的准确识别对于智能交通系统,智能城市监控系统和自动驾驶技术的集成至关重要。现代的自动解决方案受到缺乏大规模的策划数据集以及最先进的声音事件检测模型的计算需求的限制。这项工作介绍了E2PANNs(高效紧急预训练音频神经网络),这是一种源自PANNs框架的轻量级卷积神经网络架构,专门针对二进制EV警报器检测进行了优化。利用我们专用的AudioSet子集(AudioSet EV),我们在多个参考数据集上微调和评估E2PANN,并在嵌入式硬件上测试其可行性。实验活动包括消融研究,跨域基准测试和边缘设备上的实时推理部署。利用引导反向传播和ScoreCAM算法的可解释性分析提供了对模型内部表示的见解,并验证了其捕获与不同类型EV警报相关的不同光谱时间模式的能力。实时性能通过逐帧和基于事件的检测指标以及对假阳性激活的详细分析进行评估。结果表明,E2PANNs建立了一个新的国家的艺术在这个研究领域,具有高计算效率,并适用于基于边缘的音频监测和安全关键应用。
摘要:Accurate recognition of Emergency Vehicle (EV) sirens is critical for the integration of intelligent transportation systems, smart city monitoring systems, and autonomous driving technologies. Modern automatic solutions are limited by the lack of large scale, curated datasets and by the computational demands of state of the art sound event detection models. This work introduces E2PANNs (Efficient Emergency Pre trained Audio Neural Networks), a lightweight Convolutional Neural Network architecture derived from the PANNs framework, specifically optimized for binary EV siren detection. Leveraging our dedicated subset of AudioSet (AudioSet EV) we fine-tune and evaluate E2PANNs across multiple reference datasets and test its viability on embedded hardware. The experimental campaign includes ablation studies, cross-domain benchmarking, and real-time inference deployment on edge device. Interpretability analyses exploiting Guided Backpropagation and ScoreCAM algorithms provide insights into the model internal representations and validate its ability to capture distinct spectrotemporal patterns associated with different types of EV sirens. Real time performance is assessed through frame wise and event based detection metrics, as well as a detailed analysis of false positive activations. Results demonstrate that E2PANNs establish a new state of the art in this research domain, with high computational efficiency, and suitability for edge-based audio monitoring and safety-critical applications.
【14】You Sound a Little Tense: L2 Tailored Clear TTS Using Durational Vowel Properties
链接:https://arxiv.org/abs/2506.23367
备注:Accepted to ISCA Speech Synthesis Workshop, 2025
摘要:我们提出了第一个文本到语音(TTS)系统量身定制的第二语言(L2)扬声器。我们使用美国英语紧张(长)和宽松(短)元音之间的持续时间差异,以创建一个“清晰度模式”的Matcha-TTS。我们的感知研究表明,当使用我们的清晰度模式时,法语第一语言、英语第二语言听众的转录错误较少(至少9.15%),并且发现它比整体放慢的语音更令人鼓舞和尊重。值得注意的是,听众并没有意识到这些影响:尽管在清晰度模式下单词错误率降低,但听众仍然认为放慢所有目标单词是最容易理解的,这表明实际的可理解性与感知的可理解性无关。此外,我们发现,耳语ASR没有使用相同的线索,L2扬声器区分困难的元音,是不足以评估这些人的TTS系统的可懂度。
摘要:We present the first text-to-speech (TTS) system tailored to second language (L2) speakers. We use duration differences between American English tense (longer) and lax (shorter) vowels to create a "clarity mode" for Matcha-TTS. Our perception studies showed that French-L1, English-L2 listeners had fewer (at least 9.15%) transcription errors when using our clarity mode, and found it more encouraging and respectful than overall slowed down speech. Remarkably, listeners were not aware of these effects: despite the decreased word error rate in clarity mode, listeners still believed that slowing all target words was the most intelligible, suggesting that actual intelligibility does not correlate with perceived intelligibility. Additionally, we found that Whisper-ASR did not use the same cues as L2 speakers to differentiate difficult vowels and is not sufficient to assess the intelligibility of TTS systems for these individuals.
【15】XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
链接:https://arxiv.org/abs/2506.23325
摘要:语音编解码器是语音信号和大型语言模型之间的桥梁。一个理想的语音语言模型编解码器不仅要保留声学信息,而且要捕捉丰富的语义信息。然而,现有的语音编解码器难以平衡高质量的音频重建与易于通过语言模型建模。在这项研究中,我们分析了以前的编解码器在平衡语义丰富性和声学保真度方面的局限性。我们提出了XY-Tokenizer,一种新的编解码器,通过多阶段,多任务学习减轻语义和声学能力之间的冲突。实验结果表明,XY-Tokenizer在语义和声学任务中的性能与以类似比特率操作的最先进的编解码器相当,即使现有的编解码器通常只在一个方面表现出色。具体来说,XY-Tokenizer实现了强大的文本对齐,超过了基于蒸馏的语义建模方法,如SpeechTokenizer和Mimi,同时在重建和原始音频之间保持了0.83的说话人相似性得分。XY-Tokenizer的重建性能与BigCodec相当,BigCodec是目前最先进的纯声学编解码器,在类似的比特率下实现了0.84的扬声器相似性得分。代码和型号可在https://github.com/gyt1145028706/XY-Tokenizer上获得。
摘要:Speech codecs serve as bridges between speech signals and large language models. An ideal codec for speech language models should not only preserve acoustic information but also capture rich semantic information. However, existing speech codecs struggle to balance high-quality audio reconstruction with ease of modeling by language models. In this study, we analyze the limitations of previous codecs in balancing semantic richness and acoustic fidelity. We propose XY-Tokenizer, a novel codec that mitigates the conflict between semantic and acoustic capabilities through multi-stage, multi-task learning. Experimental results demonstrate that XY-Tokenizer achieves performance in both semantic and acoustic tasks comparable to that of state-of-the-art codecs operating at similar bitrates, even though those existing codecs typically excel in only one aspect. Specifically, XY-Tokenizer achieves strong text alignment, surpassing distillation-based semantic modeling methods such as SpeechTokenizer and Mimi, while maintaining a speaker similarity score of 0.83 between reconstructed and original audio. The reconstruction performance of XY-Tokenizer is comparable to that of BigCodec, the current state-of-the-art among acoustic-only codecs, which achieves a speaker similarity score of 0.84 at a similar bitrate. Code and models are available at https://github.com/gyt1145028706/XY-Tokenizer.
【16】The Florence Price Art Song Dataset and Piano Accompaniment Generator
链接:https://arxiv.org/abs/2506.23130
备注:8 pages, 4 figures. To appear in the proceedings of ISMIR 2025
摘要:佛罗伦萨湾普赖斯是20世纪初的一位作曲家,她的音乐反映了她在美国南部的成长经历、她的非洲遗产和她的西方古典训练。她被认为是第一位由大型管弦乐队演奏交响乐的非洲裔美国女性。在她去世几十年后,她的音乐最近再次受到公众和研究界的关注。除了其他流派,普莱斯还是一位多产的独唱和钢琴作曲家。音乐历史学家记录了普赖斯创作的134首艺术歌曲和钢琴/人声编曲的灵歌和民歌。我们以MuseScore、MusicXML、XML和PDF格式发布了其中112件作品的数字目录。我们还使用这个数据集微调符号音乐生成模型,以生成旋律的替代品,我们进行了盲听实验,表明我们的模型生成的替代品被认为是反映佛罗伦萨价格的风格比基线模型生成的替代品更频繁。我们将我们的模型作为Florence Price钢琴伴奏生成器与我们的数据集一起发布。
摘要:Florence B. Price was a composer in the early 20th century whose music reflects her upbringing in the American South, her African heritage, and her Western classical training. She is noted as the first African-American woman to have a symphony performed by a major orchestra. Her music has recently received renewed attention from both the public and the research community, decades after her death. In addition to other genres, Price was a prolific composer for solo voice and piano. Music historians have documented the existence of 134 art songs and piano/voice arrangements for spirituals and folk songs written by Price. We release a digital catalog of 112 of these works in MuseScore, MusicXML, MIDI, and PDF format. We also use this dataset to fine-tune a symbolic music generation model to generate accompaniments to melodies, and we conduct a blind listening experiment that shows that accompaniments generated by our model are perceived as being reflective of Florence Price's style more frequently than accompaniments generated by a baseline model. We release our model as the Florence Price Piano Accompaniment Generator alongside our dataset.
【17】TOMI: Transforming and Organizing Music Ideas for Multi-Track Compositions with Full-Song Structure
链接:https://arxiv.org/abs/2506.23094
备注:9 pages, 4 figures, 2 tables. To be published in ISMIR 2025
摘要:分层规划是一种从结构上对长序列建模的有效方法。除了考虑音乐的时间结构的层次,本文探讨了一个更重要的方面:概念层次,这涉及到产生音乐的想法,转换它们,并最终组织它们-跨越音乐的时间和空间-成为一个完整的组成。为此,我们引入了TOMI(转换和组织音乐思想)作为深度音乐生成的一种新方法,并通过调整基础LLM开发了一个基于TOMI的模型。形式上,我们通过一个稀疏的四维空间来表示多轨作曲过程,该四维空间的特征在于剪辑(短音频或音频片段),部分(时间位置),轨道(乐器层)和转换(精心制作方法)。我们的模型能够生成具有完整歌曲结构的多音轨电子音乐,并且我们进一步将基于TOMI的模型与REAPER数字音频工作站集成,实现交互式的人-AI共同创作。实验结果表明,我们的方法产生更高质量的电子音乐与更强的结构一致性相比,基线。
摘要:Hierarchical planning is a powerful approach to model long sequences structurally. Aside from considering hierarchies in the temporal structure of music, this paper explores an even more important aspect: concept hierarchy, which involves generating music ideas, transforming them, and ultimately organizing them--across musical time and space--into a complete composition. To this end, we introduce TOMI (Transforming and Organizing Music Ideas) as a novel approach in deep music generation and develop a TOMI-based model via instruction-tuned foundation LLM. Formally, we represent a multi-track composition process via a sparse, four-dimensional space characterized by clips (short audio or MIDI segments), sections (temporal positions), tracks (instrument layers), and transformations (elaboration methods). Our model is capable of generating multi-track electronic music with full-song structure, and we further integrate the TOMI-based model with the REAPER digital audio workstation, enabling interactive human-AI co-creation. Experimental results demonstrate that our approach produces higher-quality electronic music with stronger structural coherence compared to baselines.
【18】AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
链接:https://arxiv.org/abs/2506.23049
摘要:尽管语言和语音技术取得了进步,但没有开源系统能够实现完整的语音到语音,多轮对话,集成工具使用和代理推理。AURA(Agent for Understanding,Reasoning,and Automated Tool Use)是第一个开源的语音助手,能够通过动态工具调用和多轮对话完成复杂的目标驱动任务。AURA将开放权重的ASR、TTS和LLM组合在一个级联管道中,并支持日历预订、联系人查找、Web搜索和电子邮件等工具。其模块化设计允许使用自然语言提示和操作类轻松集成新工具。在VoiceBench上,AURA在OpenBookQA上的得分为92.75%,优于所有开放重量系统,接近GPT-40,在AlpacaEval上为4.39,与其他开放重量系统竞争。人工评估显示,在复杂的多回合语音任务中,任务成功率为90%。
摘要:Despite advances in language and speech technologies, no open-source system enables full speech-to-speech, multi-turn dialogue with integrated tool use and agentic reasoning. We introduce AURA (Agent for Understanding, Reasoning, and Automated Tool Use), the first open-source, speech-native assistant capable of completing complex, goal-driven tasks through dynamic tool invocation and multi-turn conversation. AURA combines open-weight ASR, TTS, and LLMs in a cascaded pipeline and supports tools such as calendar booking, contact lookup, web search, and email. Its modular design allows easy integration of new tools using natural language prompts and action classes. On VoiceBench, AURA scores 92.75% on OpenBookQA-outperforming all open-weight systems and nearing GPT-4o-and 4.39 on AlpacaEval, competitive with other open-weight systems. Human evaluation shows 90% task success on complex, multi-turn speech tasks.
【19】VisionScores -- A system-segmented image score dataset for deep learning tasks
链接:https://arxiv.org/abs/2506.23030
备注:Comments: 5 pages, 3 figures. Accepted for presentation at the 2025 IEEE International Conference on Image Processing (ICIP). \c{opyright} 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for any other use
摘要:VisionScores提出了一个新的建议,作为第一个系统分割的图像分数数据集,旨在为机器和深度学习任务提供结构丰富的高信息密度图像。限定为双手钢琴作品,它不仅要考虑某些图形相似性,还要考虑构图模式,因为这个创作过程高度依赖于乐器。它提供了两个与作曲家和作曲类型有关的场景。第一个由14 k个样本组成,考虑来自不同作者但相同作曲类型的作品,特别是奏鸣曲。后者由10.8K样本组成,呈现相反的情况,来自同一作者的各种作曲类型,是弗朗茨·李斯特的一个。所有24.8k样本都被格式化为$128 \times 512$像素的灰度jpg图像。VisionScores不仅为用户提供格式化的样本,还提供系统的顺序和片段的元数据。此外,未分割的整版分数和预先格式化的图像包括进一步分析。
摘要:VisionScores presents a novel proposal being the first system-segmented image score dataset, aiming to offer structure-rich, high information-density images for machine and deep learning tasks. Delimited to two-handed piano pieces, it was built to consider not only certain graphic similarity but also composition patterns, as this creative process is highly instrument-dependent. It provides two scenarios in relation to composer and composition type. The first, formed by 14k samples, considers works from different authors but the same composition type, specifically, Sonatinas. The latter, consisting of 10.8K samples, presents the opposite case, various composition types from the same author, being the one selected Franz Liszt. All of the 24.8k samples are formatted as grayscale jpg images of $128 \times 512$ pixels. VisionScores supplies the users not only the formatted samples but the systems' order and pieces' metadata. Moreover, unsegmented full-page scores and the pre-formatted images are included for further analysis.
【20】Feasibility of spectral-element modeling of wave propagation through the anatomy of marine mammals
链接:https://arxiv.org/abs/2506.22944
摘要:本文首次采用三维谱元法(SEM)模拟了超声波在短吻海豚(Tursiops truncatus)头部的传播。与传统的有限元方法(FEM)不同,由于昂贵的线性系统反演和较慢的收敛速度,SEM提供了指数收敛和高效的并行计算。使用计算机断层扫描(CT)扫描数据,我们开发了一个详细的六面体网格,捕获复杂的解剖特征,例如声学脂肪和颌骨。我们的平面波和球面波的模拟证实了SEM的有效性超声时域建模。这种方法为海洋生物学开辟了新的途径,有助于回声定位、人为海洋噪音污染的影响以及海洋哺乳动物听觉和咔哒声产生的生物物理学研究。通过克服FEM的局限性,SEM提供了一个强大的可扩展的工具来测试海豚生物声学的假设,具有重要意义的保护和理解海洋哺乳动物的听觉系统下日益增加的环境挑战。
摘要:This study introduces the first 3D spectral-element method (SEM) simulation of ultrasonic wave propagation in a bottlenose dolphin (Tursiops truncatus) head. Unlike traditional finite-element methods (FEM), which struggle with high-frequency simulations due to costly linear-system inversions and slower convergence, SEM offers exponential convergence and efficient parallel computation. Using Computed Tomography (CT) scan data, we developed a detailed hexahedral mesh capturing complex anatomical features, such as acoustic fats and jaws. Our simulations of plane and spherical waves confirm SEM's effectiveness for ultrasonic time-domain modeling. This approach opens new avenues for marine biology, contributing to research in echolocation, the impacts of anthropogenic marine noise pollution and the biophysics of hearing and click generation in marine mammals. By overcoming FEM's limitations, SEM provides a powerful scalable tool to test hypotheses about dolphin bioacoustics, with significant implications for conservation and understanding marine mammal auditory systems under increasing environmental challenges.
【21】Mind the Gap: Entity-Preserved Context-Aware ASR Structured Transcriptions
链接:https://arxiv.org/abs/2506.22858
备注:This is the accepted version of an article accepted to the TSD 2025 conference, published in Springer Lecture Notes in Artificial Intelligence (LNAI). The final authenticated version is available online at SpringerLink
摘要:自动语音识别(ASR)系统(例如Whisper)可以实现高转录准确性,但在命名实体和数字数据方面遇到了困难,尤其是在需要正确格式化时。这些问题增加了单词错误率(WER),并损害了法律,金融和医疗应用等关键领域的语义理解。我们提出了一种新的训练方法,通过在训练过程中添加重叠的上下文窗口来扩展ASR模型的语义上下文。通过在30秒块的两侧滑动5秒重叠,我们创建了一个40秒的“有效语义窗口”,提高实体识别和格式,同时将预测集中在中间30秒。为了解决跨越块边界的实体,我们将此类实体完全重新分配给右侧块,以确保正确的格式。此外,具有嵌入式实体标签的丰富训练数据使模型能够学习识别和特定类型的格式。在Spoken Wikipedia数据集上进行评估,我们的方法提高了语义任务的性能,包括命名实体识别(NER)和实体格式化。这些结果突出了上下文感知训练在解决长形式转录和复杂实体识别任务的ASR限制方面的有效性。
摘要:Automatic Speech Recognition (ASR) systems, such as Whisper, achieve high transcription accuracy but struggle with named entities and numerical data, especially when proper formatting is required. These issues increase word error rate (WER) and impair semantic understanding in critical domains like legal, financial, and medical applications. We propose a novel training approach that extends the semantic context of ASR models by adding overlapping context windows during training. By sliding 5-second overlaps on both sides of 30-second chunks, we create a 40-second "effective semantic window," improving entity recognition and formatting while focusing predictions on the central 30 seconds. To address entities spanning chunk boundaries, we reassign such entities entirely to the right-hand chunk, ensuring proper formatting. Additionally, enriched training data with embedded entity labels enables the model to learn both recognition and type-specific formatting. Evaluated on the Spoken Wikipedia dataset, our method improves performance across semantic tasks, including named entity recognition (NER) and entity formatting. These results highlight the effectiveness of context-aware training in addressing ASR limitations for long-form transcription and complex entity recognition tasks.
【22】Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization
链接:https://arxiv.org/abs/2506.22846
备注:This is the accepted version of an article accepted to the TSD 2025 conference, published in Springer Lecture Notes in Artificial Intelligence (LNAI). The final authenticated version is available online at SpringerLink
摘要:端到端(E2 E)自动语音识别(ASR)系统通过将所有组件集成到单个神经网络中,使该领域发生了革命性变化,基于注意力的编码器-解码器模型实现了最先进的性能。然而,它们的自回归解码过程限制了推理速度,使其不适合实时应用。相比之下,基于CTC的模型提供更快的非自回归解码,但难以有效地对语言依赖性进行建模。为了应对这一挑战,我们提出了一种新的辅助损失框架,称为存储感知中间损失(LAIL),以增强基于CTC的ASR使用大型语言模型(LLM)的语言知识。通过将连接器层连接到中间编码器层,LAIL将输出映射到LLM的嵌入空间,并在训练期间计算因果语言建模损失。这种方法增强了语言建模,同时保持CTC解码的计算效率。使用Conformer架构和各种LLaMA模型,我们在LibriSpeech,TEDLIUM 2和WSJ语料库上展示了字错误率(WER)的显着改善,以最小的计算开销实现了基于CTC的ASR的最先进的性能。
摘要:End-to-end (E2E) automatic speech recognition (ASR) systems have revolutionized the field by integrating all components into a single neural network, with attention-based encoder-decoder models achieving state-of-the-art performance. However, their autoregressive decoding process limits inference speed, making them unsuitable for real-time applications. In contrast, CTC-based models offer faster, non-autoregressive decoding but struggle to model linguistic dependencies effectively. Addressing this challenge, we propose a novel auxiliary loss framework called Language-Aware Intermediate Loss (LAIL) to enhance CTC-based ASR using the linguistic knowledge of large language models (LLMs). By attaching connector layers to intermediate encoder layers, LAIL maps outputs to the embedding space of an LLM and computes a causal language modeling loss during training. This approach enhances linguistic modeling while preserving the computational efficiency of CTC decoding. Using the Conformer architecture and various LLaMA models, we demonstrate significant improvements in Word Error Rate (WER) on the LibriSpeech, TEDLIUM2, and WSJ corpora, achieving state-of-the-art performance for CTC-based ASR with minimal computational overhead.
【23】A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
链接:https://arxiv.org/abs/2506.22810
备注:accepted by Interspeech 2025
摘要:构音障碍语音识别(DSR)增强了移动性受限的构音障碍说话者对智能设备的可访问性。以前,DSR研究受到现有数据集通常由孤立的单词,命令短语和少数人所说的有限数量的句子组成的事实的限制。这限制了对命令交互系统和说话人自适应的研究。Speech Accessibility Project(SAP)通过发布一个大型且多样化的英语构音障碍数据集改变了这一点,导致SAP挑战赛构建说话者和文本独立的DSR系统。我们通过一种新的自我训练方法增强了Whisper模型在长时间构音障碍语音上的性能。这种方法增加了训练数据,并调整了模型,以处理在推理过程中遇到的可能不完整的语音片段。我们的系统在SAP Challenge中的单词错误率和语义得分均获得第二名。
摘要:Dysarthric speech recognition (DSR) enhances the accessibility of smart devices for dysarthric speakers with limited mobility. Previously, DSR research was constrained by the fact that existing datasets typically consisted of isolated words, command phrases, and a limited number of sentences spoken by a few individuals. This constrained research to command-interaction systems and speaker adaptation. The Speech Accessibility Project (SAP) changed this by releasing a large and diverse English dysarthric dataset, leading to the SAP Challenge to build speaker- and text-independent DSR systems. We enhanced the Whisper model's performance on long dysarthric speech via a novel self-training method. This method increased training data and adapted the model to handle potentially incomplete speech segments encountered during inference. Our system achieved second place in both Word Error Rate and Semantic Score in the SAP Challenge.
【24】WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
链接:https://arxiv.org/abs/2506.22789
备注:5 pages, 4 figures, Published at The Proceedings of Interspeech 2025, code is available at http://www.github.com/UTAustin-SwarmLab/WavShape
摘要:语音嵌入通常会保留敏感属性,如说话者身份、口音或人口统计信息,从而带来有偏见的模型训练和隐私泄露的风险。我们提出了WavShape,这是一个信息论语音表示学习框架,它优化了嵌入的公平性和隐私性,同时保留了任务相关的信息。我们利用互信息(MI)估计使用Donsker-Varadhan公式来指导基于MI的编码器,该编码器系统地过滤敏感属性,同时保持下游任务所必需的语音内容。在三个已知数据集上的实验结果表明,WavShape将嵌入和敏感属性之间的MI降低了81%,同时保留了97%的任务相关信息。通过将信息理论与自监督语音模型相结合,这项工作推进了公平,隐私感知和资源高效的语音系统的发展。
摘要:Speech embeddings often retain sensitive attributes such as speaker identity, accent, or demographic information, posing risks in biased model training and privacy leakage. We propose WavShape, an information-theoretic speech representation learning framework that optimizes embeddings for fairness and privacy while preserving task-relevant information. We leverage mutual information (MI) estimation using the Donsker-Varadhan formulation to guide an MI-based encoder that systematically filters sensitive attributes while maintaining speech content essential for downstream tasks. Experimental results on three known datasets show that WavShape reduces MI between embeddings and sensitive attributes by up to 81% while retaining 97% of task-relevant information. By integrating information theory with self-supervised speech models, this work advances the development of fair, privacy-aware, and resource-efficient speech systems.
【25】Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
链接:https://arxiv.org/abs/2506.22661
备注:Accepted to ISMIR2025
摘要:音频指纹识别(AFP)允许通过提取被称为音频指纹的紧凑表示来识别未知音频内容,所述音频指纹被设计成针对常见音频降级保持鲁棒性。神经AFP方法通常采用度量学习,其中表示质量受到监督的性质和所使用的损失函数的影响。然而,最近的工作不切实际地模拟训练期间的真实音频降级,导致次优监督。此外,尽管已经提出了几种现代度量学习方法,但当前的神经AFP方法仍然依赖于NT-Xent损失,而没有探索最新的进展或经典的替代方案。在这项工作中,我们提出了一系列的最佳实践,以加强自我监督,利用音乐信号的属性和现实的房间声学。然后,我们提出了第一个系统的评估各种度量学习方法的背景下,AFP,证明了自我监督适应的三重损失产生卓越的性能。我们的研究结果还表明,每个锚点使用多个阳性样本进行训练,在损失函数中具有截然不同的效果。我们的方法建立在这些见解的基础上,并在一个大型的综合退化数据集和一个在不同音乐场所使用麦克风记录的真实数据集上实现了最先进的性能。
摘要:Audio fingerprinting (AFP) allows the identification of unknown audio content by extracting compact representations, termed audio fingerprints, that are designed to remain robust against common audio degradations. Neural AFP methods often employ metric learning, where representation quality is influenced by the nature of the supervision and the utilized loss function. However, recent work unrealistically simulates real-life audio degradation during training, resulting in sub-optimal supervision. Additionally, although several modern metric learning approaches have been proposed, current neural AFP methods continue to rely on the NT-Xent loss without exploring the recent advances or classical alternatives. In this work, we propose a series of best practices to enhance the self-supervision by leveraging musical signal properties and realistic room acoustics. We then present the first systematic evaluation of various metric learning approaches in the context of AFP, demonstrating that a self-supervised adaptation of the triplet loss yields superior performance. Our results also reveal that training with multiple positive samples per anchor has critically different effects across loss functions. Our approach is built upon these insights and achieves state-of-the-art performance on both a large, synthetically degraded dataset and a real-world dataset recorded using microphones in diverse music venues.
【26】Evaluating Sound Similarity Metrics for Differentiable, Iterative Sound-Matching
链接:https://arxiv.org/abs/2506.22628
摘要:使用合成器的手动声音设计本质上是迭代的:艺术家将合成的输出与心理目标进行比较,调整参数,重复直到满意。迭代声音匹配通过在损失函数(或相似性度量)的指导下不断地对合成器进行编程来自动化该工作流程。先前对损失函数的比较通常倾向于一个度量而不是另一个度量,但仅限于狭窄的设置:有限的合成方法,很少的损失类型,通常没有盲听测试。这留下了一个开放的问题,是否存在普遍的最佳损失,或损失的选择仍然是一个创造性的决定条件的合成方法和声音设计师的偏好。我们提出可微迭代声音匹配作为现有文献的自然延伸,因为它结合了声音设计的手动方法和机器学习的现代进步。为了分析不同合成器之间损失函数性能的变化,我们实现了四种新颖的和已建立的可微分损失函数的混合,并将它们与可微分减法,加法和AM合成器配对。对于16种合成器损失组合中的每一种,我们进行了300次随机声音匹配试验。性能进行了测量,使用参数差异,频谱距离指标,并手动分配的听力分数。我们观察到的三个性能指标之间的一致性中等水平。我们的事后分析表明,损失函数的性能是高度依赖于合成器。这些发现强调了扩大声音匹配实验范围和开发针对特定合成技术的新相似性度量的价值,而不是追求一刀切的解决方案。
摘要:Manual sound design with a synthesizer is inherently iterative: an artist compares the synthesized output to a mental target, adjusts parameters, and repeats until satisfied. Iterative sound-matching automates this workflow by continually programming a synthesizer under the guidance of a loss function (or similarity measure) toward a target sound. Prior comparisons of loss functions have typically favored one metric over another, but only within narrow settings: limited synthesis methods, few loss types, often without blind listening tests. This leaves open the question of whether a universally optimal loss exists, or the choice of loss remains a creative decision conditioned on the synthesis method and the sound designer's preference. We propose differentiable iterative sound-matching as the natural extension of the available literature, since it combines the manual approach to sound design with modern advances in machine learning. To analyze the variability of loss function performance across synthesizers, we implemented a mix of four novel and established differentiable loss functions, and paired them with differentiable subtractive, additive, and AM synthesizers. For each of the sixteen synthesizer--loss combinations, we ran 300 randomized sound-matching trials. Performance was measured using parameter differences, spectrogram-distance metrics, and manually assigned listening scores. We observed a moderate level of consistency among the three performance measures. Our post-hoc analysis shows that the loss function performance is highly dependent on the synthesizer. These findings underscore the value of expanding the scope of sound-matching experiments and developing new similarity metrics tailored to specific synthesis techniques rather than pursuing one-size-fits-all solutions.
机器翻译由腾讯交互翻译提供,仅供参考
