微信公众号:arXiv_Daily
cs.SD语音
标题: TalkingMachines:通过自回归扩散模型的实时音频驱动FaceTime风格视频
链接:https://arxiv.org/abs/2506.03099
摘要:在本文中,我们提出了TalkingMachines -一个有效的框架,将预训练的视频生成模型转换为实时,音频驱动的角色动画。TalkingMachines通过将音频大语言模型(LLM)与我们的视频生成基础模型集成,实现自然的对话体验。我们的主要贡献包括:(1)我们将预训练的SOTA图像到视频DiT适配为具有180亿个参数的音频驱动的化身生成模型;(2)我们通过从双向教师模型到稀疏因果自回归学生模型的不对称知识蒸馏来实现无限视频流,而不会累积错误;(3)我们设计了一个高吞吐量,低延迟的推理管道,其中包含几个关键的工程优化,例如:(a)DiT和VAE解码器跨单独设备的分解,(b)使用CUDA流的设备间通信和计算的有效重叠,(c)消除冗余重新计算以最大化帧生成吞吐量。请在此查看演示视频-https://aaxwaz.github.io/TalkingMachines/
摘要:In this paper, we present TalkingMachines -- an efficient framework that transforms pretrained video generation models into real-time, audio-driven character animators. TalkingMachines enables natural conversational experiences by integrating an audio large language model (LLM) with our video generation foundation model. Our primary contributions include: (1) We adapt a pretrained SOTA image-to-video DiT into an audio-driven avatar generation model of 18 billion parameters; (2) We enable infinite video streaming without error accumulation through asymmetric knowledge distillation from a bidirectional teacher model into a sparse causal, autoregressive student model; (3) We design a high-throughput, low-latency inference pipeline incorporating several key engineering optimizations such as: (a) disaggregation of the DiT and VAE decoder across separate devices, (b) efficient overlap of inter-device communication and computation using CUDA streams, (c) elimination of redundant recomputations to maximize frame-generation throughput. Please see demo videos here - https://aaxwaz.github.io/TalkingMachines/
【2】 DGMO: Training-Free Audio Source Separation through Diffusion-Guided Mask Optimization
标题: DGMO:通过扩散引导屏蔽优化实现免训练音频源分离链接:https://arxiv.org/abs/2506.02858
备注:Interspeech 2025
摘要:语音查询音频源分离(LASS)通过自然语言查询实现开放词汇的声音分离。虽然现有方法依赖于特定于任务的训练,但我们探索最初为音频生成而设计的预训练扩散模型是否可以在无需进一步训练的情况下本质上执行分离。在这项研究中,我们引入了一个无训练的框架,利用生成先验的zero-shot LASS。为了解决这些问题,我们提出了扩散引导掩模优化(DGMO),这是一个测试时优化框架,可以细化频谱图掩模,以实现精确的输入对齐分离。我们的方法有效地将预训练的扩散模型重新用于源分离,在没有特定任务监督的情况下实现有竞争力的性能。这项工作扩展了扩散模型的应用,超越了生成,建立了一个新的范例zero-shot音频分离。该代码可从以下网址获得:https://wltschmrz.github.io/DGMO/
摘要:Language-queried Audio Source Separation (LASS) enables open-vocabulary sound separation via natural language queries. While existing methods rely on task-specific training, we explore whether pretrained diffusion models, originally designed for audio generation, can inherently perform separation without further training. In this study, we introduce a training-free framework leveraging generative priors for zero-shot LASS. Analyzing na\"ive adaptations, we identify key limitations arising from modality-specific challenges.To address these issues, we propose Diffusion-Guided Mask Optimization (DGMO), a test-time optimization framework that refines spectrogram masks for precise, input-aligned separation. Our approach effectively repurposes pretrained diffusion models for source separation, achieving competitive performance without task-specific supervision. This work expands the application of diffusion models beyond generation, establishing a new paradigm for zero-shot audio separation. The code is available at: https://wltschmrz.github.io/DGMO/
【3】 UltrasonicSpheres: Localized, Multi-Channel Sound Spheres Using Off-the-Shelf Speakers and Earables
标题: UltrasonicSpheres:使用现成扬声器和耳机的本地化多通道声音球体链接:https://arxiv.org/abs/2506.02715
摘要:我们提出了一个演示的UltrasonicSpheres,一个新的系统,用于特定位置的音频传输使用可穿戴耳机,解码成可听的声音超声波信号。与传统的波束形成设置不同,UltrasonicSpheres依赖于单个超声波扬声器来广播具有多个通道的本地化音频,每个通道都编码在不同的超声波载波频率上。佩戴我们的声学透明耳机的用户可以解调他们选择的流,例如以选定的语言展示叙述,同时保持对周围环境声音的完全感知。这种体验保留了空间音频感知,给人的印象是声音直接来自声源的物理位置。这可以实现个性化、本地化的音频,而无需配对、跟踪或额外的基础设施。重要的是,没有配备耳机的游客不受影响,因为超声波信号是人耳听不见的。我们的演示邀请参与者探索多个协同定位的音频区域,并体验UltrasonicSpheres如何支持在公共空间中不显眼地传递个性化声音。
摘要:We present a demo ofUltrasonicSpheres, a novel system for location-specific audio delivery using wearable earphones that decode ultrasonic signals into audible sound. Unlike conventional beamforming setups, UltrasonicSpheres relies on single ultrasonic speakers to broadcast localized audio with multiple channels, each encoded on a distinct ultrasonic carrier frequency. Users wearing our acoustically transparent earphones can demodulate their selected stream, such as exhibit narrations in a chosen language, while remaining fully aware of ambient environmental sounds. The experience preserves spatial audio perception, giving the impression that the sound originates directly from the physical location of the source. This enables personalized, localized audio without requiring pairing, tracking, or additional infrastructure. Importantly, visitors not equipped with the earphones are unaffected, as the ultrasonic signals are inaudible to the human ear. Our demo invites participants to explore multiple co-located audio zones and experience how UltrasonicSpheres supports unobtrusive delivery of personalized sound in public spaces.
【4】 MotionRAG-Diff: A Retrieval-Augmented Diffusion Framework for Long-Term Music-to-Dance Generation
标题: MotionRAG-Diff:长期音乐到舞蹈一代的检索增强传播框架链接:https://arxiv.org/abs/2506.02661
备注:12 pages, 5 figures
摘要:生成长期的,连贯的,逼真的音乐条件下的舞蹈序列仍然是一个具有挑战性的任务,在人体运动合成。现有的方法表现出严重的局限性:运动图方法依赖于固定的模板库,限制创造性的产生;扩散模型,而能够产生新的运动,往往缺乏时间的连贯性和音乐对齐。为了解决这些挑战,我们提出了$\textbf{MotionRAG-Diff}$,一个混合框架,集成检索增强生成(RAG)与基于扩散的细化,使高品质,音乐连贯的舞蹈生成任意长期的音乐输入。我们的方法引入了三个核心创新:(1)跨模态对比学习架构,在共享的潜在空间中对齐异构的音乐和舞蹈表示,在没有配对数据的情况下建立无监督的语义对应;(2)优化的运动图系统,用于高效检索和无缝拼接运动片段,确保长序列的真实感和时间一致性;(3)多条件扩散模型,其对原始音乐信号和对比特征进行联合条件化以增强运动质量和全局同步。大量的实验表明,MotionRAG-Diff在运动质量、多样性和音乐运动同步精度方面达到了最先进的性能。这项工作建立了一个新的范例,音乐驱动的舞蹈生成协同检索为基础的模板保真度与扩散为基础的创意增强。
摘要:Generating long-term, coherent, and realistic music-conditioned dance sequences remains a challenging task in human motion synthesis. Existing approaches exhibit critical limitations: motion graph methods rely on fixed template libraries, restricting creative generation; diffusion models, while capable of producing novel motions, often lack temporal coherence and musical alignment. To address these challenges, we propose $\textbf{MotionRAG-Diff}$, a hybrid framework that integrates Retrieval-Augmented Generation (RAG) with diffusion-based refinement to enable high-quality, musically coherent dance generation for arbitrary long-term music inputs. Our method introduces three core innovations: (1) A cross-modal contrastive learning architecture that aligns heterogeneous music and dance representations in a shared latent space, establishing unsupervised semantic correspondence without paired data; (2) An optimized motion graph system for efficient retrieval and seamless concatenation of motion segments, ensuring realism and temporal coherence across long sequences; (3) A multi-condition diffusion model that jointly conditions on raw music signals and contrastive features to enhance motion quality and global synchronization. Extensive experiments demonstrate that MotionRAG-Diff achieves state-of-the-art performance in motion quality, diversity, and music-motion synchronization accuracy. This work establishes a new paradigm for music-driven dance generation by synergizing retrieval-based template fidelity with diffusion-based creative enhancement.
【5】 Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
标题: 通过Whisper微调克服多方言阿拉伯语ASB中的数据稀缺性链接:https://arxiv.org/abs/2506.02627
备注:Accepted at Interspeech 2025
摘要:虽然商用阿拉伯语自动语音识别(ASR)系统支持现代标准阿拉伯语(MSA),但它们难以处理方言语音。我们使用Mozilla Common Voice for MSA和MASC dataset for dialectal speech研究了OpenAI的Whisper对五种主要阿拉伯方言(海湾,黎凡特,伊拉克,埃及,马格里布)的微调效果。我们评估了MSA训练规模的影响,MSA数据预训练的好处,以及方言特异性与方言合并模型。我们发现,少量的MSA微调数据产生实质性的改善较小的模型,匹配较大的非微调模型。虽然MSA预训练显示出最小的好处,这表明MSA和方言之间的共享功能有限,但我们的方言池模型对方言特定的模型进行了验证。这表明,在适当平衡的情况下,汇集方言数据可以帮助解决低资源ASR中的数据稀缺问题,而不会造成显著的性能损失。
摘要:Although commercial Arabic automatic speech recognition (ASR) systems support Modern Standard Arabic (MSA), they struggle with dialectal speech. We investigate the effect of fine-tuning OpenAI's Whisper on five major Arabic dialects (Gulf, Levantine, Iraqi, Egyptian, Maghrebi) using Mozilla Common Voice for MSA and the MASC dataset for dialectal speech. We evaluate MSA training size effects, benefits of pre-training on MSA data, and dialect-specific versus dialect-pooled models. We find that small amounts of MSA fine-tuning data yield substantial improvements for smaller models, matching larger non-fine-tuned models. While MSA pre-training shows minimal benefit, suggesting limited shared features between MSA and dialects, our dialect-pooled models perform comparably to dialect-specific ones. This indicates that pooling dialectal data, when properly balanced, can help address data scarcity in low-resource ASR without significant performance loss.
【6】 Cross-attention and Self-attention for Audio-visual Speaker Diarization in MISP-Meeting Challenge
标题: MISP会议挑战中视听演讲者对话的交叉注意和自我注意链接:https://arxiv.org/abs/2506.02621
摘要:本文介绍了为多模态基于信息的语音处理(MISP)2025挑战赛任务1开发的系统。我们介绍了CASA-Net,一种嵌入式融合方法设计的端到端的视听扬声器日记(AVSD)系统。CASA-Net采用了交叉注意(CA)模块,以有效地捕捉视听信号中的跨模态交互,并采用了自我注意(SA)模块来学习视听帧之间的上下文关系。为了进一步提高性能,我们采用了一种集成伪标签细化和再训练的训练策略,提高了时间戳预测的准确性。此外,中值滤波和重叠平均作为后处理技术应用于消除离群值和平滑预测标签。我们的系统在评估集上实现了8.18%的日志化错误率(DER),比基线DER 15.52%相对改善了47.3%。
摘要:This paper presents the system developed for Task 1 of the Multi-modal Information-based Speech Processing (MISP) 2025 Challenge. We introduce CASA-Net, an embedding fusion method designed for end-to-end audio-visual speaker diarization (AVSD) systems. CASA-Net incorporates a cross-attention (CA) module to effectively capture cross-modal interactions in audio-visual signals and employs a self-attention (SA) module to learn contextual relationships among audio-visual frames. To further enhance performance, we adopt a training strategy that integrates pseudo-label refinement and retraining, improving the accuracy of timestamp predictions. Additionally, median filtering and overlap averaging are applied as post-processing techniques to eliminate outliers and smooth prediction labels. Our system achieved a diarization error rate (DER) of 8.18% on the evaluation set, representing a relative improvement of 47.3% over the baseline DER of 15.52%.
【7】 Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm
标题: 基于图注意网络和标签传播算法的说话人重叠社区检测链接:https://arxiv.org/abs/2506.02610
摘要:在说话人日志化中,传统的基于聚类的方法仍然广泛应用于现实世界的应用中。然而,这些方法与说话人嵌入和重叠语音段的复杂分布作斗争。为了解决这些问题,我们提出了一种基于图注意力网络和标签传播算法(OCDGALP)的重叠社区检测方法。所提出的框架包括两个关键组件:(1)一个图注意力网络,通过聚合来自相邻节点的信息来细化说话人嵌入和节点连接,以及(2)一个标签传播算法,为每个节点分配多个社区标签,从而实现同时聚类和重叠社区检测。实验结果表明,该方法显着降低了日志错误率(DER),实现了最先进的15.94%的DER在DIHARD-III数据集上没有oracle语音活动检测(VAD),和一个令人印象深刻的11.07%与oracle VAD。
摘要:In speaker diarization, traditional clustering-based methods remain widely used in real-world applications. However, these methods struggle with the complex distribution of speaker embeddings and overlapping speech segments. To address these limitations, we propose an Overlapping Community Detection method based on Graph Attention networks and the Label Propagation Algorithm (OCDGALP). The proposed framework comprises two key components: (1) a graph attention network that refines speaker embeddings and node connections by aggregating information from neighboring nodes, and (2) a label propagation algorithm that assigns multiple community labels to each node, enabling simultaneous clustering and overlapping community detection. Experimental results show that the proposed method significantly reduces the Diarization Error Rate (DER), achieving a state-of-the-art 15.94% DER on the DIHARD-III dataset without oracle Voice Activity Detection (VAD), and an impressive 11.07% with oracle VAD.
【8】 Synthetic Speech Source Tracing using Metric Learning
标题: 使用度量学习的合成语音源跟踪链接:https://arxiv.org/abs/2506.02590
备注:Submitted to Interspeech 2025
摘要:本文讨论了在合成语音识别生成系统背后的操纵音频通过扬声器发音启发的管道源跟踪。虽然以前的工作集中在欺骗检测,源跟踪缺乏强大的解决方案。我们评估两种方法:基于分类和度量学习。我们使用ResNet和自监督学习(SSL)主干在MLAADv5基准上测试了我们的方法。结果表明,ResNet通过度量学习方法实现了具有竞争力的性能,匹配甚至超过了基于SSL的系统。我们的工作证明了ResNet在源跟踪方面的可行性,同时强调了优化SSL表示的必要性。我们的工作将说话人识别方法与音频取证挑战联系起来,为打击合成媒体操纵提供了新的方向。
摘要:This paper addresses source tracing in synthetic speech-identifying generative systems behind manipulated audio via speaker recognition-inspired pipelines. While prior work focuses on spoofing detection, source tracing lacks robust solutions. We evaluate two approaches: classification-based and metric-learning. We tested our methods on the MLAADv5 benchmark using ResNet and self-supervised learning (SSL) backbones. The results show that ResNet achieves competitive performance with the metric learning approach, matching and even exceeding SSL-based systems. Our work demonstrates ResNet's viability for source tracing while underscoring the need to optimize SSL representations for this task. Our work bridges speaker recognition methodologies with audio forensic challenges, offering new directions for combating synthetic media manipulation.
【9】 On the Language and Gender Biases in PSTN, VoIP and Neural Audio Codecs
标题: 关于公共电话网、语音电话和神经音频编解码器中的语言和性别偏见链接:https://arxiv.org/abs/2506.02545
备注:Submitted to INTERSPEECH2025
摘要:近年来,人们越来越关注语音技术中的公平性和包容性,特别是在自动语音识别和语音情感分析等领域。当音频在处理之前被转码时,如在流或实时应用中的情况,编码机制中的任何固有偏差都可能导致不一致。这不仅影响用户体验,而且还可能通过使陈规定型观念和排斥永久化而产生更广泛的社会影响。因此,音频编码机制的公正性非常重要。在这项工作中,我们对音频编解码器的语言和性别偏见方面的稀缺研究做出了贡献。通过分析超过200万个多语言音频文件的语音质量后,通过一个代表性的编解码器子集(PSTN,VoIP和神经)转码,我们的研究结果表明,PSTN编解码器在性别方面有很强的偏见,神经编解码器引入语言偏见。
摘要:In recent years, there has been a growing focus on fairness and inclusivity within speech technology, particularly in areas such as automatic speech recognition and speech sentiment analysis. When audio is transcoded prior to processing, as is the case in streaming or real-time applications, any inherent bias in the coding mechanism may result in disparities. This not only affects user experience but can also have broader societal implications by perpetuating stereotypes and exclusion. Thus, it is important that audio coding mechanisms are unbiased. In this work, we contribute towards the scarce research with respect to language and gender biases of audio codecs. By analyzing the speech quality of over 2 million multilingual audio files after transcoding through a representative subset of codecs (PSTN, VoIP and neural), our results indicate that PSTN codecs are strongly biased in terms of gender and that neural codecs introduce language biases.
【10】 DnR-nonverbal: Cinematic Audio Source Separation Dataset Containing Non-Verbal Sounds
标题: DnR-非言语:包含非言语声音的电影音频源分离数据集链接:https://arxiv.org/abs/2506.02499
备注:Accepted to Interspeech 2025, 5 pages, 3 figures, dataset is available at this https URL
摘要:我们提出了一个新的数据集的电影音频源分离(CASS),处理非语言的声音。现有的CASS数据集只包含阅读风格的声音作为语音干。这些数据集与实际的电影音频不同,实际的电影音频更可能包括表演出来的声音。因此,在传统数据集上训练的模型往往存在这样的问题,即情绪化的声音,如笑声和尖叫声,更容易被分离为效果,而不是语音。为了解决这个问题,我们建立了一个新的数据集,DnR-nonverbal。拟议的数据集包括语音干中的笑声和尖叫声等非语言声音。通过实验,我们揭示了当前CASS模型在非言语声音提取方面的问题,并表明我们的数据集可以有效地解决合成和实际电影音频中的问题。我们的数据集可在https://zenodo.org/records/15470640上获得。
摘要:We propose a new dataset for cinematic audio source separation (CASS) that handles non-verbal sounds. Existing CASS datasets only contain reading-style sounds as a speech stem. These datasets differ from actual movie audio, which is more likely to include acted-out voices. Consequently, models trained on conventional datasets tend to have issues where emotionally heightened voices, such as laughter and screams, are more easily separated as an effect, not speech. To address this problem, we build a new dataset, DnR-nonverbal. The proposed dataset includes non-verbal sounds like laughter and screams in the speech stem. From the experiments, we reveal the issue of non-verbal sound extraction by the current CASS model and show that our dataset can effectively address the issue in the synthetic and actual movie audio. Our dataset is available at https://zenodo.org/records/15470640.
【11】 SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
标题: SOVA-Bench:基于LLM的语音助理的语音对话能力基准测试链接:https://arxiv.org/abs/2506.02457
摘要:由于大型语言模型(LLM),语音编码算法和声码器结构的稳步发展,最近的进步已经能够直接从用户指令生成语音响应。然而,基准生成的语音质量一直是一个被忽视的,但关键的问题,考虑到从追求语义准确性生动和自发的语音流的转变。以往的评价侧重于语音理解能力,缺乏对音质的量化。在本文中,我们提出了语音概念语音助手基准(SOVA-Bench),提供了一般知识,语音识别和理解,以及可用的语音LLM之间的语义和声学生成能力的理解比较。据我们所知,SOVA-Bench是语音LLM最系统的评估框架之一,启发了语音交互系统的发展方向。
摘要:Thanks to the steady progress of large language models (LLMs), speech encoding algorithms and vocoder structure, recent advancements have enabled generating speech response directly from a user instruction. However, benchmarking the generated speech quality has been a neglected but critical issue, considering the shift from the pursuit of semantic accuracy to vivid and spontaneous speech flow. Previous evaluation focused on the speech-understanding ability, lacking a quantification of acoustic quality. In this paper, we propose Speech cOnversational Voice Assistant Benchmark (SOVA-Bench), providing a comprehension comparison of the general knowledge, speech recognition and understanding, along with both semantic and acoustic generative ability between available speech LLMs. To the best of our knowledge, SOVA-Bench is one of the most systematic evaluation frameworks for speech LLMs, inspiring the direction of voice interaction systems.
【12】 Breaking the Barriers of Text-Hungry and Audio-Deficient AI
标题: 打破文本饥饿和音频匮乏的人工智能的障碍链接:https://arxiv.org/abs/2506.02443
备注:61 pages, 16 figures, 14 tables, 25 languages, 13 blockaudio per language. Presented at AI Mali, May 2025
摘要:虽然全球语言多样性涵盖了超过7164种公认的语言,但目前占主导地位的机器智能架构仍然从根本上偏向于书面文本。这一偏见将7亿多人排除在外,特别是在农村和偏远地区,他们是听识字的。在这项工作中,我们介绍了一个完全无文本的,音频到音频的机器智能框架,旨在为这一服务不足的人群,以及所有喜欢音频效率的人。我们的贡献包括完全绕过文本的新颖音频到音频翻译架构,包括频谱图,尺度图,小波和基于单元的模型。我们的方法的核心是多尺度音频语义转换(MAST),表示编码音调,韵律,扬声器和表达功能。我们进一步将MAST集成到由分数布朗运动驱动的平均场型框架的分数扩散中。它能够在不依赖文本监督的情况下生成高保真、语义一致的语音。其结果是一个强大且可扩展的系统,能够直接从原始音频中学习,即使是不成文或很少数字化的语言。这项工作代表了向音频原生机器智能系统的根本转变,为历史上被排除在当前机器智能生态系统之外的社区扩大了对语言技术的访问。
摘要:While global linguistic diversity spans more than 7164 recognized languages, the current dominant architecture of machine intelligence remains fundamentally biased toward written text. This bias excludes over 700 million people particularly in rural and remote regions who are audio-literate. In this work, we introduce a fully textless, audio-to-audio machine intelligence framework designed to serve this underserved population, and all the people who prefer audio-efficiency. Our contributions include novel Audio-to-Audio translation architectures that bypass text entirely, including spectrogram-, scalogram-, wavelet-, and unit-based models. Central to our approach is the Multiscale Audio-Semantic Transform (MAST), a representation that encodes tonal, prosodic, speaker, and expressive features. We further integrate MAST into a fractional diffusion of mean-field-type framework powered by fractional Brownian motion. It enables the generation of high-fidelity, semantically consistent speech without reliance on textual supervision. The result is a robust and scalable system capable of learning directly from raw audio, even in languages that are unwritten or rarely digitized. This work represents a fundamental shift toward audio-native machine intelligence systems, expanding access to language technologies for communities historically left out of the current machine intelligence ecosystem.
【13】 StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
标题: StarVC:语音转换中联合生成文本和语音的统一自回归框架链接:https://arxiv.org/abs/2506.02414
备注:5 pages, 2 figures, Accepted by Interspeech 2025, Demo: this https URL
摘要:语音转换(VC)修改语音以匹配目标说话者,同时保留语言内容。传统的方法通常直接从语音中提取说话人信息,而忽略了对语言内容的显式利用。由于VC从根本上涉及将说话者身份从语言内容中分离出来,因此利用结构化语义特征可以提高转换性能。然而,以前的尝试,将语义特征到VC显示有限的有效性,激励显式文本建模的集成。我们提出了StarVC,一个统一的自回归VC框架,首先预测文本标记合成声学特征之前。实验表明,StarVC在保留语言内容(即,WER和CER)和扬声器特性(即,SECS和MOS)。音频演示可以在https://thuhcsi.github.io/StarVC/上找到。
摘要:Voice Conversion (VC) modifies speech to match a target speaker while preserving linguistic content. Traditional methods usually extract speaker information directly from speech while neglecting the explicit utilization of linguistic content. Since VC fundamentally involves disentangling speaker identity from linguistic content, leveraging structured semantic features could enhance conversion performance. However, previous attempts to incorporate semantic features into VC have shown limited effectiveness, motivating the integration of explicit text modeling. We propose StarVC, a unified autoregressive VC framework that first predicts text tokens before synthesizing acoustic features. The experiments demonstrate that StarVC outperforms conventional VC methods in preserving both linguistic content (i.e., WER and CER) and speaker characteristics (i.e., SECS and MOS). Audio demo can be found at: https://thuhcsi.github.io/StarVC/.
【14】 Trusted Fake Audio Detection Based on Dirichlet Distribution
标题: 基于Dirichlet分布的可信假音频检测链接:https://arxiv.org/abs/2506.02401
摘要:随着基于深度学习的语音转换和语音合成技术的不断发展,虚假音频带来的网络安全问题日益严重。先前提出的用于防御假音频的模型已经取得了显着的性能。然而,它们都未能对模型本身所做决策的可信度进行建模。在此基础上,提出了一种基于Dirichlet分布的虚假音频检测方法,旨在提高虚假音频检测的可靠性。具体来说,我们首先通过神经网络生成证据。然后使用狄利克雷分布对不确定性进行建模。通过用狄利克雷分布的参数对置信分布建模,可以获得每个决策的不确定性估计。最后,预测的概率和相应的不确定性估计相结合,形成最终意见。在ASVspoof系列数据集上(即,ASVspoof 2019 LA,ASVspoof 2021 LA和DF),我们进行了一些比较实验,以验证所提出的模型在准确性,鲁棒性和可信度方面的优异性能。
摘要:With the continuous development of deep learning-based speech conversion and speech synthesis technologies, the cybersecurity problem posed by fake audio has become increasingly serious. Previously proposed models for defending against fake audio have attained remarkable performance. However, they all fall short in modeling the trustworthiness of the decisions made by the models themselves. Based on this, we put forward a plausible fake audio detection approach based on the Dirichlet distribution with the aim of enhancing the reliability of fake audio detection. Specifically, we first generate evidence through a neural network. Uncertainty is then modeled using the Dirichlet distribution. By modeling the belief distribution with the parameters of the Dirichlet distribution, an estimate of uncertainty can be obtained for each decision. Finally, the predicted probabilities and corresponding uncertainty estimates are combined to form the final opinion. On the ASVspoof series dataset (i.e., ASVspoof 2019 LA, ASVspoof 2021 LA, and DF), we conduct a number of comparison experiments to verify the excellent performance of the proposed model in terms of accuracy, robustness, and trustworthiness.
【15】 Cocktail-Party Audio-Visual Speech Recognition
标题: 鸡尾酒会视听语音识别链接:https://arxiv.org/abs/2506.02178
备注:Accepted at Interspeech 2025
摘要:视听语音识别(AVSR)为具有挑战性的环境中的语音识别提供了一个强大的解决方案,例如鸡尾酒会场景,在这些场景中,仅仅依靠音频是不够的。然而,当前的AVSR模型通常针对具有持续活跃扬声器的理想化场景进行优化,忽略了包括说话和无声面部片段的真实世界设置的复杂性。本研究通过引入一种新的视听鸡尾酒会数据集来解决这一差距,该数据集旨在对当前的AVSR系统进行基准测试,并强调现有方法在现实噪声条件下的局限性。此外,我们贡献了一个1526小时的AVSR数据集,包括说话的脸和沉默的脸段,使鸡尾酒会环境中的显着性能增益。我们的方法相对于最先进的方法将WER降低了67%,在极端噪声中将WER从119%降低到39.2%,而不依赖于显式分割线索。
摘要:Audio-Visual Speech Recognition (AVSR) offers a robust solution for speech recognition in challenging environments, such as cocktail-party scenarios, where relying solely on audio proves insufficient. However, current AVSR models are often optimized for idealized scenarios with consistently active speakers, overlooking the complexities of real-world settings that include both speaking and silent facial segments. This study addresses this gap by introducing a novel audio-visual cocktail-party dataset designed to benchmark current AVSR systems and highlight the limitations of prior approaches in realistic noisy conditions. Additionally, we contribute a 1526-hour AVSR dataset comprising both talking-face and silent-face segments, enabling significant performance gains in cocktail-party environments. Our approach reduces WER by 67% relative to the state-of-the-art, reducing WER from 119% to 39.2% in extreme noise, without relying on explicit segmentation cues.
【16】 Comparison of spectrogram scaling in multi-label Music Genre Recognition
标题: 多标签音乐流派识别中谱图缩放的比较链接:https://arxiv.org/abs/2506.02091
备注:14 pages, 10 figures
摘要:随着数字音频工作站的可访问性和易用性的增加,普通听众可以获得的音乐数量也在增加;此外,流派之间的差异并不总是很好地定义,并且可能是抽象的,各个唱片中的流派组合差异很大。在这篇文章中,多种预处理方法和模型训练方法进行了描述和比较,占今天的专辑折衷的性质。一个自定义的,手动标记的数据集超过18000个条目已被用来执行实验。
摘要:As the accessibility and ease-of-use of digital audio workstations increases, so does the quantity of music available to the average listener; additionally, differences between genres are not always well defined and can be abstract, with widely varying combinations of genres across individual records. In this article, multiple preprocessing methods and approaches to model training are described and compared, accounting for the eclectic nature of today's albums. A custom, manually labeled dataset of more than 18000 entries has been used to perform the experiments.
【17】 Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025
标题: 利用基于图形的多模式融合和韵律特征增强语音情感识别,以应对Interspeech 2025年自然条件下的语音情感识别挑战赛链接:https://arxiv.org/abs/2506.02088
摘要:由于情感的微妙表达和现实世界音频的不可预测性,在自然、自发的语音中训练SER模型尤其具有挑战性。在本文中,我们提出了一个强大的系统,在自然条件下的INTERSPEECH 2025语音情感识别的挑战,专注于分类情感识别。我们的方法结合了最先进的音频模型与文本特征丰富的韵律和频谱线索。特别是,我们研究了基频(F0)量化的有效性和使用预训练的音频标记模型。我们还采用了集成模型,以提高鲁棒性。在官方测试集上,我们的系统获得了39.79%的宏观F1分数(验证时为42.20%)。我们的研究结果强调了这些方法的潜力,融合技术的分析证实了图注意力网络的有效性。我们的源代码是公开的。
摘要:Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the INTERSPEECH 2025 Speech Emotion Recognition in Naturalistic Conditions Challenge, focusing on categorical emotion recognition. Our method combines state-of-the-art audio models with text features enriched by prosodic and spectral cues. In particular, we investigate the effectiveness of Fundamental Frequency (F0) quantization and the use of a pretrained audio tagging model. We also employ an ensemble model to improve robustness. On the official test set, our system achieved a Macro F1-score of 39.79% (42.20% on validation). Our results underscore the potential of these methods, and analysis of fusion techniques confirmed the effectiveness of Graph Attention Networks. Our source code is publicly available.
【18】 Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
标题: 揭开音频Deepfake起源:采用Ensemble Fusion的深度指标学习和Conformer网络方法链接:https://arxiv.org/abs/2506.02085
备注:Accepted at Interspeech 2025, Netherlands
摘要:音频deepfake正在通过先进的人工智能获得前所未有的真实感。虽然目前的研究重点是从欺骗性语音中识别真实语音,但追踪源系统同样至关重要。本文提出了一种新的音频源跟踪系统,该系统将深度度量多类N对损失与实强调和假分散框架、一致性分类网络和集成分数嵌入融合相结合。N对损失提高了区分能力,而真实强调和虚假分散通过专注于区分真实和虚假语音模式来增强鲁棒性。Conformer网络捕获音频信号中的全局和局部依赖性,这对于源跟踪至关重要。所提出的集成分数嵌入融合显示了域内和域外源跟踪场景之间的最佳权衡。我们使用弗雷歇距离和标准度量评估我们的方法,在源跟踪中表现出优于基线系统的性能。
摘要:Audio deepfakes are acquiring an unprecedented level of realism with advanced AI. While current research focuses on discerning real speech from spoofed speech, tracing the source system is equally crucial. This work proposes a novel audio source tracing system combining deep metric multi-class N-pair loss with Real Emphasis and Fake Dispersion framework, a Conformer classification network, and ensemble score-embedding fusion. The N-pair loss improves discriminative ability, while Real Emphasis and Fake Dispersion enhance robustness by focusing on differentiating real and fake speech patterns. The Conformer network captures both global and local dependencies in the audio signal, crucial for source tracing. The proposed ensemble score-embedding fusion shows an optimal trade-off between in-domain and out-of-domain source tracing scenarios. We evaluate our method using Frechet Distance and standard metrics, demonstrating superior performance in source tracing over the baseline system.
【19】 LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
标题: LASPA:基于前缀调谐交叉注意的语言无关说话人解纠缠链接:https://arxiv.org/abs/2506.02083
备注:Accepted at Interspeech 2025, Netherlands
摘要:说话人识别模型在多语言环境中面临着挑战,因为说话人嵌入中语言信息的纠缠。语音特征(如口音、语音解剖学和语言的语音结构)之间的重叠使分离语言和说话者信息变得复杂。解开这些组件可以显着提高说话人识别的准确性。为此,我们提出了一种新型的解纠缠学习策略,该策略通过前缀调整的交叉注意来集成联合学习。这种方法在说话者在不同语言之间切换时特别有效。实验结果表明,该模型适用于单语和多语言环境,包括看不见的语言。值得注意的是,该模型提高了多个数据集的相等错误率,突出了其将语言信息与说话人嵌入分离并增强不同语言条件下识别的能力。
摘要:Speaker recognition models face challenges in multi-lingual settings due to the entanglement of linguistic information within speaker embeddings. The overlap between vocal traits such as accent, vocal anatomy, and a language's phonetic structure complicates separating linguistic and speaker information. Disentangling these components can significantly improve speaker recognition accuracy. To this end, we propose a novel disentanglement learning strategy that integrates joint learning through prefix-tuned cross-attention. This approach is particularly effective when speakers switch between languages. Experimental results show the model generalizes across monolingual and multi-lingual settings, including unseen languages. Notably, the proposed model improves the equal error rate across multiple datasets, highlighting its ability to separate language information from speaker embeddings and enhance recognition in diverse linguistic conditions.
【20】 SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction
标题: SALF-MOS:扬声器不可知的潜在特征进行下采样以进行MOS预测链接:https://arxiv.org/abs/2506.02082
备注:None
摘要:语音质量评估是选择文本到语音合成(TTS)或语音转换模型的关键过程。语音合成的评估可以使用客观度量或主观度量来完成。虽然有许多客观的指标,如语音质量的感知评估(PESQ),感知客观听力质量评估(POLQA)或短时客观可懂度(STOI),但没有一个是可行的选择最佳模型。另一方面,像平均意见得分这样的主观指标是高度可靠的,但它需要大量的人工努力并且耗时。为了解决MOS评估中的问题,我们开发了一种新的模型,Speaker Agnostic Latent Features(SALF)-Mean Opinion Score(MOS),这是一种小型,端到端,高度通用和可扩展的模型,用于预测MOS评分,范围为5。我们使用卷积序列并将其堆叠以获得音频样本的潜在特征,以获得基于均方误差(MSE)、线性一致性相关系数(LCC)、斯皮尔曼秩相关系数(SRCC)和肯德尔秩相关系数(KTAU)的最佳状态的最先进的结果。
摘要:Speech quality assessment is a critical process in selecting text-to-speech synthesis (TTS) or voice conversion models. Evaluation of voice synthesis can be done using objective metrics or subjective metrics. Although there are many objective metrics like the Perceptual Evaluation of Speech Quality (PESQ), Perceptual Objective Listening Quality Assessment (POLQA) or Short-Time Objective Intelligibility (STOI) but none of them is feasible in selecting the best model. On the other hand subjective metric like Mean Opinion Score is highly reliable but it requires a lot of manual efforts and are time-consuming. To counter the issues in MOS Evaluation, we have developed a novel model, Speaker Agnostic Latent Features (SALF)-Mean Opinion Score (MOS) which is a small-sized, end-to-end, highly generalized and scalable model for predicting MOS score on a scale of 5. We use the sequences of convolutions and stack them to get the latent features of the audio samples to get the best state-of-the-art results based on mean squared error (MSE), Linear Concordance Correlation coefficient (LCC), Spearman Rank Correlation Coefficient (SRCC) and Kendall Rank Correlation Coefficient (KTAU).
【21】 Learning More with Less: Self-Supervised Approaches for Low-Resource Speech Emotion Recognition
标题: 少花钱多学:低资源语音情感识别的自我监督方法链接:https://arxiv.org/abs/2506.02059
备注:Accepted at Interspeech 2025
摘要:语音情感识别(SER)在深度学习方面取得了重大进展,但由于注释数据的稀缺性,低资源语言(LRL)仍然面临挑战。在这项工作中,我们探索无监督学习,以提高低资源环境中的SER。具体来说,我们研究了对比学习(CL)和Bootstrap Your Own Latent(BYOL)作为增强跨语言泛化的自我监督方法。我们的方法在乌尔都语中实现了10.6%的显著F1分数改善,在德语中为15.2%,在孟加拉语中为13.9%,证明了它们在LRL中的有效性。此外,我们分析模型行为,以提供影响跨语言性能的关键因素的见解,并强调在低资源SER的挑战。这项工作提供了一个基础,为代表性不足的语言开发更具包容性,可解释性和强大的情感识别系统。
摘要:Speech Emotion Recognition (SER) has seen significant progress with deep learning, yet remains challenging for Low-Resource Languages (LRLs) due to the scarcity of annotated data. In this work, we explore unsupervised learning to improve SER in low-resource settings. Specifically, we investigate contrastive learning (CL) and Bootstrap Your Own Latent (BYOL) as self-supervised approaches to enhance cross-lingual generalization. Our methods achieve notable F1 score improvements of 10.6% in Urdu, 15.2% in German, and 13.9% in Bangla, demonstrating their effectiveness in LRLs. Additionally, we analyze model behavior to provide insights on key factors influencing performance across languages, and also highlighting challenges in low-resource SER. This work provides a foundation for developing more inclusive, explainable, and robust emotion recognition systems for underrepresented languages.
【22】 Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
标题: 在视觉语音识别中利用大型语言模型:模型缩放、上下文感知解码和迭代抛光链接:https://arxiv.org/abs/2506.02012
摘要:视觉语音识别(VSR)通过分析嘴唇运动来转录语音。最近,大型语言模型(LLM)已被集成到VSR系统,导致显着的性能改善。然而,LLM的潜力还没有得到广泛的研究,如何有效地利用LLM在VSR任务仍然没有探索。本文系统地探讨了如何更好地利用LLM进行VSR任务,并提供了三个关键贡献:(1)缩放测试:我们研究了LLM大小如何影响VSR性能,确认了VSR任务中的缩放律。(2)上下文感知解码:我们添加上下文文本来指导LLM解码,提高识别精度。(3)迭代抛光:我们建议迭代地改进LLM输出,逐步减少识别错误。大量的实验表明,通过这些设计,LLM的巨大潜力可以很大程度上利用,导致显着的VSR性能改善。
摘要:Visual Speech Recognition (VSR) transcribes speech by analyzing lip movements. Recently, Large Language Models (LLMs) have been integrated into VSR systems, leading to notable performance improvements. However, the potential of LLMs has not been extensively studied, and how to effectively utilize LLMs in VSR tasks remains unexplored. This paper systematically explores how to better leverage LLMs for VSR tasks and provides three key contributions: (1) Scaling Test: We study how the LLM size affects VSR performance, confirming a scaling law in the VSR task. (2) Context-Aware Decoding: We add contextual text to guide the LLM decoding, improving recognition accuracy. (3) Iterative Polishing: We propose iteratively refining LLM outputs, progressively reducing recognition errors. Extensive experiments demonstrate that by these designs, the great potential of LLMs can be largely harnessed, leading to significant VSR performance improvement.
【23】 CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge
标题: CNNSRC 2024:第二届中文连续视觉语音识别挑战赛链接:https://arxiv.org/abs/2506.02010
备注:to be published in INTERSPEECH 2025
摘要:本文介绍了第二届汉语连续视觉语音识别挑战赛(CNVSRC 2024),该挑战赛以CNVSRC 2023为基础,旨在推进汉语大词汇量连续视觉语音识别(LVC-VSR)的研究。该挑战评估了两个测试场景:在录音室和互联网语音阅读。CNVSRC 2024使用与其前身CNVSRC 2023相同的数据集,其中包括用于培训的CN-CVS和用于开发和评估的CNVSRC-单/多。然而,CNVSRC 2024引入了两项关键改进:(1)更强大的基线系统,以及(2)额外的数据集CN-CVS 2-P1,用于开放跟踪,以提高数据量和多样性。新的挑战展示了数据预处理,特征提取,模型设计和训练策略方面的几项重要创新,进一步推动了中国LVC-VSR的最新发展。更多信息和资源可在官方网站上找到。
摘要:This paper presents the second Chinese Continuous Visual Speech Recognition Challenge (CNVSRC 2024), which builds on CNVSRC 2023 to advance research in Chinese Large Vocabulary Continuous Visual Speech Recognition (LVC-VSR). The challenge evaluates two test scenarios: reading in recording studios and Internet speech. CNVSRC 2024 uses the same datasets as its predecessor CNVSRC 2023, which involves CN-CVS for training and CNVSRC-Single/Multi for development and evaluation. However, CNVSRC 2024 introduced two key improvements: (1) a stronger baseline system, and (2) an additional dataset, CN-CVS2-P1, for open tracks to improve data volume and diversity. The new challenge has demonstrated several important innovations in data preprocessing, feature extraction, model design, and training strategies, further pushing the state-of-the-art in Chinese LVC-VSR. More details and resources are available at the official website.
【24】 Inter(sectional) Alia(s): Ambiguity in Voice Agent Identity via Intersectional Japanese Self-Referents
标题: 跨(部分)别名:通过交叉日本自我指涉的语音代理身份的模糊性链接:https://arxiv.org/abs/2506.01998
备注:CHI '25
摘要:模仿人类的会话代理人提出了关于将具有人类社会身份线索的机器拟人化的伦理问题。批评者还质疑类人代理的身份中立性假设。最近的研究表明,交叉日语代词可以引起复杂的,有时回避的印象代理身份。然而,其他“中性”的非代词自我指涉(NPSR)和语音作为一种社会表达媒介的作用仍然没有被探索。在一项众包研究中,日本参与者(N = 204)使用七个自我参照来评估三种ChatGPT声音(Juniper,Breeze和Ember)。我们发现了声音性别化的强有力证据,以及交叉自我指涉逃避性别化的潜力,即,通过中立性和难以捉摸的模糊性。值得注意的是,根据社会语言学理论,特别是boku和watakushi,年龄和正式的看法与性别有关。这项工作对代理人身份认知进行了细致入微的研究,并支持语音代理人的交叉和文化敏感工作。
摘要:Conversational agents that mimic people have raised questions about the ethics of anthropomorphizing machines with human social identity cues. Critics have also questioned assumptions of identity neutrality in humanlike agents. Recent work has revealed that intersectional Japanese pronouns can elicit complex and sometimes evasive impressions of agent identity. Yet, the role of other "neutral" non-pronominal self-referents (NPSR) and voice as a socially expressive medium remains unexplored. In a crowdsourcing study, Japanese participants (N = 204) evaluated three ChatGPT voices (Juniper, Breeze, and Ember) using seven self-referents. We found strong evidence of voice gendering alongside the potential of intersectional self-referents to evade gendering, i.e., ambiguity through neutrality and elusiveness. Notably, perceptions of age and formality intersected with gendering as per sociolinguistic theories, especially boku and watakushi. This work provides a nuanced take on agent identity perceptions and champions intersectional and culturally-sensitive work on voice agents.
【25】 PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
标题: PartialEdit:识别神经语音编辑时代的部分Deepfakes链接:https://arxiv.org/abs/2506.02958
备注:Interspeech 2025 camera ready. Project page: this https URL
摘要:神经语音编辑能够对语音话语进行无缝的部分编辑,允许修改选定的内容,同时保持音频的其余部分不变。然而,这种有用的技术也带来了deepfake的新风险。为了鼓励对检测这种部分编辑的deepfake语音的研究,我们引入了PartialEdit,这是一个使用高级神经编辑技术管理的deepfake语音数据集。我们将在PartialEdit上探索检测和定位任务。我们的实验表明,在现有PartialSpoof数据集上训练的模型无法检测到神经语音编辑模型生成的部分编辑语音。由于最近的语音编辑模型几乎都涉及神经音频编解码器,因此我们还提供了对模型在检测这些deepfake时所学到的伪影的见解。有关PartialEdit数据集和音频样本的更多信息,请访问项目页面:https://yzyouzhang.com/PartialEdit/index.html。
摘要:Neural speech editing enables seamless partial edits to speech utterances, allowing modifications to selected content while preserving the rest of the audio unchanged. This useful technique, however, also poses new risks of deepfakes. To encourage research on detecting such partially edited deepfake speech, we introduce PartialEdit, a deepfake speech dataset curated using advanced neural editing techniques. We explore both detection and localization tasks on PartialEdit. Our experiments reveal that models trained on the existing PartialSpoof dataset fail to detect partially edited speech generated by neural speech editing models. As recent speech editing models almost all involve neural audio codecs, we also provide insights into the artifacts the model learned on detecting these deepfakes. Further information about the PartialEdit dataset and audio samples can be found on the project page: https://yzyouzhang.com/PartialEdit/index.html.
【26】 CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
标题: CapSpeech:以样式标题文本转语音方式启用下游应用程序链接:https://arxiv.org/abs/2506.02863
摘要:生成式人工智能的最新进展显著改变了风格字幕文本到语音合成(CapTTS)领域。然而,由于缺乏标准化的,全面的数据集和有限的研究建立在CapTTS的下游任务,使CapTTS适应现实世界的应用仍然具有挑战性。为了解决这些差距,我们介绍CapSpeech,一个新的基准设计了一系列的CapTTS相关的任务,包括风格字幕的文本到语音合成与声音事件(CapTTS-SE),口音字幕的TTS(AccCapTTS),情感字幕的TTS(ACCAPTTTS),和文本到语音合成聊天代理(AgentTTS)。CapSpeech包含超过1000万个机器注释的音频字幕对和近36万个人工注释的音频字幕对。此外,我们还介绍了由专业配音演员和经验丰富的音频工程师收集和录制的两个新数据集,专门用于AgentTTS和CapTTS-SE任务。除了数据集之外,我们还使用CapSpeech上的自回归和非自回归模型进行了全面的实验。我们的研究结果表明,高保真度和高度可理解的语音合成在各种各样的说话风格。据我们所知,CapSpeech是最大的可用数据集,为CapTTS相关任务提供全面的注释。这些实验和发现进一步为开发CapTTS系统的挑战提供了有价值的见解。
摘要:Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to real-world applications remains challenging due to the lack of standardized, comprehensive datasets and limited research on downstream tasks built upon CapTTS. To address these gaps, we introduce CapSpeech, a new benchmark designed for a series of CapTTS-related tasks, including style-captioned text-to-speech synthesis with sound events (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS), and text-to-speech synthesis for chat agent (AgentTTS). CapSpeech comprises over 10 million machine-annotated audio-caption pairs and nearly 0.36 million human-annotated audio-caption pairs. In addition, we introduce two new datasets collected and recorded by a professional voice actor and experienced audio engineers, specifically for the AgentTTS and CapTTS-SE tasks. Alongside the datasets, we conduct comprehensive experiments using both autoregressive and non-autoregressive models on CapSpeech. Our results demonstrate high-fidelity and highly intelligible speech synthesis across a diverse range of speaking styles. To the best of our knowledge, CapSpeech is the largest available dataset offering comprehensive annotations for CapTTS-related tasks. The experiments and findings further provide valuable insights into the challenges of developing CapTTS systems.
【27】 Fast-Converging Distributed Signal Estimation in Topology-Unconstrained Wireless Acoustic Sensor Networks
标题: 不受布局约束的无线声学传感器网络中的快速收敛分布式信号估计链接:https://arxiv.org/abs/2506.02797
摘要:本文主要研究拓扑无约束的无线声传感器网络(WASNs)中的分布式信号估计问题,其中传感器节点只传输本地传感器信号的融合版本。对于这个任务,拓扑无关(TI)的分布式自适应节点特定的信号估计(DANSE)算法(TI-DANSE)先前已被提出。它收敛于非完全连接和时变网络拓扑结构中的集中式信号估计解决方案。然而,TI-DANSE在现实世界的场景中的适用性是有限的,由于其缓慢的收敛。后者是由于在TI-DANSE中,节点只能访问WANSE中所有融合信号的网络内总和。我们通过引入改进的TI-DANSE算法(称为TI-DANSE+)来解决这种低收敛速度问题,其中更新节点分别使用来自每个邻居的融合信号的部分网络内和。节点可以在其局部优化问题中最大化可用自由度的数量,从而加快收敛速度。通过将TI-DANSE+与最大化更新节点处邻居数量的树修剪策略相结合,进一步利用了这一点。在完全连接的WASNs中,TI-DANSE+的收敛速度与原始DANSE算法(后者仅为完全连接的WASNs定义)一样快,同时使用点对点数据传输而不是广播,从而节省通信带宽。如果发生链路故障,TI-DANSE+向集中式解决方案的收敛将保持不变,而不会对其公式进行任何更改。总而言之,所提出的TI-DANSE+算法可以被视为DANSE和TI-DANSE的全面替代方案,其(i)合并了两者的优点,(ii)将它们的差异调和成单个公式,并且(iii)在通信带宽使用方面显示出其自身的优势。
摘要:This paper focuses on distributed signal estimation in topology-unconstrained wireless acoustic sensor networks (WASNs) where sensor nodes only transmit fused versions of their local sensor signals. For this task, the topology-independent (TI) distributed adaptive node-specific signal estimation (DANSE) algorithm (TI-DANSE) has previously been proposed. It converges towards the centralized signal estimation solution in non-fully connected and time-varying network topologies. However, the applicability of TI-DANSE in real-world scenarios is limited due to its slow convergence. The latter results from the fact that, in TI-DANSE, nodes only have access to the in-network sum of all fused signals in the WASN. We address this low convergence speed by introducing an improved TI-DANSE algorithm, referred to as TI-DANSE+, in which updating nodes separately use the partial in-network sums of fused signals coming from each of their neighbors. Nodes can maximize the number of available degrees of freedom in their local optimization problem, leading to faster convergence. This is further exploited by combining TI-DANSE+ with a tree-pruning strategy that maximizes the number of neighbors at the updating node. In fully connected WASNs, TI-DANSE+ converges as fast as the original DANSE algorithm (the latter only defined for fully connected WASNs) while using peer-to-peer data transmission instead of broadcasting and thus saving communication bandwidth. If link failures occur, the convergence of TI-DANSE+ towards the centralized solution is preserved without any change in its formulation. Altogether, the proposed TI-DANSE+ algorithm can be viewed as an all-round alternative to DANSE and TI-DANSE which (i) merges the advantages of both, (ii) reconciliates their differences into a single formulation, and (iii) shows advantages of its own in terms of communication bandwidth usage.
【28】 On the influence of language similarity in non-target speaker verification trials
标题: 语言相似性对非目标语者确认实验的影响链接:https://arxiv.org/abs/2506.02777
备注:accepted to Interspeech 2025
摘要:在本文中,我们使用最先进的说话人验证系统ECAPA-TDNN(在VoxCeleb数据集的多语言和单语言变体上训练)研究了语言相似性在跨语言非目标说话人验证试验中的影响。我们对多语言Globalphone和LDC CTS的分数分布模式的分析揭示了涉及训练语言的扬声器比较中的聚类效应,其中比较语言的选择对分数的影响最小。相反,我们观察到的语言相似性的影响,涉及语言不包括在说话人验证系统的训练集的试验,分数与语言相似性测量的语言分类系统,特别是当使用多语言训练数据。
摘要:In this paper, we investigate the influence of language similarity in cross-lingual non-target speaker verification trials using a state-of-the-art speaker verification system, ECAPA-TDNN, trained on multilingual and monolingual variants of the VoxCeleb dataset. Our analysis of the score distribution patterns on multilingual Globalphone and LDC CTS reveals a clustering effect in speaker comparisons involving a training language, whereby the choice of comparison language only minimally impacts scores. Conversely, we observe a language similarity effect in trials involving languages not included in the training set of the speaker verification system, with scores correlating with language similarity measured by a language classification system, especially when using multilingual training data.
【29】 AuralNet: Hierarchical Attention-based 3D Binaural Localization of Overlapping Speakers
标题: AuralNet:基于层次注意的重叠说话人3D双耳定位链接:https://arxiv.org/abs/2506.02773
备注:Accepted and to appear at Interspeech 2025
摘要:我们提出了AuralNet,一种新的3D多源双耳声源定位方法,在没有源的数量的先验知识的情况下,在方位角和仰角上定位重叠的源。AuralNet采用门控的粗到细架构,将粗分类阶段与细粒度回归阶段相结合,通过扇区划分实现灵活的空间分辨率。该模型采用了多头自注意机制,以捕捉双耳信号中的空间线索,增强了在嘈杂混响环境中的鲁棒性。设计了一个掩蔽的多任务损失函数,以联合优化声音检测,方位角和仰角估计。在噪声混响条件下的大量实验证明了AuralNet优于最近的方法
摘要:We propose AuralNet, a novel 3D multi-source binaural sound source localization approach that localizes overlapping sources in both azimuth and elevation without prior knowledge of the number of sources. AuralNet employs a gated coarse-tofine architecture, combining a coarse classification stage with a fine-grained regression stage, allowing for flexible spatial resolution through sector partitioning. The model incorporates a multi-head self-attention mechanism to capture spatial cues in binaural signals, enhancing robustness in noisy-reverberant environments. A masked multi-task loss function is designed to jointly optimize sound detection, azimuth, and elevation estimation. Extensive experiments in noisy-reverberant conditions demonstrate the superiority of AuralNet over recent methods
【30】 Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
标题: 预算-隐形-情感:利用预算-LLM上下文知识进行Zero-Shot表达性语音合成,用于混合情感链接:https://arxiv.org/abs/2506.02742
摘要:现有的表达性文本到语音(TTS)系统主要对有限的分类情感进行建模,而人类对话远远超出了这些预定义的情感,因此必须探索更多样化的情感语音生成以实现更自然的交互。为了弥补这一差距,本文提出了一种新的不可见情感(PUE)的方法,通过情感引导的提示学习生成不可见的情感语音。PUE利用LLM-TTS架构进行训练,以确保分类情感相关提示和情感语音之间的情感一致性,使模型能够定量捕获每个话语的不同情感权重。在推理过程中,可以通过灵活调整情感比例和利用LLM上下文知识来生成混合情感语音,使模型能够量化不同的情感风格。我们提出的PUE成功地促进了在zero-shot设置中看不见的情感的表达性语音合成。
摘要:Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse emotional speech generation for more natural interactions. To bridge this gap, this paper proposes a novel prompt-unseen-emotion (PUE) approach to generate unseen emotional speech via emotion-guided prompt learning. PUE is trained utilizing an LLM-TTS architecture to ensure emotional consistency between categorical emotion-relevant prompts and emotional speech, allowing the model to quantitatively capture different emotion weightings per utterance. During inference, mixed emotional speech can be generated by flexibly adjusting emotion proportions and leveraging LLM contextual knowledge, enabling the model to quantify different emotional styles. Our proposed PUE successfully facilitates expressive speech synthesis of unseen emotions in a zero-shot setting.
【31】 Adaptive Differential Denoising for Respiratory Sounds Classification
标题: 呼吸音分类的自适应差异去噪链接:https://arxiv.org/abs/2506.02505
备注:accepted at Interspeech2025
摘要:自动呼吸声分类面临来自背景噪声和现有系统中去噪不足的实际挑战。 我们提出了自适应差分去噪网络,它通过三个创新集成了噪声抑制和病理特征保留: 1)自适应频率滤波器,具有可学习的频谱掩模和软收缩,以消除噪声,同时保留诊断高频分量; 2)一个差分去噪层,使用差分注意力,通过增强的样本比较来减少噪声引起的变化; 3)一个偏见去噪损失联合优化分类和鲁棒性没有干净的标签。 在ICBHI 2017数据集上的实验表明,该方法获得了65.53%的Score,比之前的sota方法提高了1.99%. 该代码可在https://github.com/deegy666/ADD-RSC上获得
摘要:Automated respiratory sound classification faces practical challenges from background noise and insufficient denoising in existing systems. We propose Adaptive Differential Denoising network, that integrates noise suppression and pathological feature preservation via three innovations: 1) Adaptive Frequency Filter with learnable spectral masks and soft shrink to eliminate noise while retaining diagnostic high-frequency components; 2) A Differential Denoise Layer using differential attention to reduce noise-induced variations through augmented sample comparisons; 3) A bias denoising loss jointly optimizing classification and robustness without clean labels. Experiments on the ICBHI2017 dataset show that our method achieves 65.53\% of the Score, which is improved by 1.99\% over the previous sota method. The code is available in https://github.com/deegy666/ADD-RSC
【32】 Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss
标题: 增强一致性损失的音乐混合物的歌词转录链接:https://arxiv.org/abs/2506.02339
备注:submitted to Interspeech
摘要:自动歌词转录(ALT)旨在从歌声中识别歌词,类似于口语的自动语音识别(ASR),但由于歌声的特定领域属性而面临更大的复杂性。虽然基础ASR模型在各种语音任务中表现出鲁棒性,但它们的性能在唱歌时会下降,特别是在有音乐伴奏的情况下。这项工作的重点是这个性能差距,并探讨低秩适应(LoRA)ALT,调查单域和双域微调策略。我们建议使用一致性损失来更好地对齐声乐和混合编码器表示,在不依赖于歌声分离的情况下改善混合的转录。我们的研究结果表明,虽然天真的双域微调表现不佳,但具有一致性损失的结构化训练会产生适度但一致的收益,这表明了将ASR基础模型应用于音乐的潜力。
摘要:Automatic Lyrics Transcription (ALT) aims to recognize lyrics from singing voices, similar to Automatic Speech Recognition (ASR) for spoken language, but faces added complexity due to domain-specific properties of the singing voice. While foundation ASR models show robustness in various speech tasks, their performance degrades on singing voice, especially in the presence of musical accompaniment. This work focuses on this performance gap and explores Low-Rank Adaptation (LoRA) for ALT, investigating both single-domain and dual-domain fine-tuning strategies. We propose using a consistency loss to better align vocal and mixture encoder representations, improving transcription on mixture without relying on singing voice separation. Our results show that while na\"ive dual-domain fine-tuning underperforms, structured training with consistency loss yields modest but consistent gains, demonstrating the potential of adapting ASR foundation models for music.
【33】 Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?
标题: 基于Mamba的音频基金会模型最适合非语言情感识别吗?链接:https://arxiv.org/abs/2506.02258
备注:Accepted to EUSIPCO 2025
摘要:在这项工作中,我们专注于非语言的声音情感识别(NVER)。我们调查基于曼巴的音频基础模型(MAFM)的第一次为NVER和假设MAFM将优于基于注意力的音频基础模型(AAFMs)的NVER利用其状态空间建模,以更有效地捕捉内在的情感结构。与AAFMs不同,由于其注意力机制,AAFMs可能会放大不相关的模式,MAFM将提取更稳定和上下文感知的表征,从而更好地区分微妙的非语言情感线索。我们的实验与国家的最先进的(SOTA)AAFM和MAFM验证了我们的假设。此外,从相关的研究,如语音情感识别,合成语音检测,其中融合的基础模型(FM)表现出更好的性能,我们还探讨融合的FM的NVER。为此,我们建议,RENO,使用雷尼发散作为一种新的损失函数的FM的有效对齐。它还利用自我注意力,更好地代表内的相互作用的FM。与RENO,通过MAFM和AAFMs的异构融合,我们显示了最高的性能相比,个别FM,其融合,并设置SOTA相比,以前的SOTA工作。
摘要:In this work, we focus on non-verbal vocal sounds emotion recognition (NVER). We investigate mamba-based audio foundation models (MAFMs) for the first time for NVER and hypothesize that MAFMs will outperform attention-based audio foundation models (AAFMs) for NVER by leveraging its state-space modeling to capture intrinsic emotional structures more effectively. Unlike AAFMs, which may amplify irrelevant patterns due to their attention mechanisms, MAFMs will extract more stable and context-aware representations, enabling better differentiation of subtle non-verbal emotional cues. Our experiments with state-of-the-art (SOTA) AAFMs and MAFMs validates our hypothesis. Further, motivated from related research such as speech emotion recognition, synthetic speech detection, where fusion of foundation models (FMs) have showed improved performance, we also explore fusion of FMs for NVER. To this end, we propose, RENO, that uses renyi-divergence as a novel loss function for effective alignment of the FMs. It also makes use of self-attention for better intra-representation interaction of the FMs. With RENO, through the heterogeneous fusion of MAFMs and AAFMs, we show the topmost performance in comparison to individual FMs, its fusion and also setting SOTA in comparison to previous SOTA work.
【34】 Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction
标题: 调查说话者预训练模型的合理有效性及其对SingMOS预测的协同能力链接:https://arxiv.org/abs/2506.02232
备注:Accepted to INTERSPEECH 2025
摘要:在这项研究中,我们专注于歌唱声音平均意见得分(SingMOS)预测。以前的研究已经表明使用最先进的(SOTA)预训练模型(PTM)的性能优势。然而,他们还没有探索说话人识别语音PTM(SPTM),如x-vector,ECAPA,我们假设它将是最有效的SingMOS预测。我们认为,由于他们的说话人识别预训练,它使他们能够捕捉细粒度的声音特征(例如,音高,音调,强度)从合成的歌声在更好的方式比其他PTM。我们用SOTA PTM(包括SPTM和音乐PTM)进行的实验验证了这一假设。此外,我们介绍了一种新的融合框架,批,使用巴塔查亚距离融合的PTM。通过与说话人识别SPTM融合的BATCH,我们报告了与所有单个PTM和基线融合技术以及设置SOTA的最高性能比较。
摘要:In this study, we focus on Singing Voice Mean Opinion Score (SingMOS) prediction. Previous research have shown the performance benefit with the use of state-of-the-art (SOTA) pre-trained models (PTMs). However, they haven't explored speaker recognition speech PTMs (SPTMs) such as x-vector, ECAPA and we hypothesize that it will be the most effective for SingMOS prediction. We believe that due to their speaker recognition pre-training, it equips them to capture fine-grained vocal features (e.g., pitch, tone, intensity) from synthesized singing voices in a much more better way than other PTMs. Our experiments with SOTA PTMs including SPTMs and music PTMs validates the hypothesis. Additionally, we introduce a novel fusion framework, BATCH that uses Bhattacharya Distance for fusion of PTMs. Through BATCH with the fusion of speaker recognition SPTMs, we report the topmost performance comparison to all the individual PTMs and baseline fusion techniques as well as setting SOTA.
【35】 Towards Machine Unlearning for Paralinguistic Speech Processing
标题: 迈向副语言语音处理的机器去学习链接:https://arxiv.org/abs/2506.02230
备注:Accepted to INTERSPEECH 2025
摘要:在这项工作中,我们开创了机器非学习(MU)的副语言语音处理(PSP)的研究。我们专注于两个关键的PSP任务:语音情感识别(SER)和抑郁症检测(DD)。为此,我们提出了SISA++,这是对先前最先进的(SOTA)MU方法SISA的一种新的扩展,通过合并在不同分片上训练的模型并进行加权平均。通过这样的修改,我们表明,SISA++在基准SER(CREMA-D)和DD(E-DAIC)数据集的学习后比SISA更能保持性能。此外,为了指导未来的研究,更容易采用MU的PSP,我们提出了“食谱食谱”-可操作的建议,选择最佳的功能表示和下游架构,可以减轻性能下降后的学习过程。
摘要:In this work, we pioneer the study of Machine Unlearning (MU) for Paralinguistic Speech Processing (PSP). We focus on two key PSP tasks: Speech Emotion Recognition (SER) and Depression Detection (DD). To this end, we propose, SISA++, a novel extension to previous state-of-the-art (SOTA) MU method, SISA by merging models trained on different shards with weight-averaging. With such modifications, we show that SISA++ preserves performance more in comparison to SISA after unlearning in benchmark SER (CREMA-D) and DD (E-DAIC) datasets. Also, to guide future research for easier adoption of MU for PSP, we present ``cookbook recipes'' - actionable recommendations for selecting optimal feature representations and downstream architectures that can mitigate performance degradation after the unlearning process.
【36】 No Audiogram: Leveraging Existing Scores for Personalized Speech Intelligibility Prediction
标题: No Audiogram:利用现有分数进行个性化语音清晰度预测链接:https://arxiv.org/abs/2506.02039
备注:Accepted at Interspeech 2025
摘要:个性化的语音清晰度预测是一个挑战。以前的方法主要依赖于听力图,这是固有的准确性有限,因为它们只捕捉听众的纯音听力阈值。我们提出了一种新的方法,利用个人现有的可懂度数据来预测他们在新音频上的表现,而不是结合额外的听众功能。我们引入了基于支持样本的可理解性预测网络(SSIPNet),这是一种深度学习模型,它利用语音基础模型从多个支持(音频,分数)对中构建收听者语音识别能力的高维表示,从而实现对不可见音频的准确预测。Clarity Prediction Challenge数据集上的结果表明,即使有少量的支持(音频,分数)对,我们的方法也优于基于音频的预测。我们的工作提出了一个新的模式,个性化的语音可懂度预测。
摘要:Personalized speech intelligibility prediction is challenging. Previous approaches have mainly relied on audiograms, which are inherently limited in accuracy as they only capture a listener's hearing threshold for pure tones. Rather than incorporating additional listener features, we propose a novel approach that leverages an individual's existing intelligibility data to predict their performance on new audio. We introduce the Support Sample-Based Intelligibility Prediction Network (SSIPNet), a deep learning model that leverages speech foundation models to build a high-dimensional representation of a listener's speech recognition ability from multiple support (audio, score) pairs, enabling accurate predictions for unseen audio. Results on the Clarity Prediction Challenge dataset show that, even with a small number of support (audio, score) pairs, our method outperforms audiogram-based predictions. Our work presents a new paradigm for personalized speech intelligibility prediction.
【37】 Singing Voice Graph Modeling for SingFake Detection
标题: 歌唱声图建模及其假唱检测链接:https://arxiv.org/abs/2406.03111
备注:Accepted by Interspeech 2024; Our code is available at this https URL
摘要:检测歌声深度伪造(SingFake)涉及确定歌声的真实性和版权。现有的语音deepfake检测模型一直在努力适应人类发声这一独特的歌唱声音领域中的不可见攻击。为了弥合这一差距,我们提出了一个突破性的SingGraph模型。该模型协同MERT声学音乐理解模型的音高和节奏分析与wav2vec2.0模型的歌词语言分析的能力。此外,我们提倡使用基于音乐领域知识的RawBoost和节拍匹配技术来增强歌声,从而提高SingFake检测性能。我们提出的方法在SingFake数据集内实现了新的最先进(SOTA)结果,在三种不同的场景中超过了以前的SOTA模型:它将看到的歌手的EER相对提高了13.2%,未看到的歌手提高了24.3%,使用不同编解码器的未看到的歌手提高了37.1%。
摘要:Detecting singing voice deepfakes, or SingFake, involves determining the authenticity and copyright of a singing voice. Existing models for speech deepfake detection have struggled to adapt to unseen attacks in this unique singing voice domain of human vocalization. To bridge the gap, we present a groundbreaking SingGraph model. The model synergizes the capabilities of the MERT acoustic music understanding model for pitch and rhythm analysis with the wav2vec2.0 model for linguistic analysis of lyrics. Additionally, we advocate for using RawBoost and beat matching techniques grounded in music domain knowledge for singing voice augmentation, thereby enhancing SingFake detection performance. Our proposed method achieves new state-of-the-art (SOTA) results within the SingFake dataset, surpassing the previous SOTA model across three distinct scenarios: it improves EER relatively for seen singers by 13.2%, for unseen singers by 24.3%, and unseen singers using different codecs by 37.1%.
【1】 InfiniteAudio: Infinite-Length Audio Generation with Consistency
标题: InfiniteAudio:具有一致性的无限长度音频生成链接:https://arxiv.org/abs/2506.03020
摘要:本文介绍了InfiniteAudio,一个简单而有效的策略,用于使用基于扩散的文本到音频方法生成无限长度的音频。当前的方法面临内存限制,因为输出大小随输入长度增加,使得长持续时间生成具有挑战性。一种常见的解决方法是连接短的音频片段,但由于缺乏共享的时间上下文,这通常会导致不一致。为了解决这个问题,InfiniteAudio无缝集成到现有的管道中,无需额外的培训。它介绍了两个关键技术:FIFO采样,具有固定大小输入的先进先出推理策略,以及弯曲去噪,它选择性地优先考虑关键扩散步骤以提高效率。实验表明,InfiniteAudio在所有指标上都达到了相当或更高的性能。音频样本可以在我们的项目页面上找到。
摘要:This paper presents InfiniteAudio, a simple yet effective strategy for generating infinite-length audio using diffusion-based text-to-audio methods. Current approaches face memory constraints because the output size increases with input length, making long duration generation challenging. A common workaround is to concatenate short audio segments, but this often leads to inconsistencies due to the lack of shared temporal context. To address this, InfiniteAudio integrates seamlessly into existing pipelines without additional training. It introduces two key techniques: FIFO sampling, a first-in, first-out inference strategy with fixed-size inputs, and curved denoising, which selectively prioritizes key diffusion steps for efficiency. Experiments show that InfiniteAudio achieves comparable or superior performance across all metrics. Audio samples are available on our project page.
【2】 PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
标题: PartialEdit:识别神经语音编辑时代的部分Deepfakes链接:https://arxiv.org/abs/2506.02958
备注:Interspeech 2025 camera ready. Project page: this https URL
摘要:神经语音编辑能够对语音话语进行无缝的部分编辑,允许修改选定的内容,同时保持音频的其余部分不变。然而,这种有用的技术也带来了deepfake的新风险。为了鼓励对检测这种部分编辑的deepfake语音的研究,我们引入了PartialEdit,这是一个使用高级神经编辑技术管理的deepfake语音数据集。我们将在PartialEdit上探索检测和定位任务。我们的实验表明,在现有PartialSpoof数据集上训练的模型无法检测到神经语音编辑模型生成的部分编辑语音。由于最近的语音编辑模型几乎都涉及神经音频编解码器,因此我们还提供了对模型在检测这些deepfake时所学到的伪影的见解。有关PartialEdit数据集和音频样本的更多信息,请访问项目页面:https://yzyouzhang.com/PartialEdit/index.html。
摘要:Neural speech editing enables seamless partial edits to speech utterances, allowing modifications to selected content while preserving the rest of the audio unchanged. This useful technique, however, also poses new risks of deepfakes. To encourage research on detecting such partially edited deepfake speech, we introduce PartialEdit, a deepfake speech dataset curated using advanced neural editing techniques. We explore both detection and localization tasks on PartialEdit. Our experiments reveal that models trained on the existing PartialSpoof dataset fail to detect partially edited speech generated by neural speech editing models. As recent speech editing models almost all involve neural audio codecs, we also provide insights into the artifacts the model learned on detecting these deepfakes. Further information about the PartialEdit dataset and audio samples can be found on the project page: https://yzyouzhang.com/PartialEdit/index.html.
【3】 Diffusion Buffer: Online Diffusion-based Speech Enhancement with Sub-Second Latency
标题: 扩散缓冲区:具有亚秒延迟的在线基于扩散的语音增强链接:https://arxiv.org/abs/2506.02908
备注:5 pages, 2 figures, Accepted to Interspeech 2025
摘要:扩散模型是一类最近已被用于语音增强的生成模型,取得了显着的成功,但在推理时间的计算昂贵。因此,这些模型对于实时处理流数据是不切实际的。在这项工作中,我们适应滑动窗口扩散框架的语音增强任务。我们的方法随着时间的推移逐渐破坏语音信号,将更多的噪声分配给缓冲区中接近当前的帧。这种方法输出去噪帧,延迟与所选缓冲区大小成比例,从而实现性能和延迟之间的权衡。实证结果表明,我们的方法优于标准的扩散模型,并在GPU上有效地运行,实现了0.3到1秒的顺序的输入输出延迟。这标志着第一个实用的基于扩散的在线语音增强解决方案。
摘要:Diffusion models are a class of generative models that have been recently used for speech enhancement with remarkable success but are computationally expensive at inference time. Therefore, these models are impractical for processing streaming data in real-time. In this work, we adapt a sliding window diffusion framework to the speech enhancement task. Our approach progressively corrupts speech signals through time, assigning more noise to frames close to the present in a buffer. This approach outputs denoised frames with a delay proportional to the chosen buffer size, enabling a trade-off between performance and latency. Empirical results demonstrate that our method outperforms standard diffusion models and runs efficiently on a GPU, achieving an input-output latency in the order of 0.3 to 1 seconds. This marks the first practical diffusion-based solution for online speech enhancement.
【4】 CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
标题: CapSpeech:以样式标题文本转语音方式启用下游应用程序链接:https://arxiv.org/abs/2506.02863
摘要:生成式人工智能的最新进展显著改变了风格字幕文本到语音合成(CapTTS)领域。然而,由于缺乏标准化的,全面的数据集和有限的研究建立在CapTTS的下游任务,使CapTTS适应现实世界的应用仍然具有挑战性。为了解决这些差距,我们介绍CapSpeech,一个新的基准设计了一系列的CapTTS相关的任务,包括风格字幕的文本到语音合成与声音事件(CapTTS-SE),口音字幕的TTS(AccCapTTS),情感字幕的TTS(ACCAPTTTS),和文本到语音合成聊天代理(AgentTTS)。CapSpeech包含超过1000万个机器注释的音频字幕对和近36万个人工注释的音频字幕对。此外,我们还介绍了由专业配音演员和经验丰富的音频工程师收集和录制的两个新数据集,专门用于AgentTTS和CapTTS-SE任务。除了数据集之外,我们还使用CapSpeech上的自回归和非自回归模型进行了全面的实验。我们的研究结果表明,高保真度和高度可理解的语音合成在各种各样的说话风格。据我们所知,CapSpeech是最大的可用数据集,为CapTTS相关任务提供全面的注释。这些实验和发现进一步为开发CapTTS系统的挑战提供了有价值的见解。
摘要:Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to real-world applications remains challenging due to the lack of standardized, comprehensive datasets and limited research on downstream tasks built upon CapTTS. To address these gaps, we introduce CapSpeech, a new benchmark designed for a series of CapTTS-related tasks, including style-captioned text-to-speech synthesis with sound events (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS), and text-to-speech synthesis for chat agent (AgentTTS). CapSpeech comprises over 10 million machine-annotated audio-caption pairs and nearly 0.36 million human-annotated audio-caption pairs. In addition, we introduce two new datasets collected and recorded by a professional voice actor and experienced audio engineers, specifically for the AgentTTS and CapTTS-SE tasks. Alongside the datasets, we conduct comprehensive experiments using both autoregressive and non-autoregressive models on CapSpeech. Our results demonstrate high-fidelity and highly intelligible speech synthesis across a diverse range of speaking styles. To the best of our knowledge, CapSpeech is the largest available dataset offering comprehensive annotations for CapTTS-related tasks. The experiments and findings further provide valuable insights into the challenges of developing CapTTS systems.
【5】 Fast-Converging Distributed Signal Estimation in Topology-Unconstrained Wireless Acoustic Sensor Networks
标题: 不受布局约束的无线声学传感器网络中的快速收敛分布式信号估计链接:https://arxiv.org/abs/2506.02797
摘要:本文主要研究拓扑无约束的无线声传感器网络(WASNs)中的分布式信号估计问题,其中传感器节点只传输本地传感器信号的融合版本。对于这个任务,拓扑无关(TI)的分布式自适应节点特定的信号估计(DANSE)算法(TI-DANSE)先前已被提出。它收敛于非完全连接和时变网络拓扑结构中的集中式信号估计解决方案。然而,TI-DANSE在现实世界的场景中的适用性是有限的,由于其缓慢的收敛。后者是由于在TI-DANSE中,节点只能访问WANSE中所有融合信号的网络内总和。我们通过引入改进的TI-DANSE算法(称为TI-DANSE+)来解决这种低收敛速度问题,其中更新节点分别使用来自每个邻居的融合信号的部分网络内和。节点可以在其局部优化问题中最大化可用自由度的数量,从而加快收敛速度。通过将TI-DANSE+与最大化更新节点处邻居数量的树修剪策略相结合,进一步利用了这一点。在完全连接的WASNs中,TI-DANSE+的收敛速度与原始DANSE算法(后者仅为完全连接的WASNs定义)一样快,同时使用点对点数据传输而不是广播,从而节省通信带宽。如果发生链路故障,TI-DANSE+向集中式解决方案的收敛将保持不变,而不会对其公式进行任何更改。总而言之,所提出的TI-DANSE+算法可以被视为DANSE和TI-DANSE的全面替代方案,其(i)合并了两者的优点,(ii)将它们的差异调和成单个公式,并且(iii)在通信带宽使用方面显示出其自身的优势。
摘要:This paper focuses on distributed signal estimation in topology-unconstrained wireless acoustic sensor networks (WASNs) where sensor nodes only transmit fused versions of their local sensor signals. For this task, the topology-independent (TI) distributed adaptive node-specific signal estimation (DANSE) algorithm (TI-DANSE) has previously been proposed. It converges towards the centralized signal estimation solution in non-fully connected and time-varying network topologies. However, the applicability of TI-DANSE in real-world scenarios is limited due to its slow convergence. The latter results from the fact that, in TI-DANSE, nodes only have access to the in-network sum of all fused signals in the WASN. We address this low convergence speed by introducing an improved TI-DANSE algorithm, referred to as TI-DANSE+, in which updating nodes separately use the partial in-network sums of fused signals coming from each of their neighbors. Nodes can maximize the number of available degrees of freedom in their local optimization problem, leading to faster convergence. This is further exploited by combining TI-DANSE+ with a tree-pruning strategy that maximizes the number of neighbors at the updating node. In fully connected WASNs, TI-DANSE+ converges as fast as the original DANSE algorithm (the latter only defined for fully connected WASNs) while using peer-to-peer data transmission instead of broadcasting and thus saving communication bandwidth. If link failures occur, the convergence of TI-DANSE+ towards the centralized solution is preserved without any change in its formulation. Altogether, the proposed TI-DANSE+ algorithm can be viewed as an all-round alternative to DANSE and TI-DANSE which (i) merges the advantages of both, (ii) reconciliates their differences into a single formulation, and (iii) shows advantages of its own in terms of communication bandwidth usage.
【6】 On the influence of language similarity in non-target speaker verification trials
标题: 语言相似性对非目标语者确认实验的影响链接:https://arxiv.org/abs/2506.02777
备注:accepted to Interspeech 2025
摘要:在本文中,我们使用最先进的说话人验证系统ECAPA-TDNN(在VoxCeleb数据集的多语言和单语言变体上训练)研究了语言相似性在跨语言非目标说话人验证试验中的影响。我们对多语言Globalphone和LDC CTS的分数分布模式的分析揭示了涉及训练语言的扬声器比较中的聚类效应,其中比较语言的选择对分数的影响最小。相反,我们观察到的语言相似性的影响,涉及语言不包括在说话人验证系统的训练集的试验,分数与语言相似性测量的语言分类系统,特别是当使用多语言训练数据。
摘要:In this paper, we investigate the influence of language similarity in cross-lingual non-target speaker verification trials using a state-of-the-art speaker verification system, ECAPA-TDNN, trained on multilingual and monolingual variants of the VoxCeleb dataset. Our analysis of the score distribution patterns on multilingual Globalphone and LDC CTS reveals a clustering effect in speaker comparisons involving a training language, whereby the choice of comparison language only minimally impacts scores. Conversely, we observe a language similarity effect in trials involving languages not included in the training set of the speaker verification system, with scores correlating with language similarity measured by a language classification system, especially when using multilingual training data.
【7】 AuralNet: Hierarchical Attention-based 3D Binaural Localization of Overlapping Speakers
标题: AuralNet:基于层次注意的重叠说话人3D双耳定位链接:https://arxiv.org/abs/2506.02773
备注:Accepted and to appear at Interspeech 2025
摘要:我们提出了AuralNet,一种新的3D多源双耳声源定位方法,在没有源的数量的先验知识的情况下,在方位角和仰角上定位重叠的源。AuralNet采用门控的粗到细架构,将粗分类阶段与细粒度回归阶段相结合,通过扇区划分实现灵活的空间分辨率。该模型采用了多头自注意机制,以捕捉双耳信号中的空间线索,增强了在嘈杂混响环境中的鲁棒性。设计了一个掩蔽的多任务损失函数,以联合优化声音检测,方位角和仰角估计。在噪声混响条件下的大量实验证明了AuralNet优于最近的方法
摘要:We propose AuralNet, a novel 3D multi-source binaural sound source localization approach that localizes overlapping sources in both azimuth and elevation without prior knowledge of the number of sources. AuralNet employs a gated coarse-tofine architecture, combining a coarse classification stage with a fine-grained regression stage, allowing for flexible spatial resolution through sector partitioning. The model incorporates a multi-head self-attention mechanism to capture spatial cues in binaural signals, enhancing robustness in noisy-reverberant environments. A masked multi-task loss function is designed to jointly optimize sound detection, azimuth, and elevation estimation. Extensive experiments in noisy-reverberant conditions demonstrate the superiority of AuralNet over recent methods
【8】 Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
标题: 预算-隐形-情感:利用预算-LLM上下文知识进行Zero-Shot表达性语音合成,用于混合情感链接:https://arxiv.org/abs/2506.02742
摘要:现有的表达性文本到语音(TTS)系统主要对有限的分类情感进行建模,而人类对话远远超出了这些预定义的情感,因此必须探索更多样化的情感语音生成以实现更自然的交互。为了弥补这一差距,本文提出了一种新的不可见情感(PUE)的方法,通过情感引导的提示学习生成不可见的情感语音。PUE利用LLM-TTS架构进行训练,以确保分类情感相关提示和情感语音之间的情感一致性,从而使模型能够定量捕获每个话语的不同情感权重。在推理过程中,可以通过灵活调整情感比例和利用LLM上下文知识来生成混合情感语音,使模型能够量化不同的情感风格。我们提出的PUE成功地促进了在zero-shot设置中看不见的情感的表达性语音合成。
摘要:Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse emotional speech generation for more natural interactions. To bridge this gap, this paper proposes a novel prompt-unseen-emotion (PUE) approach to generate unseen emotional speech via emotion-guided prompt learning. PUE is trained utilizing an LLM-TTS architecture to ensure emotional consistency between categorical emotion-relevant prompts and emotional speech, allowing the model to quantitatively capture different emotion weightings per utterance. During inference, mixed emotional speech can be generated by flexibly adjusting emotion proportions and leveraging LLM contextual knowledge, enabling the model to quantify different emotional styles. Our proposed PUE successfully facilitates expressive speech synthesis of unseen emotions in a zero-shot setting.
【9】 Adaptive Differential Denoising for Respiratory Sounds Classification
标题: 呼吸音分类的自适应差异去噪链接:https://arxiv.org/abs/2506.02505
备注:accepted at Interspeech2025
摘要:自动呼吸声分类面临来自背景噪声和现有系统中去噪不足的实际挑战。 我们提出了自适应差分去噪网络,它通过三项创新集成了噪声抑制和病理特征保留: 1)自适应频率滤波器,具有可学习的频谱掩模和软收缩,以消除噪声,同时保留诊断高频分量; 2)一个差分去噪层,使用差分注意力,通过增强的样本比较来减少噪声引起的变化; 3)一个偏见去噪损失联合优化分类和鲁棒性没有干净的标签。 在ICBHI 2017数据集上的实验表明,该方法获得了65.53%的Score,比之前的sota方法提高了1.99%. 该代码可在https://github.com/deegy666/ADD-RSC上获得
摘要:Automated respiratory sound classification faces practical challenges from background noise and insufficient denoising in existing systems. We propose Adaptive Differential Denoising network, that integrates noise suppression and pathological feature preservation via three innovations: 1) Adaptive Frequency Filter with learnable spectral masks and soft shrink to eliminate noise while retaining diagnostic high-frequency components; 2) A Differential Denoise Layer using differential attention to reduce noise-induced variations through augmented sample comparisons; 3) A bias denoising loss jointly optimizing classification and robustness without clean labels. Experiments on the ICBHI2017 dataset show that our method achieves 65.53\% of the Score, which is improved by 1.99\% over the previous sota method. The code is available in https://github.com/deegy666/ADD-RSC
【10】 Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss
标题: 增强一致性损失的音乐混合物的歌词转录链接:https://arxiv.org/abs/2506.02339
备注:submitted to Interspeech
摘要:自动歌词转录(ALT)旨在从歌声中识别歌词,类似于口语的自动语音识别(ASR),但由于歌声的特定领域属性而面临更大的复杂性。虽然基础ASR模型在各种语音任务中表现出鲁棒性,但它们的性能在唱歌时会下降,特别是在有音乐伴奏的情况下。这项工作的重点是这个性能差距,并探讨低秩适应(LoRA)ALT,调查单域和双域微调策略。我们建议使用一致性损失来更好地对齐声乐和混合编码器表示,在不依赖于歌声分离的情况下改善混合的转录。我们的研究结果表明,虽然天真的双域微调表现不佳,但具有一致性损失的结构化训练会产生适度但一致的收益,这表明了将ASR基础模型应用于音乐的潜力。
摘要:Automatic Lyrics Transcription (ALT) aims to recognize lyrics from singing voices, similar to Automatic Speech Recognition (ASR) for spoken language, but faces added complexity due to domain-specific properties of the singing voice. While foundation ASR models show robustness in various speech tasks, their performance degrades on singing voice, especially in the presence of musical accompaniment. This work focuses on this performance gap and explores Low-Rank Adaptation (LoRA) for ALT, investigating both single-domain and dual-domain fine-tuning strategies. We propose using a consistency loss to better align vocal and mixture encoder representations, improving transcription on mixture without relying on singing voice separation. Our results show that while na\"ive dual-domain fine-tuning underperforms, structured training with consistency loss yields modest but consistent gains, demonstrating the potential of adapting ASR foundation models for music.
【11】 Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?
标题: 基于Mamba的音频基金会模型最适合非语言情感识别吗?链接:https://arxiv.org/abs/2506.02258
备注:Accepted to EUSIPCO 2025
摘要:在这项工作中,我们专注于非语言的声音情感识别(NVER)。我们调查基于曼巴的音频基础模型(MAFM)的第一次为NVER和假设MAFM将优于基于注意力的音频基础模型(AAFMs)的NVER利用其状态空间建模,以更有效地捕捉内在的情感结构。与AAFMs不同,由于其注意力机制,AAFMs可能会放大不相关的模式,MAFM将提取更稳定和上下文感知的表征,从而更好地区分微妙的非语言情感线索。我们的实验与国家的最先进的(SOTA)AAFM和MAFM验证了我们的假设。此外,从相关的研究,如语音情感识别,合成语音检测,其中融合的基础模型(FM)表现出更好的性能,我们还探讨融合的FM的NVER。为此,我们建议,RENO,使用雷尼发散作为一种新的损失函数的FM的有效对齐。它还利用自我注意力,更好地代表内的相互作用的FM。与RENO,通过MAFM和AAFMs的异构融合,我们显示了最高的性能相比,个别FM,其融合,并设置SOTA相比,以前的SOTA工作。
摘要:In this work, we focus on non-verbal vocal sounds emotion recognition (NVER). We investigate mamba-based audio foundation models (MAFMs) for the first time for NVER and hypothesize that MAFMs will outperform attention-based audio foundation models (AAFMs) for NVER by leveraging its state-space modeling to capture intrinsic emotional structures more effectively. Unlike AAFMs, which may amplify irrelevant patterns due to their attention mechanisms, MAFMs will extract more stable and context-aware representations, enabling better differentiation of subtle non-verbal emotional cues. Our experiments with state-of-the-art (SOTA) AAFMs and MAFMs validates our hypothesis. Further, motivated from related research such as speech emotion recognition, synthetic speech detection, where fusion of foundation models (FMs) have showed improved performance, we also explore fusion of FMs for NVER. To this end, we propose, RENO, that uses renyi-divergence as a novel loss function for effective alignment of the FMs. It also makes use of self-attention for better intra-representation interaction of the FMs. With RENO, through the heterogeneous fusion of MAFMs and AAFMs, we show the topmost performance in comparison to individual FMs, its fusion and also setting SOTA in comparison to previous SOTA work.
【12】 Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction
标题: 调查说话者预训练模型的合理有效性及其对SingMOS预测的协同能力链接:https://arxiv.org/abs/2506.02232
备注:Accepted to INTERSPEECH 2025
摘要:在这项研究中,我们专注于歌唱声音平均意见得分(SingMOS)预测。以前的研究已经表明使用最先进的(SOTA)预训练模型(PTM)的性能优势。然而,他们还没有探索说话人识别语音PTM(SPTM),如x-vector,ECAPA,我们假设它将是最有效的SingMOS预测。我们认为,由于他们的说话人识别预训练,它使他们能够捕捉细粒度的声音特征(例如,音高,音调,强度)从合成的歌声在更好的方式比其他PTM。我们用SOTA PTM(包括SPTM和音乐PTM)进行的实验验证了这一假设。此外,我们介绍了一种新的融合框架,批,使用巴塔查亚距离融合的PTM。通过与说话人识别SPTM融合的BATCH,我们报告了与所有单个PTM和基线融合技术以及设置SOTA的最高性能比较。
摘要:In this study, we focus on Singing Voice Mean Opinion Score (SingMOS) prediction. Previous research have shown the performance benefit with the use of state-of-the-art (SOTA) pre-trained models (PTMs). However, they haven't explored speaker recognition speech PTMs (SPTMs) such as x-vector, ECAPA and we hypothesize that it will be the most effective for SingMOS prediction. We believe that due to their speaker recognition pre-training, it equips them to capture fine-grained vocal features (e.g., pitch, tone, intensity) from synthesized singing voices in a much more better way than other PTMs. Our experiments with SOTA PTMs including SPTMs and music PTMs validates the hypothesis. Additionally, we introduce a novel fusion framework, BATCH that uses Bhattacharya Distance for fusion of PTMs. Through BATCH with the fusion of speaker recognition SPTMs, we report the topmost performance comparison to all the individual PTMs and baseline fusion techniques as well as setting SOTA.
【13】 Towards Machine Unlearning for Paralinguistic Speech Processing
标题: 迈向副语言语音处理的机器去学习链接:https://arxiv.org/abs/2506.02230
备注:Accepted to INTERSPEECH 2025
摘要:在这项工作中,我们开创了机器非学习(MU)的副语言语音处理(PSP)的研究。我们专注于两个关键的PSP任务:语音情感识别(SER)和抑郁症检测(DD)。为此,我们提出了SISA++,这是对先前最先进的(SOTA)MU方法SISA的一种新的扩展,通过合并在不同分片上训练的模型并进行加权平均。通过这样的修改,我们表明,SISA++在基准SER(CREMA-D)和DD(E-DAIC)数据集的学习后比SISA更能保持性能。此外,为了指导未来的研究,更容易采用MU的PSP,我们提出了“食谱食谱”-可操作的建议,选择最佳的功能表示和下游架构,可以减轻性能下降后的学习过程。
摘要:In this work, we pioneer the study of Machine Unlearning (MU) for Paralinguistic Speech Processing (PSP). We focus on two key PSP tasks: Speech Emotion Recognition (SER) and Depression Detection (DD). To this end, we propose, SISA++, a novel extension to previous state-of-the-art (SOTA) MU method, SISA by merging models trained on different shards with weight-averaging. With such modifications, we show that SISA++ preserves performance more in comparison to SISA after unlearning in benchmark SER (CREMA-D) and DD (E-DAIC) datasets. Also, to guide future research for easier adoption of MU for PSP, we present ``cookbook recipes'' - actionable recommendations for selecting optimal feature representations and downstream architectures that can mitigate performance degradation after the unlearning process.
【14】 Dhvani: A Weakly-supervised Phonemic Error Detection and Personalized Feedback System for Hindi
标题: Dhvani:一个弱监督的印地语音素错误检测和个性化反馈系统链接:https://arxiv.org/abs/2506.02166
备注:Accepted for publication at Interspeech 2025 to be held in Rotterdam, the Netherlands
摘要:计算机辅助发音培训(CAPT)已被广泛研究用于英语。然而,在将其应用于拥有15亿人口的印度语言方面仍然存在重大差距。为印度语言量身定制的发音工具非常缺乏,尽管每年有数百万人学习它们。印地语拥有超过6亿的使用者,是全球第四大使用语言,改善印地语发音是解决这一差距的重要第一步。本文提出1)Dhvani -一种新的印地语CAPT系统,2)印地语发音错误的合成语音生成,3)一种新的方法,为学习者提供个性化的反馈。虽然该系统经常与使用梵文字形的学习者进行交互,但其核心分析目标是音素区别,利用印地语的高度语音正字法来分析发音错误的语音并提供有针对性的反馈。
摘要:Computer-Assisted Pronunciation Training (CAPT) has been extensively studied for English. However, there remains a critical gap in its application to Indian languages with a base of 1.5 billion speakers. Pronunciation tools tailored to Indian languages are strikingly lacking despite the fact that millions learn them every year. With over 600 million speakers and being the fourth most-spoken language worldwide, improving Hindi pronunciation is a vital first step toward addressing this gap. This paper proposes 1) Dhvani -- a novel CAPT system for Hindi, 2) synthetic speech generation for Hindi mispronunciations, and 3) a novel methodology for providing personalized feedback to learners. While the system often interacts with learners using Devanagari graphemes, its core analysis targets phonemic distinctions, leveraging Hindi's highly phonetic orthography to analyze mispronounced speech and provide targeted feedback.
【15】 Enhancing GOP in CTC-Based Mispronunciation Detection with Phonological Knowledge
标题: 利用音素知识增强基于ATC的发音错误检测中GOP链接:https://arxiv.org/abs/2506.02080
备注:Accepted to Interspeech 2025. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)
摘要:计算机辅助发音训练(CAPT)系统采用发音质量的自动测量,例如发音良好度(GOP)度量。GOP依赖于强制对齐,由于声学可变性,这容易导致标记和分割错误。虽然无噪声方法解决了这些挑战,但它们在计算上是昂贵的,并且随着音素序列长度和库存大小的扩展性很差。为了提高效率,我们引入了一个替代意识的语音免费GOP,限制基于音素集群和常见的学习者错误的音素替换。我们在两个L2英语语音数据集上评估了我们的GOP,一个是儿童语音,我的发音教练(MPC)和SpeechOcean762,其中包括儿童和成人语音。我们比较了RPS(限制性音素替换)和UPS(不受限制的音素替换)设置在无干扰的方法,这优于基线。我们讨论了我们的结果,并概述了未来研究的途径。
摘要:Computer-Assisted Pronunciation Training (CAPT) systems employ automatic measures of pronunciation quality, such as the goodness of pronunciation (GOP) metric. GOP relies on forced alignments, which are prone to labeling and segmentation errors due to acoustic variability. While alignment-free methods address these challenges, they are computationally expensive and scale poorly with phoneme sequence length and inventory size. To enhance efficiency, we introduce a substitution-aware alignment-free GOP that restricts phoneme substitutions based on phoneme clusters and common learner errors. We evaluated our GOP on two L2 English speech datasets, one with child speech, My Pronunciation Coach (MPC), and SpeechOcean762, which includes child and adult speech. We compared RPS (restricted phoneme substitutions) and UPS (unrestricted phoneme substitutions) setups within alignment-free methods, which outperformed the baseline. We discuss our results and outline avenues for future research.
【16】 Evaluating the Effectiveness of Pre-Trained Audio Embeddings for Classification of Parkinson's Disease Speech Data
标题: 评估预训练音频嵌入对帕金森病语音数据分类的有效性链接:https://arxiv.org/abs/2506.02078
备注:Accepted to Interspeech 2025. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)
摘要:语言障碍是帕金森病(PD)的普遍生物标志物,促进了使用语音数据进行临床应用的诊断技术的发展。虽然深层声学特征已显示出PD分类的前景,但其有效性往往因个体扬声器差异而异,这是现有文献中尚未彻底探索的因素。该研究调查了三种预训练的音频嵌入(OpenL3,VGGish和Wav2Vec2.0模型)用于PD分类的有效性。使用NeuroVoz数据集,OpenL3在舒张运动(DDK)和听和重复(LR)任务中优于其他人,捕获了PD检测的关键声学特征。只有Wav2Vec2.0显示出显着的性别偏见,男性发言人,在DDK任务中取得了更有利的结果。错误分类的情况下,揭示了非典型语音模式的挑战,强调需要改进的特征提取和模型鲁棒性PD检测。
摘要:Speech impairments are prevalent biomarkers for Parkinson's Disease (PD), motivating the development of diagnostic techniques using speech data for clinical applications. Although deep acoustic features have shown promise for PD classification, their effectiveness often varies due to individual speaker differences, a factor that has not been thoroughly explored in the existing literature. This study investigates the effectiveness of three pre-trained audio embeddings (OpenL3, VGGish and Wav2Vec2.0 models) for PD classification. Using the NeuroVoz dataset, OpenL3 outperforms others in diadochokinesis (DDK) and listen and repeat (LR) tasks, capturing critical acoustic features for PD detection. Only Wav2Vec2.0 shows significant gender bias, achieving more favorable results for male speakers, in DDK tasks. The misclassified cases reveal challenges with atypical speech patterns, highlighting the need for improved feature extraction and model robustness in PD detection.
【17】 No Audiogram: Leveraging Existing Scores for Personalized Speech Intelligibility Prediction
标题: No Audiogram:利用现有分数进行个性化语音清晰度预测链接:https://arxiv.org/abs/2506.02039
备注:Accepted at Interspeech 2025
摘要:个性化的语音清晰度预测是一个挑战。以前的方法主要依赖于听力图,这是固有的准确性有限,因为它们只捕捉听众的纯音听力阈值。我们提出了一种新的方法,利用个人现有的可懂度数据来预测他们在新音频上的表现,而不是结合额外的听众功能。我们引入了基于支持样本的可理解性预测网络(SSIPNet),这是一种深度学习模型,它利用语音基础模型从多个支持(音频,分数)对中构建收听者语音识别能力的高维表示,从而实现对不可见音频的准确预测。Clarity Prediction Challenge数据集上的结果表明,即使有少量的支持(音频,分数)对,我们的方法也优于基于音频的预测。我们的工作提出了一个新的模式,个性化的语音可懂度预测。
摘要:Personalized speech intelligibility prediction is challenging. Previous approaches have mainly relied on audiograms, which are inherently limited in accuracy as they only capture a listener's hearing threshold for pure tones. Rather than incorporating additional listener features, we propose a novel approach that leverages an individual's existing intelligibility data to predict their performance on new audio. We introduce the Support Sample-Based Intelligibility Prediction Network (SSIPNet), a deep learning model that leverages speech foundation models to build a high-dimensional representation of a listener's speech recognition ability from multiple support (audio, score) pairs, enabling accurate predictions for unseen audio. Results on the Clarity Prediction Challenge dataset show that, even with a small number of support (audio, score) pairs, our method outperforms audiogram-based predictions. Our work presents a new paradigm for personalized speech intelligibility prediction.
【18】 Towards a Japanese Full-duplex Spoken Dialogue System
标题: 迈向日本人的全日制口语对话系统链接:https://arxiv.org/abs/2506.02979
备注:Accepted to Interspeech 2025
摘要:全双工口语对话系统,它可以同时模拟人类对话的双向特征,如语音重叠和反向通道,最近引起了极大的关注。然而,对于日语的全双工口语对话系统的研究一直是有限的,并且其在日语中的发展的研究仍然很少。在本文中,我们提出了第一个公开可用的全双工口语对话模型在日本,这是建立在Moshi,一个全双工对话模型在英语。我们的模型是通过两个阶段的过程进行训练的:对日语大规模口语对话数据进行预训练,然后对高质量的立体声口语对话数据进行微调。我们进一步提高模型的性能,将合成的对话数据生成的多流文本到语音系统。评价实验表明,训练后的模型在自然性和有意义性方面优于日本基线模型。
摘要:Full-duplex spoken dialogue systems, which can model simultaneous bidirectional features of human conversations such as speech overlaps and backchannels, have attracted significant attention recently. However, the study of full-duplex spoken dialogue systems for the Japanese language has been limited, and the research on their development in Japanese remains scarce. In this paper, we present the first publicly available full-duplex spoken dialogue model in Japanese, which is built upon Moshi, a full-duplex dialogue model in English. Our model is trained through a two-stage process: pre-training on a large-scale spoken dialogue data in Japanese, followed by fine-tuning on high-quality stereo spoken dialogue data. We further enhance the model's performance by incorporating synthetic dialogue data generated by a multi-stream text-to-speech system. Evaluation experiments demonstrate that the trained model outperforms Japanese baseline models in both naturalness and meaningfulness.
【19】 A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation
标题: 用于德语方言ASR和方言到标准语音翻译的多方言数据集链接:https://arxiv.org/abs/2506.02894
备注:Accepted to Interspeech 2025
摘要:尽管德国有着多样化的方言,但它们在当前的自动语音识别(ASR)研究中的代表性不足。为了研究模型对方言变异的鲁棒性,我们提出了Betthupferl,一个评估数据集,包含四个小时的阅读语音在三个方言组在德国东南部(弗兰肯,巴伐利亚,阿勒曼尼),和半个小时的标准德语语音。我们提供了方言和标准德语transmittance,并分析它们之间的语言差异。我们对几种最先进的多语言ASR模型进行了标准德语语音翻译的基准测试,并发现了输出与方言和标准化翻译之间的差异。最佳ASR模型的定性错误分析表明,它有时正常化的语法差异,但往往保持更接近方言结构。
摘要:Although Germany has a diverse landscape of dialects, they are underrepresented in current automatic speech recognition (ASR) research. To enable studies of how robust models are towards dialectal variation, we present Betthupferl, an evaluation dataset containing four hours of read speech in three dialect groups spoken in Southeast Germany (Franconian, Bavarian, Alemannic), and half an hour of Standard German speech. We provide both dialectal and Standard German transcriptions, and analyze the linguistic differences between them. We benchmark several multilingual state-of-the-art ASR models on speech translation into Standard German, and find differences between how much the output resembles the dialectal vs. standardized transcriptions. Qualitative error analyses of the best ASR model reveal that it sometimes normalizes grammatical differences, but often stays closer to the dialectal constructions.
【20】 UltrasonicSpheres: Localized, Multi-Channel Sound Spheres Using Off-the-Shelf Speakers and Earables
标题: UltrasonicSpheres:使用现成扬声器和耳机的本地化多通道声音球体链接:https://arxiv.org/abs/2506.02715
摘要:我们提出了一个演示的UltrasonicSpheres,一个新的系统,用于特定位置的音频传输使用可穿戴耳机,解码成可听的声音超声波信号。与传统的波束形成设置不同,UltrasonicSpheres依赖于单个超声波扬声器来广播具有多个通道的本地化音频,每个通道都编码在不同的超声波载波频率上。佩戴我们的声学透明耳机的用户可以解调他们选择的流,例如以选定的语言展示叙述,同时保持对周围环境声音的完全感知。这种体验保留了空间音频感知,给人的印象是声音直接来自声源的物理位置。这可以实现个性化、本地化的音频,而无需配对、跟踪或额外的基础设施。重要的是,没有配备耳机的游客不受影响,因为超声波信号是人耳听不见的。我们的演示邀请参与者探索多个协同定位的音频区域,并体验UltrasonicSpheres如何支持在公共空间中不显眼地传递个性化声音。
摘要:We present a demo ofUltrasonicSpheres, a novel system for location-specific audio delivery using wearable earphones that decode ultrasonic signals into audible sound. Unlike conventional beamforming setups, UltrasonicSpheres relies on single ultrasonic speakers to broadcast localized audio with multiple channels, each encoded on a distinct ultrasonic carrier frequency. Users wearing our acoustically transparent earphones can demodulate their selected stream, such as exhibit narrations in a chosen language, while remaining fully aware of ambient environmental sounds. The experience preserves spatial audio perception, giving the impression that the sound originates directly from the physical location of the source. This enables personalized, localized audio without requiring pairing, tracking, or additional infrastructure. Importantly, visitors not equipped with the earphones are unaffected, as the ultrasonic signals are inaudible to the human ear. Our demo invites participants to explore multiple co-located audio zones and experience how UltrasonicSpheres supports unobtrusive delivery of personalized sound in public spaces.
【21】 MotionRAG-Diff: A Retrieval-Augmented Diffusion Framework for Long-Term Music-to-Dance Generation
标题: MotionRAG-Diff:长期音乐到舞蹈一代的检索增强传播框架链接:https://arxiv.org/abs/2506.02661
备注:12 pages, 5 figures
摘要:生成长期的,连贯的,逼真的音乐条件下的舞蹈序列仍然是一个具有挑战性的任务,在人体运动合成。现有的方法表现出严重的局限性:运动图方法依赖于固定的模板库,限制创造性的产生;扩散模型,而能够产生新的运动,往往缺乏时间的连贯性和音乐对齐。为了解决这些挑战,我们提出了$\textbf{MotionRAG-Diff}$,一个混合框架,集成检索增强生成(RAG)与基于扩散的细化,使高品质,音乐连贯的舞蹈生成任意长期的音乐输入。我们的方法引入了三个核心创新:(1)跨模态对比学习架构,在共享的潜在空间中对齐异构的音乐和舞蹈表示,在没有配对数据的情况下建立无监督的语义对应;(2)优化的运动图系统,用于高效检索和无缝拼接运动片段,确保长序列的真实感和时间一致性;(3)多条件扩散模型,其对原始音乐信号和对比特征进行联合条件化以增强运动质量和全局同步。大量的实验表明,MotionRAG-Diff在运动质量、多样性和音乐运动同步精度方面达到了最先进的性能。这项工作建立了一个新的范例,音乐驱动的舞蹈生成协同检索为基础的模板保真度与扩散为基础的创意增强。
摘要:Generating long-term, coherent, and realistic music-conditioned dance sequences remains a challenging task in human motion synthesis. Existing approaches exhibit critical limitations: motion graph methods rely on fixed template libraries, restricting creative generation; diffusion models, while capable of producing novel motions, often lack temporal coherence and musical alignment. To address these challenges, we propose $\textbf{MotionRAG-Diff}$, a hybrid framework that integrates Retrieval-Augmented Generation (RAG) with diffusion-based refinement to enable high-quality, musically coherent dance generation for arbitrary long-term music inputs. Our method introduces three core innovations: (1) A cross-modal contrastive learning architecture that aligns heterogeneous music and dance representations in a shared latent space, establishing unsupervised semantic correspondence without paired data; (2) An optimized motion graph system for efficient retrieval and seamless concatenation of motion segments, ensuring realism and temporal coherence across long sequences; (3) A multi-condition diffusion model that jointly conditions on raw music signals and contrastive features to enhance motion quality and global synchronization. Extensive experiments demonstrate that MotionRAG-Diff achieves state-of-the-art performance in motion quality, diversity, and music-motion synchronization accuracy. This work establishes a new paradigm for music-driven dance generation by synergizing retrieval-based template fidelity with diffusion-based creative enhancement.
【22】 Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
标题: 通过Whisper微调克服多方言阿拉伯语ASB中的数据稀缺性链接:https://arxiv.org/abs/2506.02627
备注:Accepted at Interspeech 2025
摘要:虽然商用阿拉伯语自动语音识别(ASR)系统支持现代标准阿拉伯语(MSA),但它们难以处理方言语音。我们使用Mozilla Common Voice for MSA和MASC dataset for dialectal speech研究了OpenAI的Whisper对五种主要阿拉伯方言(海湾,黎凡特,伊拉克,埃及,马格里布)的微调效果。我们评估了MSA训练规模的影响,MSA数据预训练的好处,以及方言特异性与方言合并模型。我们发现,少量的MSA微调数据产生实质性的改善较小的模型,匹配较大的非微调模型。虽然MSA预训练显示出最小的好处,这表明MSA和方言之间的共享功能有限,但我们的方言池模型对方言特定的模型进行了验证。这表明,在适当平衡的情况下,汇集方言数据可以帮助解决低资源ASR中的数据稀缺问题,而不会造成显著的性能损失。
摘要:Although commercial Arabic automatic speech recognition (ASR) systems support Modern Standard Arabic (MSA), they struggle with dialectal speech. We investigate the effect of fine-tuning OpenAI's Whisper on five major Arabic dialects (Gulf, Levantine, Iraqi, Egyptian, Maghrebi) using Mozilla Common Voice for MSA and the MASC dataset for dialectal speech. We evaluate MSA training size effects, benefits of pre-training on MSA data, and dialect-specific versus dialect-pooled models. We find that small amounts of MSA fine-tuning data yield substantial improvements for smaller models, matching larger non-fine-tuned models. While MSA pre-training shows minimal benefit, suggesting limited shared features between MSA and dialects, our dialect-pooled models perform comparably to dialect-specific ones. This indicates that pooling dialectal data, when properly balanced, can help address data scarcity in low-resource ASR without significant performance loss.
【23】 Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm
标题: 基于图注意网络和标签传播算法的说话人重叠社区检测链接:https://arxiv.org/abs/2506.02610
摘要:在说话人日志化中,传统的基于聚类的方法仍然广泛应用于现实世界的应用中。然而,这些方法与说话人嵌入和重叠语音段的复杂分布作斗争。为了解决这些问题,我们提出了一种基于图注意力网络和标签传播算法(OCDGALP)的重叠社区检测方法。所提出的框架包括两个关键组件:(1)一个图注意力网络,通过聚合来自相邻节点的信息来细化说话人嵌入和节点连接,以及(2)一个标签传播算法,为每个节点分配多个社区标签,从而实现同时聚类和重叠社区检测。实验结果表明,该方法显着降低了日志错误率(DER),实现了最先进的15.94%的DER在DIHARD-III数据集上没有oracle语音活动检测(VAD),和一个令人印象深刻的11.07%与oracle VAD。
摘要:In speaker diarization, traditional clustering-based methods remain widely used in real-world applications. However, these methods struggle with the complex distribution of speaker embeddings and overlapping speech segments. To address these limitations, we propose an Overlapping Community Detection method based on Graph Attention networks and the Label Propagation Algorithm (OCDGALP). The proposed framework comprises two key components: (1) a graph attention network that refines speaker embeddings and node connections by aggregating information from neighboring nodes, and (2) a label propagation algorithm that assigns multiple community labels to each node, enabling simultaneous clustering and overlapping community detection. Experimental results show that the proposed method significantly reduces the Diarization Error Rate (DER), achieving a state-of-the-art 15.94% DER on the DIHARD-III dataset without oracle Voice Activity Detection (VAD), and an impressive 11.07% with oracle VAD.
【24】 Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
标题: 超越词汇内容的韵律结构:自我监督学习的研究链接:https://arxiv.org/abs/2506.02584
备注:Accepted at INTERSPEECH 2025
摘要:在语篇理解过程中,人们利用词汇结构的可预测性。虽然可预测的结构也存在于言语中,但韵律(例如语调、节奏和响度)在多大程度上独立于词汇内容而对这种结构做出贡献尚不清楚。本研究利用自监督学习(SSL)来研究韵律声学相关结构的时间粒度。我们提出的Masked Prosody Model的表示可以预测依赖于本地信息(如单词边界)的感知标签,但为涉及长期结构(如情感识别)的标签提供最大价值。各种感知标签的探测实验显示出比未转换的音高、能量和语音活动特征更强的相对增益。我们的研究结果揭示了SSL训练目标时间尺度的重要性,并强调了复杂SSL编码结构与更受约束的经典结构相比的价值。
摘要:People exploit the predictability of lexical structures during text comprehension. Though predictable structure is also present in speech, the degree to which prosody, e.g. intonation, tempo, and loudness, contributes to such structure independently of the lexical content is unclear. This study leverages self-supervised learning (SSL) to examine the temporal granularity of structures in the acoustic correlates of prosody. Representations from our proposed Masked Prosody Model can predict perceptual labels dependent on local information, such as word boundaries, but provide the most value for labels involving longer-term structures, like emotion recognition. Probing experiments across various perceptual labels show strong relative gains over untransformed pitch, energy, and voice activity features. Our results reveal the importance of SSL training objective timescale and highlight the value of complex SSL-encoded structures compared to more constrained classical structures.
【25】 On the Language and Gender Biases in PSTN, VoIP and Neural Audio Codecs
标题: 关于公共电话网、语音电话和神经音频编解码器中的语言和性别偏见链接:https://arxiv.org/abs/2506.02545
备注:Submitted to INTERSPEECH2025
摘要:近年来,人们越来越关注语音技术中的公平性和包容性,特别是在自动语音识别和语音情感分析等领域。当音频在处理之前被转码时,如在流或实时应用中的情况,编码机制中的任何固有偏差都可能导致不一致。这不仅影响用户体验,而且还可能通过使陈规定型观念和排斥永久化而产生更广泛的社会影响。因此,音频编码机制的公正性非常重要。在这项工作中,我们对音频编解码器的语言和性别偏见方面的稀缺研究做出了贡献。通过分析超过200万个多语言音频文件的语音质量后,通过一个代表性的编解码器子集(PSTN,VoIP和神经)转码,我们的研究结果表明,PSTN编解码器在性别方面有很强的偏见,神经编解码器引入语言偏见。
摘要:In recent years, there has been a growing focus on fairness and inclusivity within speech technology, particularly in areas such as automatic speech recognition and speech sentiment analysis. When audio is transcoded prior to processing, as is the case in streaming or real-time applications, any inherent bias in the coding mechanism may result in disparities. This not only affects user experience but can also have broader societal implications by perpetuating stereotypes and exclusion. Thus, it is important that audio coding mechanisms are unbiased. In this work, we contribute towards the scarce research with respect to language and gender biases of audio codecs. By analyzing the speech quality of over 2 million multilingual audio files after transcoding through a representative subset of codecs (PSTN, VoIP and neural), our results indicate that PSTN codecs are strongly biased in terms of gender and that neural codecs introduce language biases.
【26】 DnR-nonverbal: Cinematic Audio Source Separation Dataset Containing Non-Verbal Sounds
标题: DnR-非言语:包含非言语声音的电影音频源分离数据集链接:https://arxiv.org/abs/2506.02499
备注:Accepted to Interspeech 2025, 5 pages, 3 figures, dataset is available at this https URL
摘要:我们提出了一个新的数据集的电影音频源分离(CASS),处理非语言的声音。现有的CASS数据集只包含阅读风格的声音作为语音干。这些数据集与实际的电影音频不同,实际的电影音频更可能包括表演出来的声音。因此,在传统数据集上训练的模型往往存在这样的问题,即情绪化的声音,如笑声和尖叫声,更容易被分离为效果,而不是语音。为了解决这个问题,我们建立了一个新的数据集,DnR-nonverbal。拟议的数据集包括语音干中的笑声和尖叫声等非语言声音。通过实验,我们揭示了当前CASS模型在非言语声音提取方面的问题,并表明我们的数据集可以有效地解决合成和实际电影音频中的问题。我们的数据集可以在https://zenodo.org/records/15470640上找到。
摘要:We propose a new dataset for cinematic audio source separation (CASS) that handles non-verbal sounds. Existing CASS datasets only contain reading-style sounds as a speech stem. These datasets differ from actual movie audio, which is more likely to include acted-out voices. Consequently, models trained on conventional datasets tend to have issues where emotionally heightened voices, such as laughter and screams, are more easily separated as an effect, not speech. To address this problem, we build a new dataset, DnR-nonverbal. The proposed dataset includes non-verbal sounds like laughter and screams in the speech stem. From the experiments, we reveal the issue of non-verbal sound extraction by the current CASS model and show that our dataset can effectively address the issue in the synthetic and actual movie audio. Our dataset is available at https://zenodo.org/records/15470640.
【27】 SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
标题: SOVA-Bench:基于LLM的语音助理的语音对话能力基准测试链接:https://arxiv.org/abs/2506.02457
摘要:由于大型语言模型(LLM),语音编码算法和声码器结构的稳步发展,最近的进步已经能够直接从用户指令生成语音响应。然而,基准生成的语音质量一直是一个被忽视的,但关键的问题,考虑到从追求语义准确性生动和自发的语音流的转变。以往的评价侧重于语音理解能力,缺乏对音质的量化。在本文中,我们提出了语音概念语音助手基准(SOVA-Bench),提供了一般知识,语音识别和理解,以及可用的语音LLM之间的语义和声学生成能力的理解比较。据我们所知,SOVA-Bench是语音LLM最系统的评估框架之一,启发了语音交互系统的发展方向。
摘要:Thanks to the steady progress of large language models (LLMs), speech encoding algorithms and vocoder structure, recent advancements have enabled generating speech response directly from a user instruction. However, benchmarking the generated speech quality has been a neglected but critical issue, considering the shift from the pursuit of semantic accuracy to vivid and spontaneous speech flow. Previous evaluation focused on the speech-understanding ability, lacking a quantification of acoustic quality. In this paper, we propose Speech cOnversational Voice Assistant Benchmark (SOVA-Bench), providing a comprehension comparison of the general knowledge, speech recognition and understanding, along with both semantic and acoustic generative ability between available speech LLMs. To the best of our knowledge, SOVA-Bench is one of the most systematic evaluation frameworks for speech LLMs, inspiring the direction of voice interaction systems.
【28】 Breaking the Barriers of Text-Hungry and Audio-Deficient AI
标题: 打破文本饥饿和音频匮乏的人工智能的障碍链接:https://arxiv.org/abs/2506.02443
备注:61 pages, 16 figures, 14 tables, 25 languages, 13 blockaudio per language. Presented at AI Mali, May 2025
摘要:虽然全球语言多样性涵盖了超过7164种公认的语言,但目前占主导地位的机器智能架构仍然从根本上偏向于书面文本。这一偏见将7亿多人排除在外,特别是在农村和偏远地区,他们是听识字的。在这项工作中,我们介绍了一个完全无文本的,音频到音频的机器智能框架,旨在为这一服务不足的人群,以及所有喜欢音频效率的人。我们的贡献包括完全绕过文本的新颖音频到音频翻译架构,包括频谱图,尺度图,小波和基于单元的模型。我们的方法的核心是多尺度音频语义转换(MAST),表示编码音调,韵律,扬声器和表达功能。我们进一步将MAST集成到由分数布朗运动驱动的平均场型框架的分数扩散中。它能够在不依赖文本监督的情况下生成高保真、语义一致的语音。其结果是一个强大且可扩展的系统,能够直接从原始音频中学习,即使是不成文或很少数字化的语言。这项工作代表了向音频原生机器智能系统的根本转变,为历史上被排除在当前机器智能生态系统之外的社区扩大了对语言技术的访问。
摘要:While global linguistic diversity spans more than 7164 recognized languages, the current dominant architecture of machine intelligence remains fundamentally biased toward written text. This bias excludes over 700 million people particularly in rural and remote regions who are audio-literate. In this work, we introduce a fully textless, audio-to-audio machine intelligence framework designed to serve this underserved population, and all the people who prefer audio-efficiency. Our contributions include novel Audio-to-Audio translation architectures that bypass text entirely, including spectrogram-, scalogram-, wavelet-, and unit-based models. Central to our approach is the Multiscale Audio-Semantic Transform (MAST), a representation that encodes tonal, prosodic, speaker, and expressive features. We further integrate MAST into a fractional diffusion of mean-field-type framework powered by fractional Brownian motion. It enables the generation of high-fidelity, semantically consistent speech without reliance on textual supervision. The result is a robust and scalable system capable of learning directly from raw audio, even in languages that are unwritten or rarely digitized. This work represents a fundamental shift toward audio-native machine intelligence systems, expanding access to language technologies for communities historically left out of the current machine intelligence ecosystem.
【29】 StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
标题: StarVC:语音转换中联合生成文本和语音的统一自回归框架链接:https://arxiv.org/abs/2506.02414
备注:5 pages, 2 figures, Accepted by Interspeech 2025, Demo: this https URL
摘要:语音转换(VC)修改语音以匹配目标说话者,同时保留语言内容。传统的方法通常直接从语音中提取说话人信息,而忽略了对语言内容的显式利用。由于VC从根本上涉及将说话者身份从语言内容中分离出来,因此利用结构化语义特征可以提高转换性能。然而,以前的尝试,将语义特征到VC显示有限的有效性,激励显式文本建模的集成。我们提出了StarVC,一个统一的自回归VC框架,首先预测文本标记合成声学特征之前。实验表明,StarVC在保留语言内容(即,WER和CER)和扬声器特性(即,SECS和MOS)。音频演示可以在https://thuhcsi.github.io/StarVC/上找到。
摘要:Voice Conversion (VC) modifies speech to match a target speaker while preserving linguistic content. Traditional methods usually extract speaker information directly from speech while neglecting the explicit utilization of linguistic content. Since VC fundamentally involves disentangling speaker identity from linguistic content, leveraging structured semantic features could enhance conversion performance. However, previous attempts to incorporate semantic features into VC have shown limited effectiveness, motivating the integration of explicit text modeling. We propose StarVC, a unified autoregressive VC framework that first predicts text tokens before synthesizing acoustic features. The experiments demonstrate that StarVC outperforms conventional VC methods in preserving both linguistic content (i.e., WER and CER) and speaker characteristics (i.e., SECS and MOS). Audio demo can be found at: https://thuhcsi.github.io/StarVC/.
【30】 Trusted Fake Audio Detection Based on Dirichlet Distribution
标题: 基于Dirichlet分布的可信假音频检测链接:https://arxiv.org/abs/2506.02401
摘要:随着基于深度学习的语音转换和语音合成技术的不断发展,虚假音频带来的网络安全问题日益严重。先前提出的用于防御假音频的模型已经取得了显着的性能。然而,它们都未能对模型本身所做决策的可信度进行建模。在此基础上,提出了一种基于Dirichlet分布的虚假音频检测方法,旨在提高虚假音频检测的可靠性。具体来说,我们首先通过神经网络生成证据。然后使用狄利克雷分布对不确定性进行建模。通过用狄利克雷分布的参数对置信分布建模,可以获得每个决策的不确定性估计。最后,预测的概率和相应的不确定性估计相结合,形成最终意见。在ASVspoof系列数据集上(即,ASVspoof 2019 LA,ASVspoof 2021 LA和DF),我们进行了一些比较实验,以验证所提出的模型在准确性,鲁棒性和可信度方面的优异性能。
摘要:With the continuous development of deep learning-based speech conversion and speech synthesis technologies, the cybersecurity problem posed by fake audio has become increasingly serious. Previously proposed models for defending against fake audio have attained remarkable performance. However, they all fall short in modeling the trustworthiness of the decisions made by the models themselves. Based on this, we put forward a plausible fake audio detection approach based on the Dirichlet distribution with the aim of enhancing the reliability of fake audio detection. Specifically, we first generate evidence through a neural network. Uncertainty is then modeled using the Dirichlet distribution. By modeling the belief distribution with the parameters of the Dirichlet distribution, an estimate of uncertainty can be obtained for each decision. Finally, the predicted probabilities and corresponding uncertainty estimates are combined to form the final opinion. On the ASVspoof series dataset (i.e., ASVspoof 2019 LA, ASVspoof 2021 LA, and DF), we conduct a number of comparison experiments to verify the excellent performance of the proposed model in terms of accuracy, robustness, and trustworthiness.
【31】 Sounding Like a Winner? Prosodic Differences in Post-Match Interviews
标题: 听起来像是赢家?赛后采访中的韵律差异链接:https://arxiv.org/abs/2506.02283
备注:Accepted to Interspeech 2025
摘要:本研究旨在探讨网球赛后访谈中与输赢相关的韵律特征。此外,本研究探讨了使用韵律特征和自我监督学习(SSL)表征仅基于赛后采访录音对比赛结果进行分类的可能性。通过分析音高和强度等韵律元素,以及Wav 2 Vec 2.0和HuBERT等SSL模型,目的是确定运动员是否赢得了比赛。从数据中提取传统的声学特征和深度语音表示,并使用机器学习分类器来区分获胜和失败的球员。结果表明,SSL表示有效地区分输赢的结果,捕捉微妙的语音模式与情绪状态。与此同时,韵律线索--如音高的变化--仍然是胜利的有力指标。
摘要:This study examines the prosodic characteristics associated with winning and losing in post-match tennis interviews. Additionally, this research explores the potential to classify match outcomes solely based on post-match interview recordings using prosodic features and self-supervised learning (SSL) representations. By analyzing prosodic elements such as pitch and intensity, alongside SSL models like Wav2Vec 2.0 and HuBERT, the aim is to determine whether an athlete has won or lost their match. Traditional acoustic features and deep speech representations are extracted from the data, and machine learning classifiers are employed to distinguish between winning and losing players. Results indicate that SSL representations effectively differentiate between winning and losing outcomes, capturing subtle speech patterns linked to emotional states. At the same time, prosodic cues -- such as pitch variability -- remain strong indicators of victory.
【32】 Investigating the Impact of Word Informativeness on Speech Emotion Recognition
标题: 调查词语信息量对言语情感识别的影响链接:https://arxiv.org/abs/2506.02239
备注:Accepted to Interspeech 2025
摘要:在语音情感识别中,一个关键的挑战在于识别携带最相关的声学变化的语音信号段,以辨别特定的情感。传统方法计算整个句子或较长语音部分的能量和F0等特征的泛函,可能会错过长形式统计数据中重要的细粒度变化。本研究探讨使用词的信息量,来自一个预先训练的语言模型,以确定语义上重要的部分。然后专门为这些识别的片段计算声学特征,提高情感识别的准确性。该方法利用标准的声学韵律特征,其功能,和自我监督的表示。结果表明,识别性能的显着改善时,功能计算的基础上选择的单词信息量的片段,强调这种方法的有效性。
摘要:In emotion recognition from speech, a key challenge lies in identifying speech signal segments that carry the most relevant acoustic variations for discerning specific emotions. Traditional approaches compute functionals for features such as energy and F0 over entire sentences or longer speech portions, potentially missing essential fine-grained variation in the long-form statistics. This research investigates the use of word informativeness, derived from a pre-trained language model, to identify semantically important segments. Acoustic features are then computed exclusively for these identified segments, enhancing emotion recognition accuracy. The methodology utilizes standard acoustic prosodic features, their functionals, and self-supervised representations. Results indicate a notable improvement in recognition performance when features are computed on segments selected based on word informativeness, underscoring the effectiveness of this approach.
【33】 HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation
标题: HENT-SRT:具有自蒸馏功能的分层高效神经传感器,用于联合语音识别和翻译链接:https://arxiv.org/abs/2506.02157
摘要:神经传感器(NT)为语音流提供了一个有效的框架,在自动语音识别(ASR)中表现出强大的性能。然而,NT到语音翻译(ST)的应用仍然具有挑战性,因为现有的方法在对ASR和ST联合建模时会遇到单词重新排序和性能下降的问题,从而导致与基于注意力的编码器-解码器(AED)模型之间的差距。现有的基于NT的ST方法也遭受高计算训练成本。为了解决这些问题,我们提出了HENT-SRT(用于语音识别和翻译的分层高效神经传感器),这是一种新的框架,可以分解ASR和翻译任务,以更好地处理重新排序。为了确保强大的ST,同时保持ASR性能,我们使用自蒸馏与CTC一致性正则化。此外,我们通过结合ASR传感器的最佳实践来提高计算效率,包括下采样分层编码器,无状态预测器和修剪的传感器损耗以降低训练复杂性。最后,我们在解码过程中引入了空白惩罚,减少了删除,提高了翻译质量。我们的方法在三个会话数据集上进行评估-阿拉伯语,西班牙语和普通话,在NT模型中实现了新的最先进的性能,并大大缩小了与基于AED的系统的差距。
摘要:Neural transducers (NT) provide an effective framework for speech streaming, demonstrating strong performance in automatic speech recognition (ASR). However, the application of NT to speech translation (ST) remains challenging, as existing approaches struggle with word reordering and performance degradation when jointly modeling ASR and ST, resulting in a gap with attention-based encoder-decoder (AED) models. Existing NT-based ST approaches also suffer from high computational training costs. To address these issues, we propose HENT-SRT (Hierarchical Efficient Neural Transducer for Speech Recognition and Translation), a novel framework that factorizes ASR and translation tasks to better handle reordering. To ensure robust ST while preserving ASR performance, we use self-distillation with CTC consistency regularization. Moreover, we improve computational efficiency by incorporating best practices from ASR transducers, including a down-sampled hierarchical encoder, a stateless predictor, and a pruned transducer loss to reduce training complexity. Finally, we introduce a blank penalty during decoding, reducing deletions and improving translation quality. Our approach is evaluated on three conversational datasets Arabic, Spanish, and Mandarin achieving new state-of-the-art performance among NT models and substantially narrowing the gap with AED-based systems.
【34】 Comparison of spectrogram scaling in multi-label Music Genre Recognition
标题: 多标签音乐流派识别中谱图缩放的比较链接:https://arxiv.org/abs/2506.02091
备注:14 pages, 10 figures
摘要:随着数字音频工作站的可访问性和易用性的增加,普通听众可以获得的音乐数量也在增加;此外,流派之间的差异并不总是很好地定义,并且可能是抽象的,各个唱片中的流派组合差异很大。在这篇文章中,多种预处理方法和模型训练方法进行了描述和比较,占今天的专辑折衷的性质。一个自定义的,手动标记的数据集超过18000个条目已被用来执行实验。
摘要:As the accessibility and ease-of-use of digital audio workstations increases, so does the quantity of music available to the average listener; additionally, differences between genres are not always well defined and can be abstract, with widely varying combinations of genres across individual records. In this article, multiple preprocessing methods and approaches to model training are described and compared, accounting for the eclectic nature of today's albums. A custom, manually labeled dataset of more than 18000 entries has been used to perform the experiments.
【35】 Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
标题: 揭开音频Deepfake起源:采用Ensemble Fusion的深度指标学习和Conformer网络方法链接:https://arxiv.org/abs/2506.02085
备注:Accepted at Interspeech 2025, Netherlands
摘要:音频Deepfake正在通过先进的AI获得前所未有的真实感。虽然目前的研究重点是从欺骗性语音中识别真实语音,但追踪源系统同样至关重要。本文提出了一种新的音频源跟踪系统,该系统将深度度量多类N对损失与实强调和假分散框架、一致性分类网络和集成分数嵌入融合相结合。N对损失提高了区分能力,而真实强调和虚假分散通过专注于区分真实和虚假语音模式来增强鲁棒性。Conformer网络捕获音频信号中的全局和局部依赖性,这对于源跟踪至关重要。所提出的集成分数嵌入融合显示了域内和域外源跟踪场景之间的最佳权衡。我们使用弗雷歇距离和标准度量评估我们的方法,在源跟踪中表现出优于基线系统的性能。
摘要:Audio deepfakes are acquiring an unprecedented level of realism with advanced AI. While current research focuses on discerning real speech from spoofed speech, tracing the source system is equally crucial. This work proposes a novel audio source tracing system combining deep metric multi-class N-pair loss with Real Emphasis and Fake Dispersion framework, a Conformer classification network, and ensemble score-embedding fusion. The N-pair loss improves discriminative ability, while Real Emphasis and Fake Dispersion enhance robustness by focusing on differentiating real and fake speech patterns. The Conformer network captures both global and local dependencies in the audio signal, crucial for source tracing. The proposed ensemble score-embedding fusion shows an optimal trade-off between in-domain and out-of-domain source tracing scenarios. We evaluate our method using Frechet Distance and standard metrics, demonstrating superior performance in source tracing over the baseline system.
【36】 Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
标题: 在视觉语音识别中利用大型语言模型:模型缩放、上下文感知解码和迭代抛光链接:https://arxiv.org/abs/2506.02012
摘要:视觉语音识别(VSR)通过分析嘴唇运动来转录语音。最近,大型语言模型(LLM)已被集成到VSR系统,导致显着的性能改善。然而,LLM的潜力还没有得到广泛的研究,如何有效地利用LLM在VSR任务仍然没有探索。本文系统地探讨了如何更好地利用LLM进行VSR任务,并提供了三个关键贡献:(1)缩放测试:我们研究了LLM大小如何影响VSR性能,确认了VSR任务中的缩放律。(2)上下文感知解码:我们添加上下文文本来指导LLM解码,提高识别精度。(3)迭代抛光:我们建议迭代地改进LLM输出,逐步减少识别错误。大量的实验表明,通过这些设计,LLM的巨大潜力可以很大程度上利用,导致显着的VSR性能改善。
摘要:Visual Speech Recognition (VSR) transcribes speech by analyzing lip movements. Recently, Large Language Models (LLMs) have been integrated into VSR systems, leading to notable performance improvements. However, the potential of LLMs has not been extensively studied, and how to effectively utilize LLMs in VSR tasks remains unexplored. This paper systematically explores how to better leverage LLMs for VSR tasks and provides three key contributions: (1) Scaling Test: We study how the LLM size affects VSR performance, confirming a scaling law in the VSR task. (2) Context-Aware Decoding: We add contextual text to guide the LLM decoding, improving recognition accuracy. (3) Iterative Polishing: We propose iteratively refining LLM outputs, progressively reducing recognition errors. Extensive experiments demonstrate that by these designs, the great potential of LLMs can be largely harnessed, leading to significant VSR performance improvement.
【37】 CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge
标题: CNNSRC 2024:第二届中文连续视觉语音识别挑战赛链接:https://arxiv.org/abs/2506.02010
备注:to be published in INTERSPEECH 2025
摘要:本文介绍了第二届汉语连续视觉语音识别挑战赛(CNVSRC 2024),该挑战赛以CNVSRC 2023为基础,旨在推进汉语大词汇量连续视觉语音识别(LVC-VSR)的研究。该挑战评估了两个测试场景:在录音室和互联网语音阅读。CNVSRC 2024使用与其前身CNVSRC 2023相同的数据集,其中包括用于培训的CN-CVS和用于开发和评估的CNVSRC-单/多。然而,CNVSRC 2024引入了两项关键改进:(1)更强大的基线系统,以及(2)额外的数据集CN-CVS 2-P1,用于开放跟踪,以提高数据量和多样性。新的挑战展示了数据预处理,特征提取,模型设计和训练策略方面的几项重要创新,进一步推动了中国LVC-VSR的最新发展。更多信息和资源可在官方网站上找到。
摘要:This paper presents the second Chinese Continuous Visual Speech Recognition Challenge (CNVSRC 2024), which builds on CNVSRC 2023 to advance research in Chinese Large Vocabulary Continuous Visual Speech Recognition (LVC-VSR). The challenge evaluates two test scenarios: reading in recording studios and Internet speech. CNVSRC 2024 uses the same datasets as its predecessor CNVSRC 2023, which involves CN-CVS for training and CNVSRC-Single/Multi for development and evaluation. However, CNVSRC 2024 introduced two key improvements: (1) a stronger baseline system, and (2) an additional dataset, CN-CVS2-P1, for open tracks to improve data volume and diversity. The new challenge has demonstrated several important innovations in data preprocessing, feature extraction, model design, and training strategies, further pushing the state-of-the-art in Chinese LVC-VSR. More details and resources are available at the official website.
【38】 Inter(sectional) Alia(s): Ambiguity in Voice Agent Identity via Intersectional Japanese Self-Referents
标题: 跨(部分)别名:通过交叉日本自我指涉的语音代理身份的模糊性链接:https://arxiv.org/abs/2506.01998
备注:CHI '25
摘要:模仿人类的会话代理人提出了关于将具有人类社会身份线索的机器拟人化的伦理问题。批评者还质疑类人代理的身份中立性假设。最近的研究表明,交叉日语代词可以引起复杂的,有时回避的印象代理身份。然而,其他“中性”的非代词自我指涉(NPSR)和语音作为一种社会表达媒介的作用仍然没有被探索。在一项众包研究中,日本参与者(N = 204)使用七个自我参照来评估三种ChatGPT声音(Juniper,Breeze和Ember)。我们发现了声音性别化的强有力证据,以及交叉自我指涉逃避性别化的潜力,即,通过中立性和难以捉摸的模糊性。值得注意的是,根据社会语言学理论,特别是boku和watakushi,年龄和正式的看法与性别有关。这项工作提供了一个细致入微的代理身份的看法和冠军的声音代理的跨部门和文化敏感的工作。
摘要:Conversational agents that mimic people have raised questions about the ethics of anthropomorphizing machines with human social identity cues. Critics have also questioned assumptions of identity neutrality in humanlike agents. Recent work has revealed that intersectional Japanese pronouns can elicit complex and sometimes evasive impressions of agent identity. Yet, the role of other "neutral" non-pronominal self-referents (NPSR) and voice as a socially expressive medium remains unexplored. In a crowdsourcing study, Japanese participants (N = 204) evaluated three ChatGPT voices (Juniper, Breeze, and Ember) using seven self-referents. We found strong evidence of voice gendering alongside the potential of intersectional self-referents to evade gendering, i.e., ambiguity through neutrality and elusiveness. Notably, perceptions of age and formality intersected with gendering as per sociolinguistic theories, especially boku and watakushi. This work provides a nuanced take on agent identity perceptions and champions intersectional and culturally-sensitive work on voice agents.
【39】 DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
标题: DYNEC:基于动态词汇的语音识别非自回归语境化链接:https://arxiv.org/abs/2506.00422
备注:Accepted to Interspeech 2025
摘要:上下文偏置(CB)提高了对罕见和未见过短语的自动语音识别。最近的研究引入了动态词汇表,它将上下文短语表示为自回归(AR)模型中的可扩展令牌。这种方法提高了CB的准确性,但推理速度较慢。虽然动态词汇表可以应用于非自回归(NAR)模型,如连接主义时间分类(CTC),但条件独立性假设无法捕获静态和动态标记之间的依赖关系。提出了一种基于动态词汇的NAR上下文化方法DYNAC,它是一种自适应的CTC方法,将动态词汇集成到中间层中。DYNAC根据动态词汇调节编码器,有效地捕获静态和动态令牌之间的依赖关系,同时减少实时因素(RTF)。实验结果表明,DYNAC减少了RTF的81%,在LibriSpeech 960测试干净集上的字错误率下降了0.1个点。
摘要:Contextual biasing (CB) improves automatic speech recognition for rare and unseen phrases. Recent studies have introduced dynamic vocabulary, which represents context phrases as expandable tokens in autoregressive (AR) models. This method improves CB accuracy but with slow inference speed. While dynamic vocabulary can be applied to non-autoregressive (NAR) models, such as connectionist temporal classification (CTC), the conditional independence assumption fails to capture dependencies between static and dynamic tokens. This paper proposes DYNAC (Dynamic Vocabulary-based NAR Contextualization), a self-conditioned CTC method that integrates dynamic vocabulary into intermediate layers. Conditioning the encoder on dynamic vocabulary, DYNAC effectively captures dependencies between static and dynamic tokens while reducing the real-time factor (RTF). Experimental results show that DYNAC reduces RTF by 81% with a 0.1-point degradation in word error rate on the LibriSpeech 960 test-clean set.
机器翻译由腾讯交互翻译提供,仅供参考
