今日论文合集:cs.SD语音13篇,eess.AS音频处理21篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
标题:SPGISpeech 2.0:转录多扬声器金融音频,用于扬声器标记的转录
链接:https://arxiv.org/abs/2508.05554

作者:Raymond Grossman, Taejin Park, Kunal Dhawan, Andrew Titus, Sophia Zhi, Yulia Shchadilova, Weiqing Wang, Jagadeesh Balam, Boris Ginsburg
备注:To be presented at Interspeech 2025
摘要:我们介绍SPGISpeech 2.0,一个适合在金融领域的说话人标记的转录数据集。SPGISpeech 2.0提高了适用建模任务的多样性,同时保持了原始SPGISpeech数据集的核心特征:音频片段及其相应的完全格式化的文本翻译,可用于端到端自动语音识别(ASR)。SPGISpeech 2.0包括3,780个额外小时的专业转录收入电话。此外,数据集包含用于促进多说话者ASR的每个音频片段的呼叫和说话者信息。我们验证了实用程序的SPGISpeech 2.0通过改进的扬声器标记的ASR性能的流行的语音识别模型SPGISpeech 2.0微调后。SPGISpeech 2.0免费发布供非商业使用,我们希望它能促进语音识别技术的进步,并激发广泛的研究应用。
摘要:We introduce SPGISpeech 2.0, a dataset suitable for speaker-tagged transcription in the financial domain. SPGISpeech 2.0 improves the diversity of applicable modeling tasks while maintaining the core characteristic of the original SPGISpeech dataset: audio snippets and their corresponding fully formatted text transcriptions, usable for end-to-end automatic speech recognition (ASR). SPGISpeech 2.0 consists of 3,780 additional hours of professionally transcribed earnings calls. Furthermore, the dataset contains call and speaker information for each audio snippet facilitating multi-talker ASR. We validate the utility of SPGISpeech 2.0 through improvements in speaker-tagged ASR performance of popular speech recognition models after fine-tuning on SPGISpeech 2.0. Released free for non-commercial use, we expect SPGISpeech 2.0 to foster advancements in speech recognition technologies and inspire a wide range of research applications.


【2】Embedding Alignment in Code Generation for Audio
标题:音频代码生成中嵌入对齐
链接:https://arxiv.org/abs/2508.05473

作者:Sam Kouteili, Hiren Madhu, George Typaldos, Mark Santolucito
摘要:LLM驱动的代码生成有可能彻底改变创造性的编码工作,例如实时编码,使用户能够专注于语法细节的结构基元。在这样的领域中,当提示LLM时,用户可以受益于考虑多个不同的代码候选以更好地实现他们的音乐意图。然而,代码生成模型难以呈现独特且多样的代码候选,而无法直接洞察代码的音频输出。为了更好地建立候选代码和生成的音频之间的关系,我们研究了代码和音频嵌入空间之间的映射的拓扑结构。我们发现,代码和音频嵌入并不表现出简单的线性关系,而是用一个构建的预测模型来补充,该模型表明可以学习嵌入对齐图。作为对音乐多样性输出目标的补充,我们提出了一个模型,该模型由给定的代码预测输出音频嵌入,构建代码-音频嵌入对齐图。
摘要:LLM-powered code generation has the potential to revolutionize creative coding endeavors, such as live-coding, by enabling users to focus on structural motifs over syntactic details. In such domains, when prompting an LLM, users may benefit from considering multiple varied code candidates to better realize their musical intentions. Code generation models, however, struggle to present unique and diverse code candidates, with no direct insight into the code's audio output. To better establish a relationship between code candidates and produced audio, we investigate the topology of the mapping between code and audio embedding spaces. We find that code and audio embeddings do not exhibit a simple linear relationship, but supplement this with a constructed predictive model that shows an embedding alignment map could be learned. Supplementing the aim for musically diverse output, we present a model that given code predicts output audio embedding, constructing a code-audio embedding alignment map.


【3】From Detection to Correction: Backdoor-Resilient Face Recognition via Vision-Language Trigger Detection and Noise-Based Neutralization
标题:从检测到纠正:通过视觉语言触发检测和基于噪音的中和来实现后门弹性人脸识别
链接:https://arxiv.org/abs/2508.05409

作者:Farah Wahida, M.A.P. Chamikara, Yashothara Shanmugarasa, Mohan Baruwal Chhetri, Thilina Ranbaduge, Ibrahim Khalil
备注:19 Pages, 24 Figures
摘要:生物识别系统,例如由深度神经网络(DNN)驱动的人脸识别系统,依赖于大型且高度敏感的数据集。后门攻击可以通过操纵训练过程来破坏这些系统。通过在一些训练图像中插入小的触发器,例如贴纸,化妆品或图案化的面具,对手可以在稍后的认证期间呈现相同的触发器,以被错误地识别为另一个人,从而获得未经授权的访问。针对后门攻击的现有防御机制在精确识别和缓解中毒图像而不影响数据实用性方面仍然面临挑战,这会破坏系统的整体可靠性。我们提出了一种新颖的和可推广的方法,TrueBiometric:值得信赖的生物识别技术,它使用多数投票机制,利用多个国家的最先进的大型视觉语言模型,准确地检测中毒图像。一旦被识别,中毒的样品使用目标和校准的校正噪声进行校正。我们广泛的实证结果表明,TrueBiometric检测和纠正中毒的图像与100\%的准确性,而不影响准确的干净的图像。与现有的最先进的方法相比,TrueBiometric提供了一种更实用,更准确,更有效的解决方案,用于减轻人脸识别系统中的后门攻击。
摘要:Biometric systems, such as face recognition systems powered by deep neural networks (DNNs), rely on large and highly sensitive datasets. Backdoor attacks can subvert these systems by manipulating the training process. By inserting a small trigger, such as a sticker, make-up, or patterned mask, into a few training images, an adversary can later present the same trigger during authentication to be falsely recognized as another individual, thereby gaining unauthorized access. Existing defense mechanisms against backdoor attacks still face challenges in precisely identifying and mitigating poisoned images without compromising data utility, which undermines the overall reliability of the system. We propose a novel and generalizable approach, TrueBiometric: Trustworthy Biometrics, which accurately detects poisoned images using a majority voting mechanism leveraging multiple state-of-the-art large vision language models. Once identified, poisoned samples are corrected using targeted and calibrated corrective noise. Our extensive empirical results demonstrate that TrueBiometric detects and corrects poisoned images with 100\% accuracy without compromising accuracy on clean images. Compared to existing state-of-the-art approaches, TrueBiometric offers a more practical, accurate, and effective solution for mitigating backdoor attacks in face recognition systems.


【4】A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
标题:用于实现非言语语音生成和理解的可扩展管道
链接:https://arxiv.org/abs/2508.05385

作者:Runchuan Ye, Yixuan Zhou, Renjie Yu, Zijian Lin, Kehan Li, Xiang Li, Xin Liu, Guoyang Zeng, Zhiyong Wu
摘要:人类的口语交流不仅涉及词汇内容,还包括非言语发声(NV),如笑声,叹息和咳嗽,这些发声传达情感,意图和社会信号。然而,大多数现有的语音系统仅关注语言内容,缺乏理解和生成这种非语言提示的能力,从而降低了口语界面的情商和交流丰富性。在这项工作中,我们介绍了$\textbf{NonVerbalSpeech-38 K}$,一个用于非语言语音生成和理解的大型且多样化的数据集,从真实世界的媒体中收集并使用自动管道进行注释。该数据集包含38,718个样本(约131小时),包含10类非语言线索,如笑声,嗅探和清嗓子。我们通过微调最先进的模型(包括F5-TTS和Qwen 2-Audio)进一步验证了数据集,证明了其在非言语语音生成和理解任务中的有效性。我们的贡献有三个方面:(1)我们提出了一个实用的管道来构建自然和多样化的非语言语音数据集;(2)我们发布了一个大规模的数据集,以推进非语言语音生成和理解的研究;(3)我们通过展示非语言语音合成和字幕的改进来验证数据集的有效性,从而促进更丰富的人机交互。
摘要:Human spoken communication involves not only lexical content but also non-verbal vocalizations (NVs) such as laughter, sighs, and coughs, which convey emotions, intentions, and social signals. However, most existing speech systems focus solely on verbal content and lack the ability to understand and generate such non-verbal cues, reducing the emotional intelligence and communicative richness of spoken interfaces. In this work, we introduce $\textbf{NonVerbalSpeech-38K}$, a large and diverse dataset for non-verbal speech generation and understanding, collected from real-world media and annotated using an automatic pipeline. The dataset contains 38,718 samples (about 131 hours) with 10 categories of non-verbal cues, such as laughter, sniff, and throat clearing. We further validate the dataset by fine-tuning state-of-the-art models, including F5-TTS and Qwen2-Audio, demonstrating its effectiveness in non-verbal speech generation and understanding tasks. Our contributions are threefold: (1) We propose a practical pipeline for building natural and diverse non-verbal speech datasets; (2) We release a large-scale dataset to advance research on non-verbal speech generation and understanding; (3) We validate the dataset's effectiveness by demonstrating improvements in both non-verbal speech synthesis and captioning, thereby facilitating richer human-computer interaction.


【5】Estimating Musical Surprisal from Audio in Autoregressive Diffusion Model Noise Spaces
标题:自回归扩散模型噪音空间中的音频估计音乐惊喜
链接:https://arxiv.org/abs/2508.05306

作者:Mathias Rose Bjare, Stefan Lattner, Gerhard Widmer
备注:9 pages, 1 figure, 5 tables. Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR), Daejeon, South Korea, 2025 2025
摘要:最近,信息内容(IC)的预测,从一个生成的无限词汇Transformer(GIVT)已被用来模拟音乐的预期和音频中的发音。我们调查这种建模的有效性,使用自回归扩散模型(ADM)计算的IC。我们的经验表明,IC估计模型的基础上,两个不同的扩散常微分方程(ODE)描述不同的数据更好,在负对数似然,比GIVT。我们通过检查两个任务来评估扩散模型IC在捕获音调方面的有效性:(1)捕获单声道音高音调,以及(2)检测多声道音频中的段边界。在这两项任务中,扩散模型匹配或超过GIVT的性能。我们假设,在不同的扩散过程噪声水平估计对应于音乐和音频特征存在于不同的音频粒度上的噪声。为了验证我们的假设,我们发现,对于适当的噪声水平,所研究的音乐模拟任务的结果有所改善。代码在github.com/SonyCSLParis/audioic上提供。
摘要:Recently, the information content (IC) of predictions from a Generative Infinite-Vocabulary Transformer (GIVT) has been used to model musical expectancy and surprisal in audio. We investigate the effectiveness of such modelling using IC calculated with autoregressive diffusion models (ADMs). We empirically show that IC estimates of models based on two different diffusion ordinary differential equations (ODEs) describe diverse data better, in terms of negative log-likelihood, than a GIVT. We evaluate diffusion model IC's effectiveness in capturing surprisal aspects by examining two tasks: (1) capturing monophonic pitch surprisal, and (2) detecting segment boundaries in multi-track audio. In both tasks, the diffusion models match or exceed the performance of a GIVT. We hypothesize that the surprisal estimated at different diffusion process noise levels corresponds to the surprisal of music and audio features present at different audio granularities. Testing our hypothesis, we find that, for appropriate noise levels, the studied musical surprisal tasks' results improve. Code is provided on github.com/SonyCSLParis/audioic.


【6】SpectroStream: A Versatile Neural Codec for General Audio
标题:SpectroStream:通用音频的通用神经编解码器
链接:https://arxiv.org/abs/2508.05207

作者:Yunpeng Li, Kehang Han, Brian McWilliams, Zalan Borsos, Marco Tagliasacchi
摘要:我们提出了SpectroStream,一个全频带多通道神经音频编解码器。SpectroStream是成熟的SoundStream的继承者,它将其功能扩展到24 kHz单声道音频之外,并能够以4- 16 kbps的比特率高质量地重建48 kHz立体声音乐。这是通过一种新的神经架构来实现的,该架构利用了时频域中的音频表示,从而带来了更好的音频质量,特别是在更高的采样率下。该模型还使用延迟融合策略来处理多声道音频,这对于平衡每个声道的声学质量和跨声道相位一致性至关重要。
摘要:We propose SpectroStream, a full-band multi-channel neural audio codec. Successor to the well-established SoundStream, SpectroStream extends its capability beyond 24 kHz monophonic audio and enables high-quality reconstruction of 48 kHz stereo music at bit rates of 4--16 kbps. This is accomplished with a new neural architecture that leverages audio representation in the time-frequency domain, which leads to better audio quality especially at higher sample rate. The model also uses a delayed-fusion strategy to handle multi-channel audio, which is crucial in balancing per-channel acoustic quality and cross-channel phase consistency.


【7】RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
标题:RAP:基于视频扩散Transformer的实时音频驱动人像动画
链接:https://arxiv.org/abs/2508.05115

作者:Fangyu Du, Taiqing Li, Ziwei Zhang, Qian Qiao, Tan Yu, Dingcheng Zhen, Xu Jia, Yang Yang, Shunshun Yin, Siyuan Liu
备注:11 pages, 9 figures
摘要:音频驱动的人像动画旨在从输入音频信号和单个参考图像合成逼真和自然的说话头部视频。虽然现有方法通过利用高维中间表示和显式建模运动动力学来实现高质量的结果,但其计算复杂性使其不适合实时部署。实时推理施加了严格的延迟和内存限制,通常需要使用高度压缩的潜在表示。然而,在这样紧凑的空间中操作阻碍了细粒度时空细节的保存,从而使视听同步RAP(实时音频驱动的肖像动画)复杂化,RAP是用于在实时约束下生成高质量说话肖像的统一框架。具体来说,RAP引入了一种混合注意力机制,用于细粒度的音频控制,以及一种静态-动态训练-推理范式,避免了显式的运动监督。通过这些技术,RAP实现了精确的音频驱动控制,减轻了长期的时间漂移,并保持了较高的视觉保真度。大量的实验表明,RAP实现国家的最先进的性能,同时在实时约束下运行。
摘要:Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional intermediate representations and explicitly modeling motion dynamics, their computational complexity renders them unsuitable for real-time deployment. Real-time inference imposes stringent latency and memory constraints, often necessitating the use of highly compressed latent representations. However, operating in such compact spaces hinders the preservation of fine-grained spatiotemporal details, thereby complicating audio-visual synchronization RAP (Real-time Audio-driven Portrait animation), a unified framework for generating high-quality talking portraits under real-time constraints. Specifically, RAP introduces a hybrid attention mechanism for fine-grained audio control, and a static-dynamic training-inference paradigm that avoids explicit motion supervision. Through these techniques, RAP achieves precise audio-driven control, mitigates long-term temporal drift, and maintains high visual fidelity. Extensive experiments demonstrate that RAP achieves state-of-the-art performance while operating under real-time constraints.


【8】Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation
标题:迈向无幻觉音乐:可靠歌曲生成的强化学习偏好优化框架
链接:https://arxiv.org/abs/2508.05011

作者:Huaicheng Zhang, Wei Tan, Guangzheng Li, Yixuan Zhang, Hangting Chen, Shun Lei, Chenyu Yang, Zhiyong Wu, Shuai Wang, Qijun Huang, Dong Yu
摘要:基于音频的生成语言模型的最新进展加速了AI驱动的歌词到歌曲生成。然而,这些模型经常遭受内容幻觉,产生与输入歌词不一致的输出,破坏音乐的连贯性。目前的监督微调(SFT)的方法,被动标签拟合的限制,表现出有限的自我改善和不良的幻觉缓解。为了解决这一核心挑战,我们提出了一种新的强化学习(RL)框架,利用偏好优化进行幻觉控制。我们的主要贡献包括:(1)开发一个强大的幻觉偏好数据集,该数据集通过音素错误率(PER)计算和基于规则的过滤来构建,以捕获与人类期望的一致性;(2)在RL框架内实现和评估三种不同的偏好优化策略:直接偏好优化(DPO),邻近策略优化(PPO)和组相对策略优化(GRPO)。DPO在政策外运行,以提高正令牌的可能性,实现了7.4%的PER显著降低。PPO和GRPO采用基于策略的方法,训练基于PER的奖励模型,通过奖励最大化和KL正则化迭代优化序列,分别产生4.9%和4.7%的PER降低。全面的客观和主观评估证实,我们的方法有效地抑制幻觉,同时保持音乐质量。至关重要的是,这项工作提出了一个系统的,基于RL的解决方案,幻觉控制在歌词到歌曲的一代。该框架的可移植性也解锁了音乐风格的坚持和音乐性增强的潜力,为未来的生成歌曲研究开辟了新的途径。
摘要:Recent advances in audio-based generative language models have accelerated AI-driven lyric-to-song generation. However, these models frequently suffer from content hallucination, producing outputs misaligned with the input lyrics and undermining musical coherence. Current supervised fine-tuning (SFT) approaches, limited by passive label-fitting, exhibit constrained self-improvement and poor hallucination mitigation. To address this core challenge, we propose a novel reinforcement learning (RL) framework leveraging preference optimization for hallucination control. Our key contributions include: (1) Developing a robust hallucination preference dataset constructed via phoneme error rate (PER) computation and rule-based filtering to capture alignment with human expectations; (2) Implementing and evaluating three distinct preference optimization strategies within the RL framework: Direct Preference Optimization (DPO), Proximal Policy Optimization (PPO), and Group Relative Policy Optimization (GRPO). DPO operates off-policy to enhance positive token likelihood, achieving a significant 7.4% PER reduction. PPO and GRPO employ an on-policy approach, training a PER-based reward model to iteratively optimize sequences via reward maximization and KL-regularization, yielding PER reductions of 4.9% and 4.7%, respectively. Comprehensive objective and subjective evaluations confirm that our methods effectively suppress hallucinations while preserving musical quality. Crucially, this work presents a systematic, RL-based solution to hallucination control in lyric-to-song generation. The framework's transferability also unlocks potential for music style adherence and musicality enhancement, opening new avenues for future generative song research.


【9】Pitch Accent Detection improves Pretrained Automatic Speech Recognition
标题:音调口音检测改进了预训练的自动语音识别
链接:https://arxiv.org/abs/2508.04814

作者:David Sasu, Natalie Schluter
摘要:我们展示了使用半监督语音表示的自动语音识别(ASR)系统的性能可以通过引入联合ASR和音高重音检测模型来提高一个互补的音高重音检测模块。我们模型的音高重音检测组件在最先进的任务上实现了显着的改进,将F1分数的差距缩小了41%。此外,在有限的资源微调下,联合训练中的ASR性能将LibriSpeech上的WER降低了28.3%。通过这些结果,我们展示了扩展预训练语音模型以保留或重新学习重要韵律线索(如音高口音)的重要性。
摘要:We show the performance of Automatic Speech Recognition (ASR) systems that use semi-supervised speech representations can be boosted by a complimentary pitch accent detection module, by introducing a joint ASR and pitch accent detection model. The pitch accent detection component of our model achieves a significant improvement on the state-of-the-art for the task, closing the gap in F1-score by 41%. Additionally, the ASR performance in joint training decreases WER by 28.3% on LibriSpeech, under limited resource fine-tuning. With these results, we show the importance of extending pretrained speech models to retain or re-learn important prosodic cues such as pitch accent.


【10】Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM
标题:利用冻结LLM的说话人特征增强对话注释
链接:https://arxiv.org/abs/2508.04795

作者:Thomas Thebaud, Yen-Ju Lu, Matthew Wiesner, Peter Viechnicki, Najim Dehak
备注:Accepted in the 2025 IEEE Automatic Speech Recognition and Understanding Workshop
摘要:在对话转录管道中,大型语言模型(LLM)经常用于后处理以改善语法,标点符号和可读性。我们探索了一个补充的后处理步骤:通过添加年龄、性别和情感等说话者特征的元数据标签来丰富转录的对话。有些标签对于整个对话是全局的,而有些则是随时间变化的。我们的方法将冻结的音频基础模型(如Whisper或WavLM)与冻结的LLAMA语言模型相结合,以推断这些扬声器属性,而无需对任何模型进行特定于任务的微调。使用轻量级,高效的连接器来桥接音频和语言表示,我们在保持模块化和速度的同时,在扬声器分析任务上实现了有竞争力的性能。此外,我们证明了冻结LLAMA模型可以直接比较x向量,在某些情况下实现8.8%的等错误率。
摘要:In dialogue transcription pipelines, Large Language Models (LLMs) are frequently employed in post-processing to improve grammar, punctuation, and readability. We explore a complementary post-processing step: enriching transcribed dialogues by adding metadata tags for speaker characteristics such as age, gender, and emotion. Some of the tags are global to the entire dialogue, while some are time-variant. Our approach couples frozen audio foundation models, such as Whisper or WavLM, with a frozen LLAMA language model to infer these speaker attributes, without requiring task-specific fine-tuning of either model. Using lightweight, efficient connectors to bridge audio and language representations, we achieve competitive performance on speaker profiling tasks while preserving modularity and speed. Additionally, we demonstrate that a frozen LLAMA model can compare x-vectors directly, achieving an Equal Error Rate of 8.8% in some scenarios.


【11】Wearable Music2Emotion : Assessing Emotions Induced by AI-Generated Music through Portable EEG-fNIRS Fusion
标题:可穿戴音乐2情感:通过便携式EEG-fNIRS融合评估人工智能生成的音乐引发的情感
链接:https://arxiv.org/abs/2508.04723

作者: Sha Zhao, Song Yi, Yangxuan Zhou, Jiadong Pan, Jiquan Wang, Jie Xia, Shijian Li, Shurong Dong, Gang Pan
备注:Accepted by ACM MM 2025
摘要:情绪严重影响心理健康,通过脑机接口技术的神经生理信号驱动对基于音乐的情感计算的兴趣。虽然以前的研究利用音乐的情感诱导的可访问性,三个关键的限制仍然存在:\textbf{(1)刺激约束}:音乐刺激仅限于小语料库,由于版权和策展成本,与启发式情感音乐映射的选择偏见,忽略了个人的情感档案。\textbf{(2)模态特异性}:过度依赖单峰神经数据(例如,EEG)忽略了来自跨模态信号融合的互补见解。textbf{(3)可移植性限制}:繁琐的设置(例如,64+通道基于凝胶的EEG帽)由于程序复杂性和便携性障碍而阻碍了现实世界的适用性。为了解决这些局限性,我们提出了MEEtBrain,一种用于情绪分析(效价/唤醒)的便携式多模态框架,通过无线头带将AI生成的音乐刺激与同步的EEG-fNIRS采集集成在一起。通过MEEtBrain,音乐刺激可以由AI大规模自动生成,消除主观选择偏见,同时确保音乐的多样性。我们使用我们开发的便携式设备,该设备设计为轻型头带式,使用干电极,同时收集EEG和fNIRS记录。在第一次招募中收集了来自20名参与者的14小时数据集,以验证该框架的有效性,人工智能生成的音乐引发目标情绪(效价/唤醒)。我们正在积极扩展我们的多模态数据集(最新数据集有44个参与者),并将其公开,以促进进一步的研究和实际应用。\textbf{该数据集可在https://zju-bmi-lab.github.io/ZBra上获得。
摘要:Emotions critically influence mental health, driving interest in music-based affective computing via neurophysiological signals with Brain-computer Interface techniques. While prior studies leverage music's accessibility for emotion induction, three key limitations persist: \textbf{(1) Stimulus Constraints}: Music stimuli are confined to small corpora due to copyright and curation costs, with selection biases from heuristic emotion-music mappings that ignore individual affective profiles. \textbf{(2) Modality Specificity}: Overreliance on unimodal neural data (e.g., EEG) ignores complementary insights from cross-modal signal fusion.\textbf{ (3) Portability Limitation}: Cumbersome setups (e.g., 64+ channel gel-based EEG caps) hinder real-world applicability due to procedural complexity and portability barriers. To address these limitations, we propose MEEtBrain, a portable and multimodal framework for emotion analysis (valence/arousal), integrating AI-generated music stimuli with synchronized EEG-fNIRS acquisition via a wireless headband. By MEEtBrain, the music stimuli can be automatically generated by AI on a large scale, eliminating subjective selection biases while ensuring music diversity. We use our developed portable device that is designed in a lightweight headband-style and uses dry electrodes, to simultaneously collect EEG and fNIRS recordings. A 14-hour dataset from 20 participants was collected in the first recruitment to validate the framework's efficacy, with AI-generated music eliciting target emotions (valence/arousal). We are actively expanding our multimodal dataset (44 participants in the latest dataset) and make it publicly available to promote further research and practical applications. \textbf{The dataset is available at https://zju-bmi-lab.github.io/ZBra.


【12】Toward Low-Latency End-to-End Voice Agents for Telecommunications Using Streaming ASR, Quantized LLMs, and Real-Time TTS
标题:使用流媒体ZR、量化LLM和实时TTC实现电信低延迟端到端语音代理
链接:https://arxiv.org/abs/2508.04721

作者:Vignesh Ethiraj, Ashwath David, Sidhanth Menon, Divya Vijay
摘要:我们推出了一个低延迟的电信AI语音代理管道,用于实时、交互式电信用途,为呼叫中心自动化、智能IVR(交互式语音响应)和AI驱动的客户支持提供先进的语音AI。该解决方案是为电信构建的,结合了NetoAI的四个专用模型:TLAM,一个4位量化的特定于语音的大型语言模型(LLM); T-VEC,一个特定于语音的嵌入模型; TTE,一个特定于语音的自动语音识别(ASR)模型;和T-Synth,一个特定于语音的文本到语音(TTS)模型。这些模型能够实现高度响应、适应领域的语音AI代理,支持以知识为基础的低延迟语音交互。该管道集成了流ASR(TTE)、会话智能(TLAM)、电信文档检索增强生成(RAG)和实时TTS(T-Synth),为电信语音助理树立了新的基准。为了评估该系统,我们建立了一个数据集的500人记录的电信问题的RFC,模拟真实的电信代理查询。该框架允许跨堆栈分析延迟、域相关性和实时性能。结果表明,TLAM、TTE和T-Synth提供低于1.0的实时因子(RTF),支持企业低延迟电信部署。这些人工智能代理-由TLAM,TTE和T-Synth提供支持-为下一代电信人工智能提供基础,实现自动化客户支持,诊断等。
摘要:We introduce a low-latency telecom AI voice agent pipeline for real-time, interactive telecommunications use, enabling advanced voice AI for call center automation, intelligent IVR (Interactive Voice Response), and AI-driven customer support. The solution is built for telecom, combining four specialized models by NetoAI: TSLAM, a 4-bit quantized Telecom-Specific Large Language Model (LLM); T-VEC, a Telecom-Specific Embedding Model; TTE, a Telecom-Specific Automatic Speech Recognition (ASR) model; and T-Synth, a Telecom-Specific Text-to-Speech (TTS) model. These models enable highly responsive, domain-adapted voice AI agents supporting knowledge-grounded spoken interactions with low latency. The pipeline integrates streaming ASR (TTE), conversational intelligence (TSLAM), retrieval augmented generation (RAG) over telecom documents, and real-time TTS (T-Synth), setting a new benchmark for telecom voice assistants. To evaluate the system, we built a dataset of 500 human-recorded telecom questions from RFCs, simulating real telecom agent queries. This framework allows analysis of latency, domain relevance, and real-time performance across the stack. Results show that TSLAM, TTE, and T-Synth deliver real-time factors (RTF) below 1.0, supporting enterprise, low-latency telecom deployments. These AI agents -- powered by TSLAM, TTE, and T-Synth -- provide a foundation for next-generation telecom AI, enabling automated customer support, diagnostics, and more.


【13】Keyword Spotting with Hyper-Matched Filters for Small Footprint Devices
标题:使用超匹配过滤器针对小尺寸设备进行关键字定位
链接:https://arxiv.org/abs/2508.04857

作者:Yael Segal-Feldman, Ann R. Bradlow, Matthew Goldrick, Joseph Keshet
备注:pre-print
摘要:开放词汇关键词识别(KWS)是指在语音记录中检测单词或术语的任务,无论它们是否包含在训练数据中。本文介绍了一种开放式词汇表的关键字定位模型与国家的最先进的检测精度的小尺寸设备。该模型由语音编码器、目标关键词编码器和检测网络组成。语音编码器是一个微型Whisper或微型Conformer。目标关键字编码器被实现为一个超网络,该超网络将所需的关键字作为字符串,并为卷积层生成一组唯一的权重,卷积层可以被认为是关键字特定的匹配滤波器。检测网络使用匹配的过滤器权重来执行关键字特定的卷积,该卷积引导感知器模块的交叉注意机制来确定目标术语是否出现在记录中。结果表明,我们的系统达到了最先进的检测性能,并有效地推广到域外条件,包括第二语言(L2)的语音。值得注意的是,我们最小的模型只有420万个参数,可以与大几倍的模型相匹配或优于后者,从而证明了效率和鲁棒性。
摘要:Open-vocabulary keyword spotting (KWS) refers to the task of detecting words or terms within speech recordings, regardless of whether they were included in the training data. This paper introduces an open-vocabulary keyword spotting model with state-of-the-art detection accuracy for small-footprint devices. The model is composed of a speech encoder, a target keyword encoder, and a detection network. The speech encoder is either a tiny Whisper or a tiny Conformer. The target keyword encoder is implemented as a hyper-network that takes the desired keyword as a character string and generates a unique set of weights for a convolutional layer, which can be considered as a keyword-specific matched filter. The detection network uses the matched-filter weights to perform a keyword-specific convolution, which guides the cross-attention mechanism of a Perceiver module in determining whether the target term appears in the recording. The results indicate that our system achieves state-of-the-art detection performance and generalizes effectively to out-of-domain conditions, including second-language (L2) speech. Notably, our smallest model, with just 4.2 million parameters, matches or outperforms models that are several times larger, demonstrating both efficiency and robustness.


eess.AS音频处理


【1】Investigation of Speech and Noise Latent Representations in Single-channel VAE-based Speech Enhancement
标题:单通道基于VAE的语音增强中语音和噪音潜在表示的研究
链接:https://arxiv.org/abs/2508.05293

作者:Jiatong Li, Simon Doclo
备注:5 pages, 5 figures
摘要:最近,已经提出了一种基于变分自编码器(VAE)的单通道语音增强系统,使用贝叶斯排列训练,它使用两个预训练的VAE来获得语音和噪声的潜在表示。基于这些预训练的VAE,带噪VAE学习从带噪语音生成语音和噪声潜在表示以用于语音增强。修改预训练的VAE损失项会影响预训练的语音和噪声潜在表示。在本文中,我们研究这些不同的表示如何影响语音增强性能。在DNS 3、WSJ 0-QUT和VoiceBank-DEMAND数据集上的实验表明,语音和噪声表示清晰分离的潜在空间比产生重叠语音和噪声表示的标准VAE显著提高了性能。
摘要:Recently, a variational autoencoder (VAE)-based single-channel speech enhancement system using Bayesian permutation training has been proposed, which uses two pretrained VAEs to obtain latent representations for speech and noise. Based on these pretrained VAEs, a noisy VAE learns to generate speech and noise latent representations from noisy speech for speech enhancement. Modifying the pretrained VAE loss terms affects the pretrained speech and noise latent representations. In this paper, we investigate how these different representations affect speech enhancement performance. Experiments on the DNS3, WSJ0-QUT, and VoiceBank-DEMAND datasets show that a latent space where speech and noise representations are clearly separated significantly improves performance over standard VAEs, which produce overlapping speech and noise representations.


【2】Privacy Disclosure of Similarity in Speech and Language Processing
标题:语音和语言处理中相似性的隐私披露
链接:https://arxiv.org/abs/2508.05250

作者:Tom Backstr, Mohammad Hassan Vali, My Nguyen, Silas Rech
摘要:说话者、作者和其他生物特征识别应用程序经常将样本的相似性与模板数据库进行比较以确定身份。考虑到数据可能是嘈杂的,并且相似性度量可能是不准确的,这样的比较可能无法可靠地将真实身份识别为最相似的。尽管如此,即使是基于不准确的相似性度量的相似性排名也可能泄露关于真实身份的私人信息。我们提出了一种方法来量化这样的相似性排名的隐私披露估计其概率分布。它基于确定真实说话人的相似性等级的直方图,或者当数据稀缺时,用β-二项分布建模直方图。我们用熵(比特)来表达本公开,使得来自独立特征的公开是加性的。我们的实验表明,所有测试的扬声器和作者的特征包含个人识别信息(PII),可以帮助识别,与嵌入从扬声器识别算法包含最多的信息,其次是电话嵌入,语言嵌入,和基频。我们的初步实验表明,PII的披露随着测试样本的长度而增加,但它受到数据库模板长度的限制。所提供的度量(相似性排名公开)提供了一种比较生物特征之间的PII公开并将其合并以帮助识别的方式。因此,它可以帮助全面评估语音和其他生物识别技术对隐私的威胁。
摘要:Speaker, author, and other biometric identification applications often compare a sample's similarity to a database of templates to determine the identity. Given that data may be noisy and similarity measures can be inaccurate, such a comparison may not reliably identify the true identity as the most similar. Still, even the similarity rank based on an inaccurate similarity measure can disclose private information about the true identity. We propose a methodology for quantifying the privacy disclosure of such a similarity rank by estimating its probability distribution. It is based on determining the histogram of the similarity rank of the true speaker, or when data is scarce, modeling the histogram with the beta-binomial distribution. We express the disclosure in terms of entropy (bits), such that the disclosure from independent features are additive. Our experiments demonstrate that all tested speaker and author characterizations contain personally identifying information (PII) that can aid in identification, with embeddings from speaker recognition algorithms containing the most information, followed by phone embeddings, linguistic embeddings, and fundamental frequency. Our initial experiments show that the disclosure of PII increases with the length of test samples, but it is bounded by the length of database templates. The provided metric, similarity rank disclosure, provides a way to compare the disclosure of PII between biometric features and merge them to aid identification. It can thus aid in the holistic evaluation of threats to privacy in speech and other biometric technologies.


【3】Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
标题:低资源场景下的语音LLM:数据量要求和预训练对高资源语言的影响
链接:https://arxiv.org/abs/2508.05149

作者:Seraphina Fong, Marco Matassoni, Alessio Brutti
备注:Accepted at Interspeech 2025. 5 pages, 2 figures, 3 tables
摘要:大型语言模型(LLM)在处理高资源语言的口语输入方面表现出了潜力,在各种任务中达到了最先进的性能。然而,在资源匮乏的环境中,对这些方法的适用性探讨得仍然较少。这项工作研究了使用SLAM-ASR框架将语音LLM用于低资源自动语音识别,其中可训练的轻量级投影仪连接语音编码器和LLM。首先,我们评估训练数据量要求,以匹配仅Whisper性能,再次强调有限数据的挑战。其次,我们证明了利用在高资源语言上预先训练的单语言或多语言投影仪可以减少数据稀缺的影响,特别是在小训练集的情况下。使用多语言LLM(EuroLLM,Salamandra)与whisper-large-v3-turbo,我们评估了几个公共基准测试的性能,为未来研究优化低资源语言和多语言的语音LLM提供了见解。
摘要:Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art performance in various tasks. However, their applicability is still less explored in low-resource settings. This work investigates the use of Speech LLMs for low-resource Automatic Speech Recognition using the SLAM-ASR framework, where a trainable lightweight projector connects a speech encoder and a LLM. Firstly, we assess training data volume requirements to match Whisper-only performance, re-emphasizing the challenges of limited data. Secondly, we show that leveraging mono- or multilingual projectors pretrained on high-resource languages reduces the impact of data scarcity, especially with small training sets. Using multilingual LLMs (EuroLLM, Salamandra) with whisper-large-v3-turbo, we evaluate performance on several public benchmarks, providing insights for future research on optimizing Speech LLMs for low-resource languages and multilinguality.


【4】Fairness in Dysarthric Speech Synthesis: Understanding Intrinsic Bias in Dysarthric Speech Cloning using F5-TTS
标题:构音障碍语音合成中的公平性:使用F5-TTS理解构音障碍语音克隆中的内在偏差
链接:https://arxiv.org/abs/2508.05102

作者:Anuprabha M, Krishna Gurugubelli, Anil Kumar Vuppala
备注:Accepted at Interspeech 2025
摘要:构音障碍性言语在开发辅助技术方面构成了重大挑战,主要是由于数据的可用性有限。神经语音合成方面的最新进展,尤其是零射击语音克隆(zero-shot voice cloning),促进了用于数据增强的合成语音生成;然而,它们可能会引入对构音障碍语音的偏见。在本文中,我们调查的有效性,最先进的F5-TTS克隆构音障碍的语音使用TORGO数据集,专注于可懂度,说话人相似性和韵律保存。我们还使用公平性指标(如差异影响和奇偶性差异)分析潜在的偏差,以评估构音障碍严重程度之间的差异。结果表明,在构音障碍语音合成中,F5-TTS对说话人的语音可懂度和韵律保留表现出强烈的偏好。这项研究的见解可以帮助整合公平意识的构音障碍语音合成,促进更具包容性的语音技术的进步。
摘要:Dysarthric speech poses significant challenges in developing assistive technologies, primarily due to the limited availability of data. Recent advances in neural speech synthesis, especially zero-shot voice cloning, facilitate synthetic speech generation for data augmentation; however, they may introduce biases towards dysarthric speech. In this paper, we investigate the effectiveness of state-of-the-art F5-TTS in cloning dysarthric speech using TORGO dataset, focusing on intelligibility, speaker similarity, and prosody preservation. We also analyze potential biases using fairness metrics like Disparate Impact and Parity Difference to assess disparities across dysarthric severity levels. Results show that F5-TTS exhibits a strong bias toward speech intelligibility over speaker and prosody preservation in dysarthric speech synthesis. Insights from this study can help integrate fairness-aware dysarthric speech synthesis, fostering the advancement of more inclusive speech technologies.


【5】MOVER: Combining Multiple Meeting Recognition Systems
标题:MOPER:结合多个会议识别系统
链接:https://arxiv.org/abs/2508.05055

作者:Naoyuki Kamo, Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani
摘要:在本文中,我们提出了会议识别器输出投票错误减少(MOVER),一种新的系统组合方法,会议识别任务。虽然有方法来组合日记化的输出(例如,DOVER)或自动语音识别(ASR)系统(例如,ROVER),MOVER是第一种可以结合会议识别系统的输出的方法,这些会议识别系统在日志化和ASR方面都不同。MOVER通过说话人对齐、分段分组、单词和时间组合等五个阶段将具有不同时间间隔和说话人标签的假设组合在一起。CHiME-8 DASR任务和NOTSOFAR-1任务的多通道跟踪上的实验结果表明,MOVER可以成功地将具有不同日记和识别输出的多个会议识别系统组合在一起,对于这两项任务,相对于现有技术的系统,实现了9.55%和8.51%的相对TcpWER改进。
摘要:In this paper, we propose Meeting recognizer Output Voting Error Reduction (MOVER), a novel system combination method for meeting recognition tasks. Although there are methods to combine the output of diarization (e.g., DOVER) or automatic speech recognition (ASR) systems (e.g., ROVER), MOVER is the first approach that can combine the outputs of meeting recognition systems that differ in terms of both diarization and ASR. MOVER combines hypotheses with different time intervals and speaker labels through a five-stage process that includes speaker alignment, segment grouping, word and timing combination, etc. Experimental results on the CHiME-8 DASR task and the multi-channel track of the NOTSOFAR-1 task demonstrate that MOVER can successfully combine multiple meeting recognition systems with diverse diarization and recognition outputs, achieving relative tcpWER improvements of 9.55 % and 8.51 % over the state-of-the-art systems for both tasks.


【6】REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers
标题:REC-VC:采用扩散Transformer的稳健、表现力和快速的Zero-Shot语音转换
链接:https://arxiv.org/abs/2508.04996

作者:Yuepeng Jiang, Ziqian Ning, Shuai Wang, Chengjia Wang, Mengxiao Bi, Pengcheng Zhu, Lei Xie, Zhonghua Fu
摘要:在现实世界的语音转换应用中,源语音中的环境噪声和用户对表达性输出的需求构成了关键挑战。传统的基于ASR的方法确保了噪声鲁棒性,但抑制了韵律,而基于SSL的模型提高了表现力,但遭受音色泄漏和噪声敏感性。本文提出了REF-VC,一个抗噪声的表达语音转换系统。主要创新包括:(1)随机擦除策略,以减轻SSL特征中固有的信息冗余,增强噪声鲁棒性和表达能力;(2)受E2 TTS启发的隐式对齐,以抑制非必要特征重建;(3)集成可重构模型,以加速流匹配推理,显著减少到4个步骤。实验结果表明,我们的模型优于基线,如种子VC在zero-shot场景中的噪声集,同时也执行种子VC的清洁集。此外,REF-VC可以在一个模型中兼容歌声转换。
摘要:In real-world voice conversion applications, environmental noise in source speech and user demands for expressive output pose critical challenges. Traditional ASR-based methods ensure noise robustness but suppress prosody, while SSL-based models improve expressiveness but suffer from timbre leakage and noise sensitivity. This paper proposes REF-VC, a noise-robust expressive voice conversion system. Key innovations include: (1) A random erasing strategy to mitigate the information redundancy inherent in SSL feature, enhancing noise robustness and expressiveness; (2) Implicit alignment inspired by E2TTS to suppress non-essential feature reconstruction; (3) Integration of Shortcut Models to accelerate flow matching inference, significantly reducing to 4 steps. Experimental results demonstrate that our model outperforms baselines such as Seed-VC in zero-shot scenarios on the noisy set, while also performing comparably to Seed-VC on the clean set. In addition, REF-VC can be compatible with singing voice conversion within one model.


【7】Closed-Form Successive Relative Transfer Function Vector Estimation based on Blind Oblique Projection Incorporating Noise Whitening
标题:基于盲斜投影消除噪音白化的封闭式连续相对传递函数载体估计
链接:https://arxiv.org/abs/2508.04887

作者:Henri Gode, Simon Doclo
摘要:声源的相对传递函数(RTF)在波束形成中起着至关重要的作用,可以有效地抑制噪声和干扰。本文解决了在线估计的RTF向量的多个声源在嘈杂和混响的环境中,为特定的场景,源激活连续的挑战。虽然可以直接估计第一个源的RTF向量,但在多个源同时活动的段期间估计后续源的RTF向量时出现了主要挑战。盲斜投影(BOP)方法已被提出来估计RTF矢量的一个新的激活源,通过最佳地阻止这个源。然而,这种方法面临着几个限制:由于其依赖于迭代梯度下降优化而导致的高计算复杂度,引入随机附加向量,这可能会对性能产生负面影响,以及假设高信噪比(SNR)。为了克服这些限制,在本文中,我们提出了三个扩展的BOP方法。首先,我们推导出一个封闭形式的解决方案,优化BOP成本函数,显着降低计算复杂性。其次,我们引入正交附加向量代替随机向量,提高RTF向量估计精度。第三,我们将噪声处理技术的灵感来自协方差减法和白化,在低信噪比条件下增加鲁棒性。为了提供一个逐帧的源活动模式的估计,所需的传统BOP方法和所提出的方法,我们提出了一个基于空间相干性的在线源计数方法。模拟与现实世界的混响嘈杂的录音,具有3个连续激活扬声器,有和没有先验知识的源活动模式。
摘要:Relative transfer functions (RTFs) of sound sources play a crucial role in beamforming, enabling effective noise and interference suppression. This paper addresses the challenge of online estimating the RTF vectors of multiple sound sources in noisy and reverberant environments, for the specific scenario where sources activate successively. While the RTF vector of the first source can be estimated straightforwardly, the main challenge arises in estimating the RTF vectors of subsequent sources during segments where multiple sources are simultaneously active. The blind oblique projection (BOP) method has been proposed to estimate the RTF vector of a newly activating source by optimally blocking this source. However, this method faces several limitations: high computational complexity due to its reliance on iterative gradient descent optimization, the introduction of random additional vectors, which can negatively impact performance, and the assumption of high signal-to-noise ratio (SNR). To overcome these limitations, in this paper we propose three extensions to the BOP method. First, we derive a closed-form solution for optimizing the BOP cost function, significantly reducing computational complexity. Second, we introduce orthogonal additional vectors instead of random vectors, enhancing RTF vector estimation accuracy. Third, we incorporate noise handling techniques inspired by covariance subtraction and whitening, increasing robustness in low SNR conditions. To provide a frame-by-frame estimate of the source activity pattern, required by both the conventional BOP method and the proposed method, we propose a spatial-coherence-based online source counting method. Simulations are performed with real-world reverberant noisy recordings featuring 3 successively activating speakers, with and without a-priori knowledge of the source activity pattern.


【8】Keyword Spotting with Hyper-Matched Filters for Small Footprint Devices
标题:使用超匹配过滤器针对小尺寸设备进行关键字定位
链接:https://arxiv.org/abs/2508.04857

作者:Yael Segal-Feldman, Ann R. Bradlow, Matthew Goldrick, Joseph Keshet
备注:pre-print
摘要:开放词汇关键词识别(KWS)是指在语音记录中检测单词或术语的任务,无论它们是否包含在训练数据中。本文介绍了一种开放式词汇表的关键字定位模型与国家的最先进的检测精度的小尺寸设备。该模型由语音编码器、目标关键词编码器和检测网络组成。语音编码器是一个微型Whisper或微型Conformer。目标关键字编码器被实现为一个超网络,该超网络将所需的关键字作为字符串,并为卷积层生成一组唯一的权重,卷积层可以被认为是关键字特定的匹配滤波器。检测网络使用匹配的过滤器权重来执行关键字特定的卷积,该卷积引导感知器模块的交叉注意机制来确定目标术语是否出现在记录中。结果表明,我们的系统达到了最先进的检测性能,并有效地推广到域外条件,包括第二语言(L2)的语音。值得注意的是,我们最小的模型,只有420万个参数,匹配或优于几倍大的模型,证明了效率和鲁棒性。
摘要:Open-vocabulary keyword spotting (KWS) refers to the task of detecting words or terms within speech recordings, regardless of whether they were included in the training data. This paper introduces an open-vocabulary keyword spotting model with state-of-the-art detection accuracy for small-footprint devices. The model is composed of a speech encoder, a target keyword encoder, and a detection network. The speech encoder is either a tiny Whisper or a tiny Conformer. The target keyword encoder is implemented as a hyper-network that takes the desired keyword as a character string and generates a unique set of weights for a convolutional layer, which can be considered as a keyword-specific matched filter. The detection network uses the matched-filter weights to perform a keyword-specific convolution, which guides the cross-attention mechanism of a Perceiver module in determining whether the target term appears in the recording. The results indicate that our system achieves state-of-the-art detection performance and generalizes effectively to out-of-domain conditions, including second-language (L2) speech. Notably, our smallest model, with just 4.2 million parameters, matches or outperforms models that are several times larger, demonstrating both efficiency and robustness.


【9】SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
标题:SPGISpeech 2.0:转录多扬声器金融音频,用于扬声器标记的转录
链接:https://arxiv.org/abs/2508.05554

作者:Raymond Grossman, Taejin Park, Kunal Dhawan, Andrew Titus, Sophia Zhi, Yulia Shchadilova, Weiqing Wang, Jagadeesh Balam, Boris Ginsburg
备注:To be presented at Interspeech 2025
摘要:我们介绍SPGISpeech 2.0,一个适合在金融领域的说话人标记的转录数据集。SPGISpeech 2.0提高了适用建模任务的多样性,同时保持了原始SPGISpeech数据集的核心特征:音频片段及其相应的完全格式化的文本翻译,可用于端到端自动语音识别(ASR)。SPGISpeech 2.0包括3,780个额外小时的专业转录收入电话。此外,数据集包含用于促进多说话者ASR的每个音频片段的呼叫和说话者信息。在对SPGISpeech 2.0进行微调后,我们通过改进流行语音识别模型的说话者标记的ASR性能来验证SPGISpeech 2.0的实用性。SPGISpeech 2.0免费发布供非商业使用,我们希望它能促进语音识别技术的进步,并激发广泛的研究应用。
摘要:We introduce SPGISpeech 2.0, a dataset suitable for speaker-tagged transcription in the financial domain. SPGISpeech 2.0 improves the diversity of applicable modeling tasks while maintaining the core characteristic of the original SPGISpeech dataset: audio snippets and their corresponding fully formatted text transcriptions, usable for end-to-end automatic speech recognition (ASR). SPGISpeech 2.0 consists of 3,780 additional hours of professionally transcribed earnings calls. Furthermore, the dataset contains call and speaker information for each audio snippet facilitating multi-talker ASR. We validate the utility of SPGISpeech 2.0 through improvements in speaker-tagged ASR performance of popular speech recognition models after fine-tuning on SPGISpeech 2.0. Released free for non-commercial use, we expect SPGISpeech 2.0 to foster advancements in speech recognition technologies and inspire a wide range of research applications.


【10】Embedding Alignment in Code Generation for Audio
标题:音频代码生成中嵌入对齐
链接:https://arxiv.org/abs/2508.05473

作者:Sam Kouteili, Hiren Madhu, George Typaldos, Mark Santolucito
摘要:LLM驱动的代码生成有可能彻底改变创造性的编码工作,例如实时编码,使用户能够专注于语法细节的结构基元。在这样的领域中,当提示LLM时,用户可以受益于考虑多个不同的代码候选以更好地实现他们的音乐意图。然而,代码生成模型难以呈现独特且多样的代码候选,而无法直接洞察代码的音频输出。为了更好地建立候选代码和生成的音频之间的关系,我们研究了代码和音频嵌入空间之间的映射的拓扑结构。我们发现,代码和音频嵌入并不表现出简单的线性关系,而是用一个构建的预测模型来补充,该模型表明可以学习嵌入对齐图。作为对音乐多样性输出目标的补充,我们提出了一个模型,该模型由给定的代码预测输出音频嵌入,构建代码-音频嵌入对齐图。
摘要:LLM-powered code generation has the potential to revolutionize creative coding endeavors, such as live-coding, by enabling users to focus on structural motifs over syntactic details. In such domains, when prompting an LLM, users may benefit from considering multiple varied code candidates to better realize their musical intentions. Code generation models, however, struggle to present unique and diverse code candidates, with no direct insight into the code's audio output. To better establish a relationship between code candidates and produced audio, we investigate the topology of the mapping between code and audio embedding spaces. We find that code and audio embeddings do not exhibit a simple linear relationship, but supplement this with a constructed predictive model that shows an embedding alignment map could be learned. Supplementing the aim for musically diverse output, we present a model that given code predicts output audio embedding, constructing a code-audio embedding alignment map.


【11】From Detection to Correction: Backdoor-Resilient Face Recognition via Vision-Language Trigger Detection and Noise-Based Neutralization
标题:从检测到纠正:通过视觉语言触发检测和基于噪音的中和来实现后门弹性人脸识别
链接:https://arxiv.org/abs/2508.05409

作者:Farah Wahida, M.A.P. Chamikara, Yashothara Shanmugarasa, Mohan Baruwal Chhetri, Thilina Ranbaduge, Ibrahim Khalil
备注:19 Pages, 24 Figures
摘要:生物识别系统,例如由深度神经网络(DNN)驱动的人脸识别系统,依赖于大型且高度敏感的数据集。后门攻击可以通过操纵训练过程来破坏这些系统。通过在一些训练图像中插入小的触发器,例如贴纸,化妆品或图案化的面具,对手可以在稍后的认证期间呈现相同的触发器,以被错误地识别为另一个人,从而获得未经授权的访问。针对后门攻击的现有防御机制在精确识别和缓解中毒图像而不影响数据实用性方面仍然面临挑战,这会破坏系统的整体可靠性。我们提出了一种新颖的和可推广的方法,TrueBiometric:值得信赖的生物识别技术,它使用多数投票机制,利用多个国家的最先进的大型视觉语言模型,准确地检测中毒图像。一旦被识别,中毒的样品使用目标和校准的校正噪声进行校正。我们广泛的实证结果表明,TrueBiometric检测和纠正中毒的图像与100\ %的准确性,而不影响准确的干净的图像。与现有的最先进的方法相比,TrueBiometric提供了一种更实用,更准确,更有效的解决方案,用于减轻人脸识别系统中的后门攻击。
摘要:Biometric systems, such as face recognition systems powered by deep neural networks (DNNs), rely on large and highly sensitive datasets. Backdoor attacks can subvert these systems by manipulating the training process. By inserting a small trigger, such as a sticker, make-up, or patterned mask, into a few training images, an adversary can later present the same trigger during authentication to be falsely recognized as another individual, thereby gaining unauthorized access. Existing defense mechanisms against backdoor attacks still face challenges in precisely identifying and mitigating poisoned images without compromising data utility, which undermines the overall reliability of the system. We propose a novel and generalizable approach, TrueBiometric: Trustworthy Biometrics, which accurately detects poisoned images using a majority voting mechanism leveraging multiple state-of-the-art large vision language models. Once identified, poisoned samples are corrected using targeted and calibrated corrective noise. Our extensive empirical results demonstrate that TrueBiometric detects and corrects poisoned images with 100\% accuracy without compromising accuracy on clean images. Compared to existing state-of-the-art approaches, TrueBiometric offers a more practical, accurate, and effective solution for mitigating backdoor attacks in face recognition systems.


【12】A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
标题:用于实现非言语语音生成和理解的可扩展管道
链接:https://arxiv.org/abs/2508.05385

作者:Runchuan Ye, Yixuan Zhou, Renjie Yu, Zijian Lin, Kehan Li, Xiang Li, Xin Liu, Guoyang Zeng, Zhiyong Wu
摘要:人类的口语交流不仅涉及词汇内容,还包括非言语发声(NV),如笑声,叹息和咳嗽,这些发声传达情感,意图和社会信号。然而,大多数现有的语音系统仅关注语言内容,缺乏理解和生成这种非语言提示的能力,从而降低了口语界面的情商和交流丰富性。在这项工作中,我们介绍了$\textbf{NonVerbalSpeech-38 K}$,一个用于非语言语音生成和理解的大型且多样化的数据集,从真实世界的媒体中收集并使用自动管道进行注释。该数据集包含38,718个样本(约131小时),包含10类非语言线索,如笑声,嗅探和清嗓子。我们通过微调最先进的模型(包括F5-TTS和Qwen 2-Audio)进一步验证了数据集,证明了其在非言语语音生成和理解任务中的有效性。我们的贡献有三个方面:(1)我们提出了一个实用的管道来构建自然和多样化的非语言语音数据集;(2)我们发布了一个大规模的数据集,以推进非语言语音生成和理解的研究;(3)我们通过展示非语言语音合成和字幕的改进来验证数据集的有效性,从而促进更丰富的人机交互。
摘要:Human spoken communication involves not only lexical content but also non-verbal vocalizations (NVs) such as laughter, sighs, and coughs, which convey emotions, intentions, and social signals. However, most existing speech systems focus solely on verbal content and lack the ability to understand and generate such non-verbal cues, reducing the emotional intelligence and communicative richness of spoken interfaces. In this work, we introduce $\textbf{NonVerbalSpeech-38K}$, a large and diverse dataset for non-verbal speech generation and understanding, collected from real-world media and annotated using an automatic pipeline. The dataset contains 38,718 samples (about 131 hours) with 10 categories of non-verbal cues, such as laughter, sniff, and throat clearing. We further validate the dataset by fine-tuning state-of-the-art models, including F5-TTS and Qwen2-Audio, demonstrating its effectiveness in non-verbal speech generation and understanding tasks. Our contributions are threefold: (1) We propose a practical pipeline for building natural and diverse non-verbal speech datasets; (2) We release a large-scale dataset to advance research on non-verbal speech generation and understanding; (3) We validate the dataset's effectiveness by demonstrating improvements in both non-verbal speech synthesis and captioning, thereby facilitating richer human-computer interaction.


【13】Estimating Musical Surprisal from Audio in Autoregressive Diffusion Model Noise Spaces
标题:自回归扩散模型噪音空间中的音频估计音乐惊喜
链接:https://arxiv.org/abs/2508.05306

作者:Mathias Rose Bjare, Stefan Lattner, Gerhard Widmer
备注:9 pages, 1 figure, 5 tables. Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR), Daejeon, South Korea, 2025 2025
摘要:最近,信息内容(IC)的预测,从一个生成的无限词汇Transformer(GIVT)已被用来模拟音乐的预期和音频中的发音。我们调查这种建模的有效性,使用自回归扩散模型(ADM)计算的IC。我们的经验表明,IC估计模型的基础上,两个不同的扩散常微分方程(ODE)描述不同的数据更好,在负对数似然,比GIVT。我们通过检查两个任务来评估扩散模型IC在捕获音调方面的有效性:(1)捕获单声道音高音调,以及(2)检测多声道音频中的段边界。在这两项任务中,扩散模型匹配或超过GIVT的性能。我们假设,在不同的扩散过程噪声水平估计对应于音乐和音频特征存在于不同的音频粒度上的噪声。检验我们的假设,我们发现,对于适当的噪音水平,所研究的音乐惊喜任务的结果会改善。代码在github.com/SonyCSLParis/audioic上提供。
摘要:Recently, the information content (IC) of predictions from a Generative Infinite-Vocabulary Transformer (GIVT) has been used to model musical expectancy and surprisal in audio. We investigate the effectiveness of such modelling using IC calculated with autoregressive diffusion models (ADMs). We empirically show that IC estimates of models based on two different diffusion ordinary differential equations (ODEs) describe diverse data better, in terms of negative log-likelihood, than a GIVT. We evaluate diffusion model IC's effectiveness in capturing surprisal aspects by examining two tasks: (1) capturing monophonic pitch surprisal, and (2) detecting segment boundaries in multi-track audio. In both tasks, the diffusion models match or exceed the performance of a GIVT. We hypothesize that the surprisal estimated at different diffusion process noise levels corresponds to the surprisal of music and audio features present at different audio granularities. Testing our hypothesis, we find that, for appropriate noise levels, the studied musical surprisal tasks' results improve. Code is provided on github.com/SonyCSLParis/audioic.


【14】SpectroStream: A Versatile Neural Codec for General Audio
标题:SpectroStream:通用音频的通用神经编解码器
链接:https://arxiv.org/abs/2508.05207

作者:Yunpeng Li, Kehang Han, Brian McWilliams, Zalan Borsos, Marco Tagliasacchi
摘要:我们提出了SpectroStream,一个全频带多通道神经音频编解码器。SpectroStream是成熟的SoundStream的继承者,它将其功能扩展到24 kHz单声道音频之外,并能够以4- 16 kbps的比特率高质量地重建48 kHz立体声音乐。这是通过一种新的神经架构来实现的,该架构利用了时频域中的音频表示,从而带来了更好的音频质量,特别是在更高的采样率下。该模型还使用延迟融合策略来处理多声道音频,这对于平衡每个声道的声学质量和跨声道相位一致性至关重要。
摘要:We propose SpectroStream, a full-band multi-channel neural audio codec. Successor to the well-established SoundStream, SpectroStream extends its capability beyond 24 kHz monophonic audio and enables high-quality reconstruction of 48 kHz stereo music at bit rates of 4--16 kbps. This is accomplished with a new neural architecture that leverages audio representation in the time-frequency domain, which leads to better audio quality especially at higher sample rate. The model also uses a delayed-fusion strategy to handle multi-channel audio, which is crucial in balancing per-channel acoustic quality and cross-channel phase consistency.


【15】RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
标题:RAP:基于视频扩散Transformer的实时音频驱动人像动画
链接:https://arxiv.org/abs/2508.05115

作者:Fangyu Du, Taiqing Li, Ziwei Zhang, Qian Qiao, Tan Yu, Dingcheng Zhen, Xu Jia, Yang Yang, Shunshun Yin, Siyuan Liu
备注:11 pages, 9 figures
摘要:音频驱动的人像动画旨在从输入音频信号和单个参考图像合成逼真和自然的说话头部视频。虽然现有方法通过利用高维中间表示和显式建模运动动力学来实现高质量的结果,但其计算复杂性使其不适合实时部署。实时推理施加了严格的延迟和内存限制,通常需要使用高度压缩的潜在表示。然而,在这样紧凑的空间中操作阻碍了细粒度时空细节的保存,从而使视听同步RAP(实时音频驱动的肖像动画)复杂化,RAP是用于在实时约束下生成高质量说话肖像的统一框架。具体来说,RAP引入了一种混合注意力机制,用于细粒度的音频控制,以及一种静态-动态训练-推理范式,避免了显式的运动监督。通过这些技术,RAP实现了精确的音频驱动控制,减轻了长期的时间漂移,并保持了较高的视觉保真度。大量的实验表明,RAP实现国家的最先进的性能,同时在实时约束下运行。
摘要:Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional intermediate representations and explicitly modeling motion dynamics, their computational complexity renders them unsuitable for real-time deployment. Real-time inference imposes stringent latency and memory constraints, often necessitating the use of highly compressed latent representations. However, operating in such compact spaces hinders the preservation of fine-grained spatiotemporal details, thereby complicating audio-visual synchronization RAP (Real-time Audio-driven Portrait animation), a unified framework for generating high-quality talking portraits under real-time constraints. Specifically, RAP introduces a hybrid attention mechanism for fine-grained audio control, and a static-dynamic training-inference paradigm that avoids explicit motion supervision. Through these techniques, RAP achieves precise audio-driven control, mitigates long-term temporal drift, and maintains high visual fidelity. Extensive experiments demonstrate that RAP achieves state-of-the-art performance while operating under real-time constraints.


【16】Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation
标题:迈向无幻觉音乐:可靠歌曲生成的强化学习偏好优化框架
链接:https://arxiv.org/abs/2508.05011

作者:Huaicheng Zhang, Wei Tan, Guangzheng Li, Yixuan Zhang, Hangting Chen, Shun Lei, Chenyu Yang, Zhiyong Wu, Shuai Wang, Qijun Huang, Dong Yu
摘要:基于音频的生成语言模型的最新进展加速了AI驱动的歌词到歌曲生成。然而,这些模型经常遭受内容幻觉,产生与输入歌词不一致的输出,破坏音乐的连贯性。目前的监督微调(SFT)的方法,被动标签拟合的限制,表现出有限的自我改善和不良的幻觉缓解。为了解决这一核心挑战,我们提出了一种新的强化学习(RL)框架,利用偏好优化进行幻觉控制。我们的主要贡献包括:(1)开发一个强大的幻觉偏好数据集,该数据集通过音素错误率(PER)计算和基于规则的过滤来构建,以捕获与人类期望的一致性;(2)在RL框架内实现和评估三种不同的偏好优化策略:直接偏好优化(DPO),邻近策略优化(PPO)和组相对策略优化(GRPO)。DPO在政策外运行,以提高正令牌的可能性,实现了7.4%的PER显著降低。PPO和GRPO采用基于策略的方法,训练基于PER的奖励模型,通过奖励最大化和KL正则化迭代优化序列,分别产生4.9%和4.7%的PER降低。全面的客观和主观评估证实,我们的方法有效地抑制幻觉,同时保持音乐质量。至关重要的是,这项工作提出了一个系统的,基于RL的解决方案,幻觉控制在歌词到歌曲的一代。该框架的可移植性也解锁了音乐风格的坚持和音乐性增强的潜力,为未来的生成歌曲研究开辟了新的途径。
摘要:Recent advances in audio-based generative language models have accelerated AI-driven lyric-to-song generation. However, these models frequently suffer from content hallucination, producing outputs misaligned with the input lyrics and undermining musical coherence. Current supervised fine-tuning (SFT) approaches, limited by passive label-fitting, exhibit constrained self-improvement and poor hallucination mitigation. To address this core challenge, we propose a novel reinforcement learning (RL) framework leveraging preference optimization for hallucination control. Our key contributions include: (1) Developing a robust hallucination preference dataset constructed via phoneme error rate (PER) computation and rule-based filtering to capture alignment with human expectations; (2) Implementing and evaluating three distinct preference optimization strategies within the RL framework: Direct Preference Optimization (DPO), Proximal Policy Optimization (PPO), and Group Relative Policy Optimization (GRPO). DPO operates off-policy to enhance positive token likelihood, achieving a significant 7.4% PER reduction. PPO and GRPO employ an on-policy approach, training a PER-based reward model to iteratively optimize sequences via reward maximization and KL-regularization, yielding PER reductions of 4.9% and 4.7%, respectively. Comprehensive objective and subjective evaluations confirm that our methods effectively suppress hallucinations while preserving musical quality. Crucially, this work presents a systematic, RL-based solution to hallucination control in lyric-to-song generation. The framework's transferability also unlocks potential for music style adherence and musicality enhancement, opening new avenues for future generative song research.


【17】REINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech Translation
标题:REINA:基于规则化信息的丢失,用于高效语音同步翻译
链接:https://arxiv.org/abs/2508.04946

作者:Nameer Hirschkind, Joseph Liu, Mahesh Kumar Nandwana, Xiao Yu
摘要:同时语音翻译(SimulST)系统在音频流中同时发出翻译的文本或语音。这样的系统面临着平衡翻译质量和延迟的重大挑战。我们引入了一种策略来优化这种权衡:只有在您通过这样做获得信息时才等待更多的输入。基于这种策略,我们提出了正则熵信息自适应(REINA),一种新的损失训练自适应策略使用现有的非流翻译模型。我们从信息理论的原则,REINA,REINA有助于推动报告的帕累托前沿的延迟/质量权衡比以前的作品。利用REINA,我们训练了一个SimulST模型的法语,西班牙语和德语,既从英语和英语。仅在开源或综合生成的数据上进行训练,我们为类似大小的模型实现了最先进的(SOTA)流结果。我们还引入了一个衡量流媒体效率的指标,定量显示REINA与以前的方法相比,将延迟/质量权衡提高了21%,对非流媒体基线BLEU分数进行了归一化。
摘要:Simultaneous Speech Translation (SimulST) systems stream in audio while simultaneously emitting translated text or speech. Such systems face the significant challenge of balancing translation quality and latency. We introduce a strategy to optimize this tradeoff: wait for more input only if you gain information by doing so. Based on this strategy, we present Regularized Entropy INformation Adaptation (REINA), a novel loss to train an adaptive policy using an existing non-streaming translation model. We derive REINA from information theory principles and show that REINA helps push the reported Pareto frontier of the latency/quality tradeoff over prior works. Utilizing REINA, we train a SimulST model on French, Spanish and German, both from and into English. Training on only open source or synthetically generated data, we achieve state-of-the-art (SOTA) streaming results for models of comparable size. We also introduce a metric for streaming efficiency, quantitatively showing REINA improves the latency/quality trade-off by as much as 21% compared to prior approaches, normalized against non-streaming baseline BLEU scores.


【18】Pitch Accent Detection improves Pretrained Automatic Speech Recognition
标题:音调口音检测改进了预训练的自动语音识别
链接:https://arxiv.org/abs/2508.04814

作者:David Sasu, Natalie Schluter
摘要:我们展示了使用半监督语音表示的自动语音识别(ASR)系统的性能可以通过引入联合ASR和音高重音检测模型来提高一个互补的音高重音检测模块。我们模型的音高重音检测组件在最先进的任务上实现了显着的改进,将F1分数的差距缩小了41%。此外,在有限的资源微调下,联合训练中的ASR性能将LibriSpeech上的WER降低了28.3%。通过这些结果,我们展示了扩展预训练语音模型以保留或重新学习重要韵律线索(如音高口音)的重要性。
摘要:We show the performance of Automatic Speech Recognition (ASR) systems that use semi-supervised speech representations can be boosted by a complimentary pitch accent detection module, by introducing a joint ASR and pitch accent detection model. The pitch accent detection component of our model achieves a significant improvement on the state-of-the-art for the task, closing the gap in F1-score by 41%. Additionally, the ASR performance in joint training decreases WER by 28.3% on LibriSpeech, under limited resource fine-tuning. With these results, we show the importance of extending pretrained speech models to retain or re-learn important prosodic cues such as pitch accent.


【19】Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM
标题:利用冻结LLM的说话人特征增强对话注释
链接:https://arxiv.org/abs/2508.04795

作者:Thomas Thebaud, Yen-Ju Lu, Matthew Wiesner, Peter Viechnicki, Najim Dehak
备注:Accepted in the 2025 IEEE Automatic Speech Recognition and Understanding Workshop
摘要:在对话转录管道中,大型语言模型(LLM)经常用于后处理以改善语法,标点符号和可读性。我们探索了一个互补的后处理步骤:通过添加元数据标签,如年龄,性别和情感的扬声器特性,丰富转录的对话。有些标签对于整个对话是全局的,而有些则是随时间变化的。我们的方法将冻结的音频基础模型(如Whisper或WavLM)与冻结的LLAMA语言模型相结合,以推断这些扬声器属性,而无需对任何模型进行特定于任务的微调。使用轻量级,高效的连接器来桥接音频和语言表示,我们在保持模块化和速度的同时,在扬声器分析任务上实现了有竞争力的性能。此外,我们证明了冻结LLAMA模型可以直接比较x向量,在某些情况下实现8.8%的等错误率。
摘要:In dialogue transcription pipelines, Large Language Models (LLMs) are frequently employed in post-processing to improve grammar, punctuation, and readability. We explore a complementary post-processing step: enriching transcribed dialogues by adding metadata tags for speaker characteristics such as age, gender, and emotion. Some of the tags are global to the entire dialogue, while some are time-variant. Our approach couples frozen audio foundation models, such as Whisper or WavLM, with a frozen LLAMA language model to infer these speaker attributes, without requiring task-specific fine-tuning of either model. Using lightweight, efficient connectors to bridge audio and language representations, we achieve competitive performance on speaker profiling tasks while preserving modularity and speed. Additionally, we demonstrate that a frozen LLAMA model can compare x-vectors directly, achieving an Equal Error Rate of 8.8% in some scenarios.


【20】Wearable Music2Emotion : Assessing Emotions Induced by AI-Generated Music through Portable EEG-fNIRS Fusion
标题:可穿戴音乐2情感:通过便携式EEG-fNIRS融合评估人工智能生成的音乐引发的情感
链接:https://arxiv.org/abs/2508.04723

作者: Sha Zhao, Song Yi, Yangxuan Zhou, Jiadong Pan, Jiquan Wang, Jie Xia, Shijian Li, Shurong Dong, Gang Pan
备注:Accepted by ACM MM 2025
摘要:情绪严重影响心理健康,通过脑机接口技术的神经生理信号驱动对基于音乐的情感计算的兴趣。虽然以前的研究利用音乐的情感诱导的可访问性,三个关键的限制仍然存在:\textbf{(1)刺激约束}:音乐刺激仅限于小语料库,由于版权和策展成本,与启发式情感音乐映射的选择偏见,忽略了个人的情感档案。\textbf{(2)模态特异性}:过度依赖单峰神经数据(例如,EEG)忽略了来自跨模态信号融合的互补见解。textbf{(3)可移植性限制}:繁琐的设置(例如,64+通道基于凝胶的EEG帽)由于程序复杂性和便携性障碍而阻碍了现实世界的适用性。为了解决这些局限性,我们提出了MEEtBrain,一种用于情绪分析(效价/唤醒)的便携式多模态框架,通过无线头带将AI生成的音乐刺激与同步的EEG-fNIRS采集集成在一起。通过MEEtBrain,音乐刺激可以由AI大规模自动生成,消除主观选择偏见,同时确保音乐的多样性。我们使用我们开发的便携式设备,该设备设计为轻型头带式,使用干电极,同时收集EEG和fNIRS记录。在第一次招募中收集了来自20名参与者的14小时数据集,以验证该框架的有效性,人工智能生成的音乐引发目标情绪(效价/唤醒)。我们正在积极扩展我们的多模态数据集(最新数据集有44个参与者),并将其公开,以促进进一步的研究和实际应用。\textbf{该数据集可在https://zju-bmi-lab.github.io/ZBra上获得。
摘要:Emotions critically influence mental health, driving interest in music-based affective computing via neurophysiological signals with Brain-computer Interface techniques. While prior studies leverage music's accessibility for emotion induction, three key limitations persist: \textbf{(1) Stimulus Constraints}: Music stimuli are confined to small corpora due to copyright and curation costs, with selection biases from heuristic emotion-music mappings that ignore individual affective profiles. \textbf{(2) Modality Specificity}: Overreliance on unimodal neural data (e.g., EEG) ignores complementary insights from cross-modal signal fusion.\textbf{ (3) Portability Limitation}: Cumbersome setups (e.g., 64+ channel gel-based EEG caps) hinder real-world applicability due to procedural complexity and portability barriers. To address these limitations, we propose MEEtBrain, a portable and multimodal framework for emotion analysis (valence/arousal), integrating AI-generated music stimuli with synchronized EEG-fNIRS acquisition via a wireless headband. By MEEtBrain, the music stimuli can be automatically generated by AI on a large scale, eliminating subjective selection biases while ensuring music diversity. We use our developed portable device that is designed in a lightweight headband-style and uses dry electrodes, to simultaneously collect EEG and fNIRS recordings. A 14-hour dataset from 20 participants was collected in the first recruitment to validate the framework's efficacy, with AI-generated music eliciting target emotions (valence/arousal). We are actively expanding our multimodal dataset (44 participants in the latest dataset) and make it publicly available to promote further research and practical applications. \textbf{The dataset is available at https://zju-bmi-lab.github.io/ZBra.


【21】Toward Low-Latency End-to-End Voice Agents for Telecommunications Using Streaming ASR, Quantized LLMs, and Real-Time TTS
标题:使用流媒体ZR、量化LLM和实时TTC实现电信低延迟端到端语音代理
链接:https://arxiv.org/abs/2508.04721

作者:Vignesh Ethiraj, Ashwath David, Sidhanth Menon, Divya Vijay
摘要:我们推出了一个低延迟的电信AI语音代理管道,用于实时、交互式电信用途,为呼叫中心自动化、智能IVR(交互式语音响应)和AI驱动的客户支持提供先进的语音AI。该解决方案是为电信构建的,结合了NetoAI的四个专用模型:TLAM,一个4位量化的特定于语音的大型语言模型(LLM); T-VEC,一个特定于语音的嵌入模型; TTE,一个特定于语音的自动语音识别(ASR)模型;和T-Synth,一个特定于语音的文本到语音(TTS)模型。这些模型能够实现高度响应、适应领域的语音AI代理,支持以知识为基础的低延迟语音交互。该管道集成了流ASR(TTE)、会话智能(TLAM)、电信文档检索增强生成(RAG)和实时TTS(T-Synth),为电信语音助理树立了新的基准。为了评估该系统,我们建立了一个数据集的500人记录的电信问题的RFC,模拟真实的电信代理查询。该框架允许跨堆栈分析延迟、域相关性和实时性能。结果表明,TLAM、TTE和T-Synth提供低于1.0的实时因子(RTF),支持企业低延迟电信部署。这些人工智能代理-由TLAM,TTE和T-Synth提供支持-为下一代电信人工智能提供基础,实现自动化客户支持,诊断等。
摘要:We introduce a low-latency telecom AI voice agent pipeline for real-time, interactive telecommunications use, enabling advanced voice AI for call center automation, intelligent IVR (Interactive Voice Response), and AI-driven customer support. The solution is built for telecom, combining four specialized models by NetoAI: TSLAM, a 4-bit quantized Telecom-Specific Large Language Model (LLM); T-VEC, a Telecom-Specific Embedding Model; TTE, a Telecom-Specific Automatic Speech Recognition (ASR) model; and T-Synth, a Telecom-Specific Text-to-Speech (TTS) model. These models enable highly responsive, domain-adapted voice AI agents supporting knowledge-grounded spoken interactions with low latency. The pipeline integrates streaming ASR (TTE), conversational intelligence (TSLAM), retrieval augmented generation (RAG) over telecom documents, and real-time TTS (T-Synth), setting a new benchmark for telecom voice assistants. To evaluate the system, we built a dataset of 500 human-recorded telecom questions from RFCs, simulating real telecom agent queries. This framework allows analysis of latency, domain relevance, and real-time performance across the stack. Results show that TSLAM, TTE, and T-Synth deliver real-time factors (RTF) below 1.0, supporting enterprise, low-latency telecom deployments. These AI agents -- powered by TSLAM, TTE, and T-Synth -- provide a foundation for next-generation telecom AI, enabling automated customer support, diagnostics, and more.


机器翻译由腾讯交互翻译提供,仅供参考