今日论文合集:cs.SD语音12篇,eess.AS音频处理15篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Pretrained Conformers for Audio Fingerprinting and Retrieval
标题:用于音频指纹识别和检索的预训练Conformers
链接:https://arxiv.org/abs/2508.11609

作者:Kemal Altwlkany, Elmedin Selmanovic, Sead Delalic
摘要:构象在语音处理中表现出很好的效果,因为它们能够捕获局部和全局交互。在这项工作中,我们利用一个自我监督的对比学习框架来训练基于一致性的编码器,这些编码器能够为小段音频生成独特的嵌入,从而很好地推广到以前看不见的数据。我们实现了音频检索任务的最先进的结果,同时仅使用3秒的音频来生成嵌入。我们的模型几乎完全不受时间失调的影响,并在出现其他音频失真(例如噪音、混响或极端时间拉伸)的情况下实现最先进的结果。代码和模型是公开的,当我们使用流行的和免费的不同大小的数据集进行训练和测试时,结果很容易重现。
摘要:Conformers have shown great results in speech processing due to their ability to capture both local and global interactions. In this work, we utilize a self-supervised contrastive learning framework to train conformer-based encoders that are capable of generating unique embeddings for small segments of audio, generalizing well to previously unseen data. We achieve state-of-the-art results for audio retrieval tasks while using only 3 seconds of audio to generate embeddings. Our models are almost completely immune to temporal misalignments and achieve state-of-the-art results in cases of other audio distortions such as noise, reverb or extreme temporal stretching. Code and models are made publicly available and the results are easy to reproduce as we train and test using popular and freely available datasets of different sizes.


【2】Representing Speech Through Autoregressive Prediction of Cochlear Tokens
标题:通过Cokisar代币的自回归预测来表示语音
链接:https://arxiv.org/abs/2508.11598

作者:Greta Tuckute, Klemen Kotar, Evelina Fedorenko, Daniel L.K. Yamins
摘要:我们介绍AuriStream,一个生物启发的模型,通过一个两阶段的框架启发人类听觉处理层次的语音编码。第一阶段将原始音频转换为基于人类耳蜗的时频表示,从中提取离散的\textbf{cochlear tokens}。第二阶段在耳蜗令牌上应用自回归序列模型。AuriStream学习有意义的音素和单词表示,以及最先进的词汇语义。AuriStream在不同的下游SUPERB语音任务上表现出有竞争力的性能。它补充了AuriStream强大的代表性功能,生成音频的延续,可以在频谱图空间中可视化并解码回音频,从而提供对模型预测的见解。总之,我们提出了一个两阶段的语音表示学习框架,以推进更人性化的模型的开发,从而有效地处理一系列基于语音的任务。
摘要:We introduce AuriStream, a biologically inspired model for encoding speech via a two-stage framework inspired by the human auditory processing hierarchy. The first stage transforms raw audio into a time-frequency representation based on the human cochlea, from which we extract discrete \textbf{cochlear tokens}. The second stage applies an autoregressive sequence model over the cochlear tokens. AuriStream learns meaningful phoneme and word representations, and state-of-the-art lexical semantics. AuriStream shows competitive performance on diverse downstream SUPERB speech tasks. Complementing AuriStream's strong representational capabilities, it generates continuations of audio which can be visualized in a spectrogram space and decoded back into audio, providing insights into the model's predictions. In summary, we present a two-stage framework for speech representation learning to advance the development of more human-like models that efficiently handle a range of speech-based tasks.


【3】Speech Emotion Recognition Using Fine-Tuned DWFormer:A Study on Track 1 of the IERPChallenge 2024
标题:使用微调DWFormer的语音情感识别:IERP挑战2024第一轨研究
链接:https://arxiv.org/abs/2508.11371

作者:Honghong Wang, Xupeng Jia, Jing Deng, Rong Zheng
备注:5 pages,1 figures
摘要:人工智能领域对情感识别的话题有着浓厚的兴趣。现有的情感识别模型大多致力于提高离散情感标签预测的精度。鉴于人类个性与情绪之间的直接关系,以及主观情绪表达的显著个体间差异,IERP挑战赛2024将个性特质纳入情绪识别研究。本文介绍了Fosafer提交的IERP Challenge 2024第一轨道的申请。该任务主要涉及音频中的情感识别,同时还提供文本和音频特征。在Track 1中,我们仅使用基于音频的特征,并通过数据增强和分数融合策略的整合,微调了预训练的语音情感识别模型DWFormer,从而在参赛团队中获得第一名。
摘要:The field of artificial intelligence has a strong interest in the topic of emotion recognition. The majority of extant emotion recognition models are oriented towards enhancing the precision of discrete emotion label prediction. Given the direct relationship between human personality and emotion, as well as the significant inter-individual differences in subjective emotional expression, the IERP Challenge 2024 incorporates personality traits into emotion recognition research. This paper presents the Fosafer submissions to the Track 1 of the IERP Challenge 2024. This task primarily concerns the recognition of emotions in audio, while also providing text and audio features. In Track 1, we utilized exclusively audio-based features and fine-tuned a pre-trained speech emotion recognition model, DWFormer, through the integration of data augmentation and score fusion strategies, thereby achieving the first place among the participating teams.


【4】Mitigating Category Imbalance: Fosafer System for the Multimodal Emotion and Intent Joint Understanding Challenge
标题:缓解类别不平衡:多模态情感和意图联合理解挑战的Fosafer系统
链接:https://arxiv.org/abs/2508.11362

作者:Honghong Wang, Yankai Wang, Dejun Zhang, Jing Deng, Rong Zheng
备注:2 pages. pubilshed by ICASSP2025
摘要:本文提出了Fosafer方法在多模态情感和意图联合理解挑战中的Track 2 Mandarin,其重点是实现对普通话中的情感和意图的联合识别,尽管存在类别不平衡的问题。为了缓解这个问题,我们在文本、视频和音频模态中使用了各种数据增强技术。此外,我们引入了SampleWeighted Focal对比损失,旨在解决识别少数类样本和语义相似但难以区分的样本的挑战。此外,我们微调的休伯特模型,以适应情感和意图的联合识别。为了减轻模态竞争,我们引入了模态丢弃策略。对于最终的预测,使用多数投票方法来确定结果。实验结果表明,我们的方法的有效性,它实现了第二个最好的性能在轨道2普通话的挑战。
摘要:This paper presents Fosafer approach to the Track 2 Mandarin in the Multimodal Emotion and Intent Joint Understandingchallenge, which focuses on achieving joint recognition of emotion and intent in Mandarin, despite the issue of category imbalance. To alleviate this issue, we use a variety of data augmentation techniques across text, video, and audio modalities. Additionally, we introduce the SampleWeighted Focal Contrastive loss, designed to address the challenges of recognizing minority class samples and those that are semantically similar but difficult to distinguish. Moreover, we fine-tune the Hubert model to adapt the emotion and intent joint recognition. To mitigate modal competition, we introduce a modal dropout strategy. For the final predictions, a plurality voting approach is used to determine the results. The experimental results demonstrate the effectiveness of our method, which achieves the second-best performance in the Track 2 Mandarin challenge.


【5】Benchmarking Prosody Encoding in Discrete Speech Tokens
标题:离散语音令牌中的韵律编码基准
链接:https://arxiv.org/abs/2508.11224

作者:Kentaro Onda, Satoru Fukayama, Daisuke Saito, Nobuaki Minematsu
备注:Accepted by ASRU2025
摘要:最近,通过k-均值聚类从自监督学习(SSL)模型中获得的离散令牌已被积极研究,作为语音语言模型中的伪文本和各种任务的有效中间表示。然而,这些离散的标记通常是提前学习的,与语言模型或下游任务的训练分开。因此,与离散化相关的选择(如使用的SSL模型或集群数量)必须严格地进行。特别是,语音语言模型预计将理解和生成的反应,不仅反映了语义内容,但也韵律特征。然而,已经有有限的研究离散令牌捕捉韵律信息的能力。为了解决这一差距,本研究进行了全面的分析,侧重于韵律编码的基础上,他们的敏感性,人为修改的韵律,旨在提供实用的指导方针,设计离散令牌。
摘要:Recently, discrete tokens derived from self-supervised learning (SSL) models via k-means clustering have been actively studied as pseudo-text in speech language models and as efficient intermediate representations for various tasks. However, these discrete tokens are typically learned in advance, separately from the training of language models or downstream tasks. As a result, choices related to discretization, such as the SSL model used or the number of clusters, must be made heuristically. In particular, speech language models are expected to understand and generate responses that reflect not only the semantic content but also prosodic features. Yet, there has been limited research on the ability of discrete tokens to capture prosodic information. To address this gap, this study conducts a comprehensive analysis focusing on prosodic encoding based on their sensitivity to the artificially modified prosody, aiming to provide practical guidelines for designing discrete tokens.


【6】Novel Parasitic Dual-Scale Modeling for Efficient and Accurate Multilingual Speech Translation
标题:高效准确的多语言语音翻译的新型寄生双尺度建模
链接:https://arxiv.org/abs/2508.11189

作者:Chenyang Le, Yinfeng Xia, Huiyan Li, Manhong Wang, Yutao Sun, Xingyang Ma, Yanmin Qian
备注:Interspeech 2025
摘要:语音到文本翻译的最新进展导致了能够同时处理多语言对的多语言模型的开发。然而,这些统一的模型往往受到大的参数大小,使其具有挑战性的平衡推理效率和性能,特别是在本地部署的情况下。我们提出了一种创新的寄生双尺度方法,它结合了增强的投机性采样方法与模型压缩和知识蒸馏技术。在Whisper Medium模型的基础上,我们将其增强为多语言语音翻译到whisperM 2 M,并集成了我们新颖的KVSPN模块,在六种流行语言中实现了最先进的(SOTA)性能,提高了推理效率。KVSPN实现了40\%的加速,而没有BLEU评分下降。与蒸馏方法相结合,它代表了2.6$\times$的速度比原来的耳语介质具有优越的性能。
摘要:Recent advancements in speech-to-text translation have led to the development of multilingual models capable of handling multiple language pairs simultaneously. However, these unified models often suffer from large parameter sizes, making it challenging to balance inference efficiency and performance, particularly in local deployment scenarios. We propose an innovative Parasitic Dual-Scale Approach, which combines an enhanced speculative sampling method with model compression and knowledge distillation techniques. Building on the Whisper Medium model, we enhance it for multilingual speech translation into whisperM2M, and integrate our novel KVSPN module, achieving state-of-the-art (SOTA) performance across six popular languages with improved inference efficiency. KVSPN enables a 40\% speedup with no BLEU score degradation. Combined with distillation methods, it represents a 2.6$\times$ speedup over the original Whisper Medium with superior performance.


【7】LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters
标题:LD-LAaudio-V1:具有双轻量级适配器的视频到长格式音频生成扩展
链接:https://arxiv.org/abs/2508.11074

作者:Haomin Zhang, Kristin Qi, Shuxin Yang, Zihao Chen, Chaofan Ding, Xinhan Di
备注:Gen4AVC@ICCV: 1st Workshop on Generative AI for Audio-Visual Content Creation
摘要:Generating high-quality and temporally synchronized audio from video content is essential for video editing and post-production tasks, enabling the creation of semantically aligned audio for silent videos. However, most existing approaches focus on short-form audio generation for video segments under 10 seconds or rely on noisy datasets for long-form video-to-audio zsynthesis. To address these limitations, we introduce LD-LAudio-V1, an extension of state-of-the-art video-to-audio models and it incorporates dual lightweight adapters to enable long-form audio generation. In addition, we release a clean and human-annotated video-to-audio dataset that contains pure sound effects without noise or artifacts. Our method significantly reduces splicing artifacts and temporal inconsistencies while maintaining computational efficiency. Compared to direct fine-tuning with short training videos, LD-LAudio-V1 achieves significant improvements across multiple metrics: $FD_{\text{passt}}$ 450.00 $\rightarrow$ 327.29 (+27.27%), $FD_{\text{panns}}$ 34.88 $\rightarrow$ 22.68 (+34.98%), $FD_{\text{vgg}}$ 3.75 $\rightarrow$ 1.28 (+65.87%), $KL_{\text{panns}}$ 2.49 $\rightarrow$ 2.07 (+16.87%), $KL_{\text{passt}}$ 1.78 $\rightarrow$ 1.53 (+14.04%), $IS_{\text{panns}}$ 4.17 $\rightarrow$ 4.30 (+3.12%), $IB_{\text{score}}$ 0.25 $\rightarrow$ 0.28 (+12.00%), $Energy\Delta10\text{ms}$ 0.3013 $\rightarrow$ 0.1349 (+55.23%), $Energy\Delta10\text{ms(vs.GT)}$ 0.0531 $\rightarrow$ 0.0288 (+45.76%), and $Sem.\,Rel.$ 2.73 $\rightarrow$ 3.28 (+20.15%). Our dataset aims to facilitate further research in long-form video-to-audio generation and is available at https://github.com/deepreasonings/long-form-video2audio.
摘要:Generating high-quality and temporally synchronized audio from video content is essential for video editing and post-production tasks, enabling the creation of semantically aligned audio for silent videos. However, most existing approaches focus on short-form audio generation for video segments under 10 seconds or rely on noisy datasets for long-form video-to-audio zsynthesis. To address these limitations, we introduce LD-LAudio-V1, an extension of state-of-the-art video-to-audio models and it incorporates dual lightweight adapters to enable long-form audio generation. In addition, we release a clean and human-annotated video-to-audio dataset that contains pure sound effects without noise or artifacts. Our method significantly reduces splicing artifacts and temporal inconsistencies while maintaining computational efficiency. Compared to direct fine-tuning with short training videos, LD-LAudio-V1 achieves significant improvements across multiple metrics: $FD_{\text{passt}}$ 450.00 $\rightarrow$ 327.29 (+27.27%), $FD_{\text{panns}}$ 34.88 $\rightarrow$ 22.68 (+34.98%), $FD_{\text{vgg}}$ 3.75 $\rightarrow$ 1.28 (+65.87%), $KL_{\text{panns}}$ 2.49 $\rightarrow$ 2.07 (+16.87%), $KL_{\text{passt}}$ 1.78 $\rightarrow$ 1.53 (+14.04%), $IS_{\text{panns}}$ 4.17 $\rightarrow$ 4.30 (+3.12%), $IB_{\text{score}}$ 0.25 $\rightarrow$ 0.28 (+12.00%), $Energy\Delta10\text{ms}$ 0.3013 $\rightarrow$ 0.1349 (+55.23%), $Energy\Delta10\text{ms(vs.GT)}$ 0.0531 $\rightarrow$ 0.0288 (+45.76%), and $Sem.\,Rel.$ 2.73 $\rightarrow$ 3.28 (+20.15%). Our dataset aims to facilitate further research in long-form video-to-audio generation and is available at https://github.com/deepreasonings/long-form-video2audio.


【8】Perturbed Public Voices (P$^{2}$V): A Dataset for Robust Audio Deepfake Detection
标题:受干扰的公共声音(P$^{2}$V):用于稳健音频深度伪造检测的数据集
链接:https://arxiv.org/abs/2508.10949

作者:Chongyang Gao, Marco Postiglione, Isabel Gortner, Sarit Kraus, V.S. Subrahmanian
摘要:目前的音频Deepfake检测器无法信任。虽然它们在受控基准测试中表现出色,但在现实世界中测试时却失败了。我们介绍了扰动公共声音(P$^{2}$V),这是一个IRB批准的数据集,它捕获了恶意deepfakes的三个关键方面:(1)通过LLM的身份一致性转录,(2)环境和对抗性噪声,以及(3)最先进的语音克隆(2020-2025)。实验揭示了22个最近的音频deepfake检测器的惊人漏洞:在当前数据集上训练的模型在P$^{2}$V上测试时损失了43%的性能,性能测量为deepfake音频,AUC和1-EER的F1得分的平均值。简单的对抗性扰动会导致高达16%的性能下降,而先进的克隆技术会使可检测性降低20- 30%。相比之下,经过P$^{2}$V训练的模型保持了对这些攻击的鲁棒性,同时推广到现有的数据集,为鲁棒的音频deepfake检测建立了新的基准。P$^{2}$V将在会议/期刊接受后公开发布。
摘要:Current audio deepfake detectors cannot be trusted. While they excel on controlled benchmarks, they fail when tested in the real world. We introduce Perturbed Public Voices (P$^{2}$V), an IRB-approved dataset capturing three critical aspects of malicious deepfakes: (1) identity-consistent transcripts via LLMs, (2) environmental and adversarial noise, and (3) state-of-the-art voice cloning (2020-2025). Experiments reveal alarming vulnerabilities of 22 recent audio deepfake detectors: models trained on current datasets lose 43% performance when tested on P$^{2}$V, with performance measured as the mean of F1 score on deepfake audio, AUC, and 1-EER. Simple adversarial perturbations induce up to 16% performance degradation, while advanced cloning techniques reduce detectability by 20-30%. In contrast, P$^{2}$V-trained models maintain robustness against these attacks while generalizing to existing datasets, establishing a new benchmark for robust audio deepfake detection. P$^{2}$V will be publicly released upon acceptance by a conference/journal.


【9】MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
标题:MoE-TTC:通过专家混合增强基于描述的TTC的域外文本理解
链接:https://arxiv.org/abs/2508.11326

作者:Heyang Xue, Xuchen Song, Yu Tang, Jianyu Chen, Yanru Chen, Yang Li, Yahui Zhou
摘要:基于描述的文本到语音(TTS)模型在域内文本描述上表现出很强的性能,即,在训练中遇到的。然而,在现实世界的应用程序中,用户生成的描述的多样性不可避免地引入了大量的域外输入,挑战这些系统的文本理解能力。为了解决这个问题,我们提出了MoE-TTS,一个基于解释的TTS模型,旨在增强对域外文本描述的理解。MoE-TTS采用基于模态的专家混合(MoE)方法,使用一组适应语音模态的专用权重来增强预训练的文本大语言模型(LLM),同时在训练期间保持原始LLM冻结。这种方法允许MoE-TTS有效地利用文本LLM的预训练知识和文本理解能力。我们的实验结果表明:第一,即使是最先进的闭源商业产品可以通过精心设计的域外描述测试集的挑战;第二,MoE-TTS在生成更准确地反映描述的语音方面取得了优异的性能。我们鼓励读者在https://welkinyang.github.io/MoE-TTS/上收听演示。
摘要:Description-based text-to-speech (TTS) models exhibit strong performance on in-domain text descriptions, i.e., those encountered during training. However, in real-world applications, the diverse range of user-generated descriptions inevitably introduces numerous out-of-domain inputs that challenge the text understanding capabilities of these systems. To address this issue, we propose MoE-TTS, a description-based TTS model designed to enhance the understanding of out-of-domain text descriptions. MoE-TTS employs a modality-based mixture-of-experts (MoE) approach to augment a pre-trained textual large language model (LLM) with a set of specialized weights adapted to the speech modality while maintaining the original LLM frozen during training. This approach allows MoE-TTS to effectively leverage the pre-trained knowledge and text understanding abilities of textual LLMs. Our experimental results indicate that: first, even the most advanced closed-source commercial products can be challenged by carefully designed out-of-domain description test sets; second, MoE-TTS achieves superior performance in generating speech that more accurately reflects the descriptions. We encourage readers to listen to the demos at https://welkinyang.github.io/MoE-TTS/.


【10】Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
标题:使用说话风格自然语言描述的表达性语音检索
链接:https://arxiv.org/abs/2508.11187

作者:Wonjune Kang, Deb Roy
备注:Accepted to ASRU 2025
摘要:我们介绍的任务,表达语音检索,其目标是检索语音话语在一个给定的风格的基础上,自然语言描述的风格。虽然以前的工作主要集中在执行语音检索的基础上说了什么话语,我们的目标是这样做的基础上如何说的东西。我们训练语音和文本编码器,以嵌入语音和文本描述的说话风格到一个联合的潜在空间,这使得使用自由形式的文本提示描述的情感或风格作为查询检索匹配的表达语音段。我们对我们提出的框架的各个方面进行了详细的分析,包括编码器架构,有效的跨模态对齐的训练标准,以及对任意文本查询的改进泛化的提示增强。在包含22种说话风格的多个数据集上的实验表明,我们的方法实现了强大的检索性能,通过Recall@k来衡量。
摘要:We introduce the task of expressive speech retrieval, where the goal is to retrieve speech utterances spoken in a given style based on a natural language description of that style. While prior work has primarily focused on performing speech retrieval based on what was said in an utterance, we aim to do so based on how something was said. We train speech and text encoders to embed speech and text descriptions of speaking styles into a joint latent space, which enables using free-form text prompts describing emotions or styles as queries to retrieve matching expressive speech segments. We perform detailed analyses of various aspects of our proposed framework, including encoder architectures, training criteria for effective cross-modal alignment, and prompt augmentation for improved generalization to arbitrary text queries. Experiments on multiple datasets encompassing 22 speaking styles demonstrate that our approach achieves strong retrieval performance as measured by Recall@k.


【11】CleanCTG: A Deep Learning Model for Multi-Artefact Detection and Reconstruction in Cardiotocography
标题:CleanCTG:用于子宫内膜摄影中多伪影检测和重建的深度学习模型
链接:https://arxiv.org/abs/2508.10928

作者:Sheng Wong, Beth Albert, Gabriel Davis Jones
摘要:胎儿监护(CTG)是必不可少的,但经常受到各种伪影的影响,这些伪影模糊了真实的胎儿心率(FHR)模式,并可能导致误诊或延迟干预。目前的深度学习方法通常绕过全面的噪声处理,应用最小的预处理或仅关注下游分类,而传统方法依赖于简单的插值或基于规则的过滤,仅解决丢失的样本,无法纠正复杂的伪影类型。我们提出了CleanCTG,一个端到端的双阶段模型,首先通过多尺度卷积和上下文感知的交叉注意识别多个伪影类型,然后通过伪影特定的校正分支重建损坏的片段。培训使用了超过800,000分钟的生理上真实的,合成损坏的CTG,这些CTG来自专家验证的“干净”录音。在合成数据上,CleanCTG实现了完美的伪影检测(AU-ROC = 1.00),并将损坏片段的均方误差(MSE)降低到2.74 x 10^-4(干净片段的MSE = 2.40 x 10^-6),比第二好的方法高出60%以上。对10,190分钟临床医生注释片段的外部验证产生AU-ROC = 0.95(灵敏度= 83.44%,特异性94.22%),超过了6个比较分类器。最后,当与Dawes-Redman系统集成在933个临床CTG记录上时,去噪痕迹增加了特异性(从80.70%到82.70%),并将中位决策时间缩短了33%。这些研究结果表明,明确的伪影去除和信号重建既可以保持诊断的准确性,使监测时间更短,提供了一个实用的路线,以更可靠的CTG解释。
摘要:Cardiotocography (CTG) is essential for fetal monitoring but is frequently compromised by diverse artefacts which obscure true fetal heart rate (FHR) patterns and can lead to misdiagnosis or delayed intervention. Current deep-learning approaches typically bypass comprehensive noise handling, applying minimal preprocessing or focusing solely on downstream classification, while traditional methods rely on simple interpolation or rule-based filtering that addresses only missing samples and fail to correct complex artefact types. We present CleanCTG, an end-to-end dual-stage model that first identifies multiple artefact types via multi-scale convolution and context-aware cross-attention, then reconstructs corrupted segments through artefact-specific correction branches. Training utilised over 800,000 minutes of physiologically realistic, synthetically corrupted CTGs derived from expert-verified "clean" recordings. On synthetic data, CleanCTG achieved perfect artefact detection (AU-ROC = 1.00) and reduced mean squared error (MSE) on corrupted segments to 2.74 x 10^-4 (clean-segment MSE = 2.40 x 10^-6), outperforming the next best method by more than 60%. External validation on 10,190 minutes of clinician-annotated segments yielded AU-ROC = 0.95 (sensitivity = 83.44%, specificity 94.22%), surpassing six comparator classifiers. Finally, when integrated with the Dawes-Redman system on 933 clinical CTG recordings, denoised traces increased specificity (from 80.70% to 82.70%) and shortened median time to decision by 33%. These findings suggest that explicit artefact removal and signal reconstruction can both maintain diagnostic accuracy and enable shorter monitoring sessions, offering a practical route to more reliable CTG interpretation.


【12】ASAudio: A Survey of Advanced Spatial Audio Research
标题:ASaudio:高级空间音频研究概览
链接:https://arxiv.org/abs/2508.10924

作者:Zhiyuan Zhu, Yu Zhang, Wenxiang Guo, Changhao Pan, Zhou Zhao
摘要:随着空间音频技术的快速发展,在AR、VR等场景中的应用得到了广泛的关注。与传统的单声道声音不同,空间音频提供了更逼真和身临其境的听觉体验。尽管在该领域取得了显著进展,但仍然缺乏系统地组织和分析这些方法及其基础技术的全面调查。在本文中,我们提供了一个全面的概述空间音频和系统地回顾最近的文献在该领域。为了解决这个问题,我们按时间顺序概述了现有的工作相关的空间音频和分类这些研究的基础上输入输出表示,以及生成和理解任务,从而总结了空间音频的各个研究方面。此外,我们还回顾了相关的数据集、评估指标和基准,从培训和评估的角度提供了见解。相关材料可在https://github.com/dieKarotte/ASAudio上获取。
摘要:With the rapid development of spatial audio technologies today, applications in AR, VR, and other scenarios have garnered extensive attention. Unlike traditional mono sound, spatial audio offers a more realistic and immersive auditory experience. Despite notable progress in the field, there remains a lack of comprehensive surveys that systematically organize and analyze these methods and their underlying technologies. In this paper, we provide a comprehensive overview of spatial audio and systematically review recent literature in the area. To address this, we chronologically outlining existing work related to spatial audio and categorize these studies based on input-output representations, as well as generation and understanding tasks, thereby summarizing various research aspects of spatial audio. In addition, we review related datasets, evaluation metrics, and benchmarks, offering insights from both training and evaluation perspectives. Related materials are available at https://github.com/dieKarotte/ASAudio.


eess.AS音频处理


【1】Emphasis Sensitivity in Speech Representations
标题:言语表达中的强调敏感性
链接:https://arxiv.org/abs/2508.11566

作者:Shaun Rafael Cassini, Thomas Hain, Anton Ragni
备注:Accepted to IEEE ASRU 2025
摘要:这项工作调查现代语音模型是否对韵律强调敏感-它们是否以系统不同的方式编码强调词和中性词。先前的工作通常依赖于孤立的声学相关(例如,音高、持续时间)或标签预测,这两者都错过了强调的关系结构。本文提出了一个基于残差的框架,定义强调之间的差异成对的中性和强调词表示。对自监督语音模型的分析表明,这些残差与时长变化密切相关,在词身份预测方面表现不佳,表明韵律强调的结构化,关系编码。在ASR微调模型中,残差占据的子空间比预训练模型中的子空间更紧凑50%,这进一步表明强调被编码为一致的低维变换,通过特定任务的学习变得更加结构化。
摘要:This work investigates whether modern speech models are sensitive to prosodic emphasis - whether they encode emphasized and neutral words in systematically different ways. Prior work typically relies on isolated acoustic correlates (e.g., pitch, duration) or label prediction, both of which miss the relational structure of emphasis. This paper proposes a residual-based framework, defining emphasis as the difference between paired neutral and emphasized word representations. Analysis on self-supervised speech models shows that these residuals correlate strongly with duration changes and perform poorly at word identity prediction, indicating a structured, relational encoding of prosodic emphasis. In ASR fine-tuned models, residuals occupy a subspace up to 50% more compact than in pre-trained models, further suggesting that emphasis is encoded as a consistent, low-dimensional transformation that becomes more structured with task-specific learning.


【2】Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling
标题:利用基于重新合成的持续时间建模增强野外语音情感转换
链接:https://arxiv.org/abs/2508.11535

作者:Navin Raj Prabhu, Danilo de Oliveira, Nale Lehmann-Willenbrock, Timo Gerkmann
备注:Copyright 2025 IEEE. Personal use of this material is permitted.   Permission from IEEE must be obtained for all other uses, in any current or   future media, including reprinting/republishing this material for advertising   or promotional purposes, creating new collective works, for resale or   redistribution to servers or lists, or reuse of any copyrighted component of   this work in other works
摘要:语音情感转换的目的是在保留词汇内容和说话人身份的同时,对输入语音中所表达的情感进行修改。最近,生成式建模方法在改变局部声学特性(例如基频、频谱包络和能量)方面显示出有希望的结果,但通常缺乏控制声音持续时间的能力。为了解决这个问题,我们提出了一个持续时间建模框架,使用基于再合成的离散内容表示,使修改语音持续时间,以反映目标情绪,并实现可控的语音速率,而不使用并行数据。实验结果表明,包括建议的持续时间建模框架显着提高情感表达,在野生MSP播客数据集。分析表明,低唤醒情绪与较长的持续时间和较慢的语速相关,而高唤醒情绪产生较短,较快的讲话。
摘要:Speech Emotion Conversion aims to modify the emotion expressed in input speech while preserving lexical content and speaker identity. Recently, generative modeling approaches have shown promising results in changing local acoustic properties such as fundamental frequency, spectral envelope and energy, but often lack the ability to control the duration of sounds. To address this, we propose a duration modeling framework using resynthesis-based discrete content representations, enabling modification of speech duration to reflect target emotions and achieve controllable speech rates without using parallel data. Experimental results reveal that the inclusion of the proposed duration modeling framework significantly enhances emotional expressiveness, in the in-the-wild MSP-Podcast dataset. Analyses show that low-arousal emotions correlate with longer durations and slower speech rates, while high-arousal emotions produce shorter, faster speech.


【3】MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
标题:MoE-TTC:通过专家混合增强基于描述的TTC的域外文本理解
链接:https://arxiv.org/abs/2508.11326

作者:Heyang Xue, Xuchen Song, Yu Tang, Jianyu Chen, Yanru Chen, Yang Li, Yahui Zhou
摘要:基于描述的文本到语音(TTS)模型在域内文本描述上表现出很强的性能,即,在训练中遇到的。然而,在现实世界的应用程序中,用户生成的描述的多样性不可避免地引入了大量的域外输入,挑战这些系统的文本理解能力。为了解决这个问题,我们提出了MoE-TTS,一个基于解释的TTS模型,旨在增强对域外文本描述的理解。MoE-TTS采用基于模态的专家混合(MoE)方法,使用一组适应语音模态的专用权重来增强预训练的文本大语言模型(LLM),同时在训练期间保持原始LLM冻结。这种方法允许MoE-TTS有效地利用文本LLM的预训练知识和文本理解能力。我们的实验结果表明:第一,即使是最先进的闭源商业产品可以通过精心设计的域外描述测试集的挑战;第二,MoE-TTS在生成更准确地反映描述的语音方面取得了优异的性能。我们鼓励读者在https://welkinyang.github.io/MoE-TTS/上收听演示。
摘要:Description-based text-to-speech (TTS) models exhibit strong performance on in-domain text descriptions, i.e., those encountered during training. However, in real-world applications, the diverse range of user-generated descriptions inevitably introduces numerous out-of-domain inputs that challenge the text understanding capabilities of these systems. To address this issue, we propose MoE-TTS, a description-based TTS model designed to enhance the understanding of out-of-domain text descriptions. MoE-TTS employs a modality-based mixture-of-experts (MoE) approach to augment a pre-trained textual large language model (LLM) with a set of specialized weights adapted to the speech modality while maintaining the original LLM frozen during training. This approach allows MoE-TTS to effectively leverage the pre-trained knowledge and text understanding abilities of textual LLMs. Our experimental results indicate that: first, even the most advanced closed-source commercial products can be challenged by carefully designed out-of-domain description test sets; second, MoE-TTS achieves superior performance in generating speech that more accurately reflects the descriptions. We encourage readers to listen to the demos at https://welkinyang.github.io/MoE-TTS/.


【4】EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens
标题:基于球面矢量和离散语音标记的多语言情感语音合成
链接:https://arxiv.org/abs/2508.11273

作者:Joonyong Park, Kenichi Nakamura
备注:In Proceedings of the 13th ISCA Speech Synthesis Workshop
摘要:本文介绍了一个新的框架,多语言情感的文本到语音(TTS)的合成,结合球形情感向量与离散令牌功能来自自监督学习(SSL)。通过在连续的球面坐标空间中编码情感,并利用基于SSL的表示进行语义和声学建模,MSSLSphere实现了细粒度的情感控制、有效的跨语言情感传递以及对说话者身份的鲁棒保护。我们在英语和日语语料库上评估了MSSLSphere,证明了语音清晰度,频谱保真度,韵律一致性和整体合成质量的显着改善。主观评估进一步证实,我们的方法优于基线模型的自然性和情感表达,强调其作为一个可扩展的解决方案,多语言情感TTS的潜力。
摘要:This paper introduces EmoSSLSphere, a novel framework for multilingual emotional text-to-speech (TTS) synthesis that combines spherical emotion vectors with discrete token features derived from self-supervised learning (SSL). By encoding emotions in a continuous spherical coordinate space and leveraging SSL-based representations for semantic and acoustic modeling, EmoSSLSphere enables fine-grained emotional control, effective cross-lingual emotion transfer, and robust preservation of speaker identity. We evaluate EmoSSLSphere on English and Japanese corpora, demonstrating significant improvements in speech intelligibility, spectral fidelity, prosodic consistency, and overall synthesis quality. Subjective evaluations further confirm that our method outperforms baseline models in terms of naturalness and emotional expressiveness, underscoring its potential as a scalable solution for multilingual emotional TTS.


【5】Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
标题:使用说话风格自然语言描述的表达性语音检索
链接:https://arxiv.org/abs/2508.11187

作者:Wonjune Kang, Deb Roy
备注:Accepted to ASRU 2025
摘要:我们介绍的任务,表达语音检索,其目标是检索语音话语在一个给定的风格的基础上,自然语言描述的风格。虽然以前的工作主要集中在执行语音检索的基础上说了什么话语,我们的目标是这样做的基础上如何说的东西。我们训练语音和文本编码器,以嵌入语音和文本描述的说话风格到一个联合的潜在空间,这使得使用自由形式的文本提示描述的情感或风格作为查询检索匹配的表达语音段。我们对我们提出的框架的各个方面进行了详细的分析,包括编码器架构,有效的跨模态对齐的训练标准,以及对任意文本查询的改进泛化的提示增强。在包含22种说话风格的多个数据集上的实验表明,我们的方法实现了强大的检索性能,通过Recall@k来衡量。
摘要:We introduce the task of expressive speech retrieval, where the goal is to retrieve speech utterances spoken in a given style based on a natural language description of that style. While prior work has primarily focused on performing speech retrieval based on what was said in an utterance, we aim to do so based on how something was said. We train speech and text encoders to embed speech and text descriptions of speaking styles into a joint latent space, which enables using free-form text prompts describing emotions or styles as queries to retrieve matching expressive speech segments. We perform detailed analyses of various aspects of our proposed framework, including encoder architectures, training criteria for effective cross-modal alignment, and prompt augmentation for improved generalization to arbitrary text queries. Experiments on multiple datasets encompassing 22 speaking styles demonstrate that our approach achieves strong retrieval performance as measured by Recall@k.


【6】CleanCTG: A Deep Learning Model for Multi-Artefact Detection and Reconstruction in Cardiotocography
标题:CleanCTG:用于子宫内膜摄影中多伪影检测和重建的深度学习模型
链接:https://arxiv.org/abs/2508.10928

作者:Sheng Wong, Beth Albert, Gabriel Davis Jones
摘要:胎儿监护(CTG)是必不可少的,但经常受到各种伪影的影响,这些伪影模糊了真实的胎儿心率(FHR)模式,并可能导致误诊或延迟干预。目前的深度学习方法通常绕过全面的噪声处理,应用最小的预处理或仅关注下游分类,而传统方法依赖于简单的插值或基于规则的过滤,仅解决丢失的样本,无法纠正复杂的伪影类型。我们提出了CleanCTG,一个端到端的双阶段模型,首先通过多尺度卷积和上下文感知的交叉注意识别多个伪影类型,然后通过伪影特定的校正分支重建损坏的片段。培训使用了超过800,000分钟的生理上真实的,合成损坏的CTG,这些CTG来自专家验证的“干净”录音。在合成数据上,CleanCTG实现了完美的伪影检测(AU-ROC = 1.00),并将损坏片段的均方误差(MSE)降低到2.74 x 10^-4(干净片段的MSE = 2.40 x 10^-6),比第二好的方法高出60%以上。对10,190分钟临床医生注释片段的外部验证产生AU-ROC = 0.95(灵敏度= 83.44%,特异性94.22%),超过了6个比较分类器。最后,当与Dawes-Redman系统集成在933个临床CTG记录上时,去噪痕迹增加了特异性(从80.70%到82.70%),并将中位决策时间缩短了33%。这些研究结果表明,明确的伪影去除和信号重建既可以保持诊断的准确性,使监测时间更短,提供了一个实用的路线,以更可靠的CTG解释。
摘要:Cardiotocography (CTG) is essential for fetal monitoring but is frequently compromised by diverse artefacts which obscure true fetal heart rate (FHR) patterns and can lead to misdiagnosis or delayed intervention. Current deep-learning approaches typically bypass comprehensive noise handling, applying minimal preprocessing or focusing solely on downstream classification, while traditional methods rely on simple interpolation or rule-based filtering that addresses only missing samples and fail to correct complex artefact types. We present CleanCTG, an end-to-end dual-stage model that first identifies multiple artefact types via multi-scale convolution and context-aware cross-attention, then reconstructs corrupted segments through artefact-specific correction branches. Training utilised over 800,000 minutes of physiologically realistic, synthetically corrupted CTGs derived from expert-verified "clean" recordings. On synthetic data, CleanCTG achieved perfect artefact detection (AU-ROC = 1.00) and reduced mean squared error (MSE) on corrupted segments to 2.74 x 10^-4 (clean-segment MSE = 2.40 x 10^-6), outperforming the next best method by more than 60%. External validation on 10,190 minutes of clinician-annotated segments yielded AU-ROC = 0.95 (sensitivity = 83.44%, specificity 94.22%), surpassing six comparator classifiers. Finally, when integrated with the Dawes-Redman system on 933 clinical CTG recordings, denoised traces increased specificity (from 80.70% to 82.70%) and shortened median time to decision by 33%. These findings suggest that explicit artefact removal and signal reconstruction can both maintain diagnostic accuracy and enable shorter monitoring sessions, offering a practical route to more reliable CTG interpretation.


【7】ASAudio: A Survey of Advanced Spatial Audio Research
标题:ASaudio:高级空间音频研究概览
链接:https://arxiv.org/abs/2508.10924

作者:Zhiyuan Zhu, Yu Zhang, Wenxiang Guo, Changhao Pan, Zhou Zhao
摘要:随着空间音频技术的快速发展,在AR、VR等场景中的应用得到了广泛的关注。与传统的单声道声音不同,空间音频提供了更逼真和身临其境的听觉体验。尽管在该领域取得了显著进展,但仍然缺乏系统地组织和分析这些方法及其基础技术的全面调查。在本文中,我们提供了一个全面的概述空间音频和系统地回顾最近的文献在该领域。为了解决这个问题,我们按时间顺序概述了现有的工作相关的空间音频和分类这些研究的基础上输入输出表示,以及生成和理解任务,从而总结了空间音频的各个研究方面。此外,我们还回顾了相关的数据集、评估指标和基准,从培训和评估的角度提供了见解。相关材料可在https://github.com/dieKarotte/ASAudio上查阅。
摘要:With the rapid development of spatial audio technologies today, applications in AR, VR, and other scenarios have garnered extensive attention. Unlike traditional mono sound, spatial audio offers a more realistic and immersive auditory experience. Despite notable progress in the field, there remains a lack of comprehensive surveys that systematically organize and analyze these methods and their underlying technologies. In this paper, we provide a comprehensive overview of spatial audio and systematically review recent literature in the area. To address this, we chronologically outlining existing work related to spatial audio and categorize these studies based on input-output representations, as well as generation and understanding tasks, thereby summarizing various research aspects of spatial audio. In addition, we review related datasets, evaluation metrics, and benchmarks, offering insights from both training and evaluation perspectives. Related materials are available at https://github.com/dieKarotte/ASAudio.


【8】Pretrained Conformers for Audio Fingerprinting and Retrieval
标题:用于音频指纹识别和检索的预训练Conformers
链接:https://arxiv.org/abs/2508.11609

作者:Kemal Altwlkany, Elmedin Selmanovic, Sead Delalic
摘要:构象在语音处理中表现出很好的效果,因为它们能够捕获局部和全局交互。在这项工作中,我们利用一个自我监督的对比学习框架来训练基于一致性的编码器,这些编码器能够为小段音频生成独特的嵌入,从而很好地推广到以前看不见的数据。我们实现了音频检索任务的最先进的结果,同时仅使用3秒的音频来生成嵌入。我们的模型几乎完全不受时间错位的影响,并在其他音频失真(如噪声、混响或极端时间拉伸)的情况下实现最先进的结果。代码和模型是公开的,当我们使用流行的和免费的不同大小的数据集进行训练和测试时,结果很容易重现。
摘要:Conformers have shown great results in speech processing due to their ability to capture both local and global interactions. In this work, we utilize a self-supervised contrastive learning framework to train conformer-based encoders that are capable of generating unique embeddings for small segments of audio, generalizing well to previously unseen data. We achieve state-of-the-art results for audio retrieval tasks while using only 3 seconds of audio to generate embeddings. Our models are almost completely immune to temporal misalignments and achieve state-of-the-art results in cases of other audio distortions such as noise, reverb or extreme temporal stretching. Code and models are made publicly available and the results are easy to reproduce as we train and test using popular and freely available datasets of different sizes.


【9】Representing Speech Through Autoregressive Prediction of Cochlear Tokens
标题:通过Cokisar代币的自回归预测来表示语音
链接:https://arxiv.org/abs/2508.11598

作者:Greta Tuckute, Klemen Kotar, Evelina Fedorenko, Daniel L.K. Yamins
摘要:我们介绍AuriStream,一个生物启发的模型,通过一个两阶段的框架启发人类听觉处理层次的语音编码。第一阶段将原始音频转换为基于人类耳蜗的时频表示,从中提取离散的\textbf{cochlear tokens}。第二阶段在耳蜗令牌上应用自回归序列模型。AuriStream学习有意义的音素和单词表示,以及最先进的词汇语义。AuriStream在不同的下游SUPERB语音任务上表现出有竞争力的性能。它补充了AuriStream强大的代表性功能,生成音频的延续,可以在频谱图空间中可视化并解码回音频,从而提供对模型预测的见解。总之,我们提出了一个两阶段的语音表示学习框架,以推进更人性化的模型的开发,从而有效地处理一系列基于语音的任务。
摘要:We introduce AuriStream, a biologically inspired model for encoding speech via a two-stage framework inspired by the human auditory processing hierarchy. The first stage transforms raw audio into a time-frequency representation based on the human cochlea, from which we extract discrete \textbf{cochlear tokens}. The second stage applies an autoregressive sequence model over the cochlear tokens. AuriStream learns meaningful phoneme and word representations, and state-of-the-art lexical semantics. AuriStream shows competitive performance on diverse downstream SUPERB speech tasks. Complementing AuriStream's strong representational capabilities, it generates continuations of audio which can be visualized in a spectrogram space and decoded back into audio, providing insights into the model's predictions. In summary, we present a two-stage framework for speech representation learning to advance the development of more human-like models that efficiently handle a range of speech-based tasks.


【10】Speech Emotion Recognition Using Fine-Tuned DWFormer:A Study on Track 1 of the IERPChallenge 2024
标题:使用微调DWFormer的语音情感识别:IERP挑战2024第一轨研究
链接:https://arxiv.org/abs/2508.11371

作者:Honghong Wang, Xupeng Jia, Jing Deng, Rong Zheng
备注:5 pages,1 figures
摘要:人工智能领域对情感识别的话题有着浓厚的兴趣。现有的情感识别模型大多致力于提高离散情感标签预测的精度。鉴于人类个性与情绪之间的直接关系,以及主观情绪表达的显著个体间差异,IERP挑战赛2024将个性特质纳入情绪识别研究。本文介绍了Fosafer提交的IERP Challenge 2024第一轨道的申请。该任务主要涉及音频中的情感识别,同时还提供文本和音频特征。在Track 1中,我们仅使用基于音频的特征,并通过数据增强和分数融合策略的整合,微调了预训练的语音情感识别模型DWFormer,从而在参赛团队中获得第一名。
摘要:The field of artificial intelligence has a strong interest in the topic of emotion recognition. The majority of extant emotion recognition models are oriented towards enhancing the precision of discrete emotion label prediction. Given the direct relationship between human personality and emotion, as well as the significant inter-individual differences in subjective emotional expression, the IERP Challenge 2024 incorporates personality traits into emotion recognition research. This paper presents the Fosafer submissions to the Track 1 of the IERP Challenge 2024. This task primarily concerns the recognition of emotions in audio, while also providing text and audio features. In Track 1, we utilized exclusively audio-based features and fine-tuned a pre-trained speech emotion recognition model, DWFormer, through the integration of data augmentation and score fusion strategies, thereby achieving the first place among the participating teams.


【11】Mitigating Category Imbalance: Fosafer System for the Multimodal Emotion and Intent Joint Understanding Challenge
标题:缓解类别不平衡:多模态情感和意图联合理解挑战的Fosafer系统
链接:https://arxiv.org/abs/2508.11362

作者:Honghong Wang, Yankai Wang, Dejun Zhang, Jing Deng, Rong Zheng
备注:2 pages. pubilshed by ICASSP2025
摘要:本文提出了Fosafer方法在多模态情感和意图联合理解挑战中的Track 2 Mandarin,其重点是实现对普通话中的情感和意图的联合识别,尽管存在类别不平衡的问题。为了缓解这个问题,我们在文本、视频和音频模态中使用了各种数据增强技术。此外,我们引入了SampleWeighted Focal对比损失,旨在解决识别少数类样本和语义相似但难以区分的样本的挑战。此外,我们微调的休伯特模型,以适应情感和意图的联合识别。为了减轻模态竞争,我们引入了模态丢弃策略。对于最终的预测,使用多数投票方法来确定结果。实验结果表明,我们的方法的有效性,它实现了第二个最好的性能在轨道2普通话的挑战。
摘要:This paper presents Fosafer approach to the Track 2 Mandarin in the Multimodal Emotion and Intent Joint Understandingchallenge, which focuses on achieving joint recognition of emotion and intent in Mandarin, despite the issue of category imbalance. To alleviate this issue, we use a variety of data augmentation techniques across text, video, and audio modalities. Additionally, we introduce the SampleWeighted Focal Contrastive loss, designed to address the challenges of recognizing minority class samples and those that are semantically similar but difficult to distinguish. Moreover, we fine-tune the Hubert model to adapt the emotion and intent joint recognition. To mitigate modal competition, we introduce a modal dropout strategy. For the final predictions, a plurality voting approach is used to determine the results. The experimental results demonstrate the effectiveness of our method, which achieves the second-best performance in the Track 2 Mandarin challenge.


【12】Benchmarking Prosody Encoding in Discrete Speech Tokens
标题:离散语音令牌中的韵律编码基准
链接:https://arxiv.org/abs/2508.11224

作者:Kentaro Onda, Satoru Fukayama, Daisuke Saito, Nobuaki Minematsu
备注:Accepted by ASRU2025
摘要:最近,通过k-均值聚类从自监督学习(SSL)模型中获得的离散令牌已被积极研究,作为语音语言模型中的伪文本和各种任务的有效中间表示。然而,这些离散的标记通常是提前学习的,与语言模型或下游任务的训练分开。因此,与离散化相关的选择(如使用的SSL模型或集群数量)必须严格地进行。特别是,语音语言模型预计将理解和生成的反应,不仅反映了语义内容,但也韵律特征。然而,关于离散令牌捕获韵律信息的能力的研究有限。为了解决这一差距,本研究进行了全面的分析,侧重于韵律编码的基础上,他们的敏感性,人为修改的韵律,旨在提供实用的指导方针,设计离散令牌。
摘要:Recently, discrete tokens derived from self-supervised learning (SSL) models via k-means clustering have been actively studied as pseudo-text in speech language models and as efficient intermediate representations for various tasks. However, these discrete tokens are typically learned in advance, separately from the training of language models or downstream tasks. As a result, choices related to discretization, such as the SSL model used or the number of clusters, must be made heuristically. In particular, speech language models are expected to understand and generate responses that reflect not only the semantic content but also prosodic features. Yet, there has been limited research on the ability of discrete tokens to capture prosodic information. To address this gap, this study conducts a comprehensive analysis focusing on prosodic encoding based on their sensitivity to the artificially modified prosody, aiming to provide practical guidelines for designing discrete tokens.


【13】Novel Parasitic Dual-Scale Modeling for Efficient and Accurate Multilingual Speech Translation
标题:高效准确的多语言语音翻译的新型寄生双尺度建模
链接:https://arxiv.org/abs/2508.11189

作者:Chenyang Le, Yinfeng Xia, Huiyan Li, Manhong Wang, Yutao Sun, Xingyang Ma, Yanmin Qian
备注:Interspeech 2025
摘要:语音到文本翻译的最新进展导致了能够同时处理多语言对的多语言模型的开发。然而,这些统一的模型往往受到大的参数大小,使其具有挑战性的平衡推理效率和性能,特别是在本地部署的情况下。我们提出了一种创新的寄生双尺度方法,它结合了增强的投机性采样方法与模型压缩和知识蒸馏技术。在Whisper Medium模型的基础上,我们将其增强为多语言语音翻译到whisperM 2 M,并集成了我们新颖的KVSPN模块,在六种流行语言中实现了最先进的(SOTA)性能,提高了推理效率。KVSPN实现了40\%的加速,而没有BLEU评分下降。与蒸馏方法相结合,它代表了2.6$\times$的速度比原来的耳语介质具有优越的性能。
摘要:Recent advancements in speech-to-text translation have led to the development of multilingual models capable of handling multiple language pairs simultaneously. However, these unified models often suffer from large parameter sizes, making it challenging to balance inference efficiency and performance, particularly in local deployment scenarios. We propose an innovative Parasitic Dual-Scale Approach, which combines an enhanced speculative sampling method with model compression and knowledge distillation techniques. Building on the Whisper Medium model, we enhance it for multilingual speech translation into whisperM2M, and integrate our novel KVSPN module, achieving state-of-the-art (SOTA) performance across six popular languages with improved inference efficiency. KVSPN enables a 40\% speedup with no BLEU score degradation. Combined with distillation methods, it represents a 2.6$\times$ speedup over the original Whisper Medium with superior performance.


【14】LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters
标题:LD-LAaudio-V1:具有双轻量级适配器的视频到长格式音频生成扩展
链接:https://arxiv.org/abs/2508.11074

作者:Haomin Zhang, Kristin Qi, Shuxin Yang, Zihao Chen, Chaofan Ding, Xinhan Di
备注:Gen4AVC@ICCV: 1st Workshop on Generative AI for Audio-Visual Content Creation
摘要:Generating high-quality and temporally synchronized audio from video content is essential for video editing and post-production tasks, enabling the creation of semantically aligned audio for silent videos. However, most existing approaches focus on short-form audio generation for video segments under 10 seconds or rely on noisy datasets for long-form video-to-audio zsynthesis. To address these limitations, we introduce LD-LAudio-V1, an extension of state-of-the-art video-to-audio models and it incorporates dual lightweight adapters to enable long-form audio generation. In addition, we release a clean and human-annotated video-to-audio dataset that contains pure sound effects without noise or artifacts. Our method significantly reduces splicing artifacts and temporal inconsistencies while maintaining computational efficiency. Compared to direct fine-tuning with short training videos, LD-LAudio-V1 achieves significant improvements across multiple metrics: $FD_{\text{passt}}$ 450.00 $\rightarrow$ 327.29 (+27.27%), $FD_{\text{panns}}$ 34.88 $\rightarrow$ 22.68 (+34.98%), $FD_{\text{vgg}}$ 3.75 $\rightarrow$ 1.28 (+65.87%), $KL_{\text{panns}}$ 2.49 $\rightarrow$ 2.07 (+16.87%), $KL_{\text{passt}}$ 1.78 $\rightarrow$ 1.53 (+14.04%), $IS_{\text{panns}}$ 4.17 $\rightarrow$ 4.30 (+3.12%), $IB_{\text{score}}$ 0.25 $\rightarrow$ 0.28 (+12.00%), $Energy\Delta10\text{ms}$ 0.3013 $\rightarrow$ 0.1349 (+55.23%), $Energy\Delta10\text{ms(vs.GT)}$ 0.0531 $\rightarrow$ 0.0288 (+45.76%), and $Sem.\,Rel.$ 2.73 $\rightarrow$ 3.28 (+20.15%). Our dataset aims to facilitate further research in long-form video-to-audio generation and is available at https://github.com/deepreasonings/long-form-video2audio.
摘要:Generating high-quality and temporally synchronized audio from video content is essential for video editing and post-production tasks, enabling the creation of semantically aligned audio for silent videos. However, most existing approaches focus on short-form audio generation for video segments under 10 seconds or rely on noisy datasets for long-form video-to-audio zsynthesis. To address these limitations, we introduce LD-LAudio-V1, an extension of state-of-the-art video-to-audio models and it incorporates dual lightweight adapters to enable long-form audio generation. In addition, we release a clean and human-annotated video-to-audio dataset that contains pure sound effects without noise or artifacts. Our method significantly reduces splicing artifacts and temporal inconsistencies while maintaining computational efficiency. Compared to direct fine-tuning with short training videos, LD-LAudio-V1 achieves significant improvements across multiple metrics: $FD_{\text{passt}}$ 450.00 $\rightarrow$ 327.29 (+27.27%), $FD_{\text{panns}}$ 34.88 $\rightarrow$ 22.68 (+34.98%), $FD_{\text{vgg}}$ 3.75 $\rightarrow$ 1.28 (+65.87%), $KL_{\text{panns}}$ 2.49 $\rightarrow$ 2.07 (+16.87%), $KL_{\text{passt}}$ 1.78 $\rightarrow$ 1.53 (+14.04%), $IS_{\text{panns}}$ 4.17 $\rightarrow$ 4.30 (+3.12%), $IB_{\text{score}}$ 0.25 $\rightarrow$ 0.28 (+12.00%), $Energy\Delta10\text{ms}$ 0.3013 $\rightarrow$ 0.1349 (+55.23%), $Energy\Delta10\text{ms(vs.GT)}$ 0.0531 $\rightarrow$ 0.0288 (+45.76%), and $Sem.\,Rel.$ 2.73 $\rightarrow$ 3.28 (+20.15%). Our dataset aims to facilitate further research in long-form video-to-audio generation and is available at https://github.com/deepreasonings/long-form-video2audio.


【15】Perturbed Public Voices (P$^{2}$V): A Dataset for Robust Audio Deepfake Detection
标题:受干扰的公共声音(P$^{2}$V):用于稳健音频深度伪造检测的数据集
链接:https://arxiv.org/abs/2508.10949

作者: Chongyang Gao, Marco Postiglione, Isabel Gortner, Sarit Kraus, V.S. Subrahmanian
摘要:目前的音频Deepfake检测器无法信任。虽然它们在受控基准测试中表现出色,但在现实世界中测试时却失败了。我们介绍了扰动公共声音(P $^{2}$V),这是一个IRB批准的数据集,它捕获了恶意deepfakes的三个关键方面:(1)通过LLM的身份一致性转录,(2)环境和对抗性噪声,以及(3)最先进的语音克隆(2020 - 2025)。实验揭示了22个最近的音频deepfake检测器的惊人漏洞:在当前数据集上训练的模型在P $^{2}$V上测试时损失了43%的性能,性能测量为deepfake音频,AUC和1-EER的F1得分的平均值。简单的对抗性扰动会导致高达16%的性能下降,而先进的克隆技术会使可检测性降低20 - 30%。相比之下,经过P $^{2}$V训练的模型保持了对这些攻击的鲁棒性,同时推广到现有的数据集,为鲁棒的音频deepfake检测建立了新的基准。P $^{2}$V将在会议/期刊接受后公开发布。
摘要:Current audio deepfake detectors cannot be trusted. While they excel on controlled benchmarks, they fail when tested in the real world. We introduce Perturbed Public Voices (P$^{2}$V), an IRB-approved dataset capturing three critical aspects of malicious deepfakes: (1) identity-consistent transcripts via LLMs, (2) environmental and adversarial noise, and (3) state-of-the-art voice cloning (2020-2025). Experiments reveal alarming vulnerabilities of 22 recent audio deepfake detectors: models trained on current datasets lose 43% performance when tested on P$^{2}$V, with performance measured as the mean of F1 score on deepfake audio, AUC, and 1-EER. Simple adversarial perturbations induce up to 16% performance degradation, while advanced cloning techniques reduce detectability by 20-30%. In contrast, P$^{2}$V-trained models maintain robustness against these attacks while generalizing to existing datasets, establishing a new benchmark for robust audio deepfake detection. P$^{2}$V will be publicly released upon acceptance by a conference/journal.


机器翻译由腾讯交互翻译提供,仅供参考