今日论文合集:cs.SD语音11篇,eess.AS音频处理14篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement
标题:MambAttention:具有多头注意力的Mamba,用于可推广的单通道语音增强
链接:https://arxiv.org/abs/2507.00966
作者:und Kühne, Jesper Jensen, Jan Østergaard, Zheng-Hua Tan
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing for possible publication
摘要:随着Mamba和xLSTM等新序列模型的出现,一些研究表明,这些模型在单通道语音增强、自动语音识别和自监督音频表示学习方面与最先进的模型相匹配或优于最先进的模型。然而,之前的研究已经表明,像LSTM和Mamba这样的序列模型往往会过拟合训练集。为了解决这个问题,以前的工作已经表明,添加自注意LSTM大大提高了单通道语音增强的泛化性能。然而,无论是混合曼巴和时间-频率注意力模型的概念,也没有他们的泛化性能进行了探讨语音增强。在本文中,我们提出了一种新的混合体系结构,MambAttention,它结合了Mamba和共享的时间和频率的多头注意力模块的可推广的单通道语音增强。为了训练我们的模型,我们引入了VoiceBank+Demand Extended(VB-DemandEx),这是一个受VoiceBank+Demand启发的数据集,但具有更具有挑战性的噪声类型和更低的信噪比。在VB-DemandEx上进行训练后,我们提出的MambAttention模型在两个域外数据集(DNS 2020和EARS-WHAM_v2)上的所有报告指标上都显著优于现有的最先进的基于LSTM、xLSTM、Mamba和Conformer的系统,而在域内数据集VB-DemandEx上的性能与之相匹配。消融研究强调了时间和频率多头注意模块之间的权重共享对泛化性能的作用。最后,我们探索将共享的时间和频率多头注意力模块与LSTM和xLSTM集成,这在域外数据集上产生了显着的性能改进。然而,我们的MambAttention模型在所有报告的评估指标中,在两个域外数据集上仍然具有优势。
摘要:With the advent of new sequence models like Mamba and xLSTM, several studies have shown that these models match or outperform state-of-the-art models in single-channel speech enhancement, automatic speech recognition, and self-supervised audio representation learning. However, prior research has demonstrated that sequence models like LSTM and Mamba tend to overfit to the training set. To address this issue, previous works have shown that adding self-attention to LSTMs substantially improves generalization performance for single-channel speech enhancement. Nevertheless, neither the concept of hybrid Mamba and time-frequency attention models nor their generalization performance have been explored for speech enhancement. In this paper, we propose a novel hybrid architecture, MambAttention, which combines Mamba and shared time- and frequency-multi-head attention modules for generalizable single-channel speech enhancement. To train our model, we introduce VoiceBank+Demand Extended (VB-DemandEx), a dataset inspired by VoiceBank+Demand but with more challenging noise types and lower signal-to-noise ratios. Trained on VB-DemandEx, our proposed MambAttention model significantly outperforms existing state-of-the-art LSTM-, xLSTM-, Mamba-, and Conformer-based systems of similar complexity across all reported metrics on two out-of-domain datasets: DNS 2020 and EARS-WHAM_v2, while matching their performance on the in-domain dataset VB-DemandEx. Ablation studies highlight the role of weight sharing between the time- and frequency-multi-head attention modules for generalization performance. Finally, we explore integrating the shared time- and frequency-multi-head attention modules with LSTM and xLSTM, which yields a notable performance improvement on the out-of-domain datasets. However, our MambAttention model remains superior on both out-of-domain datasets across all reported evaluation metrics.


【2】Multi-interaction TTS toward professional recording reproduction

标题:多互动TTC迈向专业录音复制
链接:https://arxiv.org/abs/2507.00808
作者:nagawa, Kenichi Fujita, Aya Watanabe, Yusuke Ijima
备注:7 pages,6 figures, Accepted to Speech Synthesis Workshop 2025 (SSW13)
摘要:配音导演经常通过提供反馈来迭代地改进配音演员的表演,以达到预期的效果。虽然这种基于迭代反馈的细化过程在实际录音中很重要,但在文本到语音合成(TTS)中却被忽视了。结果,即使合成的语音经常偏离用户的预期风格,在初始合成之后的细粒度风格细化也是不可能的。为了解决这个问题,我们提出了一个TTS方法与多步交互,让用户直观,快速地完善合成语音。我们的方法模型的TTS模型和它的用户之间的相互作用,模仿配音演员和语音导演之间的关系。实验表明,该模型及其相应的数据集,使迭代风格的改进,根据用户的方向,从而展示了其多交互能力。示例音频可在https://ntt-hilab-gensp上找到。github.io/ssw13multiinteraction_tts/
摘要:Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is important in actual recordings, it has been overlooked in text-to-speech synthesis (TTS). As a result, fine-grained style refinement after the initial synthesis is not possible, even though the synthesized speech often deviates from the user's intended style. To address this issue, we propose a TTS method with multi-step interaction that allows users to intuitively and rapidly refine synthetized speech. Our approach models the interaction between the TTS model and its user to emulate the relationship between voice actors and voice directors. Experiments show that the proposed model with its corresponding dataset enable iterative style refinements in accordance with users' directions, thus demonstrating its multi-interaction capability. Sample audios are available: https://ntt-hilab-gensp. github.io/ssw13multiinteraction_tts/


【3】Leveraging Large Language Models for Spontaneous Speech-Based Suicide Risk Detection

标题:利用大型语言模型进行基于自发言语的自杀风险检测
链接:https://arxiv.org/abs/2507.00693
作者:, Jiao Fu, Long Guo, Hong Liu
备注:Accepted to Interspeech 2025
摘要:早期识别自杀风险对预防自杀行为至关重要。因此,识别和研究与自杀风险相关的模式和标记物已成为当前研究的重点。在本文中,我们介绍了我们在第一届SpeechWellness Challenge(SW 1)中的工作结果,该挑战旨在探索语音作为识别青少年自杀风险的非侵入性和易于访问的心理健康指标。我们的方法利用大型语言模型(LLM)作为特征提取的主要工具,以及传统的声学和语义特征。该方法在测试集上达到了74%的准确率,在SW 1挑战中排名第一。这些研究结果表明,在自杀风险评估的背景下,基于LLM的方法分析语音的潜力。
摘要:Early identification of suicide risk is crucial for preventing suicidal behaviors. As a result, the identification and study of patterns and markers related to suicide risk have become a key focus of current research. In this paper, we present the results of our work in the 1st SpeechWellness Challenge (SW1), which aims to explore speech as a non-invasive and easily accessible mental health indicator for identifying adolescents at risk of suicide.Our approach leverages large language model (LLM) as the primary tool for feature extraction, alongside conventional acoustic and semantic features. The proposed method achieves an accuracy of 74\% on the test set, ranking first in the SW1 challenge. These findings demonstrate the potential of LLM-based methods for analyzing speech in the context of suicide risk assessment.


【4】MuteSwap: Silent Face-based Voice Conversion

标题:MuteSwap:基于无声面部的语音转换
链接:https://arxiv.org/abs/2507.00498
作者:, Yu Fang, Zhouhan Lin
摘要:传统的语音转换依赖于来自两侧的音频输入来修改从源说话者到目标说话者的语音特性。然而,当干净的音频不可用时,例如在无声视频或嘈杂环境中,该过程变得不可行。在这项工作中,我们专注于无声的基于面部的语音转换(SFVC),它完全从视觉输入进行语音转换的任务。也就是说,给定目标说话者的图像和包含嘴唇运动的源说话者的无声视频,SFVC生成对准目标说话者的身份的语音,同时保留源无声视频中的语音内容。由于这项任务需要生成可理解的语音并仅使用视觉线索转换身份,因此特别具有挑战性。为了解决这个问题,我们引入了MuteSwap,这是一个新的框架,它采用对比学习来对齐跨模态身份,并最小化互信息来分离共享的视觉特征。实验结果表明,MuteSwap在语音合成和身份转换方面都取得了令人印象深刻的性能,特别是在噪声条件下,依赖于音频输入的方法无法产生可理解的结果,证明了我们的训练方法的有效性和SFVC的可行性。
摘要:Conventional voice conversion modifies voice characteristics from a source speaker to a target speaker, relying on audio input from both sides. However, this process becomes infeasible when clean audio is unavailable, such as in silent videos or noisy environments. In this work, we focus on the task of Silent Face-based Voice Conversion (SFVC), which does voice conversion entirely from visual inputs. i.e., given images of a target speaker and a silent video of a source speaker containing lip motion, SFVC generates speech aligning the identity of the target speaker while preserving the speech content in the source silent video. As this task requires generating intelligible speech and converting identity using only visual cues, it is particularly challenging. To address this, we introduce MuteSwap, a novel framework that employs contrastive learning to align cross-modality identities and minimize mutual information to separate shared visual features. Experimental results show that MuteSwap achieves impressive performance in both speech synthesis and identity conversion, especially under noisy conditions where methods dependent on audio input fail to produce intelligible results, demonstrating both the effectiveness of our training approach and the feasibility of SFVC.


【5】AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences

标题:AudioBERTScore:基于音频嵌入序列相似性的环境声音合成客观评估
链接:https://arxiv.org/abs/2507.00475
作者:shi, Ryosuke Sakai, Shinnosuke Takamichi, Yusuke Kanamori, Yuki Okamoto
摘要:本文提出了一种新的文本到音频(TTA)合成音频的客观评价指标,旨在提高TTA模型的性能。在TTA中,合成声音的主观评价是重要的,但其实现需要金钱成本。因此,使用诸如梅尔倒谱失真的客观评价,但是这些客观度量与主观评价值之间的相关性弱。我们提出的客观评价指标,AudioBERTScore,计算嵌入的合成和参考声音之间的相似性。该方法不仅基于传统BERTScore中使用的最大范数,而且基于$p$范数来反映环境声音的非局部性质。实验结果表明,该方法得到的分数与主观评价值有更高的相关性比传统的度量。
摘要:We propose a novel objective evaluation metric for synthesized audio in text-to-audio (TTA), aiming to improve the performance of TTA models. In TTA, subjective evaluation of the synthesized sound is an important, but its implementation requires monetary costs. Therefore, objective evaluation such as mel-cepstral distortion are used, but the correlation between these objective metrics and subjective evaluation values is weak. Our proposed objective evaluation metric, AudioBERTScore, calculates the similarity between embedding of the synthesized and reference sounds. The method is based not only on the max-norm used in conventional BERTScore but also on the $p$-norm to reflect the non-local nature of environmental sounds. Experimental results show that scores obtained by the proposed method have a higher correlation with subjective evaluation values than conventional metrics.


【6】Beat and Downbeat Tracking in Performance MIDI Using an End-to-End Transformer Architecture

标题:使用端到端Transformer架构的性能数据库中的节拍和强拍跟踪
链接:https://arxiv.org/abs/2507.00466
作者: Murgul, Michael Heizmann
备注:Accepted to the 22nd Sound and Music Computing Conference (SMC), 2025
摘要:音乐演奏中的节拍跟踪是记谱级音乐转录和节奏分析的一项具有挑战性和重要性的任务,但现有的方法主要集中在基于音频的方法。本文提出了一种端到端的基于变换器的模型,用于在性能优化中进行节拍和强拍跟踪,利用编码器-解码器架构将节拍输入序列到序列转换为节拍注释。我们的方法引入了新颖的数据预处理技术,包括动态增强和优化的标记化策略,以提高不同数据集的准确性和可推广性。我们使用A-MAPS、ASAP、GuitarSet和Leduc数据集进行了广泛的实验,将我们的模型与最先进的隐马尔可夫模型(Hacking)和基于深度学习的节拍跟踪方法进行了比较。结果表明,我们的模型优于现有的符号音乐节拍跟踪方法,在各种音乐风格和乐器上获得了有竞争力的F1分数。我们的研究结果强调了Transformer架构在符号节拍跟踪方面的潜力,并建议未来与自动音乐转录系统集成,以增强音乐分析和乐谱生成。
摘要:Beat tracking in musical performance MIDI is a challenging and important task for notation-level music transcription and rhythmical analysis, yet existing methods primarily focus on audio-based approaches. This paper proposes an end-to-end transformer-based model for beat and downbeat tracking in performance MIDI, leveraging an encoder-decoder architecture for sequence-to-sequence translation of MIDI input to beat annotations. Our approach introduces novel data preprocessing techniques, including dynamic augmentation and optimized tokenization strategies, to improve accuracy and generalizability across different datasets. We conduct extensive experiments using the A-MAPS, ASAP, GuitarSet, and Leduc datasets, comparing our model against state-of-the-art hidden Markov models (HMMs) and deep learning-based beat tracking methods. The results demonstrate that our model outperforms existing symbolic music beat tracking approaches, achieving competitive F1-scores across various musical styles and instruments. Our findings highlight the potential of transformer architectures for symbolic beat tracking and suggest future integration with automatic music transcription systems for enhanced music analysis and score generation.


【7】A High-Fidelity Speech Super Resolution Network using a Complex Global Attention Module with Spectro-Temporal Loss

标题:使用具有谱时间损失的复杂全球注意力模块的高保真语音超分辨率网络
链接:https://arxiv.org/abs/2507.00229
作者:slam Tamiti, Biraj Joshi, Rida Hasan, Rashedul Hasan, Taieba Athay, Nursad Mamun, Anomadarshi Barua
摘要:语音超分辨率(SSR)通过提高采样率来增强低分辨率语音。虽然大多数SSR方法专注于幅度重建,但最近的研究突出了相位重建对改善感知质量的重要性。因此,我们引入了CTFT-Net,这是一种复杂的时频变换网络,可以在复杂域中重建幅度和相位,以改进SSR任务。它采用了一个复杂的全局注意力块来模拟音素间和频率间的依赖关系,并采用了一个复杂的构象来捕获长距离和局部特征,从而提高了频率重建和噪声鲁棒性。CTFT-Net采用时域和多分辨率频域损失函数,以获得更好的泛化能力。实验表明,CTFT-Net在VCTK数据集上的性能优于最先进的模型(NU-Wave,WSRGlow,NVSR,AERO),特别是对于极端上采样(2 kHz至48 kHz),有效地重建高频而没有噪声伪影。
摘要:Speech super-resolution (SSR) enhances low-resolution speech by increasing the sampling rate. While most SSR methods focus on magnitude reconstruction, recent research highlights the importance of phase reconstruction for improved perceptual quality. Therefore, we introduce CTFT-Net, a Complex Time-Frequency Transformation Network that reconstructs both magnitude and phase in complex domains for improved SSR tasks. It incorporates a complex global attention block to model inter-phoneme and inter-frequency dependencies and a complex conformer to capture long-range and local features, improving frequency reconstruction and noise robustness. CTFT-Net employs time-domain and multi-resolution frequency-domain loss functions for better generalization. Experiments show CTFT-Net outperforms state-of-the-art models (NU-Wave, WSRGlow, NVSR, AERO) on the VCTK dataset, particularly for extreme upsampling (2 kHz to 48 kHz), reconstructing high frequencies effectively without noisy artifacts.


【8】LearnAFE: Circuit-Algorithm Co-design Framework for Learnable Audio Analog Front-End

标题:LearnADE:可学习音频模拟前端的电路-算法协同设计框架
链接:https://arxiv.org/abs/2507.00755
作者:, Zhongyi Zhang, Cong Sheng Leow, Wang Ling Goh, Yuan Gao
备注:11 pages, 15 figures, accepted for publication on IEEE Transactions on Circuits and Systems I: Regular Papers
摘要:提出了一种用于音频信号分类的可学习模拟前端(AFE)的电路-算法协同设计框架。分开设计AFE和后端分类器是一种常见的做法,但并不理想,如本文所示。相反,本文提出了一种后端分类器与AFE的传递函数的联合优化,以实现系统级优化。更具体地,模拟带通滤波器(BPF)组的传递函数参数在分类器的信噪比(SNR)感知训练回路中被调谐。使用一个共同设计的损失函数LBPF,这项工作表现出优异的优化的滤波器组和分类器。在开源的SKY 130 130 nm CMOS工艺上实现,优化后的设计在输入信噪比从5 dB到20 dB的宽范围内,仅用22 k个分类器参数,就能实现10个关键词分类任务的90.5%-94.2%的准确率。与传统方法相比,所提出的音频AFE实现了8.7%和12.9%的功率和电容面积分别减少。
摘要:This paper presents a circuit-algorithm co-design framework for learnable analog front-end (AFE) in audio signal classification. Designing AFE and backend classifiers separately is a common practice but non-ideal, as shown in this paper. Instead, this paper proposes a joint optimization of the backend classifier with the AFE's transfer function to achieve system-level optimum. More specifically, the transfer function parameters of an analog bandpass filter (BPF) bank are tuned in a signal-to-noise ratio (SNR)-aware training loop for the classifier. Using a co-design loss function LBPF, this work shows superior optimization of both the filter bank and the classifier. Implemented in open-source SKY130 130nm CMOS process, the optimized design achieved 90.5%-94.2% accuracy for 10-keyword classification task across a wide range of input signal SNR from 5 dB to 20 dB, with only 22k classifier parameters. Compared to conventional approach, the proposed audio AFE achieves 8.7% and 12.9% reduction in power and capacitor area respectively.


【9】Mitigating Language Mismatch in SSL-Based Speaker Anonymization

标题:缓解基于SSL的说话人解析中的语言不匹配
链接:https://arxiv.org/abs/2507.00458
作者:, Wen-Chin Huang, Xin Wang, Xiaoxiao Miao, Junichi Yamagishi
备注:Accepted to Interspeech 2025
摘要:说话人匿名的目的是在保护说话人身份的同时,保留语音的内容信息和可懂度。然而,大多数发言人匿名系统(SAS)的开发和评估只使用英语,导致其他语言的效用下降。本文研究了日语和汉语语音合成中的语言失配问题。首先,我们用日语语音微调基于自监督学习(SSL)的内容编码器,以验证有效的语言适应。然后,我们提出了微调的多语言SSL模型与日语语音和评估的SAS在日语和普通话。下游实验表明,使用目标语言微调仅英语的SSL模型可以提高清晰度,同时维护隐私,并且多语言SSL进一步扩展了SAS在不同语言中的实用性。这些发现强调了语言适应和多语言预训练的SSL强大的多语言扬声器匿名的重要性。
摘要:Speaker anonymization aims to protect speaker identity while preserving content information and the intelligibility of speech. However, most speaker anonymization systems (SASs) are developed and evaluated using only English, resulting in degraded utility for other languages. This paper investigates language mismatch in SASs for Japanese and Mandarin speech. First, we fine-tune a self-supervised learning (SSL)-based content encoder with Japanese speech to verify effective language adaptation. Then, we propose fine-tuning a multilingual SSL model with Japanese speech and evaluating the SAS in Japanese and Mandarin. Downstream experiments show that fine-tuning an English-only SSL model with the target language enhances intelligibility while maintaining privacy and that multilingual SSL further extends SASs' utility across different languages. These findings highlight the importance of language adaptation and multilingual pre-training of SSLs for robust multilingual speaker anonymization.


【10】Do Music Source Separation Models Preserve Spatial Information in Binaural Audio?

标题:音乐源分离模型会保留双耳音频中的空间信息吗?
链接:https://arxiv.org/abs/2507.00155
作者:balla, Agnieszka Roginska, Magdalena Fuentes
备注:6 pages + references, 4 figures, 2 tables, 26th International Society for Music Information Retrieval (ISMIR) Conference
摘要:在音乐信息检索领域,双耳音频仍然是一个探索不足的领域。受虚拟和增强现实体验的日益普及以及可访问性的潜在应用的启发,我们研究了现有的音乐源分离(MSS)模型在双耳音频上的表现。虽然这些模型处理双通道输入,但尚不清楚它们如何有效地保留空间信息。在这项工作中,我们评估了几种流行的MSS模型如何在标准立体声和新颖的双耳数据集上保留空间信息。我们的双耳数据是使用MUSDB 18-HQ和开源头部相关的传递函数通过沿水平面随机定位仪器源来合成的。然后,我们使用信号处理和基于耳间提示的度量来评估分离的茎的空间质量。我们的研究结果表明,立体声MSS模型未能保留的空间信息,保持双耳音频的沉浸式质量的关键,以及退化取决于模型架构以及目标工具。最后,我们强调了MSS和沉浸式音频交叉点未来工作的宝贵机会。
摘要:Binaural audio remains underexplored within the music information retrieval community. Motivated by the rising popularity of virtual and augmented reality experiences as well as potential applications to accessibility, we investigate how well existing music source separation (MSS) models perform on binaural audio. Although these models process two-channel inputs, it is unclear how effectively they retain spatial information. In this work, we evaluate how several popular MSS models preserve spatial information on both standard stereo and novel binaural datasets. Our binaural data is synthesized using stems from MUSDB18-HQ and open-source head-related transfer functions by positioning instrument sources randomly along the horizontal plane. We then assess the spatial quality of the separated stems using signal processing and interaural cue-based metrics. Our results show that stereo MSS models fail to preserve the spatial information critical for maintaining the immersive quality of binaural audio, and that the degradation depends on model architecture as well as the target instrument. Finally, we highlight valuable opportunities for future work at the intersection of MSS and immersive audio.


【11】Musical Source Separation of Brazilian Percussion

标题:巴西打击乐的音乐源分离
链接:https://arxiv.org/abs/2503.04995
作者:balla, Giovana Morais, Magdalena Fuentes
备注:2 pages + references, 1 figure, 1 table, Extended Abstracts for the Late-Breaking Demo Session of the 25th International Society for Music Information Retrieval Conference
摘要:音乐源分离(MSS)最近在西方音乐背景下将乐器从混合物中分离出来方面取得了重大突破,但由于缺乏数据,对非西方乐器的研究仍然有限。在这个演示中,我们使用现有的巴西sama打击乐器数据集来创建人工混合物,用于训练U-Net模型来分离桑巴舞中的传统乐器surdo鼓。尽管训练数据有限,但考虑到鼓的重复模式及其特有的低音音色,该模型有效地隔离了鼓。这些结果表明,MSS系统可以成功地利用在更具有文化包容性的情况下工作,而不需要收集大量的数据。
摘要:Musical source separation (MSS) has recently seen a big breakthrough in separating instruments from a mixture in the context of Western music, but research on non-Western instruments is still limited due to a lack of data. In this demo, we use an existing dataset of Brazilian sama percussion to create artificial mixtures for training a U-Net model to separate the surdo drum, a traditional instrument in samba. Despite limited training data, the model effectively isolates the surdo, given the drum's repetitive patterns and its characteristic low-pitched timbre. These results suggest that MSS systems can be successfully harnessed to work in more culturally-inclusive scenarios without the need of collecting extensive amounts of data.


eess.AS音频处理


【1】Improving Stereo 3D Sound Event Localization and Detection: Perceptual Features, Stereo-specific Data Augmentation, and Distance Normalization

标题:改进立体声3D声音事件定位和检测:感知特征、立体声特定数据增强和距离标准化
链接:https://arxiv.org/abs/2507.00874
作者:eow, Ee-Leng Tan, Santi Peksi, Woon-Seng Gan
备注:Technical report for DCASE 2025 Challenge Task 3
摘要:本技术报告介绍了我们对DCASE 2025挑战赛任务3的提交:常规视频内容中的立体声事件定位和检测(SELD)。我们在本报告中讨论了纯音频任务,并介绍了几个关键贡献。首先,我们设计感知动机的输入功能,提高事件检测,声源定位和距离估计。其次,我们专门针对立体声音频的复杂性调整增强策略,包括通道交换和时频掩蔽。我们还将最近提出的FilterAugment技术,尚未探索SELD工作。最后,我们在训练过程中应用距离归一化方法来稳定回归目标。在立体STARSS23数据集上的实验表明,在所有SELD指标上都有一致的性能增益。复制我们工作的代码可以在以下存储库中找到:https://github.com/itsjunwei/NTU_SNTL_Task3
摘要:This technical report presents our submission to Task 3 of the DCASE 2025 Challenge: Stereo Sound Event Localization and Detection (SELD) in Regular Video Content. We address the audio-only task in this report and introduce several key contributions. First, we design perceptually-motivated input features that improve event detection, sound source localization, and distance estimation. Second, we adapt augmentation strategies specifically for the intricacies of stereo audio, including channel swapping and time-frequency masking. We also incorporate the recently proposed FilterAugment technique that has yet to be explored for SELD work. Lastly, we apply a distance normalization approach during training to stabilize regression targets. Experiments on the stereo STARSS23 dataset demonstrate consistent performance gains across all SELD metrics. Code to replicate our work is available in this repository: https://github.com/itsjunwei/NTU_SNTL_Task3


【2】LearnAFE: Circuit-Algorithm Co-design Framework for Learnable Audio Analog Front-End

标题:LearnADE:可学习音频模拟前端的电路-算法协同设计框架
链接:https://arxiv.org/abs/2507.00755
作者:, Zhongyi Zhang, Cong Sheng Leow, Wang Ling Goh, Yuan Gao
备注:11 pages, 15 figures, accepted for publication on IEEE Transactions on Circuits and Systems I: Regular Papers
摘要:提出了一种用于音频信号分类的可学习模拟前端(AFE)的电路-算法协同设计框架。分开设计AFE和后端分类器是一种常见的做法,但并不理想,如本文所示。相反,本文提出了一种后端分类器与AFE的传递函数的联合优化,以实现系统级优化。更具体地,模拟带通滤波器(BPF)组的传递函数参数在分类器的信噪比(SNR)感知训练回路中被调谐。使用一个共同设计的损失函数LBPF,这项工作表现出优异的优化的滤波器组和分类器。在开源的SKY 130 130 nm CMOS工艺上实现,优化后的设计在输入信噪比从5 dB到20 dB的宽范围内,仅用22 k个分类器参数,就能实现10个关键词分类任务的90.5%-94.2%的准确率。与传统方法相比,所提出的音频AFE实现了8.7%和12.9%的功率和电容面积分别减少。
摘要:This paper presents a circuit-algorithm co-design framework for learnable analog front-end (AFE) in audio signal classification. Designing AFE and backend classifiers separately is a common practice but non-ideal, as shown in this paper. Instead, this paper proposes a joint optimization of the backend classifier with the AFE's transfer function to achieve system-level optimum. More specifically, the transfer function parameters of an analog bandpass filter (BPF) bank are tuned in a signal-to-noise ratio (SNR)-aware training loop for the classifier. Using a co-design loss function LBPF, this work shows superior optimization of both the filter bank and the classifier. Implemented in open-source SKY130 130nm CMOS process, the optimized design achieved 90.5%-94.2% accuracy for 10-keyword classification task across a wide range of input signal SNR from 5 dB to 20 dB, with only 22k classifier parameters. Compared to conventional approach, the proposed audio AFE achieves 8.7% and 12.9% reduction in power and capacitor area respectively.


【3】Mitigating Language Mismatch in SSL-Based Speaker Anonymization

标题:缓解基于SSL的说话人解析中的语言不匹配
链接:https://arxiv.org/abs/2507.00458
作者:, Wen-Chin Huang, Xin Wang, Xiaoxiao Miao, Junichi Yamagishi
备注:Accepted to Interspeech 2025
摘要:说话人匿名的目的是在保护说话人身份的同时,保留语音的内容信息和可懂度。然而,大多数发言人匿名系统(SAS)的开发和评估只使用英语,导致其他语言的效用下降。本文研究了日语和汉语语音合成中的语言失配问题。首先,我们用日语语音微调基于自监督学习(SSL)的内容编码器,以验证有效的语言适应。然后,我们提出了微调的多语言SSL模型与日语语音和评估的SAS在日语和普通话。下游的实验表明,微调的英语SSL模型与目标语言提高了可理解性,同时保持隐私和多语言SSL进一步扩展SAS的效用在不同的语言。这些发现强调了语言适应和多语言预训练的SSL强大的多语言扬声器匿名的重要性。
摘要:Speaker anonymization aims to protect speaker identity while preserving content information and the intelligibility of speech. However, most speaker anonymization systems (SASs) are developed and evaluated using only English, resulting in degraded utility for other languages. This paper investigates language mismatch in SASs for Japanese and Mandarin speech. First, we fine-tune a self-supervised learning (SSL)-based content encoder with Japanese speech to verify effective language adaptation. Then, we propose fine-tuning a multilingual SSL model with Japanese speech and evaluating the SAS in Japanese and Mandarin. Downstream experiments show that fine-tuning an English-only SSL model with the target language enhances intelligibility while maintaining privacy and that multilingual SSL further extends SASs' utility across different languages. These findings highlight the importance of language adaptation and multilingual pre-training of SSLs for robust multilingual speaker anonymization.


【4】Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges

标题:收集、策划和注释名人的高质量语音深度伪造数据集:过程和挑战
链接:https://arxiv.org/abs/2507.00324
作者:i, Surya Subramani, Raksha Varahamurthy, Nithin Adupa, Lekha Bollinani, Hafiz Malik
摘要:语音合成的最新进展在保持语音真实性方面带来了前所未有的挑战,特别是关于经常成为模仿攻击目标的公众人物。本文提出了一种全面的方法,用于收集,策划和生成政治人物的合成语音数据,并详细分析了所遇到的挑战。我们介绍了一种系统的方法,结合自动化管道收集高质量的真正的语音样本,具有基于转录的分割,显着提高合成语音质量。我们尝试了各种合成方法;从单扬声器到zero-shot合成,并记录了我们方法的演变。由此产生的数据集包括来自10位公众人物的真实和合成语音样本,表现出卓越的质量,NISQA-TTS自然度得分为3.69,最高的人类错误分类率为61.9%。
摘要:Recent advances in speech synthesis have introduced unprecedented challenges in maintaining voice authenticity, particularly concerning public figures who are frequent targets of impersonation attacks. This paper presents a comprehensive methodology for collecting, curating, and generating synthetic speech data for political figures and a detailed analysis of challenges encountered. We introduce a systematic approach incorporating an automated pipeline for collecting high-quality bonafide speech samples, featuring transcription-based segmentation that significantly improves synthetic speech quality. We experimented with various synthesis approaches; from single-speaker to zero-shot synthesis, and documented the evolution of our methodology. The resulting dataset comprises bonafide and synthetic speech samples from ten public figures, demonstrating superior quality with a NISQA-TTS naturalness score of 3.69 and the highest human misclassification rate of 61.9\%.


【5】Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis

标题:语音合成中韵律建模的随机方法研究
链接:https://arxiv.org/abs/2507.00227
作者:r, Florian Lux, Alejandro Pérez-González-de-Martos, Angelina Elizarova, Lindsey Vanderlyn, Dirk Väth, Ngoc Thang Vu
备注:Accepted at Interspeech 2025
摘要:虽然近年来生成方法发展迅速,但生成表达性韵律仍然是文本到语音合成中具有挑战性的任务。这对于通过诸如音高、能量和持续时间等参数来明确地对韵律进行建模的系统尤其如此,这通常是为了可解释性和可控性而进行的。在这项工作中,我们研究了随机方法用于此任务的有效性,包括规范化流、条件流匹配和整流。我们将这些方法与传统的确定性基线以及真实的人类实现进行比较。我们广泛的主观和客观的评价表明,随机方法产生的自然韵律与人类说话者通过捕捉人类语音中固有的变化。此外,它们通过允许调整采样温度来打开额外的可控性选项。
摘要:While generative methods have progressed rapidly in recent years, generating expressive prosody for an utterance remains a challenging task in text-to-speech synthesis. This is particularly true for systems that model prosody explicitly through parameters such as pitch, energy, and duration, which is commonly done for the sake of interpretability and controllability. In this work, we investigate the effectiveness of stochastic methods for this task, including Normalizing Flows, Conditional Flow Matching, and Rectified Flows. We compare these methods to a traditional deterministic baseline, as well as to real human realizations. Our extensive subjective and objective evaluations demonstrate that stochastic methods produce natural prosody on par with human speakers by capturing the variability inherent in human speech. Further, they open up additional controllability options by allowing the sampling temperature to be tuned.


【6】Do Music Source Separation Models Preserve Spatial Information in Binaural Audio?

标题:音乐源分离模型会保留双耳音频中的空间信息吗?
链接:https://arxiv.org/abs/2507.00155
作者:balla, Agnieszka Roginska, Magdalena Fuentes
备注:6 pages + references, 4 figures, 2 tables, 26th International Society for Music Information Retrieval (ISMIR) Conference
摘要:在音乐信息检索领域,双耳音频仍然是一个探索不足的领域。受虚拟和增强现实体验的日益普及以及可访问性的潜在应用的启发,我们研究了现有的音乐源分离(MSS)模型在双耳音频上的表现。虽然这些模型处理双通道输入,但尚不清楚它们如何有效地保留空间信息。在这项工作中,我们评估了几种流行的MSS模型如何在标准立体声和新颖的双耳数据集上保留空间信息。我们的双耳数据是使用MUSDB 18-HQ和开源头部相关的传递函数通过沿水平面随机定位仪器源来合成的。然后,我们使用信号处理和基于耳间提示的度量来评估分离的茎的空间质量。我们的研究结果表明,立体声MSS模型无法保留的空间信息,保持双耳音频的沉浸式质量的关键,以及退化取决于模型架构以及目标工具。最后,我们强调了MSS和沉浸式音频交叉点未来工作的宝贵机会。
摘要:Binaural audio remains underexplored within the music information retrieval community. Motivated by the rising popularity of virtual and augmented reality experiences as well as potential applications to accessibility, we investigate how well existing music source separation (MSS) models perform on binaural audio. Although these models process two-channel inputs, it is unclear how effectively they retain spatial information. In this work, we evaluate how several popular MSS models preserve spatial information on both standard stereo and novel binaural datasets. Our binaural data is synthesized using stems from MUSDB18-HQ and open-source head-related transfer functions by positioning instrument sources randomly along the horizontal plane. We then assess the spatial quality of the separated stems using signal processing and interaural cue-based metrics. Our results show that stereo MSS models fail to preserve the spatial information critical for maintaining the immersive quality of binaural audio, and that the degradation depends on model architecture as well as the target instrument. Finally, we highlight valuable opportunities for future work at the intersection of MSS and immersive audio.


【7】MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement

标题:MambAttention:具有多头注意力的Mamba,用于可推广的单通道语音增强
链接:https://arxiv.org/abs/2507.00966
作者:und Kühne, Jesper Jensen, Jan Østergaard, Zheng-Hua Tan
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing for possible publication
摘要:随着Mamba和xLSTM等新序列模型的出现,一些研究表明,这些模型在单通道语音增强、自动语音识别和自监督音频表示学习方面与最先进的模型相匹配或优于最先进的模型。然而,之前的研究已经表明,像LSTM和Mamba这样的序列模型往往会过拟合训练集。为了解决这个问题,以前的工作已经表明,添加自注意LSTM大大提高了单通道语音增强的泛化性能。然而,无论是混合曼巴和时间-频率注意力模型的概念,也没有他们的泛化性能进行了探讨语音增强。在本文中,我们提出了一种新的混合体系结构,MambAttention,它结合了Mamba和共享的时间和频率的多头注意力模块的可推广的单通道语音增强。为了训练我们的模型,我们引入了VoiceBank+Demand Extended(VB-DemandEx),这是一个受VoiceBank+Demand启发的数据集,但具有更具有挑战性的噪声类型和更低的信噪比。在VB-DemandEx上进行训练后,我们提出的MambAttention模型在两个域外数据集(DNS 2020和EARS-WHAM_v2)上的所有报告指标上都显著优于现有的最先进的基于LSTM、xLSTM、Mamba和Conformer的系统,而在域内数据集VB-DemandEx上的性能与之相匹配。消融研究强调了时间和频率多头注意模块之间的权重共享对泛化性能的作用。最后,我们探索将共享的时间和频率多头注意力模块与LSTM和xLSTM集成,这在域外数据集上产生了显着的性能改进。然而,我们的MambAttention模型在所有报告的评估指标中,在两个域外数据集上仍然具有优势。
摘要:With the advent of new sequence models like Mamba and xLSTM, several studies have shown that these models match or outperform state-of-the-art models in single-channel speech enhancement, automatic speech recognition, and self-supervised audio representation learning. However, prior research has demonstrated that sequence models like LSTM and Mamba tend to overfit to the training set. To address this issue, previous works have shown that adding self-attention to LSTMs substantially improves generalization performance for single-channel speech enhancement. Nevertheless, neither the concept of hybrid Mamba and time-frequency attention models nor their generalization performance have been explored for speech enhancement. In this paper, we propose a novel hybrid architecture, MambAttention, which combines Mamba and shared time- and frequency-multi-head attention modules for generalizable single-channel speech enhancement. To train our model, we introduce VoiceBank+Demand Extended (VB-DemandEx), a dataset inspired by VoiceBank+Demand but with more challenging noise types and lower signal-to-noise ratios. Trained on VB-DemandEx, our proposed MambAttention model significantly outperforms existing state-of-the-art LSTM-, xLSTM-, Mamba-, and Conformer-based systems of similar complexity across all reported metrics on two out-of-domain datasets: DNS 2020 and EARS-WHAM_v2, while matching their performance on the in-domain dataset VB-DemandEx. Ablation studies highlight the role of weight sharing between the time- and frequency-multi-head attention modules for generalization performance. Finally, we explore integrating the shared time- and frequency-multi-head attention modules with LSTM and xLSTM, which yields a notable performance improvement on the out-of-domain datasets. However, our MambAttention model remains superior on both out-of-domain datasets across all reported evaluation metrics.


【8】Multi-interaction TTS toward professional recording reproduction

标题:多互动TTC迈向专业录音复制
链接:https://arxiv.org/abs/2507.00808
作者:nagawa, Kenichi Fujita, Aya Watanabe, Yusuke Ijima
备注:7 pages,6 figures, Accepted to Speech Synthesis Workshop 2025 (SSW13)
摘要:配音导演经常通过提供反馈来迭代地改进配音演员的表演,以达到预期的效果。虽然这种基于迭代反馈的细化过程在实际录音中很重要,但在文本到语音合成(TTS)中却被忽视了。结果,即使合成的语音经常偏离用户的预期风格,在初始合成之后的细粒度风格细化也是不可能的。为了解决这个问题,我们提出了一个TTS方法与多步交互,让用户直观,快速地完善合成语音。我们的方法模型的TTS模型和它的用户之间的相互作用,模仿配音演员和语音导演之间的关系。实验表明,该模型及其相应的数据集,使迭代风格的改进,根据用户的方向,从而展示了其多交互能力。示例音频可在https://ntt-hilab-gensp上找到。github.io/ssw13multiinteraction_tts/
摘要:Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is important in actual recordings, it has been overlooked in text-to-speech synthesis (TTS). As a result, fine-grained style refinement after the initial synthesis is not possible, even though the synthesized speech often deviates from the user's intended style. To address this issue, we propose a TTS method with multi-step interaction that allows users to intuitively and rapidly refine synthetized speech. Our approach models the interaction between the TTS model and its user to emulate the relationship between voice actors and voice directors. Experiments show that the proposed model with its corresponding dataset enable iterative style refinements in accordance with users' directions, thus demonstrating its multi-interaction capability. Sample audios are available: https://ntt-hilab-gensp. github.io/ssw13multiinteraction_tts/


【9】Leveraging Large Language Models for Spontaneous Speech-Based Suicide Risk Detection

标题:利用大型语言模型进行基于自发言语的自杀风险检测
链接:https://arxiv.org/abs/2507.00693
作者:, Jiao Fu, Long Guo, Hong Liu
备注:Accepted to Interspeech 2025
摘要:早期识别自杀风险对预防自杀行为至关重要。因此,识别和研究与自杀风险相关的模式和标记物已成为当前研究的重点。在本文中,我们介绍了我们在第一届SpeechWellness Challenge(SW 1)中的工作结果,该挑战旨在探索语音作为识别青少年自杀风险的非侵入性和易于访问的心理健康指标。我们的方法利用大型语言模型(LLM)作为特征提取的主要工具,以及传统的声学和语义特征。该方法在测试集上达到了74%的准确率,在SW 1挑战中排名第一。这些研究结果表明,在自杀风险评估的背景下,基于LLM的方法分析语音的潜力。
摘要:Early identification of suicide risk is crucial for preventing suicidal behaviors. As a result, the identification and study of patterns and markers related to suicide risk have become a key focus of current research. In this paper, we present the results of our work in the 1st SpeechWellness Challenge (SW1), which aims to explore speech as a non-invasive and easily accessible mental health indicator for identifying adolescents at risk of suicide.Our approach leverages large language model (LLM) as the primary tool for feature extraction, alongside conventional acoustic and semantic features. The proposed method achieves an accuracy of 74\% on the test set, ranking first in the SW1 challenge. These findings demonstrate the potential of LLM-based methods for analyzing speech in the context of suicide risk assessment.


【10】MuteSwap: Silent Face-based Voice Conversion

标题:MuteSwap:基于无声面部的语音转换
链接:https://arxiv.org/abs/2507.00498
作者:, Yu Fang, Zhouhan Lin
摘要:传统的语音转换依赖于来自两侧的音频输入来修改从源说话者到目标说话者的语音特性。然而,当干净的音频不可用时,例如在无声视频或嘈杂环境中,该过程变得不可行。在这项工作中,我们专注于无声的基于面部的语音转换(SFVC),它完全从视觉输入进行语音转换的任务。也就是说,给定目标说话者的图像和包含嘴唇运动的源说话者的无声视频,SFVC生成对准目标说话者的身份的语音,同时保留源无声视频中的语音内容。由于这项任务需要生成可理解的语音并仅使用视觉线索转换身份,因此特别具有挑战性。为了解决这个问题,我们引入了MuteSwap,这是一个新的框架,它采用对比学习来对齐跨模态身份,并最小化互信息来分离共享的视觉特征。实验结果表明,MuteSwap在语音合成和身份转换方面都取得了令人印象深刻的性能,特别是在噪声条件下,依赖于音频输入的方法无法产生可理解的结果,证明了我们的训练方法的有效性和SFVC的可行性。
摘要:Conventional voice conversion modifies voice characteristics from a source speaker to a target speaker, relying on audio input from both sides. However, this process becomes infeasible when clean audio is unavailable, such as in silent videos or noisy environments. In this work, we focus on the task of Silent Face-based Voice Conversion (SFVC), which does voice conversion entirely from visual inputs. i.e., given images of a target speaker and a silent video of a source speaker containing lip motion, SFVC generates speech aligning the identity of the target speaker while preserving the speech content in the source silent video. As this task requires generating intelligible speech and converting identity using only visual cues, it is particularly challenging. To address this, we introduce MuteSwap, a novel framework that employs contrastive learning to align cross-modality identities and minimize mutual information to separate shared visual features. Experimental results show that MuteSwap achieves impressive performance in both speech synthesis and identity conversion, especially under noisy conditions where methods dependent on audio input fail to produce intelligible results, demonstrating both the effectiveness of our training approach and the feasibility of SFVC.


【11】AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences

标题:AudioBERTScore:基于音频嵌入序列相似性的环境声音合成客观评估
链接:https://arxiv.org/abs/2507.00475
作者:shi, Ryosuke Sakai, Shinnosuke Takamichi, Yusuke Kanamori, Yuki Okamoto
摘要:本文提出了一种新的文本到音频(TTA)合成音频的客观评价指标,旨在提高TTA模型的性能。在TTA中,合成声音的主观评价是重要的,但其实现需要金钱成本。因此,使用诸如梅尔倒谱失真的客观评价,但是这些客观度量与主观评价值之间的相关性弱。我们提出的客观评价指标,AudioBERTScore,计算嵌入的合成和参考声音之间的相似性。该方法不仅基于传统BERTScore中使用的最大范数,而且基于$p$范数来反映环境声音的非局部性质。实验结果表明,该方法得到的分数与主观评价值有更高的相关性比传统的度量。
摘要:We propose a novel objective evaluation metric for synthesized audio in text-to-audio (TTA), aiming to improve the performance of TTA models. In TTA, subjective evaluation of the synthesized sound is an important, but its implementation requires monetary costs. Therefore, objective evaluation such as mel-cepstral distortion are used, but the correlation between these objective metrics and subjective evaluation values is weak. Our proposed objective evaluation metric, AudioBERTScore, calculates the similarity between embedding of the synthesized and reference sounds. The method is based not only on the max-norm used in conventional BERTScore but also on the $p$-norm to reflect the non-local nature of environmental sounds. Experimental results show that scores obtained by the proposed method have a higher correlation with subjective evaluation values than conventional metrics.


【12】Beat and Downbeat Tracking in Performance MIDI Using an End-to-End Transformer Architecture

标题:使用端到端Transformer架构的性能数据库中的节拍和强拍跟踪
链接:https://arxiv.org/abs/2507.00466
作者: Murgul, Michael Heizmann
备注:Accepted to the 22nd Sound and Music Computing Conference (SMC), 2025
摘要:音乐演奏中的节拍跟踪是记谱级音乐转录和节奏分析的一项具有挑战性和重要性的任务,但现有的方法主要集中在基于音频的方法。本文提出了一种端到端的基于变换器的模型,用于在性能优化中进行节拍和强拍跟踪,利用编码器-解码器架构将节拍输入序列到序列转换为节拍注释。我们的方法引入了新颖的数据预处理技术,包括动态增强和优化的标记化策略,以提高不同数据集的准确性和可推广性。我们使用A-MAPS、ASAP、GuitarSet和Leduc数据集进行了广泛的实验,将我们的模型与最先进的隐马尔可夫模型(Hacking)和基于深度学习的节拍跟踪方法进行了比较。结果表明,我们的模型优于现有的符号音乐节拍跟踪方法,在各种音乐风格和乐器上获得了有竞争力的F1分数。我们的研究结果突出了Transformer架构在符号节拍跟踪方面的潜力,并建议将来与自动音乐转录系统集成,以增强音乐分析和乐谱生成。
摘要:Beat tracking in musical performance MIDI is a challenging and important task for notation-level music transcription and rhythmical analysis, yet existing methods primarily focus on audio-based approaches. This paper proposes an end-to-end transformer-based model for beat and downbeat tracking in performance MIDI, leveraging an encoder-decoder architecture for sequence-to-sequence translation of MIDI input to beat annotations. Our approach introduces novel data preprocessing techniques, including dynamic augmentation and optimized tokenization strategies, to improve accuracy and generalizability across different datasets. We conduct extensive experiments using the A-MAPS, ASAP, GuitarSet, and Leduc datasets, comparing our model against state-of-the-art hidden Markov models (HMMs) and deep learning-based beat tracking methods. The results demonstrate that our model outperforms existing symbolic music beat tracking approaches, achieving competitive F1-scores across various musical styles and instruments. Our findings highlight the potential of transformer architectures for symbolic beat tracking and suggest future integration with automatic music transcription systems for enhanced music analysis and score generation.


【13】A High-Fidelity Speech Super Resolution Network using a Complex Global Attention Module with Spectro-Temporal Loss

标题:使用具有谱时间损失的复杂全球注意力模块的高保真语音超分辨率网络
链接:https://arxiv.org/abs/2507.00229
作者:slam Tamiti, Biraj Joshi, Rida Hasan, Rashedul Hasan, Taieba Athay, Nursad Mamun, Anomadarshi Barua
摘要:语音超分辨率(SSR)通过提高采样率来增强低分辨率语音。虽然大多数SSR方法专注于幅度重建,但最近的研究突出了相位重建对改善感知质量的重要性。因此,我们引入了CTFT-Net,这是一种复杂的时频变换网络,可以在复杂域中重建幅度和相位,以改进SSR任务。它采用了一个复杂的全局注意力块来模拟音素间和频率间的依赖关系,并采用了一个复杂的构象来捕获长距离和局部特征,从而提高了频率重建和噪声鲁棒性。CTFT-Net采用时域和多分辨率频域损失函数,以获得更好的泛化能力。实验表明,CTFT-Net在VCTK数据集上的性能优于最先进的模型(NU-Wave,WSRGlow,NVSR,AERO),特别是对于极端上采样(2 kHz至48 kHz),有效地重建高频而没有噪声伪影。
摘要:Speech super-resolution (SSR) enhances low-resolution speech by increasing the sampling rate. While most SSR methods focus on magnitude reconstruction, recent research highlights the importance of phase reconstruction for improved perceptual quality. Therefore, we introduce CTFT-Net, a Complex Time-Frequency Transformation Network that reconstructs both magnitude and phase in complex domains for improved SSR tasks. It incorporates a complex global attention block to model inter-phoneme and inter-frequency dependencies and a complex conformer to capture long-range and local features, improving frequency reconstruction and noise robustness. CTFT-Net employs time-domain and multi-resolution frequency-domain loss functions for better generalization. Experiments show CTFT-Net outperforms state-of-the-art models (NU-Wave, WSRGlow, NVSR, AERO) on the VCTK dataset, particularly for extreme upsampling (2 kHz to 48 kHz), reconstructing high frequencies effectively without noisy artifacts.


【14】Leveraging Unlabeled Audio-Visual Data in Speech Emotion Recognition using Knowledge Distillation

标题:基于知识蒸馏的语音情感识别中未标记音视频数据的利用
链接:https://arxiv.org/abs/2507.00055
作者:ndyala, Pedro Morgado, William Sethares
备注:Accepted at INTERSPEECH 2025
摘要:集成到人机交互系统的语音接口可以受益于语音情感识别(SER)以基于用户情感定制响应。由于人类通过多模态视听线索传达情感,因此开发使用这两种模态的SER系统是有益的。然而,收集大量的标记数据用于其开发是昂贵的。本文提出了一个名为LightweightSER(LiSER)的知识蒸馏框架,该框架利用未标记的视听数据进行SER,使用基于高级语音和面部表示模型构建的大型教师模型。LiSER将有关语音情感和面部表情的知识从教师模型转移到轻量级学生模型。在两个基准数据集RAVDESS和CREMA-D上进行的实验表明,LiSER可以减少SER任务对大量标记数据集的依赖。
摘要:Voice interfaces integral to the human-computer interaction systems can benefit from speech emotion recognition (SER) to customize responses based on user emotions. Since humans convey emotions through multi-modal audio-visual cues, developing SER systems using both the modalities is beneficial. However, collecting a vast amount of labeled data for their development is expensive. This paper proposes a knowledge distillation framework called LightweightSER (LiSER) that leverages unlabeled audio-visual data for SER, using large teacher models built on advanced speech and face representation models. LiSER transfers knowledge regarding speech emotions and facial expressions from the teacher models to lightweight student models. Experiments conducted on two benchmark datasets, RAVDESS and CREMA-D, demonstrate that LiSER can reduce the dependence on extensive labeled datasets for SER tasks.


机器翻译由腾讯交互翻译提供,仅供参考