今日论文合集:cs.SD语音11篇,eess.AS音频处理11篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】 On Adversarial Attacks In Acoustic Drone Localization

标题:声学无人机定位中的对抗攻击
链接:https://arxiv.org/abs/2502.20325
作者:Tamir Shor,  Chaim Baskin,  Alex Bronstein
摘要:近年来,多旋翼航空自主飞行器(MAV,更广泛地称为“无人机”)由于其在广泛和多样化的领域(例如,农业、商业交付、搜索和救援)。基于视觉的方法对照明条件和遮挡的敏感性促使人们越来越多地研究依赖于其他方式的导航,例如声学传感。在非受控环境中大规模使用无人机执行任务的一个主要问题是其导航系统可能受到敌对攻击的威胁,使用户面临关键任务故障,安全漏洞和可能危及操作员和旁观者的安全结果。虽然以前的工作在基于声学的无人机定位方面取得了令人印象深刻的进展,但以前对无人机导航的对抗性攻击的研究只涉及基于视觉传感的系统。在这项工作中,我们的目标是通过全面分析PGD对抗性攻击对声学无人机定位的影响来弥补这一差距。我们还开发了一种对抗扰动恢复算法,能够显着减少这种攻击在我们的设置的影响。用于复制所有实验的代码将在出版时发布。
摘要:Multi-rotor aerial autonomous vehicles (MAVs, more widely known as "drones")have been generating increased interest in recent years due to their growingapplicability in a vast and diverse range of fields (e.g., agriculture,commercial delivery, search and rescue). The sensitivity of visual-basedmethods to lighting conditions and occlusions had prompted growing study ofnavigation reliant on other modalities, such as acoustic sensing. A majorconcern in using drones in scale for tasks in non-controlled environments isthe potential threat of adversarial attacks over their navigational systems,exposing users to mission-critical failures, security breaches, and compromisedsafety outcomes that can endanger operators and bystanders. While previous workshows impressive progress in acoustic-based drone localization, prior researchin adversarial attacks over drone navigation only addresses visualsensing-based systems. In this work, we aim to compensate for this gap bysupplying a comprehensive analysis of the effect of PGD adversarial attacksover acoustic drone localization. We furthermore develop an algorithm foradversarial perturbation recovery, capable of markedly diminishing the affectof such attacks in our setting. The code for reproducing all experiments willbe released upon publication.

【2】 Adapting Automatic Speech Recognition for Accented Air Traffic Control  Communications
标题:将自动语音识别适应强调的空中交通管制通信
链接:https://arxiv.org/abs/2502.20311
作者:Marcus Yu Zhe Wee,  Justin Juin Hng Wong,  Lynus Lim,  Joe Yu Wei Tan,  Prannaya Gupta,  Dillion Lim,  En Hao Tew,  Aloysius Keng Siew Han,  Yong Zhi Lim
摘要:空中交通管制(ATC)中的有效沟通对于维护航空安全至关重要,但口音英语带来的挑战在自动语音识别(ASR)系统中仍然没有得到解决。现有的模型与东南亚口音(SEA口音)语音的转录准确性斗争,特别是在嘈杂的ATC环境中。这项研究提出了使用新创建的数据集专门针对东南亚口音进行微调的ASR模型的开发。我们的研究取得了显着的改善,实现了0.0982或9.82%的字错误率(WER)的SEA口音的ATC语音。此外,本文强调了特定区域数据集和以口音为重点的培训的重要性,为在资源受限的军事行动中部署ASR系统提供了一条途径。研究结果强调,需要噪声鲁棒的训练技术和区域特定的数据集,以提高ATC通信中非西方口音的转录准确性。
摘要:Effective communication in Air Traffic Control (ATC) is critical tomaintaining aviation safety, yet the challenges posed by accented Englishremain largely unaddressed in Automatic Speech Recognition (ASR) systems.Existing models struggle with transcription accuracy for SoutheastAsian-accented (SEA-accented) speech, particularly in noisy ATC environments.This study presents the development of ASR models fine-tuned specifically forSoutheast Asian accents using a newly created dataset. Our research achievessignificant improvements, achieving a Word Error Rate (WER) of 0.0982 or 9.82%on SEA-accented ATC speech. Additionally, the paper highlights the importanceof region-specific datasets and accent-focused training, offering a pathway fordeploying ASR systems in resource-constrained military operations. The findingsemphasize the need for noise-robust training techniques and region-specificdatasets to improve transcription accuracy for non-Western accents in ATCcommunications.

【3】 DIN-CTS: Low-Complexity Depthwise-Inception Neural Network with  Contrastive Training Strategy for Deepfake Speech Detection
标题:DIN-CTS:具有对比训练策略的低复杂度深度初始神经网络,用于Deepfake语音检测
链接:https://arxiv.org/abs/2502.20225
作者:Lam Pham,  Dat Tran,  Florian Skopik,  Alexander Schindler,  Silvia Poletti,  Fischinger David,  Martin Boyer
摘要:在本文中,我们提出了一种基于低复杂度深度初始网络(DIN)的深度假语音检测(DSD)方法,该网络采用对比训练策略(CTS)进行训练。在此框架中,输入音频记录首先使用短时傅立叶变换(STFT)和线性滤波器(LF)转换成频谱图,然后用于训练DIN。一旦经过训练,DIN处理真实的话语以提取音频嵌入,这些音频嵌入用于构建代表真实语音的高斯分布。然后,通过计算测试话语与该分布之间的距离来执行Deepfake检测,以确定话语是假的还是真的。为了评估我们提出的系统,我们对ASVspoof 2019 LA的基准数据集进行了广泛的实验。实验结果表明,深度-感知网络与对比学习策略相结合,在区分真假话语方面是有效的。我们实现了等错误率(EER),准确度(Acc.),F1,AUC分数分别为4.6%,95.4%,97.3%和98.9%,使用单个低复杂度DIN,仅1.77 M参数和985 M FLOPS短音频片段(4秒)。此外,我们提出的系统在ASVspoof 2019 LA挑战赛中的表现优于单系统提交,展示了其实时应用的潜力。
摘要:In this paper, we propose a deep neural network approach for deepfake speechdetection (DSD) based on a lowcomplexity Depthwise-Inception Network (DIN)trained with a contrastive training strategy (CTS). In this framework, inputaudio recordings are first transformed into spectrograms using Short-TimeFourier Transform (STFT) and Linear Filter (LF), which are then used to trainthe DIN. Once trained, the DIN processes bonafide utterances to extract audioembeddings, which are used to construct a Gaussian distribution representinggenuine speech. Deepfake detection is then performed by computing the distancebetween a test utterance and this distribution to determine whether theutterance is fake or bonafide. To evaluate our proposed systems, we conductedextensive experiments on the benchmark dataset of ASVspoof 2019 LA. Theexperimental results demonstrate the effectiveness of combining theDepthwise-Inception Network with the contrastive learning strategy indistinguishing between fake and bonafide utterances. We achieved Equal ErrorRate (EER), Accuracy (Acc.), F1, AUC scores of 4.6%, 95.4%, 97.3%, and 98.9%respectively using a single, low-complexity DIN with just 1.77 M parameters and985 M FLOPS on short audio segments (4 seconds). Furthermore, our proposedsystem outperforms the single-system submissions in the ASVspoof 2019 LAchallenge, showcasing its potential for real-time applications.

【4】 DGFM: Full Body Dance Generation Driven by Music Foundation Models
标题:DGFM:音乐基金会模特推动的全身舞一代
链接:https://arxiv.org/abs/2502.20176
作者:Xinran Liu,  Zhenhua Feng,  Diptesh Kanojia,  Wenwu Wang
备注:Accepted to the Audio Imagination Workshop of NeurlPS 2024
摘要:在音乐驱动的舞蹈动作生成中,大多数现有方法使用手工制作的特征,而忽略了音乐基础模型对跨模态内容生成的深刻影响。为了弥合这一差距,我们提出了一个基于扩散的方法,产生舞蹈动作的文字和音乐的条件。我们的方法提取的音乐特征相结合的音乐基础模型与手工制作的功能,从而提高了生成的舞蹈序列的质量。该方法有效地利用了高层语义信息和低层时间细节的优势,提高了模型对音乐特征的理解能力。为了显示所提出的方法的优点,我们将其与四个音乐基础模型和两组手工制作的音乐特征进行比较。实验结果表明,该方法得到了最真实的舞蹈序列,并与输入音乐达到了最佳匹配。
摘要:In music-driven dance motion generation, most existing methods usehand-crafted features and neglect that music foundation models have profoundlyimpacted cross-modal content generation. To bridge this gap, we propose adiffusion-based method that generates dance movements conditioned on text andmusic. Our approach extracts music features by combining high-level featuresobtained by music foundation model with hand-crafted features, therebyenhancing the quality of generated dance sequences. This method effectivelyleverages the advantages of high-level semantic information and low-leveltemporal details to improve the model's capability in music featureunderstanding. To show the merits of the proposed method, we compare it withfour music foundation models and two sets of hand-crafted music features. Theresults demonstrate that our method obtains the most realistic dance sequencesand achieves the best match with the input music.

【5】 DiffCSS: Diverse and Expressive Conversational Speech Synthesis with  Diffusion Models
标题:DistCSS:采用扩散模型的多样化且富有表达力的对话语音合成
链接:https://arxiv.org/abs/2502.19924
作者:Weihao wu,  Zhiwei Lin,  Yixuan Zhou,  Jingbei Li,  Rui Niu,  Qinghua Wu,  Songjun Cao,  Long Ma,  Zhiyong Wu
备注:Accepted by ICASSP 2025
摘要:会话语音合成(CSS)的目标是合成既符合语境又富有表达力的语音,人们已经做出了相当大的努力来增强对会话语境的理解。然而,现有的CSS系统仅限于确定性预测,忽略了潜在响应的多样性。此外,他们很少采用基于语言模型(LM)的TTS骨干,限制了合成语音的自然度和质量。为了解决这些问题,在本文中,我们提出了DiffCSS,一个创新的CSS框架,利用扩散模型和基于LM的TTS骨干,以产生多样化的,富有表现力的,上下文连贯的语音。提出了一种基于扩散的上下文感知韵律预测器,用于对多模态会话上下文条件下的不同韵律嵌入进行采样。在此基础上,提出了一种韵律可控的基于LM的TTS主干,用于合成具有韵律嵌入的高质量语音。实验结果表明,与现有的CSS系统相比,DiffCSS系统合成的语音更加多样化,上下文连贯,表达能力更强
摘要:Conversational speech synthesis (CSS) aims to synthesize both contextuallyappropriate and expressive speech, and considerable efforts have been made toenhance the understanding of conversational context. However, existing CSSsystems are limited to deterministic prediction, overlooking the diversity ofpotential responses. Moreover, they rarely employ language model (LM)-based TTSbackbones, limiting the naturalness and quality of synthesized speech. Toaddress these issues, in this paper, we propose DiffCSS, an innovative CSSframework that leverages diffusion models and an LM-based TTS backbone togenerate diverse, expressive, and contextually coherent speech. Adiffusion-based context-aware prosody predictor is proposed to sample diverseprosody embeddings conditioned on multimodal conversational context. Then aprosody-controllable LM-based TTS backbone is developed to synthesizehigh-quality speech with sampled prosody embeddings. Experimental resultsdemonstrate that the synthesized speech from DiffCSS is more diverse,contextually coherent, and expressive than existing CSS systems

【6】 Does Your Voice Assistant Remember? Analyzing Conversational Context  Recall and Utilization in Voice Interaction Models
标题:你的语音助手还记得吗?语音交互模型中对话上下文回忆和利用分析
链接:https://arxiv.org/abs/2502.19759
作者:Heeseung Kim,  Che Hyun Lee,  Sangkwon Park,  Jiheum Yeom,  Nohil Park,  Sangwon Yu,  Sungroh Yoon
备注:Work in Progress, Project Page: this https URL
摘要:多回合语音交互模型的最新进展已经改善了用户模型通信。然而,虽然闭源模型有效地保留和回忆过去的话语,开源模型是否也有这种能力仍然没有被探索。为了填补这一空白,我们系统地评估如何利用过去的话语使用ContextDialog,我们为此目的提出的基准开源交互模型。我们的研究结果表明,基于语音的模型比基于文本的模型更困难,特别是在回忆语音中传达的信息时,即使使用检索增强生成,模型仍然难以回答有关过去话语的问题。这些见解突出了开源模型的关键局限性,并提出了改善记忆保持和检索鲁棒性的方法。
摘要:Recent advancements in multi-turn voice interaction models have improveduser-model communication. However, while closed-source models effectivelyretain and recall past utterances, whether open-source models share thisability remains unexplored. To fill this gap, we systematically evaluate howwell open-source interaction models utilize past utterances usingContextDialog, a benchmark we proposed for this purpose. Our findings show thatspeech-based models have more difficulty than text-based ones, especially whenrecalling information conveyed in speech, and even with retrieval-augmentedgeneration, models still struggle with questions about past utterances. Theseinsights highlight key limitations in open-source models and suggest ways toimprove memory retention and retrieval robustness.

【7】 When Large Language Models Meet Speech: A Survey on Integration  Approaches
标题:当大型语言模型遇到言语时:整合方法的调查
链接:https://arxiv.org/abs/2502.19548
作者:Zhengdong Yang,  Shuichiro Shimizu,  Yahan Yu,  Chenhui Chu
摘要:大型语言模型(LLM)的最新进展激发了人们将其应用扩展到基于文本的任务之外的兴趣。大量的研究已经探索了将其他模态与LLM相结合,特别是语音模态,它与文本有着天然的联系。本文综述了语音与LLM的集成,将其分为三种主要方法:基于文本的集成、基于潜在表征的集成和基于音频令牌的集成。我们还展示了如何将这些方法应用于各种语音相关的应用程序,并强调了该领域的挑战,为
摘要:Recent advancements in large language models (LLMs) have spurred interest inexpanding their application beyond text-based tasks. A large number of studieshave explored integrating other modalities with LLMs, notably speech modality,which is naturally related to text. This paper surveys the integration ofspeech with LLMs, categorizing the methodologies into three primary approaches:text-based, latent-representation-based, and audio-token-based integration. Wealso demonstrate how these methods are applied across various speech-relatedapplications and highlight the challenges in this field to offer inspirationfor

【8】 Filtro Adaptativo y Modulo de Grabacion en Dispositivo Para Mejora en la  Calidad de Audicion
标题:Filtro Adaptativo y Modulo de Grabacion en Dispositivo Para Mejora en la Calidad de Audicion
链接:https://arxiv.org/abs/2502.19444
作者:Carlos Elihu Palomino Torres,  Francisco Claudio Chichipe Mondragon,  Frank Antonio Siesquen Rodriguez,  Mariana Alexandra Huaynate Leon
备注:in Spanish language
摘要:本计画提出利用ESP32、LMS自适应滤波器及人工智能技术,发展一个即时听觉强化系统。I2S INMP44麦克风捕捉声音,并在通过MAX98357扬声器播放之前进行动态处理以抑制噪声。该系统可持续适应不同的声学环境,确保提高语音清晰度和优化聆听体验
摘要:This project presents the development of a real-time auditory enhancementsystem utilizing an ESP32, an LMS adaptive filter, and artificial intelligencetechniques. An I2S INMP44 microphone captures the sound, which is dynamicallyprocessed to suppress noise before being played through a MAX98357 speaker. Thesystem continuously adapts to varying acoustic environments, ensuring improvedspeech clarity and an optimized listening experience

【9】 UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
标题:UniCodec:具有单域自适应码本的统一音频编解码器
链接:https://arxiv.org/abs/2502.20067
作者:Yidi Jiang,  Qian Chen,  Shengpeng Ji,  Yu Xi,  Wen Wang,  Chong Zhang,  Xianghu Yue,  ShiLiang Zhang,  Haizhou Li
备注:12 pages, 9 tables
摘要:音频语言模型的出现是由神经音频编解码器授权的,神经音频编解码器建立了与语言模型范例兼容的连续波形和离散标记之间的关键映射。从多层残差矢量量化器到单层量化器的演变趋势有利于语言自回归解码。然而,通过单个码本处理多域音频信号的能力仍然受到域间分布差异的限制。在这项工作中,我们介绍了UniCodec,一个统一的音频编解码器与一个单一的码本,以支持多域音频数据,包括语音,音乐和声音。为了实现这一点,我们提出了一个分区域自适应码本方法和域混合专家策略,以捕捉每个音频域的不同特性。此外,为了在没有辅助模块的情况下丰富编解码器的语义密度,我们提出了一种自监督掩码预测建模方法。综合的客观和主观评估表明,UniCodec在三个音频域上实现了出色的音频重建性能,优于现有的单一码本的统一神经编解码器,甚至在声学和语义表示能力上超过了最先进的特定领域编解码器。
摘要:The emergence of audio language models is empowered by neural audio codecs,which establish critical mappings between continuous waveforms and discretetokens compatible with language model paradigms. The evolutionary trends frommulti-layer residual vector quantizer to single-layer quantizer are beneficialfor language-autoregressive decoding. However, the capability to handlemulti-domain audio signals through a single codebook remains constrained byinter-domain distribution discrepancies. In this work, we introduce UniCodec, aunified audio codec with a single codebook to support multi-domain audio data,including speech, music, and sound. To achieve this, we propose a partitioneddomain-adaptive codebook method and domain Mixture-of-Experts strategy tocapture the distinct characteristics of each audio domain. Furthermore, toenrich the semantic density of the codec without auxiliary modules, we proposea self-supervised mask prediction modeling approach. Comprehensive objectiveand subjective evaluations demonstrate that UniCodec achieves excellent audioreconstruction performance across the three audio domains, outperformingexisting unified neural codecs with a single codebook, and even surpassesstate-of-the-art domain-specific codecs on both acoustic and semanticrepresentation capabilities.

【10】 CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality  and ASR
标题:CleanMel:Mel频谱图增强可提高语音质量和ASB
链接:https://arxiv.org/abs/2502.20040
作者:Nian Shao,  Rui Zhou,  Pengyu Wang,  Xian Li,  Ying Fang,  Yujie Yang,  Xiaofei Li
备注:Submission to IEEE/ACM Trans. on TASLP
摘要:在这项工作中,我们提出了CleanMel,一个单通道梅尔频谱图去噪和去混响网络,用于提高语音质量和自动语音识别(ASR)性能。建议的网络作为输入的噪声和混响麦克风记录,并预测相应的清洁梅尔频谱图。增强后的Mel谱图既可以用神经声码器转换成语音波形,也可以直接用于ASR。建议的网络是由交错的跨带和窄带处理在梅尔频域,学习的全频带频谱模式和信号的窄带特性,分别。与线性频域或时域语音增强相比,Mel谱图增强的主要优点是Mel频率以更紧凑的方式呈现语音,因此更容易学习,这将有利于语音质量和ASR。在四个英文和一个中文数据集上的实验结果表明,该模型在语音质量和ASR性能方面都有明显的改善。我们模型的代码和音频示例可在https://audio.westlake.edu.cn/Research/CleanMel.html上在线获得。
摘要:In this work, we propose CleanMel, a single-channel Mel-spectrogram denoisingand dereverberation network for improving both speech quality and automaticspeech recognition (ASR) performance. The proposed network takes as input thenoisy and reverberant microphone recording and predicts the corresponding cleanMel-spectrogram. The enhanced Mel-spectrogram can be either transformed tospeech waveform with a neural vocoder or directly used for ASR. The proposednetwork is composed of interleaved cross-band and narrow-band processing in theMel-frequency domain, for learning the full-band spectral pattern and thenarrow-band properties of signals, respectively. Compared to linear-frequencydomain or time-domain speech enhancement, the key advantage of Mel-spectrogramenhancement is that Mel-frequency presents speech in a more compact way andthus is easier to learn, which will benefit both speech quality and ASR.Experimental results on four English and one Chinese datasets demonstrate asignificant improvement in both speech quality and ASR performance achieved bythe proposed model. Code and audio examples of our model are available onlinein https://audio.westlake.edu.cn/Research/CleanMel.html.

【11】 PrimeK-Net: Multi-scale Spectral Learning via Group Prime-Kernel  Convolutional Neural Networks for Single Channel Speech Enhancement
标题:PrimeK-Net:通过群素-核卷积神经网络进行多尺度谱学习,用于单通道语音增强
链接:https://arxiv.org/abs/2502.19906
作者:Zizhen Lin,  Junyu Wang,  Ruili Li,  Fei Shen,  Xi Xuan
备注:This paper was accepeted by ICASSP 2025
摘要:单通道语音增强是一个具有挑战性的不适定问题,其重点是从退化信号中估计干净的语音。现有的研究已经证明了在语音增强任务中将卷积神经网络(CNN)与Transformers相结合的竞争性能。然而,现有的框架没有充分解决计算效率,并忽略了自然的多尺度分布的频谱。此外,CNN在语音增强中的潜力尚未完全实现。为了解决这些问题,本研究提出了一个深度可分离的扩张密集块(DSDDB)和一个组素核前馈通道注意(GPFCA)模块。具体地,DSDDB向现有框架的编码器/解码器引入了更高的参数和计算效率。GPFCA模块取代了Conformer的位置,以线性复杂度提取频谱的深层时间和频率特征。GPFCA利用所提出的群素核前馈网络(GPFN)来集成多粒度的远程、中程和短程感受野,同时利用素数的属性来避免周期性重叠效应。实验结果表明,本研究中提出的PrimeK-Net在VoiceBank+Demand数据集上实现了最先进的(SOTA)性能,仅用1.41 M参数就达到了3.61的PESQ分数。
摘要:Single-channel speech enhancement is a challenging ill-posed problem focusedon estimating clean speech from degraded signals. Existing studies havedemonstrated the competitive performance of combining convolutional neuralnetworks (CNNs) with Transformers in speech enhancement tasks. However,existing frameworks have not sufficiently addressed computational efficiencyand have overlooked the natural multi-scale distribution of the spectrum.Additionally, the potential of CNNs in speech enhancement has yet to be fullyrealized. To address these issues, this study proposes a Deep Separable DilatedDense Block (DSDDB) and a Group Prime Kernel Feedforward Channel Attention(GPFCA) module. Specifically, the DSDDB introduces higher parameter andcomputational efficiency to the Encoder/Decoder of existing frameworks. TheGPFCA module replaces the position of the Conformer, extracting deep temporaland frequency features of the spectrum with linear complexity. The GPFCAleverages the proposed Group Prime Kernel Feedforward Network (GPFN) tointegrate multi-granularity long-range, medium-range, and short-range receptivefields, while utilizing the properties of prime numbers to avoid periodicoverlap effects. Experimental results demonstrate that PrimeK-Net, proposed inthis study, achieves state-of-the-art (SOTA) performance on theVoiceBank+Demand dataset, reaching a PESQ score of 3.61 with only 1.41Mparameters.

eess.AS音频处理

【1】 UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
标题:UniCodec:具有单域自适应码本的统一音频编解码器
链接:https://arxiv.org/abs/2502.20067
作者:Yidi Jiang,  Qian Chen,  Shengpeng Ji,  Yu Xi,  Wen Wang,  Chong Zhang,  Xianghu Yue,  ShiLiang Zhang,  Haizhou Li
备注:12 pages, 9 tables
摘要:音频语言模型的出现是由神经音频编解码器授权的,神经音频编解码器建立了与语言模型范例兼容的连续波形和离散标记之间的关键映射。从多层残差矢量量化器到单层量化器的演变趋势有利于语言自回归解码。然而,通过单个码本处理多域音频信号的能力仍然受到域间分布差异的限制。在这项工作中,我们介绍了UniCodec,一个统一的音频编解码器与一个单一的码本,以支持多域音频数据,包括语音,音乐和声音。为了实现这一点,我们提出了一个分区域自适应码本方法和域混合专家策略,以捕捉每个音频域的不同特性。此外,为了在没有辅助模块的情况下丰富编解码器的语义密度,我们提出了一种自监督掩码预测建模方法。综合的客观和主观评估表明,UniCodec在三个音频域上实现了出色的音频重建性能,优于现有的单一码本的统一神经编解码器,甚至在声学和语义表示能力上超过了最先进的特定领域编解码器。
摘要:The emergence of audio language models is empowered by neural audio codecs,which establish critical mappings between continuous waveforms and discretetokens compatible with language model paradigms. The evolutionary trends frommulti-layer residual vector quantizer to single-layer quantizer are beneficialfor language-autoregressive decoding. However, the capability to handlemulti-domain audio signals through a single codebook remains constrained byinter-domain distribution discrepancies. In this work, we introduce UniCodec, aunified audio codec with a single codebook to support multi-domain audio data,including speech, music, and sound. To achieve this, we propose a partitioneddomain-adaptive codebook method and domain Mixture-of-Experts strategy tocapture the distinct characteristics of each audio domain. Furthermore, toenrich the semantic density of the codec without auxiliary modules, we proposea self-supervised mask prediction modeling approach. Comprehensive objectiveand subjective evaluations demonstrate that UniCodec achieves excellent audioreconstruction performance across the three audio domains, outperformingexisting unified neural codecs with a single codebook, and even surpassesstate-of-the-art domain-specific codecs on both acoustic and semanticrepresentation capabilities.

【2】 CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality  and ASR
标题:CleanMel:Mel频谱图增强可提高语音质量和ASB
链接:https://arxiv.org/abs/2502.20040
作者:Nian Shao,  Rui Zhou,  Pengyu Wang,  Xian Li,  Ying Fang,  Yujie Yang,  Xiaofei Li
备注:Submission to IEEE/ACM Trans. on TASLP
摘要:在这项工作中,我们提出了CleanMel,一个单通道梅尔频谱图去噪和去混响网络,用于提高语音质量和自动语音识别(ASR)性能。建议的网络作为输入的噪声和混响麦克风记录,并预测相应的清洁梅尔频谱图。增强后的Mel谱图既可以用神经声码器转换成语音波形,也可以直接用于ASR。建议的网络是由交错的跨带和窄带处理在梅尔频域,学习的全频带频谱模式和信号的窄带特性,分别。与线性频域或时域语音增强相比,Mel谱图增强的主要优点是Mel频率以更紧凑的方式呈现语音,因此更容易学习,这将有利于语音质量和ASR。在四个英文和一个中文数据集上的实验结果表明,该模型在语音质量和ASR性能方面都有明显的改善。我们模型的代码和音频示例可在https://audio.westlake.edu.cn/Research/CleanMel.html上在线获得。
摘要:In this work, we propose CleanMel, a single-channel Mel-spectrogram denoisingand dereverberation network for improving both speech quality and automaticspeech recognition (ASR) performance. The proposed network takes as input thenoisy and reverberant microphone recording and predicts the corresponding cleanMel-spectrogram. The enhanced Mel-spectrogram can be either transformed tospeech waveform with a neural vocoder or directly used for ASR. The proposednetwork is composed of interleaved cross-band and narrow-band processing in theMel-frequency domain, for learning the full-band spectral pattern and thenarrow-band properties of signals, respectively. Compared to linear-frequencydomain or time-domain speech enhancement, the key advantage of Mel-spectrogramenhancement is that Mel-frequency presents speech in a more compact way andthus is easier to learn, which will benefit both speech quality and ASR.Experimental results on four English and one Chinese datasets demonstrate asignificant improvement in both speech quality and ASR performance achieved bythe proposed model. Code and audio examples of our model are available onlinein https://audio.westlake.edu.cn/Research/CleanMel.html.

【3】 PrimeK-Net: Multi-scale Spectral Learning via Group Prime-Kernel  Convolutional Neural Networks for Single Channel Speech Enhancement
标题:PrimeK-Net:通过群素-核卷积神经网络进行多尺度谱学习,用于单通道语音增强
链接:https://arxiv.org/abs/2502.19906
作者:Zizhen Lin,  Junyu Wang,  Ruili Li,  Fei Shen,  Xi Xuan
备注:This paper was accepeted by ICASSP 2025
摘要:单通道语音增强是一个具有挑战性的不适定问题,其重点是从退化信号中估计干净的语音。现有的研究已经证明了在语音增强任务中将卷积神经网络(CNN)与Transformers相结合的竞争性能。然而,现有的框架没有充分解决计算效率,并忽略了自然的多尺度分布的频谱。此外,CNN在语音增强中的潜力尚未完全实现。为了解决这些问题,本研究提出了一个深度可分离的扩张密集块(DSDDB)和一个组素核前馈通道注意(GPFCA)模块。具体地,DSDDB向现有框架的编码器/解码器引入了更高的参数和计算效率。GPFCA模块取代了Conformer的位置,以线性复杂度提取频谱的深层时间和频率特征。GPFCA利用所提出的群素核前馈网络(GPFN)来集成多粒度的远程、中程和短程感受野,同时利用素数的属性来避免周期性重叠效应。实验结果表明,本研究中提出的PrimeK-Net在VoiceBank+Demand数据集上实现了最先进的(SOTA)性能,仅用1.41 M参数就达到了3.61的PESQ分数。
摘要:Single-channel speech enhancement is a challenging ill-posed problem focusedon estimating clean speech from degraded signals. Existing studies havedemonstrated the competitive performance of combining convolutional neuralnetworks (CNNs) with Transformers in speech enhancement tasks. However,existing frameworks have not sufficiently addressed computational efficiencyand have overlooked the natural multi-scale distribution of the spectrum.Additionally, the potential of CNNs in speech enhancement has yet to be fullyrealized. To address these issues, this study proposes a Deep Separable DilatedDense Block (DSDDB) and a Group Prime Kernel Feedforward Channel Attention(GPFCA) module. Specifically, the DSDDB introduces higher parameter andcomputational efficiency to the Encoder/Decoder of existing frameworks. TheGPFCA module replaces the position of the Conformer, extracting deep temporaland frequency features of the spectrum with linear complexity. The GPFCAleverages the proposed Group Prime Kernel Feedforward Network (GPFN) tointegrate multi-granularity long-range, medium-range, and short-range receptivefields, while utilizing the properties of prime numbers to avoid periodicoverlap effects. Experimental results demonstrate that PrimeK-Net, proposed inthis study, achieves state-of-the-art (SOTA) performance on theVoiceBank+Demand dataset, reaching a PESQ score of 3.61 with only 1.41Mparameters.

【4】 On Adversarial Attacks In Acoustic Drone Localization
标题:声学无人机定位中的对抗攻击
链接:https://arxiv.org/abs/2502.20325
作者:Tamir Shor,  Chaim Baskin,  Alex Bronstein
摘要:近年来,多旋翼航空自主飞行器(MAV,更广泛地称为“无人机”)由于其在广泛和多样化的领域(例如,农业、商业交付、搜索和救援)。基于视觉的方法对照明条件和遮挡的敏感性促使人们越来越多地研究依赖于其他方式的导航,例如声学传感。在非受控环境中大规模使用无人机执行任务的一个主要问题是其导航系统可能受到敌对攻击的威胁,使用户面临关键任务故障,安全漏洞和可能危及操作员和旁观者的安全结果。虽然以前的工作在基于声学的无人机定位方面取得了令人印象深刻的进展,但以前对无人机导航的对抗性攻击的研究只涉及基于视觉传感的系统。在这项工作中,我们的目标是通过全面分析PGD对抗性攻击对声学无人机定位的影响来弥补这一差距。我们还开发了一种对抗扰动恢复算法,能够显着减少这种攻击在我们的设置的影响。用于复制所有实验的代码将在出版时发布。
摘要:Multi-rotor aerial autonomous vehicles (MAVs, more widely known as "drones")have been generating increased interest in recent years due to their growingapplicability in a vast and diverse range of fields (e.g., agriculture,commercial delivery, search and rescue). The sensitivity of visual-basedmethods to lighting conditions and occlusions had prompted growing study ofnavigation reliant on other modalities, such as acoustic sensing. A majorconcern in using drones in scale for tasks in non-controlled environments isthe potential threat of adversarial attacks over their navigational systems,exposing users to mission-critical failures, security breaches, and compromisedsafety outcomes that can endanger operators and bystanders. While previous workshows impressive progress in acoustic-based drone localization, prior researchin adversarial attacks over drone navigation only addresses visualsensing-based systems. In this work, we aim to compensate for this gap bysupplying a comprehensive analysis of the effect of PGD adversarial attacksover acoustic drone localization. We furthermore develop an algorithm foradversarial perturbation recovery, capable of markedly diminishing the affectof such attacks in our setting. The code for reproducing all experiments willbe released upon publication.

【5】 Adapting Automatic Speech Recognition for Accented Air Traffic Control  Communications
标题:将自动语音识别适应强调的空中交通管制通信
链接:https://arxiv.org/abs/2502.20311
作者:Marcus Yu Zhe Wee,  Justin Juin Hng Wong,  Lynus Lim,  Joe Yu Wei Tan,  Prannaya Gupta,  Dillion Lim,  En Hao Tew,  Aloysius Keng Siew Han,  Yong Zhi Lim
摘要:空中交通管制(ATC)中的有效沟通对于维护航空安全至关重要,但口音英语带来的挑战在自动语音识别(ASR)系统中仍然没有得到解决。现有的模型与东南亚口音(SEA口音)语音的转录准确性斗争,特别是在嘈杂的ATC环境中。这项研究提出了使用新创建的数据集专门针对东南亚口音进行微调的ASR模型的开发。我们的研究取得了显着的改善,实现了0.0982或9.82%的字错误率(WER)的SEA口音的ATC语音。此外,本文强调了特定区域数据集和以口音为重点的培训的重要性,为在资源受限的军事行动中部署ASR系统提供了一条途径。研究结果强调,需要噪声鲁棒的训练技术和区域特定的数据集,以提高ATC通信中非西方口音的转录准确性。
摘要:Effective communication in Air Traffic Control (ATC) is critical tomaintaining aviation safety, yet the challenges posed by accented Englishremain largely unaddressed in Automatic Speech Recognition (ASR) systems.Existing models struggle with transcription accuracy for SoutheastAsian-accented (SEA-accented) speech, particularly in noisy ATC environments.This study presents the development of ASR models fine-tuned specifically forSoutheast Asian accents using a newly created dataset. Our research achievessignificant improvements, achieving a Word Error Rate (WER) of 0.0982 or 9.82%on SEA-accented ATC speech. Additionally, the paper highlights the importanceof region-specific datasets and accent-focused training, offering a pathway fordeploying ASR systems in resource-constrained military operations. The findingsemphasize the need for noise-robust training techniques and region-specificdatasets to improve transcription accuracy for non-Western accents in ATCcommunications.

【6】 DIN-CTS: Low-Complexity Depthwise-Inception Neural Network with  Contrastive Training Strategy for Deepfake Speech Detection
标题:DIN-CTS:具有对比训练策略的低复杂度深度初始神经网络,用于Deepfake语音检测
链接:https://arxiv.org/abs/2502.20225
作者:Lam Pham,  Dat Tran,  Florian Skopik,  Alexander Schindler,  Silvia Poletti,  Fischinger David,  Martin Boyer
摘要:在本文中,我们提出了一种基于低复杂度深度初始网络(DIN)的深度假语音检测(DSD)方法,该网络采用对比训练策略(CTS)进行训练。在此框架中,输入音频记录首先使用短时傅立叶变换(STFT)和线性滤波器(LF)转换成频谱图,然后用于训练DIN。一旦经过训练,DIN处理真实的话语以提取音频嵌入,这些音频嵌入用于构建代表真实语音的高斯分布。然后,通过计算测试话语与该分布之间的距离来执行Deepfake检测,以确定话语是假的还是真的。为了评估我们提出的系统,我们对ASVspoof 2019 LA的基准数据集进行了广泛的实验。实验结果表明,深度-感知网络与对比学习策略相结合,在区分真假话语方面是有效的。我们实现了等错误率(EER),准确度(Acc.),F1,AUC分数分别为4.6%,95.4%,97.3%和98.9%,使用单个低复杂度DIN,仅1.77 M参数和985 M FLOPS短音频片段(4秒)。此外,我们提出的系统在ASVspoof 2019 LA挑战赛中的表现优于单系统提交,展示了其实时应用的潜力。
摘要:In this paper, we propose a deep neural network approach for deepfake speechdetection (DSD) based on a lowcomplexity Depthwise-Inception Network (DIN)trained with a contrastive training strategy (CTS). In this framework, inputaudio recordings are first transformed into spectrograms using Short-TimeFourier Transform (STFT) and Linear Filter (LF), which are then used to trainthe DIN. Once trained, the DIN processes bonafide utterances to extract audioembeddings, which are used to construct a Gaussian distribution representinggenuine speech. Deepfake detection is then performed by computing the distancebetween a test utterance and this distribution to determine whether theutterance is fake or bonafide. To evaluate our proposed systems, we conductedextensive experiments on the benchmark dataset of ASVspoof 2019 LA. Theexperimental results demonstrate the effectiveness of combining theDepthwise-Inception Network with the contrastive learning strategy indistinguishing between fake and bonafide utterances. We achieved Equal ErrorRate (EER), Accuracy (Acc.), F1, AUC scores of 4.6%, 95.4%, 97.3%, and 98.9%respectively using a single, low-complexity DIN with just 1.77 M parameters and985 M FLOPS on short audio segments (4 seconds). Furthermore, our proposedsystem outperforms the single-system submissions in the ASVspoof 2019 LAchallenge, showcasing its potential for real-time applications.

【7】 DGFM: Full Body Dance Generation Driven by Music Foundation Models
标题:DGFM:音乐基金会模特推动的全身舞一代
链接:https://arxiv.org/abs/2502.20176
作者:Xinran Liu,  Zhenhua Feng,  Diptesh Kanojia,  Wenwu Wang
备注:Accepted to the Audio Imagination Workshop of NeurlPS 2024
摘要:在音乐驱动的舞蹈动作生成中,大多数现有方法使用手工制作的特征,而忽略了音乐基础模型对跨模态内容生成的深刻影响。为了弥合这一差距,我们提出了一个基于扩散的方法,产生舞蹈动作的文字和音乐的条件。我们的方法提取的音乐特征相结合的音乐基础模型与手工制作的功能,从而提高了生成的舞蹈序列的质量。该方法有效地利用了高层语义信息和低层时间细节的优势,提高了模型对音乐特征的理解能力。为了显示所提出的方法的优点,我们将其与四个音乐基础模型和两组手工制作的音乐特征进行比较。实验结果表明,该方法得到了最真实的舞蹈序列,并与输入音乐达到了最佳匹配。
摘要:In music-driven dance motion generation, most existing methods usehand-crafted features and neglect that music foundation models have profoundlyimpacted cross-modal content generation. To bridge this gap, we propose adiffusion-based method that generates dance movements conditioned on text andmusic. Our approach extracts music features by combining high-level featuresobtained by music foundation model with hand-crafted features, therebyenhancing the quality of generated dance sequences. This method effectivelyleverages the advantages of high-level semantic information and low-leveltemporal details to improve the model's capability in music featureunderstanding. To show the merits of the proposed method, we compare it withfour music foundation models and two sets of hand-crafted music features. Theresults demonstrate that our method obtains the most realistic dance sequencesand achieves the best match with the input music.

【8】 DiffCSS: Diverse and Expressive Conversational Speech Synthesis with  Diffusion Models
标题:DistCSS:采用扩散模型的多样化且富有表达力的对话语音合成
链接:https://arxiv.org/abs/2502.19924
作者:Weihao wu,  Zhiwei Lin,  Yixuan Zhou,  Jingbei Li,  Rui Niu,  Qinghua Wu,  Songjun Cao,  Long Ma,  Zhiyong Wu
备注:Accepted by ICASSP 2025
摘要:会话语音合成(CSS)的目标是合成既符合语境又富有表达力的语音,人们已经做出了相当大的努力来增强对会话语境的理解。然而,现有的CSS系统仅限于确定性预测,忽略了潜在响应的多样性。此外,他们很少采用基于语言模型(LM)的TTS骨干,限制了合成语音的自然度和质量。为了解决这些问题,在本文中,我们提出了DiffCSS,一个创新的CSS框架,利用扩散模型和基于LM的TTS骨干,以产生多样化的,富有表现力的,上下文连贯的语音。提出了一种基于扩散的上下文感知韵律预测器,用于对多模态会话上下文条件下的不同韵律嵌入进行采样。在此基础上,提出了一种韵律可控的基于LM的TTS主干,用于合成具有韵律嵌入的高质量语音。实验结果表明,与现有的CSS系统相比,DiffCSS系统合成的语音更加多样化,上下文连贯,表达能力更强
摘要:Conversational speech synthesis (CSS) aims to synthesize both contextuallyappropriate and expressive speech, and considerable efforts have been made toenhance the understanding of conversational context. However, existing CSSsystems are limited to deterministic prediction, overlooking the diversity ofpotential responses. Moreover, they rarely employ language model (LM)-based TTSbackbones, limiting the naturalness and quality of synthesized speech. Toaddress these issues, in this paper, we propose DiffCSS, an innovative CSSframework that leverages diffusion models and an LM-based TTS backbone togenerate diverse, expressive, and contextually coherent speech. Adiffusion-based context-aware prosody predictor is proposed to sample diverseprosody embeddings conditioned on multimodal conversational context. Then aprosody-controllable LM-based TTS backbone is developed to synthesizehigh-quality speech with sampled prosody embeddings. Experimental resultsdemonstrate that the synthesized speech from DiffCSS is more diverse,contextually coherent, and expressive than existing CSS systems

【9】 Does Your Voice Assistant Remember? Analyzing Conversational Context  Recall and Utilization in Voice Interaction Models
标题:你的语音助手还记得吗?语音交互模型中对话上下文回忆和利用分析
链接:https://arxiv.org/abs/2502.19759
作者:Heeseung Kim,  Che Hyun Lee,  Sangkwon Park,  Jiheum Yeom,  Nohil Park,  Sangwon Yu,  Sungroh Yoon
备注:Work in Progress, Project Page: this https URL
摘要:多回合语音交互模型的最新进展已经改善了用户模型通信。然而,虽然闭源模型有效地保留和回忆过去的话语,开源模型是否也有这种能力仍然没有被探索。为了填补这一空白,我们系统地评估如何利用过去的话语使用ContextDialog,我们为此目的提出的基准开源交互模型。我们的研究结果表明,基于语音的模型比基于文本的模型更困难,特别是在回忆语音中传达的信息时,即使使用检索增强生成,模型仍然难以回答有关过去话语的问题。这些见解突出了开源模型的关键局限性,并提出了改善记忆保持和检索鲁棒性的方法。
摘要:Recent advancements in multi-turn voice interaction models have improveduser-model communication. However, while closed-source models effectivelyretain and recall past utterances, whether open-source models share thisability remains unexplored. To fill this gap, we systematically evaluate howwell open-source interaction models utilize past utterances usingContextDialog, a benchmark we proposed for this purpose. Our findings show thatspeech-based models have more difficulty than text-based ones, especially whenrecalling information conveyed in speech, and even with retrieval-augmentedgeneration, models still struggle with questions about past utterances. Theseinsights highlight key limitations in open-source models and suggest ways toimprove memory retention and retrieval robustness.

【10】 When Large Language Models Meet Speech: A Survey on Integration  Approaches
标题:当大型语言模型遇到言语时:整合方法的调查
链接:https://arxiv.org/abs/2502.19548
作者:Zhengdong Yang,  Shuichiro Shimizu,  Yahan Yu,  Chenhui Chu
摘要:大型语言模型(LLM)的最新进展激发了人们将其应用扩展到基于文本的任务之外的兴趣。大量的研究已经探索了将其他模态与LLM相结合,特别是语音模态,它与文本有着天然的联系。本文综述了语音与LLM的集成,将其分为三种主要方法:基于文本的集成、基于潜在表征的集成和基于音频令牌的集成。我们还展示了如何将这些方法应用于各种语音相关的应用程序,并强调了该领域的挑战,为
摘要:Recent advancements in large language models (LLMs) have spurred interest inexpanding their application beyond text-based tasks. A large number of studieshave explored integrating other modalities with LLMs, notably speech modality,which is naturally related to text. This paper surveys the integration ofspeech with LLMs, categorizing the methodologies into three primary approaches:text-based, latent-representation-based, and audio-token-based integration. Wealso demonstrate how these methods are applied across various speech-relatedapplications and highlight the challenges in this field to offer inspirationfor

【11】 Filtro Adaptativo y Modulo de Grabacion en Dispositivo Para Mejora en la  Calidad de Audicion
标题:Filtro Adaptativo y Modulo de Grabacion en Dispositivo Para Mejora en la Calidad de Audicion
链接:https://arxiv.org/abs/2502.19444
作者:Carlos Elihu Palomino Torres,  Francisco Claudio Chichipe Mondragon,  Frank Antonio Siesquen Rodriguez,  Mariana Alexandra Huaynate Leon
备注:in Spanish language
摘要:本计画提出利用ESP32、LMS自适应滤波器及人工智能技术,发展一个即时听觉强化系统。I2S INMP44麦克风捕捉声音,并在通过MAX98357扬声器播放之前进行动态处理以抑制噪音。该系统可持续适应不同的声学环境,确保提高语音清晰度和优化聆听体验
摘要:This project presents the development of a real-time auditory enhancementsystem utilizing an ESP32, an LMS adaptive filter, and artificial intelligencetechniques. An I2S INMP44 microphone captures the sound, which is dynamicallyprocessed to suppress noise before being played through a MAX98357 speaker. Thesystem continuously adapts to varying acoustic environments, ensuring improvedspeech clarity and an optimized listening experience

机器翻译由腾讯交互翻译提供,仅供参考