今日论文合集:cs.SD语音5篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音
【1】  WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken  Dialogue Models
标题:WavRAG:语音对话模型的音频集成检索增强生成
链接:https://arxiv.org/abs/2502.14727
作者:Yifu Chen,  Shengpeng Ji,  Haoxiao Wang,  Ziqing Wang,  Siyu Chen,  Jinzheng He,  Jin Xu,  Zhou Zhao
摘要:检索增强生成(RAG)由于其能够使大型语言模型(LLM)集成外部知识而获得广泛采用。然而,现有的RAG框架主要是为基于文本的LLM设计的,并依赖于自动语音识别来处理语音输入,这会丢弃关键的音频信息,冒着转录错误的风险,并增加计算开销。因此,我们引入了WavRAG,这是第一个具有本地端到端音频支持的检索增强生成框架。WavRAG提供了两个关键功能:1)对ASR进行分类,WavRAG直接处理原始音频进行嵌入和检索。2)WavRAG将音频和文本集成到统一的知识表示中。具体来说,我们提出了WavRetriever,以方便从文本-音频混合知识库检索,并通过集成的思想链推理,进一步提高口语对话模型的上下文能力。与最先进的ASR-Text RAG管道相比,WavRAG实现了相当的检索性能,同时提供了10倍的加速。此外,WavRAG独特的文本-音频混合检索能力将RAG的边界扩展到音频模态。
摘要:Retrieval Augmented Generation (RAG) has gained widespread adoption owing toits capacity to empower large language models (LLMs) to integrate externalknowledge. However, existing RAG frameworks are primarily designed fortext-based LLMs and rely on Automatic Speech Recognition to process speechinput, which discards crucial audio information, risks transcription errors,and increases computational overhead. Therefore, we introduce WavRAG, the firstretrieval augmented generation framework with native, end-to-end audio support.WavRAG offers two key features: 1) Bypassing ASR, WavRAG directly processes rawaudio for both embedding and retrieval. 2) WavRAG integrates audio and textinto a unified knowledge representation. Specifically, we propose theWavRetriever to facilitate the retrieval from a text-audio hybrid knowledgebase, and further enhance the in-context capabilities of spoken dialogue modelsthrough the integration of chain-of-thought reasoning. In comparison tostate-of-the-art ASR-Text RAG pipelines, WavRAG achieves comparable retrievalperformance while delivering a 10x acceleration. Furthermore, WavRAG's uniquetext-audio hybrid retrieval capability extends the boundaries of RAG to theaudio modality.

【2】 Pitch Imperfect: Detecting Audio Deepfakes Through Acoustic Prosodic  Analysis
标题:音调不完美:通过声学韵律分析检测音频Deepfake
链接:https://arxiv.org/abs/2502.14726
作者:Kevin Warren,  Daniel Olszewski,  Seth Layton,  Kevin Butler,  Carrie Gates,  Patrick Traynor
摘要:音频深度伪造与有机语音越来越难以区分,经常欺骗认证系统和人类听众。虽然许多技术使用低级音频特征或优化黑盒模型训练,但专注于人类用于识别语音的特征可能是一种更长期稳健的检测方法。我们探索韵律的使用,或人类语音的高级语言特征(例如,音高、语调、抖动)作为检测音频深度伪造的更基本的手段。我们开发了一个基于六个经典韵律特征的检测器,并证明了我们的模型与社区使用的其他基线模型一样,可以检测音频deepfake,准确率为93%,EER为24.7%。更重要的是,我们证明了使用基于语言特征的方法比现有模型的好处,通过应用自适应对手使用$L_{\infty}$ norm攻击对检测器和使用注意力机制在我们的训练可解释性。我们表明,我们可以解释的韵律功能,具有最高的影响模型的决定(抖动,微光和平均基频)和其他模型是非常容易受到简单的$L_{\infty}$范数攻击(99.3%的相对精度下降)。虽然整体性能可能相似,但我们说明了韵律特征方法对音频deepfake检测的鲁棒性和可解释性的好处。
摘要:Audio deepfakes are increasingly in-differentiable from organic speech, oftenfooling both authentication systems and human listeners. While many techniquesuse low-level audio features or optimization black-box model training, focusingon the features that humans use to recognize speech will likely be a morelong-term robust approach to detection. We explore the use of prosody, or thehigh-level linguistic features of human speech (e.g., pitch, intonation,jitter) as a more foundational means of detecting audio deepfakes. We develop adetector based on six classical prosodic features and demonstrate that ourmodel performs as well as other baseline models used by the community to detectaudio deepfakes with an accuracy of 93% and an EER of 24.7%. More importantly,we demonstrate the benefits of using a linguistic features-based approach overexisting models by applying an adaptive adversary using an $L_{\infty}$ normattack against the detectors and using attention mechanisms in our training forexplainability. We show that we can explain the prosodic features that havehighest impact on the model's decision (Jitter, Shimmer and Mean FundamentalFrequency) and that other models are extremely susceptible to simple$L_{\infty}$ norm attacks (99.3% relative degradation in accuracy). Whileoverall performance may be similar, we illustrate the robustness andexplainability benefits to a prosody feature approach to audio deepfakedetection.

【3】 SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer  Based Speech Recognition
标题:SegAug:用于基于RNN传感器的鲁棒语音识别的ATC对齐分段增强
链接:https://arxiv.org/abs/2502.14685
作者:Khanh Le,  Tuan Vu Ho,  Dung Tran,  Duc Thanh Chau
备注:Accepted to ICASSP 2025
摘要:RNN-Transducer(RNN-T)是语音识别中广泛采用的架构,在端到端框架中集成了声学和语言建模。然而,RNN-T预测器倾向于过度依赖训练数据中的连续单词依赖性,导致高删除错误率,特别是对于不太常见或域外短语。现有的解决方案,如正则化和数据增强,通常会损害性能的其他方面。我们提出了SegAug,一种基于增强的技术,生成上下文不同的音频文本对低的重复级别的语义。这种方法鼓励模型更多地关注声学特征,同时使其内部语言模型的学习文本模式多样化,从而减少删除错误并提高整体性能。对LibriSpeech和Tedlium-v3数据集的评估表明,在小规模环境下,WER相对降低了12.5%,在大规模环境下降低了6.9%。值得注意的是,大部分的改善源于减少的删除错误,相对减少分别为45.4%和18.5%。这些结果突出了SegAug在提高RNN-T鲁棒性方面的有效性,为在各种具有挑战性的场景中增强语音识别性能提供了一个有前途的解决方案。
摘要:RNN-Transducer (RNN-T) is a widely adopted architecture in speechrecognition, integrating acoustic and language modeling in an end-to-endframework. However, the RNN-T predictor tends to over-rely on consecutive worddependencies in training data, leading to high deletion error rates,particularly with less common or out-of-domain phrases. Existing solutions,such as regularization and data augmentation, often compromise other aspects ofperformance. We propose SegAug, an alignment-based augmentation technique thatgenerates contextually varied audio-text pairs with low sentence-levelsemantics. This method encourages the model to focus more on acoustic featureswhile diversifying the learned textual patterns of its internal language model,thereby reducing deletion errors and enhancing overall performance. Evaluationson the LibriSpeech and Tedlium-v3 datasets demonstrate a relative WER reductionof up to 12.5% on small-scale and 6.9% on large-scale settings. Notably, mostof the improvement stems from reduced deletion errors, with relative reductionsof 45.4% and 18.5%, respectively. These results highlight SegAug'seffectiveness in improving RNN-T's robustness, offering a promising solutionfor enhancing speech recognition performance across diverse and challengingscenarios.

【4】 ChunkFormer: Masked Chunking Conformer For Long-Form Speech  Transcription
标题:ChunkFormer:用于长形式语音转录的掩蔽分块协调器
链接:https://arxiv.org/abs/2502.14673
作者:Khanh Le,  Tuan Vu Ho,  Dung Tran,  Duc Thanh Chau
备注:Accepted to ICASSP 2025
摘要:以工业规模部署ASR模型对硬件资源管理提出了重大挑战,特别是对于音频可能持续数小时的长格式转录任务。大型Conformer型号尽管功能强大,但仅限于在80GB GPU上处理15分钟的音频。此外,可变的输入长度会使效率低下的情况恶化,因为标准的插入会导致过多的填充,从而增加资源消耗和执行时间。为了解决这个问题,我们引入了ChunkFormer,这是一种高效的ASR模型,它使用相对正确的上下文进行分块处理,从而在低内存GPU上实现长时间的音频传输。ChunkFormer在80GB GPU上处理长达16小时的音频,比当前最先进的FastConformer长1.5倍,同时还提高了长格式转录性能,与Conformer相比,单词错误率绝对降低了7.7%,并在较短的任务中保持了准确性。ChunkFormer的masked mounted m
摘要:Deploying ASR models at an industrial scale poses significant challenges inhardware resource management, especially for long-form transcription taskswhere audio may last for hours. Large Conformer models, despite theircapabilities, are limited to processing only 15 minutes of audio on an 80GBGPU. Furthermore, variable input lengths worsen inefficiencies, as standardbatching leads to excessive padding, increasing resource consumption andexecution time. To address this, we introduce ChunkFormer, an efficient ASRmodel that uses chunk-wise processing with relative right context, enablinglong audio transcriptions on low-memory GPUs. ChunkFormer handles up to 16hours of audio on an 80GB GPU, 1.5x longer than the current state-of-the-artFastConformer, while also boosting long-form transcription performance with upto 7.7% absolute reduction on word error rate and maintaining accuracy onshorter tasks compared to Conformer. By eliminating the need for padding instandard batching, ChunkFormer's masked batching technique reduces executiontime and memory usage by more than 3x in batch processing, substantiallyreducing costs for a wide range of ASR systems, particularly regarding GPUresources for models serving in real-world applications.

【5】 ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by  Reducing Data Distribution Errors
标题:ATRI:通过减少数据分布错误来缓解多语言音频文本检索不确定性
链接:https://arxiv.org/abs/2502.14627
作者:Yuguo Yin,  Yuxin Xie,  Wenyuan Yang,  Dongchao Yang,  Jinghan Ru,  Xianwei Zhuang,  Liming Liang,  Yuexian Zou
摘要:多语言音频文本检索(ML-ATR)是一项具有挑战性的任务,旨在从数据库中检索音频片段或多语言文本。然而,现有的ML-ATR方案遭受例如跨语言的相似性匹配的不一致。我们从理论上分析了多语种模态对齐方向误差和权重误差的不一致性,并提出了量化不一致性的理论权重误差上限。通过对权重误差上界的分析,我们发现不一致性问题源于语言随机抽样造成的数据分布误差。针对ML-ATR中数据分布错误对召回率和一致性的影响,提出了一种基于1-to-k对比学习和音-英语共锚对比学习的一致性ML-ATR方案。在翻译的AudioCaps和Clotho数据集上的实验结果表明,我们的方案在包括英语在内的八种主流语言的召回率和一致性指标上达到了最先进的性能。我们的代码将在https://github.com/ATRI-ACL/ATRI-ACL上提供。
摘要:Multilingual audio-text retrieval (ML-ATR) is a challenging task that aims toretrieve audio clips or multilingual texts from databases. However, existingML-ATR schemes suffer from inconsistencies for instance similarity matchingacross languages. We theoretically analyze the inconsistency in terms of bothmultilingual modal alignment direction error and weight error, and propose thetheoretical weight error upper bound for quantifying the inconsistency. Basedon the analysis of the weight error upper bound, we find that the inconsistencyproblem stems from the data distribution error caused by random sampling oflanguages. We propose a consistent ML-ATR scheme using 1-to-k contrastivelearning and audio-English co-anchor contrastive learning, aiming to mitigatethe negative impact of data distribution error on recall and consistency inML-ATR. Experimental results on the translated AudioCaps and Clotho datasetsshow that our scheme achieves state-of-the-art performance on recall andconsistency metrics for eight mainstream languages, including English. Our codewill be available at https://github.com/ATRI-ACL/ATRI-ACL.

【6】 Differentiable Black-box and Gray-box Modeling of Nonlinear Audio  Effects
标题:非线性音效的区分黑匣子和灰盒建模
链接:https://arxiv.org/abs/2502.14405
作者:Marco Comunità,  Christian J. Steinmetz,  Joshua D. Reiss
摘要:音频效果广泛用于音频和音乐内容创作的每个阶段。大多数可微分音频效果建模方法属于黑盒或灰盒范例;并且大多数模型已经被提出并应用于非线性效果,如吉他放大器,失真,失真,模糊和压缩器。尽管已经针对手头的任务引入了过多的架构,但仍然缺乏对现有技术的理解,因为大多数出版物都是用一种类型的非线性音频效果和非常少量的设备进行实验。  在这项工作中,我们的目标是通过比较大量非线性音频效果的黑盒和灰盒架构,确定最适合各种设备的音频效果建模景观。在此过程中,我们还:引入时变灰盒模型并提出压缩器、失真和模糊的模型,发布用于音频效果研究的大型数据集- ToneTwist AFx https://github.com/mcomunita/tonetwist-afx-dataset-这也是第一个对社区贡献开放的数据集,根据各种指标评估模型并进行广泛的主观评估。代码https://github.com/mcomunita/nablafx和补充材料https://github.com/mcomunita/nnlinafx-supp-material也可提供。
摘要:Audio effects are extensively used at every stage of audio and music contentcreation. The majority of differentiable audio effects modeling approaches fallinto the black-box or gray-box paradigms; and most models have been proposedand applied to nonlinear effects like guitar amplifiers, overdrive, distortion,fuzz and compressor. Although a plethora of architectures have been introducedfor the task at hand there is still lack of understanding on the state of theart, since most publications experiment with one type of nonlinear audio effectand a very small number of devices. In this work we aim to shed light on the audio effects modeling landscape bycomparing black-box and gray-box architectures on a large number of nonlinearaudio effects, identifying the most suitable for a wide range of devices. Inthe process, we also: introduce time-varying gray-box models and propose modelsfor compressor, distortion and fuzz, publish a large dataset for audio effectsresearch - ToneTwist AFx https://github.com/mcomunita/tonetwist-afx-dataset -that is also the first open to community contributions, evaluate models on avariety of metrics and conduct extensive subjective evaluation. Codehttps://github.com/mcomunita/nablafx and supplementary materialhttps://github.com/mcomunita/nnlinafx-supp-material are also available.

【7】 NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio  Disentanglement for Talking Head Synthesis
标题:NeRF-3DTalker:具有3D先验辅助音频解纠缠的神经辐射场,用于说话的头部合成
链接:https://arxiv.org/abs/2502.14178
作者:Xiaoxing Liu,  Zhilei Liu,  Chongke Bi
备注:Accepted by ICASSP 2025
摘要:讲话头部合成是使用音频合成嘴唇同步的讲话头部视频。最近,NeRF的能力,以提高合成的说话人的真实感和纹理细节引起了研究人员的注意。然而,目前大多数基于音频的NeRF方法都只关注正面人脸的渲染。这些方法无法在新的视图中生成清晰的说话人。当前3D说话头部合成中的另一个普遍挑战是难以对准声学和视觉空间,这通常导致所生成的说话头部的次优对口型。为了解决这些问题,我们提出了神经辐射场与3D事先辅助音频解纠缠的讲话头合成(NeRF-3DTalker)。具体而言,所提出的方法采用3D先验信息合成清晰的说话头与自由的意见。此外,我们提出了一个3D事先辅助的音频解开模块,它的目的是解开音频分为两个不同的类别:功能相关的3D获奖演讲运动和功能相关的说话风格。此外,重新定位所产生的帧是远离说话人的运动空间在真实空间中,我们已经设计了一个局部全球标准化空间。该方法从全局和局部语义两个角度对生成的帧中的不规则位置进行规范化。通过全面的定性和定量实验,已经证明我们的NeRF-3DTalker在合成逼真的说话头部视频方面优于最先进的技术,表现出卓越的图像质量和嘴唇同步。项目网页:https://nerf-3dtalker.github.io/NeRF-3Dtalker。
摘要:Talking head synthesis is to synthesize a lip-synchronized talking head videousing audio. Recently, the capability of NeRF to enhance the realism andtexture details of synthesized talking heads has attracted the attention ofresearchers. However, most current NeRF methods based on audio are exclusivelyconcerned with the rendering of frontal faces. These methods are unable togenerate clear talking heads in novel views. Another prevalent challenge incurrent 3D talking head synthesis is the difficulty in aligning acoustic andvisual spaces, which often results in suboptimal lip-syncing of the generatedtalking heads. To address these issues, we propose Neural Radiance Field with3D Prior Aided Audio Disentanglement for Talking Head Synthesis(NeRF-3DTalker). Specifically, the proposed method employs 3D prior informationto synthesize clear talking heads with free views. Additionally, we propose a3D Prior Aided Audio Disentanglement module, which is designed to disentanglethe audio into two distinct categories: features related to 3D awarded speechmovements and features related to speaking style. Moreover, to reposition thegenerated frames that are distant from the speaker's motion space in the realspace, we have devised a local-global Standardized Space. This methodnormalizes the irregular positions in the generated frames from both global andlocal semantic perspectives. Through comprehensive qualitative and quantitativeexperiments, it has been demonstrated that our NeRF-3DTalker outperformsstate-of-the-art in synthesizing realistic talking head videos, exhibitingsuperior image quality and lip synchronization. Project page:https://nerf-3dtalker.github.io/NeRF-3Dtalker.

【8】 On the application of Visibility Graphs in the Spectral Domain for  Speaker Recognition
标题:频谱域可见性图在说话人识别中的应用
链接:https://arxiv.org/abs/2502.14110
作者:Hernan Bocaccio,  Sergio Iglesias-Pérez,  Miguel Romance,  Regino Criado,  Gabriel B. Mindlin
备注:13 pages, 5 figures
摘要:在这项研究中,我们探讨了潜在的可见性图在频谱域的说话人识别。成年参与者被要求记录五个西班牙元音的发音。对于每一个发声,我们计算的频谱,考虑到语音生产的源滤波器模型,其中共振峰是由声道作为一个无源滤波器与谐振频率的形状。频谱配置文件表现出一致的扬声器内的特性,反映个人声道解剖结构,同时显示扬声器之间的变化。然后,我们从这些光谱配置文件中构建可见性图,并提取各种图论度量来捕获它们的拓扑特征。这些指标被组装成代表每个扬声器的五个元音的特征向量。使用在这些特征上训练的决策树的集合,我们在说话人识别中实现了高准确度。我们的分析确定了关键的拓扑特征,这些特征对区分说话者至关重要。这项研究证明了可见性图的频谱分析的有效性和它们在说话人识别中的潜力。我们还讨论了这种方法的鲁棒性,提供洞察其适用于现实世界的说话人识别系统。本研究利用语音信号在频谱域的拓扑特性,拓展了说话人识别的特征提取工具箱。
摘要:In this study, we explore the potential of visibility graphs in the spectraldomain for speaker recognition. Adult participants were instructed to recordvocalizations of the five Spanish vowels. For each vocalization, we computedthe frequency spectrum considering the source-filter model of speechproduction, where formants are shaped by the vocal tract acting as a passivefilter with resonant frequencies. Spectral profiles exhibited consistentintra-speaker characteristics, reflecting individual vocal tract anatomies,while showing variation between speakers. We then constructed visibility graphsfrom these spectral profiles and extracted various graph-theoretic metrics tocapture their topological features. These metrics were assembled into featurevectors representing the five vowels for each speaker. Using an ensemble ofdecision trees trained on these features, we achieved high accuracy in speakeridentification. Our analysis identified key topological features that werecritical in distinguishing between speakers. This study demonstrates theeffectiveness of visibility graphs for spectral analysis and their potential inspeaker recognition. We also discuss the robustness of this approach, offeringinsights into its applicability for real-world speaker recognition systems.This research contributes to expanding the feature extraction toolbox forspeaker recognition by leveraging the topological properties of speech signalsin the spectral domain.

【9】 Adaptive Convolution for CNN-based Speech Enhancement Models
标题:基于CNN的语音增强模型的自适应卷积
链接:https://arxiv.org/abs/2502.14224
作者:Dahan Wang,  Xiaobin Rong,  Shiruo Sun,  Yuxiang Hu,  Changbao Zhu,  Jing Lu
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing
摘要:基于深度学习的语音增强方法显著提高了语音质量和可懂度。卷积神经网络(CNN)已被证明是许多高性能模型的重要组成部分。在本文中,我们介绍了自适应卷积,一个有效的和通用的卷积模块,提高了模型的能力,自适应地表示语音信号。自适应卷积执行逐帧因果动态卷积,通过组装多个并行候选内核来为每个帧生成时变内核。轻量级注意机制利用当前和历史信息为每个候选内核分配自适应权重,指导它们的聚合。这使得卷积运算能够适应帧级语音频谱特征,从而实现更有效的提取和重构。在各种基于CNN的模型上的实验结果表明,自适应卷积显著提高了性能,而计算复杂度的增加可以忽略不计,特别是对于轻量级模型。此外,我们还提出了自适应卷积递归网络(AdaptCRN),这是一种超轻量模型,它结合了自适应卷积和高效的编码器-解码器设计,与具有类似甚至更高计算成本的模型相比,具有更高的性能。
摘要:Deep learning-based speech enhancement methods have significantly improvedspeech quality and intelligibility. Convolutional neural networks (CNNs) havebeen proven to be essential components of many high-performance models. In thispaper, we introduce adaptive convolution, an efficient and versatileconvolutional module that enhances the model's capability to adaptivelyrepresent speech signals. Adaptive convolution performs frame-wise causaldynamic convolution, generating time-varying kernels for each frame byassembling multiple parallel candidate kernels. A Lightweight attentionmechanism leverages both current and historical information to assign adaptiveweights to each candidate kernel, guiding their aggregation. This enables theconvolution operation to adapt to frame-level speech spectral features, leadingto more efficient extraction and reconstruction. Experimental results onvarious CNN-based models demonstrate that adaptive convolution significantlyimproves the performance with negligible increases in computational complexity,especially for lightweight models. Furthermore, we propose the adaptiveconvolutional recurrent network (AdaptCRN), an ultra-lightweight model thatincorporates adaptive convolution and an efficient encoder-decoder design,achieving superior performance compared to models with similar or even highercomputational costs.

eess.AS音频处理

【1】 Role of the Pretraining and the Adaptation data sizes for low-resource  real-time MRI video segmentation
标题:预训练和自适应数据大小在低资源实时MRI视频分割中的作用
链接:https://arxiv.org/abs/2502.14418
作者:Masoud Thajudeen Tholan,  Vinayaka Hegde,  Chetan Sharma,  Prasanta Kumar Ghosh
备注:Accepted to ICASSP 2025
摘要:实时磁共振成像(rtMRI)经常用于语音产生研究,因为它提供了发音过程中声道的完整视图。本研究探讨rtMRI在分析声道运动的有效性,采用SegNet和UNet模型的空气组织边界(ATB)分割任务。我们使用越来越多的主题和视频对一些基础模型进行了预训练,以评估两个数据集的性能。首先,由来自相同数据源的未见过的视频组成的未见过的主题,比其匹配条件好0.33%和0.91%(分别为像素分类准确度(PCA)和Dice系数)。第二,包括来自新数据源的未见过的视频,其中我们获得了99.63%和98.09%的匹配条件性能准确度(分别为PCA和Dice系数)。这里,匹配条件性能是指仅在测试对象上训练的模型的性能,该测试对象被设置为其他模型的基准。我们的研究结果强调了在有限数据下微调和调整模型的重要性。值得注意的是,我们证明了有效的模型自适应可以用来自任何新数据集的少至15个rtMRI帧来实现。
摘要:Real-time Magnetic Resonance Imaging (rtMRI) is frequently used in speechproduction studies as it provides a complete view of the vocal tract duringarticulation. This study investigates the effectiveness of rtMRI in analyzingvocal tract movements by employing the SegNet and UNet models for Air-TissueBoundary (ATB)segmentation tasks. We conducted pretraining of a few base modelsusing increasing numbers of subjects and videos, to assess performance on twodatasets. First, consisting of unseen subjects with unseen videos from the samedata source, achieving 0.33% and 0.91% (Pixel-wise Classification Accuracy(PCA) and Dice Coefficient respectively) better than its matched condition.Second, comprising unseen videos from a new data source, where we obtained anaccuracy of 99.63% and 98.09% (PCA and Dice Coefficient respectively) of itsmatched condition performance. Here, matched condition performance refers tothe performance of a model trained only on the test subjects which was set as abenchmark for the other models. Our findings highlight the significance offine-tuning and adapting models with limited data. Notably, we demonstratedthat effective model adaptation can be achieved with as few as 15 rtMRI framesfrom any new dataset.

【2】 Adaptive Convolution for CNN-based Speech Enhancement Models
标题:基于CNN的语音增强模型的自适应卷积
链接:https://arxiv.org/abs/2502.14224
作者:Dahan Wang,  Xiaobin Rong,  Shiruo Sun,  Yuxiang Hu,  Changbao Zhu,  Jing Lu
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing
摘要:基于深度学习的语音增强方法显著提高了语音质量和可懂度。卷积神经网络(CNN)已被证明是许多高性能模型的重要组成部分。在本文中,我们介绍了自适应卷积,一个有效的和通用的卷积模块,提高了模型的能力,自适应地表示语音信号。自适应卷积执行逐帧因果动态卷积,通过组装多个并行候选内核来为每个帧生成时变内核。轻量级注意机制利用当前和历史信息为每个候选内核分配自适应权重,指导它们的聚合。这使得卷积运算能够适应帧级语音频谱特征,从而实现更有效的提取和重构。在各种基于CNN的模型上的实验结果表明,自适应卷积显著提高了性能,而计算复杂度的增加可以忽略不计,特别是对于轻量级模型。此外,我们还提出了自适应卷积递归网络(AdaptCRN),这是一种超轻量模型,它结合了自适应卷积和高效的编码器-解码器设计,与具有类似甚至更高计算成本的模型相比,具有更高的性能。
摘要:Deep learning-based speech enhancement methods have significantly improvedspeech quality and intelligibility. Convolutional neural networks (CNNs) havebeen proven to be essential components of many high-performance models. In thispaper, we introduce adaptive convolution, an efficient and versatileconvolutional module that enhances the model's capability to adaptivelyrepresent speech signals. Adaptive convolution performs frame-wise causaldynamic convolution, generating time-varying kernels for each frame byassembling multiple parallel candidate kernels. A Lightweight attentionmechanism leverages both current and historical information to assign adaptiveweights to each candidate kernel, guiding their aggregation. This enables theconvolution operation to adapt to frame-level speech spectral features, leadingto more efficient extraction and reconstruction. Experimental results onvarious CNN-based models demonstrate that adaptive convolution significantlyimproves the performance with negligible increases in computational complexity,especially for lightweight models. Furthermore, we propose the adaptiveconvolutional recurrent network (AdaptCRN), an ultra-lightweight model thatincorporates adaptive convolution and an efficient encoder-decoder design,achieving superior performance compared to models with similar or even highercomputational costs.

【3】 Gesture-Aware Zero-Shot Speech Recognition for Patients with Language  Disorders
标题:语言障碍患者的手势感知Zero-Shot语音识别
链接:https://arxiv.org/abs/2502.13983
作者:Seungbae Kim,  Daeun Lee,  Brielle Stark,  Jinyoung Han
摘要:由于语言处理和理解能力有限,语言障碍患者往往面临重大的沟通挑战,这也影响了他们与主要依赖于自动语音识别(ASR)的语音辅助系统的交互。尽管ASR在解决不流利问题方面取得了进展,但很少有人关注整合非语言沟通方法,如手势,语言障碍患者基本上依赖手势来补充他们的沟通。认识到需要解释的潜在意义的视觉信息,而不是单独捕获的语音,我们提出了一个手势感知的ASR系统,利用多模态大语言模型与zero-shot学习的个人语音障碍。我们的实验结果和分析表明,包括手势信息显着提高语义理解。这项研究可以帮助开发有效的沟通技术,专门用于满足语言障碍者的独特需求。
摘要:Individuals with language disorders often face significant communicationchallenges due to their limited language processing and comprehensionabilities, which also affect their interactions with voice-assisted systemsthat mostly rely on Automatic Speech Recognition (ASR). Despite advancements inASR that address disfluencies, there has been little attention on integratingnon-verbal communication methods, such as gestures, which individuals withlanguage disorders substantially rely on to supplement their communication.Recognizing the need to interpret the latent meanings of visual information notcaptured by speech alone, we propose a gesture-aware ASR system utilizing amultimodal large language model with zero-shot learning for individuals withspeech impairments. Our experiment results and analyses show that includinggesture information significantly enhances semantic understanding. This studycan help develop effective communication technologies, specifically designed tomeet the unique needs of individuals with language impairments.

【4】 Benchmarking Automatic Speech Recognition coupled LLM Modules for  Medical Diagnostics
标题:用于医疗诊断的自动语音识别与LLM模块进行基准测试
链接:https://arxiv.org/abs/2502.13982
作者:Kabir Kumar
摘要:自然语言处理(NLP)和语音识别代理通过实现高效、可访问和专业的患者支持,同时自动化繁重的工作,正在迅速发展医疗保健。这份报告是我的自我项目,其中通过两个阶段的系统分析了医疗通话记录的模型:用于语音转录的自动语音识别(ASR)和用于上下文感知的大型语言模型(LLM),专业响应。ASR,对电话录音进行微调,提供了电话中各种患者语音的一般化转录,而LLM将转录文本与医疗诊断相匹配。一种新的音频预处理策略,被部署为提供传入的记录/呼叫数据的不变性,负载有足够的噪声/削波增强,以使管道对麦克风的类型和患者在呼叫/记录时可能具有的环境条件具有鲁棒性。
摘要:Natural Language Processing (NLP) and Voice Recognition agents are rapidlyevolving healthcare by enabling efficient, accessible, and professional patientsupport while automating grunt work. This report serves as my self projectwherein models finetuned on medical call recordings are analysed through atwo-stage system: Automatic Speech Recognition (ASR) for speech transcriptionand a Large Language Model (LLM) for context-aware, professional responses.ASR, finetuned on phone call recordings provides generalised transcription ofdiverse patient speech over call, while the LLM matches transcribed text tomedical diagnosis. A novel audio preprocessing strategy, is deployed to provideinvariance to incoming recording/call data, laden with sufficient augmentationwith noise/clipping to make the pipeline robust to the type of microphone andambient conditions the patient might have while calling/recording.

【5】 WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken  Dialogue Models
标题:WavRAG:语音对话模型的音频集成检索增强生成
链接:https://arxiv.org/abs/2502.14727
作者:Yifu Chen,  Shengpeng Ji,  Haoxiao Wang,  Ziqing Wang,  Siyu Chen,  Jinzheng He,  Jin Xu,  Zhou Zhao
摘要:检索增强生成(RAG)由于其能够使大型语言模型(LLM)集成外部知识而获得广泛采用。然而,现有的RAG框架主要是为基于文本的LLM设计的,并依赖于自动语音识别来处理语音输入,这会丢弃关键的音频信息,冒着转录错误的风险,并增加计算开销。因此,我们引入了WavRAG,这是第一个具有本地端到端音频支持的检索增强生成框架。WavRAG提供了两个关键功能:1)对ASR进行分类,WavRAG直接处理原始音频进行嵌入和检索。2)WavRAG将音频和文本集成到统一的知识表示中。具体来说,我们提出了WavRetriever,以方便从文本-音频混合知识库检索,并通过集成的思想链推理,进一步提高口语对话模型的上下文能力。与最先进的ASR-Text RAG管道相比,WavRAG实现了相当的检索性能,同时提供了10倍的加速。此外,WavRAG独特的文本-音频混合检索能力将RAG的边界扩展到音频模态。
摘要:Retrieval Augmented Generation (RAG) has gained widespread adoption owing toits capacity to empower large language models (LLMs) to integrate externalknowledge. However, existing RAG frameworks are primarily designed fortext-based LLMs and rely on Automatic Speech Recognition to process speechinput, which discards crucial audio information, risks transcription errors,and increases computational overhead. Therefore, we introduce WavRAG, the firstretrieval augmented generation framework with native, end-to-end audio support.WavRAG offers two key features: 1) Bypassing ASR, WavRAG directly processes rawaudio for both embedding and retrieval. 2) WavRAG integrates audio and textinto a unified knowledge representation. Specifically, we propose theWavRetriever to facilitate the retrieval from a text-audio hybrid knowledgebase, and further enhance the in-context capabilities of spoken dialogue modelsthrough the integration of chain-of-thought reasoning. In comparison tostate-of-the-art ASR-Text RAG pipelines, WavRAG achieves comparable retrievalperformance while delivering a 10x acceleration. Furthermore, WavRAG's uniquetext-audio hybrid retrieval capability extends the boundaries of RAG to theaudio modality.

【6】 Pitch Imperfect: Detecting Audio Deepfakes Through Acoustic Prosodic  Analysis
标题:音调不完美:通过声学韵律分析检测音频Deepfake
链接:https://arxiv.org/abs/2502.14726
作者:Kevin Warren,  Daniel Olszewski,  Seth Layton,  Kevin Butler,  Carrie Gates,  Patrick Traynor
摘要:音频deepfake与有机语音越来越难以区分,经常欺骗认证系统和人类听众。虽然许多技术使用低级音频特征或优化黑盒模型训练,但专注于人类用于识别语音的特征可能是一种更长期稳健的检测方法。我们探索韵律的使用,或人类语音的高级语言特征(例如,音高、语调、抖动)作为检测音频深度伪造的更基本的手段。我们开发了一个基于六个经典韵律特征的检测器,并证明了我们的模型与社区使用的其他基线模型一样,可以检测音频deepfake,准确率为93%,EER为24.7%。更重要的是,我们证明了使用基于语言特征的方法比现有模型的好处,通过应用自适应对手使用$L_{\infty}$ norm攻击对检测器和使用注意力机制在我们的训练可解释性。我们表明,我们可以解释的韵律功能,具有最高的影响模型的决定(抖动,微光和平均基频)和其他模型是非常容易受到简单的$L_{\infty}$范数攻击(99.3%的相对精度下降)。虽然整体性能可能相似,但我们说明了韵律特征方法用于音频深度伪造检测的鲁棒性和可解释性优势。
摘要:Audio deepfakes are increasingly in-differentiable from organic speech, oftenfooling both authentication systems and human listeners. While many techniquesuse low-level audio features or optimization black-box model training, focusingon the features that humans use to recognize speech will likely be a morelong-term robust approach to detection. We explore the use of prosody, or thehigh-level linguistic features of human speech (e.g., pitch, intonation,jitter) as a more foundational means of detecting audio deepfakes. We develop adetector based on six classical prosodic features and demonstrate that ourmodel performs as well as other baseline models used by the community to detectaudio deepfakes with an accuracy of 93% and an EER of 24.7%. More importantly,we demonstrate the benefits of using a linguistic features-based approach overexisting models by applying an adaptive adversary using an $L_{\infty}$ normattack against the detectors and using attention mechanisms in our training forexplainability. We show that we can explain the prosodic features that havehighest impact on the model's decision (Jitter, Shimmer and Mean FundamentalFrequency) and that other models are extremely susceptible to simple$L_{\infty}$ norm attacks (99.3% relative degradation in accuracy). Whileoverall performance may be similar, we illustrate the robustness andexplainability benefits to a prosody feature approach to audio deepfakedetection.

【7】 SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer  Based Speech Recognition
标题:SegAug:用于基于RNN传感器的鲁棒语音识别的ATC对齐分段增强
链接:https://arxiv.org/abs/2502.14685
作者:Khanh Le,  Tuan Vu Ho,  Dung Tran,  Duc Thanh Chau
备注:Accepted to ICASSP 2025
摘要:RNN-Transducer(RNN-T)是语音识别中广泛采用的架构,在端到端框架中集成了声学和语言建模。然而,RNN-T预测器倾向于过度依赖训练数据中的连续单词依赖性,导致高删除错误率,特别是对于不太常见或域外短语。现有的解决方案,如正则化和数据增强,通常会损害性能的其他方面。我们提出了SegAug,一种基于增强的技术,生成上下文不同的音频文本对低的重复级别的语义。这种方法鼓励模型更多地关注声学特征,同时使其内部语言模型的学习文本模式多样化,从而减少删除错误并提高整体性能。对LibriSpeech和Tedlium-v3数据集的评估表明,在小规模环境下,WER相对降低了12.5%,在大规模环境下降低了6.9%。值得注意的是,大部分的改善源于减少的删除错误,相对减少分别为45.4%和18.5%。这些结果突出了SegAug在提高RNN-T鲁棒性方面的有效性,为在各种具有挑战性的场景中增强语音识别性能提供了一个有前途的解决方案。
摘要:RNN-Transducer (RNN-T) is a widely adopted architecture in speechrecognition, integrating acoustic and language modeling in an end-to-endframework. However, the RNN-T predictor tends to over-rely on consecutive worddependencies in training data, leading to high deletion error rates,particularly with less common or out-of-domain phrases. Existing solutions,such as regularization and data augmentation, often compromise other aspects ofperformance. We propose SegAug, an alignment-based augmentation technique thatgenerates contextually varied audio-text pairs with low sentence-levelsemantics. This method encourages the model to focus more on acoustic featureswhile diversifying the learned textual patterns of its internal language model,thereby reducing deletion errors and enhancing overall performance. Evaluationson the LibriSpeech and Tedlium-v3 datasets demonstrate a relative WER reductionof up to 12.5% on small-scale and 6.9% on large-scale settings. Notably, mostof the improvement stems from reduced deletion errors, with relative reductionsof 45.4% and 18.5%, respectively. These results highlight SegAug'seffectiveness in improving RNN-T's robustness, offering a promising solutionfor enhancing speech recognition performance across diverse and challengingscenarios.

【8】 ChunkFormer: Masked Chunking Conformer For Long-Form Speech  Transcription
标题:ChunkFormer:用于长形式语音转录的掩蔽分块协调器
链接:https://arxiv.org/abs/2502.14673
作者:Khanh Le,  Tuan Vu Ho,  Dung Tran,  Duc Thanh Chau
备注:Accepted to ICASSP 2025
摘要:以工业规模部署ASR模型对硬件资源管理提出了重大挑战,特别是对于音频可能持续数小时的长格式转录任务。大型Conformer型号尽管功能强大,但仅限于在80GB GPU上处理15分钟的音频。此外,可变的输入长度会使效率低下的情况恶化,因为标准的重复操作会导致过多的填充,从而增加资源消耗和执行时间。为了解决这个问题,我们引入了ChunkFormer,这是一种高效的ASR模型,它使用相对正确的上下文进行分块处理,从而在低内存GPU上实现长时间的音频传输。ChunkFormer在80GB GPU上处理长达16小时的音频,比当前最先进的FastConformer长1.5倍,同时还提高了长格式转录性能,与Conformer相比,单词错误率绝对降低了7.7%,并在较短的任务中保持了准确性。ChunkFormer的masked mounted m
摘要:Deploying ASR models at an industrial scale poses significant challenges inhardware resource management, especially for long-form transcription taskswhere audio may last for hours. Large Conformer models, despite theircapabilities, are limited to processing only 15 minutes of audio on an 80GBGPU. Furthermore, variable input lengths worsen inefficiencies, as standardbatching leads to excessive padding, increasing resource consumption andexecution time. To address this, we introduce ChunkFormer, an efficient ASRmodel that uses chunk-wise processing with relative right context, enablinglong audio transcriptions on low-memory GPUs. ChunkFormer handles up to 16hours of audio on an 80GB GPU, 1.5x longer than the current state-of-the-artFastConformer, while also boosting long-form transcription performance with upto 7.7% absolute reduction on word error rate and maintaining accuracy onshorter tasks compared to Conformer. By eliminating the need for padding instandard batching, ChunkFormer's masked batching technique reduces executiontime and memory usage by more than 3x in batch processing, substantiallyreducing costs for a wide range of ASR systems, particularly regarding GPUresources for models serving in real-world applications.

【9】 ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by  Reducing Data Distribution Errors
标题:ATRI:通过减少数据分布错误来缓解多语言音频文本检索不确定性
链接:https://arxiv.org/abs/2502.14627
作者:Yuguo Yin,  Yuxin Xie,  Wenyuan Yang,  Dongchao Yang,  Jinghan Ru,  Xianwei Zhuang,  Liming Liang,  Yuexian Zou
摘要:多语言音频文本检索(ML-ATR)是一项具有挑战性的任务,旨在从数据库中检索音频片段或多语言文本。然而,现有的ML-ATR方案遭受例如跨语言的相似性匹配的不一致。我们从理论上分析了多语种模态对齐方向误差和权重误差的不一致性,并提出了量化不一致性的理论权重误差上限。通过对权重误差上界的分析,发现不一致性问题源于语言随机抽样造成的数据分布误差。针对ML-ATR中数据分布错误对召回率和一致性的影响,提出了一种基于1-to-k对比学习和音-英语共锚对比学习的一致性ML-ATR方案。在翻译的AudioCaps和Clotho数据集上的实验结果表明,我们的方案在包括英语在内的八种主流语言的召回率和一致性指标上达到了最先进的性能。我们的代码将在https://github.com/ATRI-ACL/ATRI-ACL上提供。
摘要:Multilingual audio-text retrieval (ML-ATR) is a challenging task that aims toretrieve audio clips or multilingual texts from databases. However, existingML-ATR schemes suffer from inconsistencies for instance similarity matchingacross languages. We theoretically analyze the inconsistency in terms of bothmultilingual modal alignment direction error and weight error, and propose thetheoretical weight error upper bound for quantifying the inconsistency. Basedon the analysis of the weight error upper bound, we find that the inconsistencyproblem stems from the data distribution error caused by random sampling oflanguages. We propose a consistent ML-ATR scheme using 1-to-k contrastivelearning and audio-English co-anchor contrastive learning, aiming to mitigatethe negative impact of data distribution error on recall and consistency inML-ATR. Experimental results on the translated AudioCaps and Clotho datasetsshow that our scheme achieves state-of-the-art performance on recall andconsistency metrics for eight mainstream languages, including English. Our codewill be available at https://github.com/ATRI-ACL/ATRI-ACL.

【10】 Differentiable Black-box and Gray-box Modeling of Nonlinear Audio  Effects
标题:非线性音效的区分黑匣子和灰盒建模
链接:https://arxiv.org/abs/2502.14405
作者:Marco Comunità,  Christian J. Steinmetz,  Joshua D. Reiss
摘要:音频效果广泛用于音频和音乐内容创作的每个阶段。大多数可微分音频效果建模方法属于黑盒或灰盒范例;并且大多数模型已经被提出并应用于非线性效果,如吉他放大器,失真,失真,模糊和压缩器。尽管已经针对手头的任务引入了过多的架构,但仍然缺乏对现有技术的理解,因为大多数出版物都是用一种类型的非线性音频效果和非常少量的设备进行实验。  在这项工作中,我们的目标是通过比较大量非线性音频效果的黑盒和灰盒架构,确定最适合各种设备的音频效果建模景观。在此过程中,我们还:引入时变灰盒模型并提出压缩器、失真和模糊的模型,发布用于音频效果研究的大型数据集- ToneTwist AFx https://github.com/mcomunita/tonetwist-afx-dataset-这也是第一个对社区贡献开放的数据集,根据各种指标评估模型并进行广泛的主观评估。代码https://github.com/mcomunita/nablafx和补充材料https://github.com/mcomunita/nnlinafx-supp-material也可提供。
摘要:Audio effects are extensively used at every stage of audio and music contentcreation. The majority of differentiable audio effects modeling approaches fallinto the black-box or gray-box paradigms; and most models have been proposedand applied to nonlinear effects like guitar amplifiers, overdrive, distortion,fuzz and compressor. Although a plethora of architectures have been introducedfor the task at hand there is still lack of understanding on the state of theart, since most publications experiment with one type of nonlinear audio effectand a very small number of devices. In this work we aim to shed light on the audio effects modeling landscape bycomparing black-box and gray-box architectures on a large number of nonlinearaudio effects, identifying the most suitable for a wide range of devices. Inthe process, we also: introduce time-varying gray-box models and propose modelsfor compressor, distortion and fuzz, publish a large dataset for audio effectsresearch - ToneTwist AFx https://github.com/mcomunita/tonetwist-afx-dataset -that is also the first open to community contributions, evaluate models on avariety of metrics and conduct extensive subjective evaluation. Codehttps://github.com/mcomunita/nablafx and supplementary materialhttps://github.com/mcomunita/nnlinafx-supp-material are also available.

【11】 NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio  Disentanglement for Talking Head Synthesis
标题:NeRF-3DTalker:具有3D先验辅助音频解纠缠的神经辐射场,用于说话的头部合成
链接:https://arxiv.org/abs/2502.14178
作者:Xiaoxing Liu,  Zhilei Liu,  Chongke Bi
备注:Accepted by ICASSP 2025
摘要:讲话头部合成是使用音频合成嘴唇同步的讲话头部视频。最近,NeRF的能力,以提高合成的说话人的真实感和纹理细节引起了研究人员的注意。然而,目前大多数基于音频的NeRF方法都只关注正面人脸的渲染。这些方法无法在新的视图中生成清晰的说话人。当前3D说话头部合成中的另一个普遍挑战是难以对准声学和视觉空间,这通常导致所生成的说话头部的次优对口型。为了解决这些问题,我们提出了神经辐射场与3D事先辅助音频解纠缠的讲话头合成(NeRF-3DTalker)。具体而言,所提出的方法采用3D先验信息合成清晰的说话头与自由的意见。此外,我们提出了一个3D事先辅助的音频解开模块,它的目的是解开音频分为两个不同的类别:功能相关的3D获奖演讲运动和功能相关的说话风格。此外,重新定位所产生的帧是远离说话人的运动空间在真实空间中,我们已经设计了一个局部全球标准化空间。该方法从全局和局部语义两个角度对生成的帧中的不规则位置进行规范化。通过全面的定性和定量实验,已经证明我们的NeRF-3DTalker在合成逼真的说话头部视频方面优于最先进的技术,表现出卓越的图像质量和嘴唇同步。项目网页:https://nerf-3dtalker.github.io/NeRF-3Dtalker。
摘要:Talking head synthesis is to synthesize a lip-synchronized talking head videousing audio. Recently, the capability of NeRF to enhance the realism andtexture details of synthesized talking heads has attracted the attention ofresearchers. However, most current NeRF methods based on audio are exclusivelyconcerned with the rendering of frontal faces. These methods are unable togenerate clear talking heads in novel views. Another prevalent challenge incurrent 3D talking head synthesis is the difficulty in aligning acoustic andvisual spaces, which often results in suboptimal lip-syncing of the generatedtalking heads. To address these issues, we propose Neural Radiance Field with3D Prior Aided Audio Disentanglement for Talking Head Synthesis(NeRF-3DTalker). Specifically, the proposed method employs 3D prior informationto synthesize clear talking heads with free views. Additionally, we propose a3D Prior Aided Audio Disentanglement module, which is designed to disentanglethe audio into two distinct categories: features related to 3D awarded speechmovements and features related to speaking style. Moreover, to reposition thegenerated frames that are distant from the speaker's motion space in the realspace, we have devised a local-global Standardized Space. This methodnormalizes the irregular positions in the generated frames from both global andlocal semantic perspectives. Through comprehensive qualitative and quantitativeexperiments, it has been demonstrated that our NeRF-3DTalker outperformsstate-of-the-art in synthesizing realistic talking head videos, exhibitingsuperior image quality and lip synchronization. Project page:https://nerf-3dtalker.github.io/NeRF-3Dtalker.

【12】 LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems
标题:适用于全日制口语对话系统的LLM增强型对话管理
链接:https://arxiv.org/abs/2502.14145
作者:Hao Zhang,  Weiwei Li,  Rilin Chen,  Vinay Kothapally,  Meng Yu,  Dong Yu
备注:In submission to INTERSPEECH 2025
摘要:在口语对话系统(SDS)中实现全双工通信需要在听、说和思考之间进行实时协调。本文提出了一种语义语音活动检测(VAD)模块作为对话管理器(DM),有效地管理轮对全双工SDS。实现为一个轻量级(0.5B)LLM微调全双工会话数据,语义VAD预测四个控制令牌,以调节轮流切换和轮流保持,区分有意和无意的闯入,同时检测查询完成处理用户暂停和犹豫。通过在短时间间隔内处理输入语音,语义VAD实现了实时决策,而核心对话引擎(CDE)仅在响应生成时激活,从而减少了计算开销。这种设计允许独立的DM优化,而无需重新训练CDE,平衡可扩展的下一代全双工SDS的交互精度和推理效率。
摘要:Achieving full-duplex communication in spoken dialogue systems (SDS) requiresreal-time coordination between listening, speaking, and thinking. This paperproposes a semantic voice activity detection (VAD) module as a dialogue manager(DM) to efficiently manage turn-taking in full-duplex SDS. Implemented as alightweight (0.5B) LLM fine-tuned on full-duplex conversation data, thesemantic VAD predicts four control tokens to regulate turn-switching andturn-keeping, distinguishing between intentional and unintentional barge-inswhile detecting query completion for handling user pauses and hesitations. Byprocessing input speech in short intervals, the semantic VAD enables real-timedecision-making, while the core dialogue engine (CDE) is only activated forresponse generation, reducing computational overhead. This design allowsindependent DM optimization without retraining the CDE, balancing interactionaccuracy and inference efficiency for scalable, next-generation full-duplexSDS.

【13】 On the application of Visibility Graphs in the Spectral Domain for  Speaker Recognition
标题:频谱域可见性图在说话人识别中的应用
链接:https://arxiv.org/abs/2502.14110
作者:Hernan Bocaccio,  Sergio Iglesias-Pérez,  Miguel Romance,  Regino Criado,  Gabriel B. Mindlin
备注:13 pages, 5 figures
摘要:在这项研究中,我们探讨了潜在的可见性图在频谱域的说话人识别。成年参与者被要求记录五个西班牙元音的发音。对于每一个发声,我们计算的频谱,考虑到语音生产的源滤波器模型,其中共振峰是由声道作为一个无源滤波器与谐振频率的形状。频谱配置文件表现出一致的扬声器内的特性,反映个人声道解剖结构,同时显示扬声器之间的变化。然后,我们从这些光谱配置文件中构建可见性图,并提取各种图论度量来捕获它们的拓扑特征。这些指标被组装成代表每个扬声器的五个元音的特征向量。使用在这些特征上训练的决策树的集合,我们在说话人识别中实现了高准确度。我们的分析确定了关键的拓扑特征,这些特征对区分说话者至关重要。这项研究证明了可见性图的频谱分析的有效性和它们在说话人识别中的潜力。我们还讨论了这种方法的鲁棒性,提供洞察其适用于现实世界的说话人识别系统。本研究利用语音信号在频谱域的拓扑特性,拓展了说话人识别的特征提取工具箱。
摘要:In this study, we explore the potential of visibility graphs in the spectraldomain for speaker recognition. Adult participants were instructed to recordvocalizations of the five Spanish vowels. For each vocalization, we computedthe frequency spectrum considering the source-filter model of speechproduction, where formants are shaped by the vocal tract acting as a passivefilter with resonant frequencies. Spectral profiles exhibited consistentintra-speaker characteristics, reflecting individual vocal tract anatomies,while showing variation between speakers. We then constructed visibility graphsfrom these spectral profiles and extracted various graph-theoretic metrics tocapture their topological features. These metrics were assembled into featurevectors representing the five vowels for each speaker. Using an ensemble ofdecision trees trained on these features, we achieved high accuracy in speakeridentification. Our analysis identified key topological features that werecritical in distinguishing between speakers. This study demonstrates theeffectiveness of visibility graphs for spectral analysis and their potential inspeaker recognition. We also discuss the robustness of this approach, offeringinsights into its applicability for real-world speaker recognition systems.This research contributes to expanding the feature extraction toolbox forspeaker recognition by leveraging the topological properties of speech signalsin the spectral domain.

机器翻译由腾讯交互翻译提供,仅供参考