本文经arXiv每日学术速递授权转载
【1】Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits
链接:https://arxiv.org/abs/2501.03805
备注:SLT 2024
摘要:神经语音编辑的进步引起了人们对它们在欺骗攻击中滥用的担忧。传统的部分编辑语音语料库主要集中在剪切和粘贴编辑,虽然保持说话人的一致性,往往会引入可检测的不连续性。最近的方法,如A\text上标{3}T和Voicebox,通过利用上下文信息来改进过渡。为了促进欺骗检测研究,我们介绍了语音填充编辑(SINE)数据集,使用Voicebox创建。我们详细介绍了重新实现Voicebox训练和数据集创建的过程。主观评价证实,使用这种新技术编辑的语音比传统的剪切和粘贴方法更具有挑战性。尽管人类的困难,实验结果表明,基于自我监督的检测器可以在不同的编辑方法的检测,定位和泛化方面取得显着的性能。数据集和相关模型将向公众开放。
摘要:Neural speech editing advancements have raised concerns about their misuse inspoofing attacks. Traditional partially edited speech corpora primarily focuson cut-and-paste edits, which, while maintaining speaker consistency, oftenintroduce detectable discontinuities. Recent methods, likeA\textsuperscript{3}T and Voicebox, improve transitions by leveragingcontextual information. To foster spoofing detection research, we introduce theSpeech INfilling Edit (SINE) dataset, created with Voicebox. We detailed theprocess of re-implementing Voicebox training and dataset creation. Subjectiveevaluations confirm that speech edited using this novel technique is morechallenging to detect than conventional cut-and-paste methods. Despite humandifficulty, experimental results demonstrate that self-supervised-baseddetectors can achieve remarkable performance in detection, localization, andgeneralization across different edit methods. The dataset and related modelswill be made publicly available.
标题:多标签跨语言自动音乐流派分类,包含句子BERT
链接:https://arxiv.org/abs/2501.03769
备注:5 pages
摘要:音乐类型是由歌曲的风格特征和艺术家观众的文化偏好形成的。使用歌词对音乐流派进行自动分类可以在诸如推荐系统、播放列表创建和库组织的若干应用中有用。我们提出了一个基于sBERT生成的多语言句子嵌入的多标签、跨语言流派分类系统。使用双语葡萄牙语-英语数据集与八个重叠的流派,我们证明了系统的能力,训练在一种语言的歌词和预测流派在另一个。我们的方法优于翻译歌词和使用词袋表示的基线方法,将体裁平均F1分数从0.35提高到0.69。分类器使用一对多架构,使其能够为单个歌词分配多个流派标签。实验结果表明,数据集中化显著提高了跨语言性能。这种方法提供了一个可扩展的解决方案,跨代表性不足的语言和文化领域的流派分类,推进音乐信息检索系统的能力。
摘要:Music genres are shaped by both the stylistic features of songs and thecultural preferences of artists' audiences. Automatic classification of musicgenres using lyrics can be useful in several applications such asrecommendation systems, playlist creation, and library organization. We presenta multi-label, cross-lingual genre classification system based on multilingualsentence embeddings generated by sBERT. Using a bilingual Portuguese-Englishdataset with eight overlapping genres, we demonstrate the system's ability totrain on lyrics in one language and predict genres in another. Our approachoutperforms the baseline approach of translating lyrics and using abag-of-words representation, improving the genrewise average F1-Score from 0.35to 0.69. The classifier uses a one-vs-all architecture, enabling it to assignmultiple genre labels to a single lyric. Experimental results reveal thatdataset centralization notably improves cross-lingual performance. Thisapproach offers a scalable solution for genre classification acrossunderrepresented languages and cultural domains, advancing the capabilities ofmusic information retrieval systems.
标题:用于从神经活动重建高保真语音的NeuIncept解码器
链接:https://arxiv.org/abs/2501.03757
摘要:本文介绍了一种新的算法设计的语音合成从神经活动记录使用侵入性脑电图(EEG)技术。该系统提供了一个有前途的通信解决方案,为个人严重的语音障碍。我们的方法的核心是集成的时间-频率功能的高伽玛波段计算从EEG记录与先进的NeuroIncept解码器架构。这种神经网络架构结合了卷积神经网络(CNN)和门控递归单元(GRU),可以从神经模式中重建音频频谱图。我们的模型显示了预测谱图和实际谱图之间稳健的平均相关系数,尽管受试者间的差异表明参与者之间存在不同的神经处理机制。总的来说,我们的研究强调了神经解码技术在恢复言语障碍患者沟通能力方面的潜力,并为脑机接口技术的未来发展铺平了道路。
摘要:This paper introduces a novel algorithm designed for speech synthesis fromneural activity recordings obtained using invasive electroencephalography (EEG)techniques. The proposed system offers a promising communication solution forindividuals with severe speech impairments. Central to our approach is theintegration of time-frequency features in the high-gamma band computed from EEGrecordings with an advanced NeuroIncept Decoder architecture. This neuralnetwork architecture combines Convolutional Neural Networks (CNNs) and GatedRecurrent Units (GRUs) to reconstruct audio spectrograms from neural patterns.Our model demonstrates robust mean correlation coefficients between predictedand actual spectrograms, though inter-subject variability indicates distinctneural processing mechanisms among participants. Overall, our study highlightsthe potential of neural decoding techniques to restore communicative abilitiesin individuals with speech disorders and paves the way for future advancementsin brain-computer interface technologies.
标题:Guitar-TECHS:一个电子吉他数据集,涵盖技术、音乐摘录、和弦和音阶,使用多种硬件阵列
链接:https://arxiv.org/abs/2501.03720
备注:None
摘要:与吉他相关的机器听力研究涉及音色转换、性能生成和自动转录等任务。然而,由于声学多样性和音乐内容不足,小数据集往往限制了模型的鲁棒性。为了解决这些问题,我们引入了Guitar-TECHS,这是一个全面的数据集,包含各种吉他技术,音乐摘录,和弦和音阶。这些元素由不同的音乐家在不同的录音环境中演奏。Guitar-TECHS集成了两个立体声麦克风的录音:一个位于表演者头部的自我中心麦克风和一个位于表演者前方的偏心麦克风。它还包括直接输入录音和麦克风放大器输出,提供广泛的音频输入和录音质量。所有信号和标签都正确同步。它的多视角和多模态内容使Guitar-TECHS成为推进数据驱动吉他研究的宝贵资源,并开发强大的吉他聆听算法。我们提供了经验数据来证明数据集在训练吉他谱转录的鲁棒模型方面的有效性。
摘要:Guitar-related machine listening research involves tasks like timbretransfer, performance generation, and automatic transcription. However, smalldatasets often limit model robustness due to insufficient acoustic diversityand musical content. To address these issues, we introduce Guitar-TECHS, acomprehensive dataset featuring a variety of guitar techniques, musicalexcerpts, chords, and scales. These elements are performed by diverse musiciansacross various recording settings. Guitar-TECHS incorporates recordings fromtwo stereo microphones: an egocentric microphone positioned on the performer'shead and an exocentric microphone placed in front of the performer. It alsoincludes direct input recordings and microphoned amplifier outputs, offering awide spectrum of audio inputs and recording qualities. All signals and MIDIlabels are properly synchronized. Its multi-perspective and multi-modal contentmakes Guitar-TECHS a valuable resource for advancing data-driven guitarresearch, and to develop robust guitar listening algorithms. We provideempirical data to demonstrate the dataset's effectiveness in training robustmodels for Guitar Tablature Transcription.
标题:无监督语音分割:使用语音语言模型的通用方法
链接:https://arxiv.org/abs/2501.03711
摘要:在本文中,我们介绍了一种无监督的语音分割方法,它建立在以前研究的方法,例如,说话人日记,同时适用于一组包含的声学语义区别,铺平了道路,走向一般的无监督语音分割方法。与传统的语音和音频分割不同,传统的语音和音频分割主要关注输入信号中的频谱变化,例如,电话分割,我们的方法试图将说出的话语分割成具有不同声学语义风格的块,专注于不能很好地翻译成文本的声学语义信息,例如,情绪或说话者。虽然大多数语音分割任务只处理一种风格的变化,例如,情感日记化,我们的方法试图处理多种声学语义风格的变化。利用语音语言模型(SLM)的最新进展,我们提出了一个简单的无监督的方法来分割一个给定的语音话语。我们经验证明了所提出的方法的有效性,考虑几个设置。结果表明,该方法是优于评估基线的边界检测,段纯度,和过分割。代码可在https://github.com/avishaiElmakies/unsupervised_speech_segmentation_using_slm上获得。
摘要:In this paper, we introduce an unsupervised approach for Speech Segmentation,which builds on previously researched approaches, e.g., Speaker Diarization,while being applicable to an inclusive set of acoustic-semantic distinctions,paving a path towards a general Unsupervised Speech Segmentation approach.Unlike traditional speech and audio segmentation, which mainly focuses onspectral changes in the input signal, e.g., phone segmentation, our approachtries to segment the spoken utterance into chunks with differingacoustic-semantic styles, focusing on acoustic-semantic information that doesnot translate well into text, e.g., emotion or speaker. While most SpeechSegmentation tasks only handle one style change, e.g., emotion diarization, ourapproach tries to handle multiple acoustic-semantic style changes. Leveragingrecent advances in Speech Language Models (SLMs), we propose a simpleunsupervised method to segment a given speech utterance. We empiricallydemonstrate the effectiveness of the proposed approach by considering severalsetups. Results suggest that the proposed method is superior to the evaluatedbaselines on boundary detection, segment purity, and over-segmentation. Code isavailable athttps://github.com/avishaiElmakies/unsupervised_speech_segmentation_using_slm.
标题:MAJL:用于音乐源分离和音调估计的模型不可知联合学习框架
链接:https://arxiv.org/abs/2501.03689
摘要:音乐源分离和音高估计是音乐信息检索中的两个重要任务。通常,音高估计的输入是从音乐源分离的输出中获得的。因此,现有方法试图同时执行这两个任务,以便利用两个任务之间的互利关系。然而,这些方法仍然面临两个关键挑战,限制了这两个任务的改进:缺乏标记数据和联合学习优化。为了解决这些挑战,我们提出了一个模型无关的联合学习(MAJL)框架,这两个任务。MAJL是一个通用框架,可以为每个任务使用变体模型。它包括一个两阶段的训练方法和一个动态加权方法,称为动态加权硬样本(DWHS),分别解决缺乏标记数据和联合学习优化。在公共音乐数据集上的实验结果表明,MAJL在这两个任务上都优于最先进的方法,音乐源分离的信号失真比(SDR)显著提高了0.92,音高估计的原始音高准确度(RPA)显著提高了2.71%。此外,全面的研究不仅验证了MAJL的各个组成部分的有效性,但也表明MAJL在适应不同的模型架构的巨大的通用性。
摘要:Music source separation and pitch estimation are two vital tasks in musicinformation retrieval. Typically, the input of pitch estimation is obtainedfrom the output of music source separation. Therefore, existing methods havetried to perform these two tasks simultaneously, so as to leverage the mutuallybeneficial relationship between both tasks. However, these methods still facetwo critical challenges that limit the improvement of both tasks: the lack oflabeled data and joint learning optimization. To address these challenges, wepropose a Model-Agnostic Joint Learning (MAJL) framework for both tasks. MAJLis a generic framework and can use variant models for each task. It includes atwo-stage training method and a dynamic weighting method named Dynamic Weightson Hard Samples (DWHS), which addresses the lack of labeled data and jointlearning optimization, respectively. Experimental results on public musicdatasets show that MAJL outperforms state-of-the-art methods on both tasks,with significant improvements of 0.92 in Signal-to-Distortion Ratio (SDR) formusic source separation and 2.71% in Raw Pitch Accuracy (RPA) for pitchestimation. Furthermore, comprehensive studies not only validate theeffectiveness of each component of MAJL, but also indicate the great generalityof MAJL in adapting to different model architectures.
标题:有效且高效的语音基础模型混合精度量化
链接:https://arxiv.org/abs/2501.03643
备注:To appear at IEEE ICASSP 2025
摘要:本文提出了一种新的混合精度量化方法的语音基础模型,紧密结合的混合精度学习和量化模型参数估计到一个单一的模型压缩阶段。在LibriSpeech数据集上使用微调的wav2vec2.0-base和HuBERT-large模型进行的实验表明,所得到的混合精度量化模型将无损压缩率提高了1.7倍和1.9倍,分别超过了在单独和不连贯的阶段中执行精度学习和模型参数量化的均匀精度和两阶段混合精度量化基线,而在32位全精度模型上不会引起统计上的字错误率(WER)增加。wav2vec2.0-base和HuBERT-large模型的系统压缩时间比两阶段混合精度基线减少了1.9倍和1.5倍,同时两者都产生了较低的WER。性能最佳的3.5位混合精度量化HuBERT-large模型产生的无损压缩比为32位全精度系统的8.6倍。
摘要:This paper presents a novel mixed-precision quantization approach for speechfoundation models that tightly integrates mixed-precision learning andquantized model parameter estimation into one single model compression stage.Experiments conducted on LibriSpeech dataset with fine-tuned wav2vec2.0-baseand HuBERT-large models suggest the resulting mixed-precision quantized modelsincreased the lossless compression ratio by factors up to 1.7x and 1.9x overthe respective uniform-precision and two-stage mixed-precision quantizedbaselines that perform precision learning and model parameters quantization inseparate and disjointed stages, while incurring no statistically word errorrate (WER) increase over the 32-bit full-precision models. The systemcompression time of wav2vec2.0-base and HuBERT-large models is reduced by up to1.9 and 1.5 times over the two-stage mixed-precision baselines, while bothproduce lower WERs. The best-performing 3.5-bit mixed-precision quantizedHuBERT-large model produces a lossless compression ratio of 8.6x over the32-bit full-precision system.
标题:AADNet:基于线索掩蔽范式探索脑电时空信息以快速准确地定位和音色检测听觉注意力
链接:https://arxiv.org/abs/2501.03571
摘要:在噪声环境中,从脑电信号中解码出的听觉注意力可以推断出使用者正在关注哪个源。解码算法和实验范式设计对于技术在实际应用中的发展至关重要。为了模拟真实场景,本研究提出了一个线索掩蔽听觉注意范式,以避免实验前的信息泄露。为了在低延迟的情况下获得高解码精度,提出了一种端到端深度学习模型AADNet,以利用来自EEG信号的短时间窗口的时空信息。结果表明,在0.5秒的EEG窗口下,AADNet对听觉方向注意(OA)和音色注意(TA)的平均解码准确率分别为93.46%和91.09%。它的性能明显优于以前的五种方法,并且不需要原始音频源的知识。这一工作证明了从脑电信号中快速准确地检测听觉注意的方向和音色是可能的。研究结果为多属性听觉注意力的实时解码提供了理论依据,为神经导航助听器和其他辅助听力设备的应用提供了参考。
摘要:Auditory attention decoding from electroencephalogram (EEG) could infer towhich source the user is attending in noisy environments. Decoding algorithmsand experimental paradigm designs are crucial for the development of technologyin practical applications. To simulate real-world scenarios, this studyproposed a cue-masked auditory attention paradigm to avoid information leakagebefore the experiment. To obtain high decoding accuracy with low latency, anend-to-end deep learning model, AADNet, was proposed to exploit thespatiotemporal information from the short time window of EEG signals. Theresults showed that with a 0.5-second EEG window, AADNet achieved an averageaccuracy of 93.46% and 91.09% in decoding auditory orientation attention (OA)and timbre attention (TA), respectively. It significantly outperformed fiveprevious methods and did not need the knowledge of the original audio source.This work demonstrated that it was possible to detect the orientation andtimbre of auditory attention from EEG signals fast and accurately. The resultsare promising for the real-time multi-property auditory attention decoding,facilitating the application of the neuro-steered hearing aids and otherassistive listening devices.
标题:用于口语关键词定位的声道长度扭曲功能
链接:https://arxiv.org/abs/2501.03523
摘要:在本文中,我们提出了几种方法,将声道长度(VAN)扭曲的功能口语关键字定位(KWS)。第一种方法是独立于VTL的KWS,涉及训练单个深度神经网络(DNN),该网络利用具有各种扭曲因子的VTL特征。在训练期间,每个时期随机选择特定的VTL功能,从而允许探索VTL变体。在测试过程中,测试话语的具有不同扭曲因子的特征被相对于DNN评分,并以相等的权重组合。在第二种方法中,针对DNN对测试话语的常规特征(没有扭曲)进行评分。第三种方法,VTL连接KWS,连接扭曲的特征以形成KWS的高维特征。在英文Google Command数据集上进行的评估表明,所提出的方法提高了KWS的准确性。
摘要:In this paper, we propose several methods that incorporate vocal tract length(VTL) warped features for spoken keyword spotting (KWS). The first method,VTL-independent KWS, involves training a single deep neural network (DNN) thatutilizes VTL features with various warping factors. During training, a specificVTL feature is randomly selected per epoch, allowing the exploration of VTLvariations. During testing, the VTL features with different warping factors ofa test utterance are scored against the DNN and combined with equal weight. Inthe second method scores the conventional features of a test utterance (withoutVTL warping) against the DNN. The third method, VTL-concatenation KWS,concatenates VTL warped features to form high-dimensional features for KWS.Evaluations carried out on the English Google Command dataset demonstrate thatthe proposed methods improve the accuracy of KWS.
标题:LHGNN:用于音频分类和标记的局部-高位图神经网络
链接:https://arxiv.org/abs/2501.03464
摘要:Transformers在音频处理任务中建立了新的基准,利用自我注意机制来捕获音频数据中的复杂模式和依赖关系。然而,它们对成对交互的关注限制了它们处理识别不同音频对象所必需的高阶关系的能力。为了解决这一限制,这项工作引入了局部高阶图神经网络(LHGNN),这是一种基于图的模型,通过将局部邻域信息与来自模糊C均值聚类的高阶数据相结合来增强特征理解,从而捕获更广泛的音频关系。在三个公开的音频数据集上对该模型进行的评估表明,它在所有基准测试中的性能都优于基于transformer的模型,同时使用的参数要少得多。此外,LHGNN在缺乏ImageNet预训练的场景中表现出明显的优势,在无法获得大量预训练数据的环境中建立了其有效性和效率。
摘要:Transformers have set new benchmarks in audio processing tasks, leveragingself-attention mechanisms to capture complex patterns and dependencies withinaudio data. However, their focus on pairwise interactions limits their abilityto process the higher-order relations essential for identifying distinct audioobjects. To address this limitation, this work introduces the Local- HigherOrder Graph Neural Network (LHGNN), a graph based model that enhances featureunderstanding by integrating local neighbourhood information with higher-orderdata from Fuzzy C-Means clusters, thereby capturing a broader spectrum of audiorelationships. Evaluation of the model on three publicly available audiodatasets shows that it outperforms Transformer-based models across allbenchmarks while operating with substantially fewer parameters. Moreover, LHGNNdemonstrates a distinct advantage in scenarios lacking ImageNet pretraining,establishing its effectiveness and efficiency in environments where extensivepretraining data is unavailable.
标题:用于说话人验证的频谱感知低等级自适应
链接:https://arxiv.org/abs/2501.03829
备注:Accepted by ICASSP 2025
摘要:先前的研究表明,预训练模型的权重矩阵的主要奇异向量捕获关键知识。相反,与小奇异值相关联的那些可能包含噪声或不太可靠的信息。因此,基于LoRA的参数高效微调(PEFT)方法,它不限制频谱空间的使用,可能是不有效的任务,需要高的表示能力。在这项研究中,我们通过将预先训练的权重矩阵的频谱信息纳入微调过程来增强现有的PEFT技术。我们调查光谱适应策略,特别侧重于顶部奇异向量的加法调整。这是通过将奇异值分解(SVD)应用于预训练的权重矩阵并限制顶部频谱空间内的微调来实现的。VoxCeleb 1和CN-Celeb 1上的大量说话人验证实验表明,所提出的方法增强了调谐性能。代码发布于https://github.com/lizhepolyu/SpectralFT。
摘要:Previous research has shown that the principal singular vectors of apre-trained model's weight matrices capture critical knowledge. In contrast,those associated with small singular values may contain noise or less reliableinformation. As a result, the LoRA-based parameter-efficient fine-tuning (PEFT)approach, which does not constrain the use of the spectral space, may not beeffective for tasks that demand high representation capacity. In this study, weenhance existing PEFT techniques by incorporating the spectral information ofpre-trained weight matrices into the fine-tuning process. We investigatespectral adaptation strategies with a particular focus on the additiveadjustment of top singular vectors. This is accomplished by applying singularvalue decomposition (SVD) to the pre-trained weight matrices and restrictingthe fine-tuning within the top spectral space. Extensive speaker verificationexperiments on VoxCeleb1 and CN-Celeb1 demonstrate enhanced tuning performancewith the proposed approach. Code is released athttps://github.com/lizhepolyu/SpectralFT.
标题:通用说话人嵌入自由目标说话人提取和个人语音活动检测
链接:https://arxiv.org/abs/2501.03612
摘要:确定“谁在什么时候说了什么”在现实世界的应用中仍然具有挑战性。在典型的场景中,说话人日记(SD)被用来解决“谁在什么时候说话”的问题,而目标说话人提取(TSE)或目标说话人自动语音识别(TSASR)技术被用来解决“谁说了什么”的问题。“虽然一些工作已经取得了可喜的成果,通过结合SD和TSE系统,SD和TSE之间的不一致仍然存在输出不一致和场景不匹配。为了解决这些限制,我们提出了一个通用的扬声器嵌入自由目标扬声器提取和个人语音活动检测(USEF-TP)模型,共同执行TSE和个人语音活动检测(PVAD)。USEF-TP利用通过交叉注意机制获得的帧级特征作为说话者相关特征,而不是像传统方法那样使用说话者嵌入。此外,多任务学习算法与一个多任务感知差分损失函数的应用,以确保在不同级别的扬声器重叠的鲁棒性能。实验结果表明,我们提出的USEF-TP模型在LibriMix和SparseLibriMix数据集上的TSE和PVAD任务中取得了优异的性能。
摘要:Determining 'who spoke what and when' remains challenging in real-worldapplications. In typical scenarios, Speaker Diarization (SD) is employed toaddress the problem of 'who spoke when,' while Target Speaker Extraction (TSE)or Target Speaker Automatic Speech Recognition (TSASR) techniques are utilizedto resolve the issue of 'who spoke what.' Although some works have achievedpromising results by combining SD and TSE systems, inconsistencies remainbetween SD and TSE regarding both output inconsistency and scenario mismatch.To address these limitations, we propose a Universal Speaker Embedding FreeTarget Speaker Extraction and Personal Voice Activity Detection (USEF-TP) modelthat jointly performs TSE and Personal Voice Activity Detection (PVAD). USEF-TPleverages frame-level features obtained through a cross-attention mechanism asspeaker-related features instead of using speaker embeddings as in traditionalapproaches. Additionally, a multi-task learning algorithm with a scenario-awaredifferentiated loss function is applied to ensure robust performance acrossvarious levels of speaker overlap. The experimental results show that ourproposed USEF-TP model achieves superior performance in TSE and PVAD tasks onthe LibriMix and SparseLibriMix datasets.
标题:开发一种用于帕金森病诊断的通用言语标记物
链接:https://arxiv.org/abs/2501.03581
摘要:帕金森氏病(PD)是一种神经退行性疾病,其特征在于运动症状,包括在早期阶段改变声音产生。早期诊断不仅对改善PD患者的生活质量至关重要,而且对增强早期神经退行性变期间潜在疾病修饰疗法的疗效也至关重要,这是当前诊断工具经常错过的窗口。在本文中,我们提出了一个更普遍的方法,通过域自适应和自监督学习的PD识别。我们展示了所提出的方法在不同语言的不同数据集上的泛化能力。我们的方法利用了HuBERT,这是一个最初为语音识别而训练的大型深度神经网络,并进一步在来自与目标群体相似的人群的未标记语音数据上训练它,即,老年人以自我监督的方式。然后对模型进行微调,使其适用于多种语言的不同数据集,包括英语,意大利语和西班牙语。对四个公开的PD数据集的评估证明了该模型的有效性,平均特异性为92.1%,平均灵敏度为91.2%。该方法在大规模人群中提供了客观和一致的评估,解决了人类评估中固有的可变性,并提供了一种非侵入性,成本效益和可访问的诊断选项。
摘要:Parkinson's Disease (PD) is a neurodegenerative disorder characterized bymotor symptoms, including altered voice production in the early stages. Earlydiagnosis is crucial not only to improve PD patients' quality of life but alsoto enhance the efficacy of potential disease-modifying therapies during earlyneurodegeneration, a window often missed by current diagnostic tools. In thispaper, we propose a more generalizable approach to PD recognition throughdomain adaptation and self-supervised learning. We demonstrate thegeneralization capabilities of the proposed approach across diverse datasets indifferent languages. Our approach leverages HuBERT, a large deep neural networkoriginally trained for speech recognition and further trains it on unlabeledspeech data from a population that is similar to the target group, i.e., theelderly, in a self-supervised manner. The model is then fine-tuned and adaptedfor use across different datasets in multiple languages, including English,Italian, and Spanish. Evaluations on four publicly available PD datasetsdemonstrate the model's efficacy, achieving an average specificity of 92.1% andan average sensitivity of 91.2%. This method offers objective and consistentevaluations across large populations, addressing the variability inherent inhuman assessments and providing a non-invasive, cost-effective and accessiblediagnostic option.
标题:突破尖峰:尖峰窗口解码,实现加速、精确的自动语音识别
链接:https://arxiv.org/abs/2501.03257
备注:Accepted by ICASSP 2025
摘要:近年来,端到端自动语音识别已成为工业界和学术界的主流方法。为了在特定场景中优化系统性能,加权语音状态转换器(WFST)被广泛用于集成声学和语言模型,利用其在静态图中隐式融合语言模型的能力,从而确保鲁棒的识别,同时也促进快速纠错。然而,WFST需要通过自回归逐帧搜索CTC后验概率,这大大阻碍了推理速度。在这项工作中,我们深入研究了CTC输出的尖峰特性,并进一步提出了猜想,即非空白尖峰的相邻帧携带有益于模型的语义信息。在此基础上,我们提出了尖峰窗口解码算法,该算法通过使WFST中解码的帧的数量与CTC输出中的尖峰帧的数量线性相关,从而大大提高了推理速度,同时保证了识别性能。我们的方法实现了SOTA识别的准确性,大大加快了解码速度,在AISHELL-1和大规模内部数据集上都得到了证明,建立了一种将CTC输出与WFST集成的开创性方法。
摘要:Recently, end-to-end automatic speech recognition has become the mainstreamapproach in both industry and academia. To optimize system performance inspecific scenarios, the Weighted Finite-State Transducer (WFST) is extensivelyused to integrate acoustic and language models, leveraging its capacity toimplicitly fuse language models within static graphs, thereby ensuring robustrecognition while also facilitating rapid error correction. However, WFSTnecessitates a frame-by-frame search of CTC posterior probabilities throughautoregression, which significantly hampers inference speed. In this work, wethoroughly investigate the spike property of CTC outputs and further proposethe conjecture that adjacent frames to non-blank spikes carry semanticinformation beneficial to the model. Building on this, we propose the SpikeWindow Decoding algorithm, which greatly improves the inference speed by makingthe number of frames decoded in WFST linearly related to the number of spikingframes in the CTC output, while guaranteeing the recognition performance. Ourmethod achieves SOTA recognition accuracy with significantly acceleratesdecoding speed, proven across both AISHELL-1 and large-scale In-House datasets,establishing a pioneering approach for integrating CTC output with WFST.
标题:通过MEG驱动编码模型架起听觉感知和语言理解的桥梁
链接:https://arxiv.org/abs/2501.03246
备注:10 pages, 4 figures, Accepted at ICLR2024 Workshop TS4H
摘要:了解听觉和语言处理背后的神经机制是推进认知神经科学的关键。在这项研究中,我们使用脑磁图(MEG)数据来分析大脑对口语刺激的反应。我们开发了两种不同的编码模型:音频到MEG编码器,它使用时频分解(TFD)和wav 2 vec 2潜在空间表示,以及文本到MEG编码器,它利用CLIP和GPT-2嵌入。这两种模型都成功地预测了神经活动,证明了估计和观察到的MEG信号之间的显着相关性。然而,文本到MEG模型优于基于音频的模型,实现了更高的皮尔逊相关(PC)分数。在空间上,我们确定,基于语义的嵌入(TFD和wav 2 vec 2)主要激活侧颞区,这是负责初级听觉处理和听觉信号的整合。相比之下,文本嵌入(CLIP和GPT-2)主要涉及额叶皮层,特别是布罗卡区,该区域与高阶语言处理相关,包括语义整合和语言产生,特别是在8-30 Hz频率范围内。这些区域的强烈参与表明,听觉刺激是通过更直接的感觉通路处理的,而语言信息是通过整合意义和认知控制的网络编码的。我们的研究结果揭示了听觉和语言信息处理的不同神经通路,在额叶区域的文本表征具有更高的编码准确性。这些见解完善了我们对大脑处理听觉和文本信息的功能结构的理解,为复杂语言刺激的神经反应建模提供了定量进步。
摘要:Understanding the neural mechanisms behind auditory and linguistic processingis key to advancing cognitive neuroscience. In this study, we useMagnetoencephalography (MEG) data to analyze brain responses to spoken languagestimuli. We develop two distinct encoding models: an audio-to-MEG encoder,which uses time-frequency decompositions (TFD) and wav2vec2 latent spacerepresentations, and a text-to-MEG encoder, which leverages CLIP and GPT-2embeddings. Both models successfully predict neural activity, demonstratingsignificant correlations between estimated and observed MEG signals. However,the text-to-MEG model outperforms the audio-based model, achieving higherPearson Correlation (PC) score. Spatially, we identify that auditory-basedembeddings (TFD and wav2vec2) predominantly activate lateral temporal regions,which are responsible for primary auditory processing and the integration ofauditory signals. In contrast, textual embeddings (CLIP and GPT-2) primarilyengage the frontal cortex, particularly Broca's area, which is associated withhigher-order language processing, including semantic integration and languageproduction, especially in the 8-30 Hz frequency range. The strong involvementof these regions suggests that auditory stimuli are processed through moredirect sensory pathways, while linguistic information is encoded via networksthat integrate meaning and cognitive control. Our results reveal distinctneural pathways for auditory and linguistic information processing, with higherencoding accuracy for text representations in the frontal regions. Theseinsights refine our understanding of the brain's functional architecture inprocessing auditory and textual information, offering quantitative advancementsin the modelling of neural responses to complex language stimuli.
标题:用于说话人验证的频谱感知低等级自适应
链接:https://arxiv.org/abs/2501.03829
备注:Accepted by ICASSP 2025
摘要:先前的研究表明,预训练模型的权重矩阵的主要奇异向量捕获关键知识。相反,与小奇异值相关联的那些可能包含噪声或不太可靠的信息。因此,基于LoRA的参数高效微调(PEFT)方法,它不限制频谱空间的使用,可能是不有效的任务,需要高的表示能力。在这项研究中,我们通过将预先训练的权重矩阵的频谱信息纳入微调过程来增强现有的PEFT技术。我们调查光谱适应策略,特别侧重于顶部奇异向量的加法调整。这是通过将奇异值分解(SVD)应用于预训练的权重矩阵并限制顶部频谱空间内的微调来实现的。在VoxCeleb 1和CN-Celeb 1上进行的广泛说话人验证实验表明,所提出的方法增强了调谐性能。代码发布于https://github.com/lizhepolyu/SpectralFT。
摘要:Previous research has shown that the principal singular vectors of apre-trained model's weight matrices capture critical knowledge. In contrast,those associated with small singular values may contain noise or less reliableinformation. As a result, the LoRA-based parameter-efficient fine-tuning (PEFT)approach, which does not constrain the use of the spectral space, may not beeffective for tasks that demand high representation capacity. In this study, weenhance existing PEFT techniques by incorporating the spectral information ofpre-trained weight matrices into the fine-tuning process. We investigatespectral adaptation strategies with a particular focus on the additiveadjustment of top singular vectors. This is accomplished by applying singularvalue decomposition (SVD) to the pre-trained weight matrices and restrictingthe fine-tuning within the top spectral space. Extensive speaker verificationexperiments on VoxCeleb1 and CN-Celeb1 demonstrate enhanced tuning performancewith the proposed approach. Code is released athttps://github.com/lizhepolyu/SpectralFT.
标题:用于弱监督声音事件检测的帧级预测的伪强标签
链接:https://arxiv.org/abs/2501.03740
备注:5 pages
摘要:弱监督声音事件检测(WSSED)依赖于没有精确起始和偏移时间的音频标签,由于缺乏包括事件的精确时间边界的强标记数据,其已经变得普遍。本研究引入帧级伪强标签(FPSL),以克服缺乏时间信息WSSED生成伪强标签从帧级预测。这增强了训练期间的时间定位,并解决了裁剪式弱监督的局限性。我们在三个基准数据集(DCASE2017任务4,DCASE2018任务4和UrbanSED)上验证了我们的方法,并在复调声音检测分数(PSDS),基于事件的F1分数和基于交叉点的F1分数等关键指标上展示了显着的改进。例如,使用FPSL训练的卷积递归神经网络(CRNN)在DCASE2017上的PSDS 1中比基线模型高出4.9%,在DCASE2018上高出7.6%,在UrbanSED上高出1.8%,证实了我们的方法在提高模型性能方面的有效性。
摘要:Weakly Supervised Sound Event Detection (WSSED), which relies on audio tagswithout precise onset and offset times, has become prevalent due to thescarcity of strongly labeled data that includes exact temporal boundaries forevents. This study introduces Frame-level Pseudo Strong Labeling (FPSL) toovercome the lack of temporal information in WSSED by generating pseudo stronglabels from frame-level predictions. This enhances temporal localization duringtraining and addresses the limitations of clip-wise weak supervision. Wevalidate our approach across three benchmark datasets (DCASE2017 Task 4,DCASE2018 Task 4, and UrbanSED) and demonstrate significant improvements in keymetrics such as the Polyphonic Sound Detection Scores (PSDS), event-based F1scores, and intersection-based F1 scores. For example, Convolutional RecurrentNeural Networks (CRNNs) trained with FPSL outperform baseline models by 4.9% inPSDS1 on DCASE2017, 7.6% on DCASE2018, and 1.8% on UrbanSED, confirming theeffectiveness of our method in enhancing model performance.
标题:通过分析视觉刺激叙事中的话题演变和跨模式一致性来检测神经认知障碍
链接:https://arxiv.org/abs/2501.03727
备注:12 pages, 8 figures
摘要:神经认知障碍(NCD)的早期发现对于及时干预和疾病管理至关重要。语音分析提供了一种非侵入性和可扩展的筛选方法,特别是通过神经心理学评估工具中的叙事任务。传统的叙事分析往往侧重于微观结构中的局部指标,如词语使用和句法。虽然这些特征提供了对语言产生能力的洞察,但它们往往无法捕捉到全球叙事模式或微观结构。宏观结构包括连贯性、主题组织和逻辑进展,反映了对识别非传染性疾病可能至关重要的基本认知技能。为了解决这一差距,我们建议通过分析主题转移,时间动态和叙事的连贯性来研究特定的认知和语言挑战,旨在通过识别叙事障碍来揭示认知缺陷,并探索其对沟通和认知的影响。这项调查是基于CU-MARVEL兔子故事语料库,其中包括来自758名老年人的讲故事任务的录音。我们开发了两种方法:动态主题模型(DTM)为基础的时间分析,以检查随着时间的推移的主题的演变,和文本图像时间对齐网络(TITAN),以评估之间的连贯性口语叙事和视觉刺激。基于DTM的方法验证了动态主题一致性作为宏观结构度量的有效性(F1=0.61,AUC=0.78)。TITAN方法实现了最高性能(F1=0.72,AUC=0.81),超过了已建立的微观结构和宏观结构特征集。交叉比较和回归任务进一步证明了所提出的用于NCD检测的动态宏观结构建模方法的有效性。
摘要:Early detection of neurocognitive disorders (NCDs) is crucial for timelyintervention and disease management. Speech analysis offers a non-intrusive andscalable screening method, particularly through narrative tasks inneuropsychological assessment tools. Traditional narrative analysis oftenfocuses on local indicators in microstructure, such as word usage and syntax.While these features provide insights into language production abilities, theyoften fail to capture global narrative patterns, or microstructures.Macrostructures include coherence, thematic organization, and logicalprogressions, reflecting essential cognitive skills potentially critical forrecognizing NCDs. Addressing this gap, we propose to investigate specificcognitive and linguistic challenges by analyzing topical shifts, temporaldynamics, and the coherence of narratives over time, aiming to reveal cognitivedeficits by identifying narrative impairments, and exploring their impact oncommunication and cognition. The investigation is based on the CU-MARVEL RabbitStory corpus, which comprises recordings of a story-telling task from 758 olderadults. We developed two approaches: the Dynamic Topic Models (DTM)-basedtemporal analysis to examine the evolution of topics over time, and theText-Image Temporal Alignment Network (TITAN) to evaluate the coherence betweenspoken narratives and visual stimuli. DTM-based approach validated theeffectiveness of dynamic topic consistency as a macrostructural metric(F1=0.61, AUC=0.78). The TITAN approach achieved the highest performance(F1=0.72, AUC=0.81), surpassing established microstructural and macrostructuralfeature sets. Cross-comparison and regression tasks further demonstrated theeffectiveness of proposed dynamic macrostructural modeling approaches for NCDdetection.
标题:通用说话人嵌入自由目标说话人提取和个人语音活动检测
链接:https://arxiv.org/abs/2501.03612
摘要:确定“谁在什么时候说了什么”在现实世界的应用中仍然具有挑战性。在典型的场景中,说话人日记(SD)被用来解决“谁在什么时候说话”的问题,而目标说话人提取(TSE)或目标说话人自动语音识别(TSASR)技术被用来解决“谁说了什么”的问题。“虽然一些工作已经取得了可喜的成果,通过结合SD和TSE系统,SD和TSE之间的不一致仍然存在输出不一致和场景不匹配。为了解决这些限制,我们提出了一个通用的扬声器嵌入自由目标扬声器提取和个人语音活动检测(USEF-TP)模型,共同执行TSE和个人语音活动检测(PVAD)。USEF-TP利用通过交叉注意机制获得的帧级特征作为说话者相关特征,而不是像传统方法那样使用说话者嵌入。此外,多任务学习算法与一个多任务感知差分损失函数的应用,以确保在不同级别的扬声器重叠的鲁棒性能。实验结果表明,我们提出的USEF-TP模型在LibriMix和SparseLibriMix数据集上的TSE和PVAD任务中取得了优异的性能。
摘要:Determining 'who spoke what and when' remains challenging in real-worldapplications. In typical scenarios, Speaker Diarization (SD) is employed toaddress the problem of 'who spoke when,' while Target Speaker Extraction (TSE)or Target Speaker Automatic Speech Recognition (TSASR) techniques are utilizedto resolve the issue of 'who spoke what.' Although some works have achievedpromising results by combining SD and TSE systems, inconsistencies remainbetween SD and TSE regarding both output inconsistency and scenario mismatch.To address these limitations, we propose a Universal Speaker Embedding FreeTarget Speaker Extraction and Personal Voice Activity Detection (USEF-TP) modelthat jointly performs TSE and Personal Voice Activity Detection (PVAD). USEF-TPleverages frame-level features obtained through a cross-attention mechanism asspeaker-related features instead of using speaker embeddings as in traditionalapproaches. Additionally, a multi-task learning algorithm with a scenario-awaredifferentiated loss function is applied to ensure robust performance acrossvarious levels of speaker overlap. The experimental results show that ourproposed USEF-TP model achieves superior performance in TSE and PVAD tasks onthe LibriMix and SparseLibriMix datasets.
标题:开发一种用于帕金森病诊断的通用言语标记物
链接:https://arxiv.org/abs/2501.03581
摘要:帕金森氏病(PD)是一种神经退行性疾病,其特征在于运动症状,包括在早期阶段改变声音产生。早期诊断不仅对改善PD患者的生活质量至关重要,而且对增强早期神经退行性变期间潜在疾病修饰疗法的疗效也至关重要,这是当前诊断工具经常错过的窗口。在本文中,我们提出了一个更普遍的方法,通过域自适应和自监督学习的PD识别。我们展示了所提出的方法在不同语言的不同数据集上的泛化能力。我们的方法利用了HuBERT,这是一个最初为语音识别而训练的大型深度神经网络,并进一步在来自与目标群体相似的人群的未标记语音数据上训练它,即,老年人以自我监督的方式。然后对模型进行微调,使其适用于多种语言的不同数据集,包括英语,意大利语和西班牙语。对四个公开的PD数据集的评估证明了该模型的有效性,平均特异性为92.1%,平均灵敏度为91.2%。该方法在大规模人群中提供客观且一致的评估,解决了人类评估固有的变异性,并提供了一种无创、具有成本效益且易于使用的诊断选择。
摘要:Parkinson's Disease (PD) is a neurodegenerative disorder characterized bymotor symptoms, including altered voice production in the early stages. Earlydiagnosis is crucial not only to improve PD patients' quality of life but alsoto enhance the efficacy of potential disease-modifying therapies during earlyneurodegeneration, a window often missed by current diagnostic tools. In thispaper, we propose a more generalizable approach to PD recognition throughdomain adaptation and self-supervised learning. We demonstrate thegeneralization capabilities of the proposed approach across diverse datasets indifferent languages. Our approach leverages HuBERT, a large deep neural networkoriginally trained for speech recognition and further trains it on unlabeledspeech data from a population that is similar to the target group, i.e., theelderly, in a self-supervised manner. The model is then fine-tuned and adaptedfor use across different datasets in multiple languages, including English,Italian, and Spanish. Evaluations on four publicly available PD datasetsdemonstrate the model's efficacy, achieving an average specificity of 92.1% andan average sensitivity of 91.2%. This method offers objective and consistentevaluations across large populations, addressing the variability inherent inhuman assessments and providing a non-invasive, cost-effective and accessiblediagnostic option.
标题:病理性言语的深度学习:一项调查
链接:https://arxiv.org/abs/2501.03536
备注:Submitted to IEEE JSTSP Special Issue on Modelling and Processing Language and Speech in Neurodegenerative Disorders
摘要:神经退行性言语障碍口语技术的进步对于满足临床和技术需求至关重要。这篇综述论文对于推进这一领域至关重要,因为它全面回顾了病理性语音检测、自动语音识别、病理性语音清晰度增强、清晰度和严重程度评估以及病理性语音数据增强方法中的最新方法。它还强调了关键挑战,例如确保鲁棒性,隐私性和可解释性。本文最后探讨了有前途的未来方向,包括采用多模态方法以及图神经网络和大型语言模型的整合,以进一步推进神经退行性语音障碍的语音技术
摘要:Advancements in spoken language technologies for neurodegenerative speechdisorders are crucial for meeting both clinical and technological needs. Thisoverview paper is vital for advancing the field, as it presents a comprehensivereview of state-of-the-art methods in pathological speech detection, automaticspeech recognition, pathological speech intelligibility enhancement,intelligibility and severity assessment, and data augmentation approaches forpathological speech. It also high-lights key challenges, such as ensuringrobustness, privacy, and interpretability. The paper concludes by exploringpromising future directions, including the adoption of multimodal approachesand the integration of graph neural networks and large language models tofurther advance speech technology for neurodegenerative speech disorders
标题:突破尖峰:尖峰窗口解码,实现加速、精确的自动语音识别
链接:https://arxiv.org/abs/2501.03257
备注:Accepted by ICASSP 2025
摘要:近年来,端到端自动语音识别已成为工业界和学术界的主流方法。为了在特定场景中优化系统性能,加权语音状态转换器(WFST)被广泛用于集成声学和语言模型,利用其在静态图中隐式融合语言模型的能力,从而确保鲁棒的识别,同时也促进快速纠错。然而,WFST需要通过自回归逐帧搜索CTC后验概率,这大大阻碍了推理速度。在这项工作中,我们深入研究了CTC输出的尖峰特性,并进一步提出了猜想,即非空白尖峰的相邻帧携带有益于模型的语义信息。在此基础上,我们提出了尖峰窗口解码算法,该算法通过使WFST中解码的帧的数量与CTC输出中的尖峰帧的数量线性相关,从而大大提高了推理速度,同时保证了识别性能。我们的方法实现了SOTA识别精度,大大加快了解码速度,在AISHELL-1和大规模内部数据集上都得到了证明,建立了一种将CTC输出与WFST集成的开创性方法。
摘要:Recently, end-to-end automatic speech recognition has become the mainstreamapproach in both industry and academia. To optimize system performance inspecific scenarios, the Weighted Finite-State Transducer (WFST) is extensivelyused to integrate acoustic and language models, leveraging its capacity toimplicitly fuse language models within static graphs, thereby ensuring robustrecognition while also facilitating rapid error correction. However, WFSTnecessitates a frame-by-frame search of CTC posterior probabilities throughautoregression, which significantly hampers inference speed. In this work, wethoroughly investigate the spike property of CTC outputs and further proposethe conjecture that adjacent frames to non-blank spikes carry semanticinformation beneficial to the model. Building on this, we propose the SpikeWindow Decoding algorithm, which greatly improves the inference speed by makingthe number of frames decoded in WFST linearly related to the number of spikingframes in the CTC output, while guaranteeing the recognition performance. Ourmethod achieves SOTA recognition accuracy with significantly acceleratesdecoding speed, proven across both AISHELL-1 and large-scale In-House datasets,establishing a pioneering approach for integrating CTC output with WFST.
标题:检测不可检测的内容:评估当前欺骗检测方法针对无缝语音编辑的有效性
链接:https://arxiv.org/abs/2501.03805
备注:SLT 2024
摘要:神经语音编辑的进步引起了人们对它们在欺骗攻击中滥用的担忧。传统的部分编辑语音语料库主要集中在剪切和粘贴编辑,虽然保持说话人的一致性,往往会引入可检测的不连续性。最近的方法,如A\text上标{3}T和Voicebox,通过利用上下文信息来改进过渡。为了促进欺骗检测研究,我们介绍了语音填充编辑(SINE)数据集,使用Voicebox创建。我们详细介绍了重新实现Voicebox训练和数据集创建的过程。主观评价证实,使用这种新技术编辑的语音比传统的剪切和粘贴方法更具有挑战性。尽管人类的困难,实验结果表明,基于自我监督的检测器可以在不同的编辑方法的检测,定位和泛化方面取得显着的性能。数据集和相关模型将向公众开放。
摘要:Neural speech editing advancements have raised concerns about their misuse inspoofing attacks. Traditional partially edited speech corpora primarily focuson cut-and-paste edits, which, while maintaining speaker consistency, oftenintroduce detectable discontinuities. Recent methods, likeA\textsuperscript{3}T and Voicebox, improve transitions by leveragingcontextual information. To foster spoofing detection research, we introduce theSpeech INfilling Edit (SINE) dataset, created with Voicebox. We detailed theprocess of re-implementing Voicebox training and dataset creation. Subjectiveevaluations confirm that speech edited using this novel technique is morechallenging to detect than conventional cut-and-paste methods. Despite humandifficulty, experimental results demonstrate that self-supervised-baseddetectors can achieve remarkable performance in detection, localization, andgeneralization across different edit methods. The dataset and related modelswill be made publicly available.
标题:多标签跨语言自动音乐流派分类,包含句子BERT
链接:https://arxiv.org/abs/2501.03769
备注:5 pages
摘要:音乐类型是由歌曲的风格特征和艺术家观众的文化偏好形成的。使用歌词对音乐流派进行自动分类可以在诸如推荐系统、播放列表创建和库组织的若干应用中有用。我们提出了一个基于sBERT生成的多语言句子嵌入的多标签、跨语言流派分类系统。使用双语葡萄牙语-英语数据集与八个重叠的流派,我们证明了系统的能力,训练在一种语言的歌词和预测流派在另一个。我们的方法优于翻译歌词和使用词袋表示的基线方法,将体裁平均F1分数从0.35提高到0.69。分类器使用一对多架构,使其能够为单个歌词分配多个流派标签。实验结果表明,数据集集中化显着提高了跨语言性能。这种方法提供了一个可扩展的解决方案,跨代表性不足的语言和文化领域的流派分类,推进音乐信息检索系统的能力。
摘要:Music genres are shaped by both the stylistic features of songs and thecultural preferences of artists' audiences. Automatic classification of musicgenres using lyrics can be useful in several applications such asrecommendation systems, playlist creation, and library organization. We presenta multi-label, cross-lingual genre classification system based on multilingualsentence embeddings generated by sBERT. Using a bilingual Portuguese-Englishdataset with eight overlapping genres, we demonstrate the system's ability totrain on lyrics in one language and predict genres in another. Our approachoutperforms the baseline approach of translating lyrics and using abag-of-words representation, improving the genrewise average F1-Score from 0.35to 0.69. The classifier uses a one-vs-all architecture, enabling it to assignmultiple genre labels to a single lyric. Experimental results reveal thatdataset centralization notably improves cross-lingual performance. Thisapproach offers a scalable solution for genre classification acrossunderrepresented languages and cultural domains, advancing the capabilities ofmusic information retrieval systems.
标题:用于从神经活动重建高保真语音的NeuIncept解码器
链接:https://arxiv.org/abs/2501.03757
摘要:本文介绍了一种新的算法设计的语音合成从神经活动记录使用侵入性脑电图(EEG)技术。该系统提供了一个有前途的通信解决方案,为个人严重的语音障碍。我们的方法的核心是集成的时间-频率功能的高伽玛波段计算从EEG记录与先进的NeuroIncept解码器架构。这种神经网络架构结合了卷积神经网络(CNN)和门控递归单元(GRU),可以从神经模式中重建音频频谱图。我们的模型显示了预测和实际频谱图之间的稳健平均相关系数,尽管受试者间的差异表明参与者之间存在不同的神经处理机制。总的来说,我们的研究强调了神经解码技术在恢复言语障碍患者沟通能力方面的潜力,并为脑机接口技术的未来发展铺平了道路。
摘要:This paper introduces a novel algorithm designed for speech synthesis fromneural activity recordings obtained using invasive electroencephalography (EEG)techniques. The proposed system offers a promising communication solution forindividuals with severe speech impairments. Central to our approach is theintegration of time-frequency features in the high-gamma band computed from EEGrecordings with an advanced NeuroIncept Decoder architecture. This neuralnetwork architecture combines Convolutional Neural Networks (CNNs) and GatedRecurrent Units (GRUs) to reconstruct audio spectrograms from neural patterns.Our model demonstrates robust mean correlation coefficients between predictedand actual spectrograms, though inter-subject variability indicates distinctneural processing mechanisms among participants. Overall, our study highlightsthe potential of neural decoding techniques to restore communicative abilitiesin individuals with speech disorders and paves the way for future advancementsin brain-computer interface technologies.
标题:Guitar-TECHS:一个电子吉他数据集,涵盖技术、音乐摘录、和弦和音阶,使用多种硬件阵列
链接:https://arxiv.org/abs/2501.03720
备注:None
摘要:与吉他相关的机器听力研究涉及音色转换、性能生成和自动转录等任务。然而,由于声学多样性和音乐内容不足,小数据集往往限制了模型的鲁棒性。为了解决这些问题,我们引入了Guitar-TECHS,这是一个全面的数据集,包含各种吉他技术,音乐摘录,和弦和音阶。这些元素由不同的音乐家在不同的录音环境中演奏。Guitar-TECHS集成了两个立体声麦克风的录音:一个位于表演者头部的自我中心麦克风和一个位于表演者前方的偏心麦克风。它还包括直接输入录音和麦克风放大器输出,提供广泛的音频输入和录音质量。所有信号和标签都正确同步。它的多视角和多模态内容使Guitar-TECHS成为推进数据驱动吉他研究的宝贵资源,并开发强大的吉他聆听算法。我们提供了经验数据来证明数据集在训练吉他谱转录的鲁棒模型方面的有效性。
摘要:Guitar-related machine listening research involves tasks like timbretransfer, performance generation, and automatic transcription. However, smalldatasets often limit model robustness due to insufficient acoustic diversityand musical content. To address these issues, we introduce Guitar-TECHS, acomprehensive dataset featuring a variety of guitar techniques, musicalexcerpts, chords, and scales. These elements are performed by diverse musiciansacross various recording settings. Guitar-TECHS incorporates recordings fromtwo stereo microphones: an egocentric microphone positioned on the performer'shead and an exocentric microphone placed in front of the performer. It alsoincludes direct input recordings and microphoned amplifier outputs, offering awide spectrum of audio inputs and recording qualities. All signals and MIDIlabels are properly synchronized. Its multi-perspective and multi-modal contentmakes Guitar-TECHS a valuable resource for advancing data-driven guitarresearch, and to develop robust guitar listening algorithms. We provideempirical data to demonstrate the dataset's effectiveness in training robustmodels for Guitar Tablature Transcription.
标题:无监督语音分割:使用语音语言模型的通用方法
链接:https://arxiv.org/abs/2501.03711
摘要:在本文中,我们介绍了一种无监督的语音分割方法,它建立在以前研究的方法,例如,说话人日记,同时适用于一组包含的声学语义区别,铺平了道路,走向一般的无监督语音分割方法。与传统的语音和音频分割不同,传统的语音和音频分割主要关注输入信号中的频谱变化,例如,电话分割,我们的方法试图将说出的话语分割成具有不同声学语义风格的块,专注于不能很好地翻译成文本的声学语义信息,例如,情绪或说话者。虽然大多数语音分割任务只处理一种风格的变化,例如,情感日记化,我们的方法试图处理多种声学语义风格的变化。利用语音语言模型(SLM)的最新进展,我们提出了一个简单的无监督的方法来分割一个给定的语音话语。我们经验证明了所提出的方法的有效性,考虑几个设置。结果表明,该方法是优于评估基线的边界检测,段纯度,和过分割。代码可在https://github.com/avishaiElmakies/unsupervised_speech_segmentation_using_slm上获得。
摘要:In this paper, we introduce an unsupervised approach for Speech Segmentation,which builds on previously researched approaches, e.g., Speaker Diarization,while being applicable to an inclusive set of acoustic-semantic distinctions,paving a path towards a general Unsupervised Speech Segmentation approach.Unlike traditional speech and audio segmentation, which mainly focuses onspectral changes in the input signal, e.g., phone segmentation, our approachtries to segment the spoken utterance into chunks with differingacoustic-semantic styles, focusing on acoustic-semantic information that doesnot translate well into text, e.g., emotion or speaker. While most SpeechSegmentation tasks only handle one style change, e.g., emotion diarization, ourapproach tries to handle multiple acoustic-semantic style changes. Leveragingrecent advances in Speech Language Models (SLMs), we propose a simpleunsupervised method to segment a given speech utterance. We empiricallydemonstrate the effectiveness of the proposed approach by considering severalsetups. Results suggest that the proposed method is superior to the evaluatedbaselines on boundary detection, segment purity, and over-segmentation. Code isavailable athttps://github.com/avishaiElmakies/unsupervised_speech_segmentation_using_slm.
标题:MAJL:用于音乐源分离和音调估计的模型不可知联合学习框架
链接:https://arxiv.org/abs/2501.03689
摘要:None
摘要:Music source separation and pitch estimation are two vital tasks in musicinformation retrieval. Typically, the input of pitch estimation is obtainedfrom the output of music source separation. Therefore, existing methods havetried to perform these two tasks simultaneously, so as to leverage the mutuallybeneficial relationship between both tasks. However, these methods still facetwo critical challenges that limit the improvement of both tasks: the lack oflabeled data and joint learning optimization. To address these challenges, wepropose a Model-Agnostic Joint Learning (MAJL) framework for both tasks. MAJLis a generic framework and can use variant models for each task. It includes atwo-stage training method and a dynamic weighting method named Dynamic Weightson Hard Samples (DWHS), which addresses the lack of labeled data and jointlearning optimization, respectively. Experimental results on public musicdatasets show that MAJL outperforms state-of-the-art methods on both tasks,with significant improvements of 0.92 in Signal-to-Distortion Ratio (SDR) formusic source separation and 2.71% in Raw Pitch Accuracy (RPA) for pitchestimation. Furthermore, comprehensive studies not only validate theeffectiveness of each component of MAJL, but also indicate the great generalityof MAJL in adapting to different model architectures.
标题:有效且高效的语音基础模型混合精度量化
链接:https://arxiv.org/abs/2501.03643
备注:To appear at IEEE ICASSP 2025
摘要:本文提出了一种新的混合精度量化方法的语音基础模型,紧密结合的混合精度学习和量化模型参数估计到一个单一的模型压缩阶段。在LibriSpeech数据集上使用微调的wav2vec2.0-base和HuBERT-large模型进行的实验表明,所得到的混合精度量化模型将无损压缩率提高了1.7倍和1.9倍,分别超过了在单独和不连贯的阶段中执行精度学习和模型参数量化的均匀精度和两阶段混合精度量化基线,而在32位全精度模型上不会引起统计上的字错误率(WER)增加。wav2vec2.0-base和HuBERT-large模型的系统压缩时间比两阶段混合精度基线减少了1.9倍和1.5倍,同时两者都产生了较低的WER。性能最佳的3.5位混合精度量化HuBERT-large模型产生的无损压缩比为32位全精度系统的8.6倍。
摘要:This paper presents a novel mixed-precision quantization approach for speechfoundation models that tightly integrates mixed-precision learning andquantized model parameter estimation into one single model compression stage.Experiments conducted on LibriSpeech dataset with fine-tuned wav2vec2.0-baseand HuBERT-large models suggest the resulting mixed-precision quantized modelsincreased the lossless compression ratio by factors up to 1.7x and 1.9x overthe respective uniform-precision and two-stage mixed-precision quantizedbaselines that perform precision learning and model parameters quantization inseparate and disjointed stages, while incurring no statistically word errorrate (WER) increase over the 32-bit full-precision models. The systemcompression time of wav2vec2.0-base and HuBERT-large models is reduced by up to1.9 and 1.5 times over the two-stage mixed-precision baselines, while bothproduce lower WERs. The best-performing 3.5-bit mixed-precision quantizedHuBERT-large model produces a lossless compression ratio of 8.6x over the32-bit full-precision system.
标题:AADNet:基于线索掩蔽范式探索脑电时空信息以快速准确地定位和音色检测听觉注意力
链接:https://arxiv.org/abs/2501.03571
摘要:在噪声环境中,从脑电信号中解码出的听觉注意力可以推断出使用者正在关注哪个源。解码算法和实验范式设计对于技术在实际应用中的发展至关重要。为了模拟真实场景,本研究提出了一个线索掩蔽听觉注意范式,以避免实验前的信息泄露。为了在低延迟的情况下获得高解码精度,提出了一种端到端深度学习模型AADNet,以利用来自EEG信号的短时间窗口的时空信息。结果表明,在0.5秒的EEG窗口下,AADNet对听觉方向注意(OA)和音色注意(TA)的平均解码准确率分别为93.46%和91.09%。它的性能明显优于以前的五种方法,并且不需要原始音频源的知识。这一工作证明了从脑电信号中快速准确地检测听觉注意的方向和音色是可能的。研究结果为多属性听觉注意力的实时解码提供了理论依据,为神经导航助听器和其他辅助听力设备的应用提供了参考。
摘要:Auditory attention decoding from electroencephalogram (EEG) could infer towhich source the user is attending in noisy environments. Decoding algorithmsand experimental paradigm designs are crucial for the development of technologyin practical applications. To simulate real-world scenarios, this studyproposed a cue-masked auditory attention paradigm to avoid information leakagebefore the experiment. To obtain high decoding accuracy with low latency, anend-to-end deep learning model, AADNet, was proposed to exploit thespatiotemporal information from the short time window of EEG signals. Theresults showed that with a 0.5-second EEG window, AADNet achieved an averageaccuracy of 93.46% and 91.09% in decoding auditory orientation attention (OA)and timbre attention (TA), respectively. It significantly outperformed fiveprevious methods and did not need the knowledge of the original audio source.This work demonstrated that it was possible to detect the orientation andtimbre of auditory attention from EEG signals fast and accurately. The resultsare promising for the real-time multi-property auditory attention decoding,facilitating the application of the neuro-steered hearing aids and otherassistive listening devices.
标题:用于口语关键词定位的声道长度扭曲功能
链接:https://arxiv.org/abs/2501.03523
摘要:在本文中,我们提出了几种方法,将声道长度(VAN)扭曲的功能口语关键字定位(KWS)。第一种方法是独立于VTL的KWS,涉及训练单个深度神经网络(DNN),该网络利用具有各种扭曲因子的特征。在训练期间,每个时期随机选择特定的VTL功能,从而允许探索VTL变体。在测试过程中,测试话语具有不同扭曲因子的VTL特征将根据DNN进行评分,并以相等的权重进行组合。在第二种方法中,针对DNN对测试话语的常规特征(没有扭曲)进行评分。第三种方法,VTL连接KWS,连接扭曲的特征以形成KWS的高维特征。在英文Google Command数据集上进行的评估表明,所提出的方法提高了KWS的准确性。
摘要:In this paper, we propose several methods that incorporate vocal tract length(VTL) warped features for spoken keyword spotting (KWS). The first method,VTL-independent KWS, involves training a single deep neural network (DNN) thatutilizes VTL features with various warping factors. During training, a specificVTL feature is randomly selected per epoch, allowing the exploration of VTLvariations. During testing, the VTL features with different warping factors ofa test utterance are scored against the DNN and combined with equal weight. Inthe second method scores the conventional features of a test utterance (withoutVTL warping) against the DNN. The third method, VTL-concatenation KWS,concatenates VTL warped features to form high-dimensional features for KWS.Evaluations carried out on the English Google Command dataset demonstrate thatthe proposed methods improve the accuracy of KWS.
标题:LHGNN:用于音频分类和标记的局部-高位图神经网络
链接:https://arxiv.org/abs/2501.03464
摘要:Transformers在音频处理任务中建立了新的基准,利用自我注意机制来捕获音频数据中的复杂模式和依赖关系。然而,它们对成对交互的关注限制了它们处理识别不同音频对象所必需的高阶关系的能力。为了解决这一限制,这项工作引入了局部高阶图神经网络(LHGNN),这是一种基于图的模型,通过将局部邻域信息与来自模糊C均值聚类的高阶数据相结合来增强特征理解,从而捕获更广泛的音频关系。在三个公开的音频数据集上对该模型进行的评估表明,它在所有基准测试中的性能都优于基于transformer的模型,同时使用的参数要少得多。此外,LHGNN在缺乏ImageNet预训练的场景中表现出明显的优势,在无法获得大量预训练数据的环境中建立了其有效性和效率。
摘要:Transformers have set new benchmarks in audio processing tasks, leveragingself-attention mechanisms to capture complex patterns and dependencies withinaudio data. However, their focus on pairwise interactions limits their abilityto process the higher-order relations essential for identifying distinct audioobjects. To address this limitation, this work introduces the Local- HigherOrder Graph Neural Network (LHGNN), a graph based model that enhances featureunderstanding by integrating local neighbourhood information with higher-orderdata from Fuzzy C-Means clusters, thereby capturing a broader spectrum of audiorelationships. Evaluation of the model on three publicly available audiodatasets shows that it outperforms Transformer-based models across allbenchmarks while operating with substantially fewer parameters. Moreover, LHGNNdemonstrates a distinct advantage in scenarios lacking ImageNet pretraining,establishing its effectiveness and efficiency in environments where extensivepretraining data is unavailable.
标题:通过MEG驱动编码模型架起听觉感知和语言理解的桥梁
链接:https://arxiv.org/abs/2501.03246
备注:10 pages, 4 figures, Accepted at ICLR2024 Workshop TS4H
摘要:理解听觉和语言处理背后的神经机制是推进认知神经科学的关键。在这项研究中,我们使用脑磁图(MEG)数据来分析大脑对口语刺激的反应。我们开发了两种不同的编码模型:音频到MEG编码器,它使用时频分解(TFD)和wav 2 vec 2潜在空间表示,以及文本到MEG编码器,它利用CLIP和GPT-2嵌入。这两种模型都成功地预测了神经活动,证明了估计和观察到的MEG信号之间的显着相关性。然而,文本到MEG模型优于基于音频的模型,实现了更高的皮尔逊相关(PC)分数。在空间上,我们确定,基于语义的嵌入(TFD和wav 2 vec 2)主要激活侧颞区,这是负责初级听觉处理和听觉信号的整合。相比之下,文本嵌入(CLIP和GPT-2)主要涉及额叶皮层,特别是布罗卡区,该区域与高阶语言处理相关,包括语义整合和语言产生,特别是在8-30 Hz频率范围内。这些区域的强烈参与表明听觉刺激是通过更直接的感觉通路处理的,而语言信息是通过整合意义和认知控制的网络编码的。我们的研究结果揭示了听觉和语言信息处理的不同神经通路,在额叶区域的文本表征具有更高的编码准确性。这些见解完善了我们对大脑处理听觉和文本信息的功能结构的理解,为复杂语言刺激的神经反应建模提供了定量进步。
摘要:Understanding the neural mechanisms behind auditory and linguistic processingis key to advancing cognitive neuroscience. In this study, we useMagnetoencephalography (MEG) data to analyze brain responses to spoken languagestimuli. We develop two distinct encoding models: an audio-to-MEG encoder,which uses time-frequency decompositions (TFD) and wav2vec2 latent spacerepresentations, and a text-to-MEG encoder, which leverages CLIP and GPT-2embeddings. Both models successfully predict neural activity, demonstratingsignificant correlations between estimated and observed MEG signals. However,the text-to-MEG model outperforms the audio-based model, achieving higherPearson Correlation (PC) score. Spatially, we identify that auditory-basedembeddings (TFD and wav2vec2) predominantly activate lateral temporal regions,which are responsible for primary auditory processing and the integration ofauditory signals. In contrast, textual embeddings (CLIP and GPT-2) primarilyengage the frontal cortex, particularly Broca's area, which is associated withhigher-order language processing, including semantic integration and languageproduction, especially in the 8-30 Hz frequency range. The strong involvementof these regions suggests that auditory stimuli are processed through moredirect sensory pathways, while linguistic information is encoded via networksthat integrate meaning and cognitive control. Our results reveal distinctneural pathways for auditory and linguistic information processing, with higherencoding accuracy for text representations in the frontal regions. Theseinsights refine our understanding of the brain's functional architecture inprocessing auditory and textual information, offering quantitative advancementsin the modelling of neural responses to complex language stimuli.
