本文经arXiv每日学术速递授权转载
链接:https://arxiv.org/abs/2501.04630
备注:Accepted at Artificial Intelligence for Music Workshop at AAAI 2025 (this https URL)
摘要:符号音乐分析任务通常由最初为自然语言处理开发的模型执行,例如Transformers。这种模型需要将输入数据表示为序列,这是通过标记化过程实现的。符号音乐的符号化策略通常依赖于绝对的音高值来表示音高信息。然而,音乐研究在很大程度上促进了更高层次的表现,如旋律轮廓和和声关系,其中音高间隔比绝对音高更具表现力。在这项工作中,我们介绍了一个通用的框架,用于构建基于区间的标记。通过对三个音乐分析任务进行评估,我们表明这种基于间隔的标记化提高了模型性能,并促进了其可解释性。
摘要:Symbolic music analysis tasks are often performed by models originallydeveloped for Natural Language Processing, such as Transformers. Such modelsrequire the input data to be represented as sequences, which is achievedthrough a process of tokenization. Tokenization strategies for symbolic musicoften rely on absolute MIDI values to represent pitch information. However,music research largely promotes the benefit of higher-level representationssuch as melodic contour and harmonic relations for which pitch intervals turnout to be more expressive than absolute pitches. In this work, we introduce ageneral framework for building interval-based tokenizations. By evaluatingthese tokenizations on three music analysis tasks, we show that suchinterval-based tokenizations improve model performances and facilitate theirexplainability.
标题:时间同步ASB模型的端到端训练中的正确标签上下文
链接:https://arxiv.org/abs/2501.04521
备注:Accepted for presentation at ICASSP 2025
摘要:当前的时间同步序列到序列自动语音识别(ASR)模型是通过使用对所有对齐求和的序列级交叉熵来训练的。由于区别性公式,将正确的标签上下文并入训练标准的梯度会导致归一化问题,并且在数学上没有明确定义。经典的混合神经网络隐马尔可夫模型(NN-HMM)具有其固有的生成公式,能够对正确的标签上下文进行调节。然而,由于HMM状态绑定,正确标签上下文的身份从未显式建模。在这项工作中,我们提出了一个因子损失与辅助左,右标签上下文,总结了所有的路线。我们表明,当训练数据资源有限时,包含正确的标签上下文特别有益。此外,我们还表明,可以完全依赖全和标准来构建分解混合隐马尔可夫模型系统。实验在Switchboard 300 h和LibriSpeech 960 h上进行。
摘要:Current time-synchronous sequence-to-sequence automatic speech recognition(ASR) models are trained by using sequence level cross-entropy that sums overall alignments. Due to the discriminative formulation, incorporating the rightlabel context into the training criterion's gradient causes normalizationproblems and is not mathematically well-defined. The classic hybrid neuralnetwork hidden Markov model (NN-HMM) with its inherent generative formulationenables conditioning on the right label context. However, due to the HMMstate-tying the identity of the right label context is never modeledexplicitly. In this work, we propose a factored loss with auxiliary left andright label contexts that sums over all alignments. We show that the inclusionof the right label context is particularly beneficial when training dataresources are limited. Moreover, we also show that it is possible to build afactored hybrid HMM system by relying exclusively on the full-sum criterion.Experiments were conducted on Switchboard 300h and LibriSpeech 960h.
标题:语音纯度引导的离散令牌用于发音障碍语音识别
链接:https://arxiv.org/abs/2501.04379
备注:ICASSP 2025
摘要:提取的离散标记提供了有效的和域自适应的语音特征。他们的应用程序,表现出清晰度不精确和大的不匹配对正常的声音混乱的讲话仍然是未知的。为了改善他们的语音歧视,削弱在无监督K-均值或矢量量化的连续特征,本文提出了新的音素纯度指导(PPG)离散令牌构音障碍语音识别。语音标签监督是用来正规化的最大似然和重建误差成本标准的K-均值和VAE-VQ为基础的离散令牌提取。在UASpeech语料库上进行的实验表明,从HuBERT中提取的PPG离散令牌特征始终优于混合TDNN和端到端(E2 E)Conformer系统,该系统使用基于非PPG的K-means或VAE-VQ令牌跨越不同码本大小,统计上显著的单词错误率(WER)降低高达0.99%和1.77%的绝对值在16名构音障碍者的UASpeech测试集上,两组间的差异分别为3.21%和4.82%。通过组合使用不同令牌特征的系统获得了最低WER 23.25%。手机纯度指标也得到了持续改善。T-SNE可视化进一步表明,在引入音素纯度指导后,K-means/VAE-VQ聚类之间产生了更清晰的决策边界。
摘要:Discrete tokens extracted provide efficient and domain adaptable speechfeatures. Their application to disordered speech that exhibits articulationimprecision and large mismatch against normal voice remains unexplored. Toimprove their phonetic discrimination that is weakened during unsupervisedK-means or vector quantization of continuous features, this paper proposesnovel phone-purity guided (PPG) discrete tokens for dysarthric speechrecognition. Phonetic label supervision is used to regularize maximumlikelihood and reconstruction error costs used in standard K-means and VAE-VQbased discrete token extraction. Experiments conducted on the UASpeech corpussuggest that the proposed PPG discrete token features extracted from HuBERTconsistently outperform hybrid TDNN and End-to-End (E2E) Conformer systemsusing non-PPG based K-means or VAE-VQ tokens across varying codebook sizes bystatistically significant word error rate (WER) reductions up to 0.99\% and1.77\% absolute (3.21\% and 4.82\% relative) respectively on the UASpeech testset of 16 dysarthric speakers. The lowest WER of 23.25\% was obtained bycombining systems using different token features. Consistent improvements onthe phone purity metric were also achieved. T-SNE visualization furtherdemonstrates sharper decision boundaries were produced between K-means/VAE-VQclusters after introducing phone-purity guidance.
标题:MAD-紫外线:第一次通过超声发声挑战检测INTERSPEECT小鼠自闭症
链接:https://arxiv.org/abs/2501.04292
备注:5 pages, 1 figure and 2 tables. For MAD-UV Challenge 2025
摘要:小鼠自闭症检测通过超声发声(MAD-UV)挑战介绍了第一个INTERSPEECH挑战,重点是通过小鼠的发声检测自闭症谱系障碍(ASD)。参与者的任务是开发模型,根据高采样率的记录将小鼠自动分类为野生型或ASD模型。我们的基线系统采用了一个简单的基于CNN的分类,使用三种不同的频谱图特征。结果表明,自动ASD检测的可行性,与所考虑的听觉范围的功能实现最佳性能(UAR为0.600的段级和0.625的主题级分类)。这一挑战将语音技术和生物医学研究联系起来,为我们通过机器学习方法提高对ASD模型的理解提供了机会。这些发现为发声分析提出了有希望的方向,并突出了听觉和超声发声在ASD检测中的潜在价值。
摘要:The Mice Autism Detection via Ultrasound Vocalization (MAD-UV) Challengeintroduces the first INTERSPEECH challenge focused on detecting autism spectrumdisorder (ASD) in mice through their vocalizations. Participants are taskedwith developing models to automatically classify mice as either wild-type orASD models based on recordings with a high sampling rate. Our baseline systememploys a simple CNN-based classification using three different spectrogramfeatures. Results demonstrate the feasibility of automated ASD detection, withthe considered audible-range features achieving the best performance (UAR of0.600 for segment-level and 0.625 for subject-level classification). Thischallenge bridges speech technology and biomedical research, offeringopportunities to advance our understanding of ASD models through machinelearning approaches. The findings suggest promising directions for vocalizationanalysis and highlight the potential value of audible and ultrasoundvocalizations in ASD detection.
标题:DrawSpeech:使用韵律草图作为控制条件的表达性语音合成
链接:https://arxiv.org/abs/2501.04256
备注:Accepted by ICASSP 2025
摘要:控制文语转换系统合成具有用户期望韵律特征的语音已引起人们的广泛关注。为了实现语音合成的可控性,目前的研究主要集中在两个方面:(1)使用参考语音作为韵律提示来指导语音合成;(2)使用自然语言描述来控制语音合成过程。然而,找到准确包含用户想要合成的韵律的参考语音需要花费大量的精力。文语转换系统中基于描述的引导只能确定整体韵律,难以实现对合成语音的细粒度韵律控制。在本文中,我们提出了DrawSpeech,一个草图条件扩散模型,能够生成语音的基础上由用户绘制的任何韵律草图。具体而言,韵律草图被馈送到DrawSpeech以提供预期韵律趋势的粗略指示。DrawSpeech然后根据粗略的草图恢复详细的音高和能量轮廓,并合成所需的语音。实验结果表明,DrawSpeech可以生成具有多种韵律的语音,并且可以以用户友好的方式精确地控制细粒度的韵律。我们的实现和音频样本是公开的。
摘要:Controlling text-to-speech (TTS) systems to synthesize speech with theprosodic characteristics expected by users has attracted much attention. Toachieve controllability, current studies focus on two main directions: (1)using reference speech as prosody prompt to guide speech synthesis, and (2)using natural language descriptions to control the generation process. However,finding reference speech that exactly contains the prosody that users want tosynthesize takes a lot of effort. Description-based guidance in TTS systems canonly determine the overall prosody, which has difficulty in achievingfine-grained prosody control over the synthesized speech. In this paper, wepropose DrawSpeech, a sketch-conditioned diffusion model capable of generatingspeech based on any prosody sketches drawn by users. Specifically, the prosodysketches are fed to DrawSpeech to provide a rough indication of the expectedprosody trends. DrawSpeech then recovers the detailed pitch and energy contoursbased on the coarse sketches and synthesizes the desired speech. Experimentalresults show that DrawSpeech can generate speech with a wide variety of prosodyand can precisely control the fine-grained prosody in a user-friendly manner.Our implementation and audio samples are publicly available.
标题:再次聆听和观看:视听语音识别的生成性错误纠正
链接:https://arxiv.org/abs/2501.04038
摘要:与传统的自动语音识别(ASR)不同,视听语音识别(AVSR)同时获取音频和视频信号来推断转录。最近的研究表明,大语言模型(LLM)可以有效地用于生成错误校正(GER)在ASR中通过预测最好的转录从ASR生成的N-最好的假设。然而,这些LLM缺乏同时理解音频和视频的能力,使得GER方法在AVSR中应用具有挑战性。在这项工作中,我们提出了一个新的GER范式AVSR,称为AVGER,遵循的概念“听和看了一遍”。具体来说,我们首先使用强大的AVSR系统读取音频和视频信号,以获得N-Best假设,然后使用基于Q-former的多模态同步编码器再次读取音频和视频信息,并将其分别转换为LLM可以理解的音频和视频压缩表示。之后,视听压缩表示和N-Best假设一起构成跨模态提示,以指导LLM产生最佳转录。此外,本文还提出了一种多级一致性约束训练准则,包括逻辑级、话语级和表示级,以提高校正精度,同时增强音视频压缩表示的可解释性。在LRS 3数据集上的实验结果表明,该方法优于当前主流的AVSR系统。建议的AVGER可以减少字错误率(WER)的24%相比,他们。代码和模型可以在https://github.com/CircleRedRain/AVGER上找到。
摘要:Unlike traditional Automatic Speech Recognition (ASR), Audio-Visual SpeechRecognition (AVSR) takes audio and visual signals simultaneously to infer thetranscription. Recent studies have shown that Large Language Models (LLMs) canbe effectively used for Generative Error Correction (GER) in ASR by predictingthe best transcription from ASR-generated N-best hypotheses. However, theseLLMs lack the ability to simultaneously understand audio and visual, making theGER approach challenging to apply in AVSR. In this work, we propose a novel GERparadigm for AVSR, termed AVGER, that follows the concept of ``listening andseeing again''. Specifically, we first use the powerful AVSR system to read theaudio and visual signals to get the N-Best hypotheses, and then use theQ-former-based Multimodal Synchronous Encoder to read the audio and visualinformation again and convert them into an audio and video compressionrepresentation respectively that can be understood by LLM. Afterward, theaudio-visual compression representation and the N-Best hypothesis togetherconstitute a Cross-modal Prompt to guide the LLM in producing the besttranscription. In addition, we also proposed a Multi-Level ConsistencyConstraint training criterion, including logits-level, utterance-level andrepresentations-level, to improve the correction accuracy while enhancing theinterpretability of audio and visual compression representations. Theexperimental results on the LRS3 dataset show that our method outperformscurrent mainstream AVSR systems. The proposed AVGER can reduce the Word ErrorRate (WER) by 24% compared to them. Code and models can be found at:https://github.com/CircleRedRain/AVGER.
标题:FleSpeech:具有各种预设的可靠可控语音生成
链接:https://arxiv.org/abs/2501.04644
备注:14 pages, 3 figures
摘要:可控的语音生成方法通常依赖于单个或固定的提示,这阻碍了创造性和灵活性。这些限制使得在某些场景下很难满足特定的用户需求,例如在保持所选扬声器音色的同时调整风格,或者选择风格并生成与角色视觉外观相匹配的声音。为了克服这些挑战,我们提出了\textit{FleSpeech},一种新的多阶段语音生成框架,通过集成各种形式的控制,可以更灵活地操纵语音属性。FleSpeech采用多模式提示编码器,将不同的文本、音频和视觉提示处理并统一到一个有凝聚力的表示中。这种方法增强了语音合成的适应性,并支持对生成的语音进行创造性和精确的控制。此外,我们开发了一个多模态数据集的数据收集管道,以促进该领域的进一步研究和应用。综合的主观和客观实验证明了FleSpeech的有效性。音频样本可在https://kkksuper.github.io/FleSpeech/上获得
摘要:Controllable speech generation methods typically rely on single or fixedprompts, hindering creativity and flexibility. These limitations make itdifficult to meet specific user needs in certain scenarios, such as adjustingthe style while preserving a selected speaker's timbre, or choosing a style andgenerating a voice that matches a character's visual appearance. To overcomethese challenges, we propose \textit{FleSpeech}, a novel multi-stage speechgeneration framework that allows for more flexible manipulation of speechattributes by integrating various forms of control. FleSpeech employs amultimodal prompt encoder that processes and unifies different text, audio, andvisual prompts into a cohesive representation. This approach enhances theadaptability of speech synthesis and supports creative and precise control overthe generated speech. Additionally, we develop a data collection pipeline formultimodal datasets to facilitate further research and applications in thisfield. Comprehensive subjective and objective experiments demonstrate theeffectiveness of FleSpeech. Audio samples are available athttps://kkksuper.github.io/FleSpeech/
标题:CLARVC:采用解开潜在扩散模型和对抗训练的Zero-Shot式语音转换
链接:https://arxiv.org/abs/2501.04416
备注:5 pages, 3 figures, accepted by ICASSP 2025
摘要:风格语音转换的目的是在保持原说话人身份的前提下,将源语音的说话风格转换为所需的说话风格。然而,以前的风格语音转换方法主要集中在定义明确的领域,如情感方面,限制了它们的实际应用。在这项研究中,我们提出了一种新的Zero-shot风格的语音转换方法,它利用语音编解码器和语音提示机制的潜在扩散模型,以促进在上下文中学习说话风格转换。为了理清说话风格和说话人音色之间的关系,本文引入信息瓶颈过滤源语音中的说话风格,并采用不确定性建模自适应实例归一化(UMAdaIN)对风格提示中的说话人音色进行扰动。此外,我们提出了一种新的对抗性训练策略,以增强上下文学习和提高风格相似性。在44,000小时的语音数据上进行的实验表明,在zero-shot场景中,CNOVC在生成具有不同说话风格的语音方面具有优越的性能。
摘要:Style voice conversion aims to transform the speaking style of source speechinto a desired style while keeping the original speaker's identity. However,previous style voice conversion approaches primarily focus on well-defineddomains such as emotional aspects, limiting their practical applications. Inthis study, we present ZSVC, a novel Zero-shot Style Voice Conversion approachthat utilizes a speech codec and a latent diffusion model with speech promptingmechanism to facilitate in-context learning for speaking style conversion. Todisentangle speaking style and speaker timbre, we introduce informationbottleneck to filter speaking style in the source speech and employ UncertaintyModeling Adaptive Instance Normalization (UMAdaIN) to perturb the speakertimbre in the style prompt. Moreover, we propose a novel adversarial trainingstrategy to enhance in-context learning and improve style similarity.Experiments conducted on 44,000 hours of speech data demonstrate the superiorperformance of ZSVC in generating speech with diverse speaking styles inzero-shot scenarios.
标题:使用Transformers和基于VAE的数据增强解码脑电语音感知
链接:https://arxiv.org/abs/2501.04359
备注:19 pages, 15 figures, 2 tables
摘要:从非侵入性脑信号(如脑电图(EEG))解码语音有可能推动脑机接口(BCI)的发展,并应用于无声通信和语音障碍患者的辅助技术。然而,基于EEG的语音解码面临着重大挑战,例如噪声数据,有限的数据集,以及在语音感知等复杂任务上的性能差。本研究试图通过采用变分自动编码器(VAE)进行EEG数据增强来解决这些挑战,以提高数据质量,并将最先进的(SOTA)序列到序列深度学习架构(最初在肌电图(EMG)任务中取得成功)应用于基于EEG的语音解码。此外,我们适应这个架构的词分类任务。使用Brennan数据集,其中包含受试者听叙述语音的EEG记录,我们对数据进行预处理,并评估EEG到单词/句子任务的分类和序列到序列模型。我们的实验表明,VAE有潜力重建人工EEG数据增强。与此同时,我们的序列到序列模型在生成句子方面比我们的分类模型表现得更有希望,尽管两者都仍然是具有挑战性的任务。这些发现为未来的脑电语音感知解码研究奠定了基础,并可能扩展到语音产生任务,例如无声或想象的语音。
摘要:Decoding speech from non-invasive brain signals, such aselectroencephalography (EEG), has the potential to advance brain-computerinterfaces (BCIs), with applications in silent communication and assistivetechnologies for individuals with speech impairments. However, EEG-based speechdecoding faces major challenges, such as noisy data, limited datasets, and poorperformance on complex tasks like speech perception. This study attempts toaddress these challenges by employing variational autoencoders (VAEs) for EEGdata augmentation to improve data quality and applying a state-of-the-art(SOTA) sequence-to-sequence deep learning architecture, originally successfulin electromyography (EMG) tasks, to EEG-based speech decoding. Additionally, weadapt this architecture for word classification tasks. Using the Brennandataset, which contains EEG recordings of subjects listening to narratedspeech, we preprocess the data and evaluate both classification andsequence-to-sequence models for EEG-to-words/sentences tasks. Our experimentsshow that VAEs have the potential to reconstruct artificial EEG data foraugmentation. Meanwhile, our sequence-to-sequence model achieves more promisingperformance in generating sentences compared to our classification model,though both remain challenging tasks. These findings lay the groundwork forfuture research on EEG speech perception decoding, with possible extensions tospeech production tasks such as silent or imagined speech.
标题:基于DNN的音频处理闭环系统中的无伪影音质
链接:https://arxiv.org/abs/2501.04116
摘要:深度神经网络(DNN)的最新进展显着改善了各种音频处理应用,包括语音增强,合成和助听器算法。基于DNN的闭环系统由于其强大的性能和适应不同条件的能力而在这些应用中受到欢迎。尽管它们的有效性,当前基于DNN的闭环系统经常遭受由次优采样方法引入的伪影引起的声音质量下降。为了应对这一挑战,我们引入了dCoNNear,这是一种新型的DNN架构,旨在无缝集成到闭环框架中。该架构专门用于防止生成虚假伪像。我们通过闭环框架内的原理验证示例证明了dCoNNear的有效性,该框架采用正常和听力受损配置文件的听觉处理的生物物理学现实模型来设计个性化助听器算法。我们的研究结果表明,dCoNNear不仅准确地模拟了现有非DNN生物物理模型的所有处理阶段,而且还消除了可听伪影,从而提高了助听器算法的音质。这项研究提出了一种新颖的,无伪影的闭环框架,提高了音频处理系统的音质,为音频和听力技术中的高保真应用提供了一个有前途的解决方案。
摘要:Recent advances in deep neural networks (DNNs) have significantly improvedvarious audio processing applications, including speech enhancement, synthesis,and hearing aid algorithms. DNN-based closed-loop systems have gainedpopularity in these applications due to their robust performance and ability toadapt to diverse conditions. Despite their effectiveness, current DNN-basedclosed-loop systems often suffer from sound quality degradation caused byartifacts introduced by suboptimal sampling methods. To address this challenge,we introduce dCoNNear, a novel DNN architecture designed for seamlessintegration into closed-loop frameworks. This architecture specifically aims toprevent the generation of spurious artifacts. We demonstrate the effectivenessof dCoNNear through a proof-of-principle example within a closed-loop frameworkthat employs biophysically realistic models of auditory processing for bothnormal and hearing-impaired profiles to design personalized hearing aidalgorithms. Our results show that dCoNNear not only accurately simulates allprocessing stages of existing non-DNN biophysical models but also eliminatesaudible artifacts, thereby enhancing the sound quality of the resulting hearingaid algorithms. This study presents a novel, artifact-free closed-loopframework that improves the sound quality of audio processing systems, offeringa promising solution for high-fidelity applications in audio and hearingtechnologies.
标题:FleSpeech:具有各种预设的可靠可控语音生成
链接:https://arxiv.org/abs/2501.04644
备注:14 pages, 3 figures
摘要:可控的语音生成方法通常依赖于单个或固定的提示,这阻碍了创造性和灵活性。这些限制使得在某些场景下很难满足特定的用户需求,例如在保持所选扬声器音色的同时调整风格,或者选择风格并生成与角色视觉外观相匹配的声音。为了克服这些挑战,我们提出了\textit{FleSpeech},一种新的多阶段语音生成框架,通过集成各种形式的控制,可以更灵活地操纵语音属性。FleSpeech采用多模式提示编码器,将不同的文本、音频和视觉提示处理并统一到一个有凝聚力的表示中。这种方法增强了语音合成的适应性,并支持对生成的语音进行创造性和精确的控制。此外,我们开发了一个多模态数据集的数据收集管道,以促进该领域的进一步研究和应用。综合的主观和客观实验证明了FleSpeech的有效性。音频样本可在https://kkksuper.github.io/FleSpeech/上获得
摘要:Controllable speech generation methods typically rely on single or fixedprompts, hindering creativity and flexibility. These limitations make itdifficult to meet specific user needs in certain scenarios, such as adjustingthe style while preserving a selected speaker's timbre, or choosing a style andgenerating a voice that matches a character's visual appearance. To overcomethese challenges, we propose \textit{FleSpeech}, a novel multi-stage speechgeneration framework that allows for more flexible manipulation of speechattributes by integrating various forms of control. FleSpeech employs amultimodal prompt encoder that processes and unifies different text, audio, andvisual prompts into a cohesive representation. This approach enhances theadaptability of speech synthesis and supports creative and precise control overthe generated speech. Additionally, we develop a data collection pipeline formultimodal datasets to facilitate further research and applications in thisfield. Comprehensive subjective and objective experiments demonstrate theeffectiveness of FleSpeech. Audio samples are available athttps://kkksuper.github.io/FleSpeech/
标题:CLARVC:采用解开潜在扩散模型和对抗训练的Zero-Shot式语音转换
链接:https://arxiv.org/abs/2501.04416
备注:5 pages, 3 figures, accepted by ICASSP 2025
摘要:风格语音转换的目的是在保持原说话人身份的前提下,将源语音的说话风格转换为所需的说话风格。然而,以前的风格语音转换方法主要集中在定义明确的领域,如情感方面,限制了它们的实际应用。在这项研究中,我们提出了一种新的Zero-shot风格的语音转换方法,它利用语音编解码器和语音提示机制的潜在扩散模型,以促进在上下文中学习说话风格转换。为了理清说话风格和说话人音色之间的关系,本文引入信息瓶颈过滤源语音中的说话风格,并采用不确定性建模自适应实例归一化(UMAdaIN)对风格提示中的说话人音色进行扰动。此外,我们提出了一种新的对抗性训练策略,以增强上下文学习和提高风格相似性。在44,000小时的语音数据上进行的实验表明,在zero-shot场景中,CNOVC在生成具有不同说话风格的语音方面具有优越的性能。
摘要:Style voice conversion aims to transform the speaking style of source speechinto a desired style while keeping the original speaker's identity. However,previous style voice conversion approaches primarily focus on well-defineddomains such as emotional aspects, limiting their practical applications. Inthis study, we present ZSVC, a novel Zero-shot Style Voice Conversion approachthat utilizes a speech codec and a latent diffusion model with speech promptingmechanism to facilitate in-context learning for speaking style conversion. Todisentangle speaking style and speaker timbre, we introduce informationbottleneck to filter speaking style in the source speech and employ UncertaintyModeling Adaptive Instance Normalization (UMAdaIN) to perturb the speakertimbre in the style prompt. Moreover, we propose a novel adversarial trainingstrategy to enhance in-context learning and improve style similarity.Experiments conducted on 44,000 hours of speech data demonstrate the superiorperformance of ZSVC in generating speech with diverse speaking styles inzero-shot scenarios.
标题:使用Transformers和基于VAE的数据增强解码脑电语音感知
链接:https://arxiv.org/abs/2501.04359
备注:19 pages, 15 figures, 2 tables
摘要:从非侵入性脑信号(如脑电图(EEG))解码语音有可能推动脑机接口(BCI)的发展,并应用于无声通信和语音障碍患者的辅助技术。然而,基于EEG的语音解码面临着重大挑战,例如噪声数据,有限的数据集,以及在语音感知等复杂任务上的性能差。本研究试图通过采用变分自动编码器(VAE)进行EEG数据增强来解决这些挑战,以提高数据质量,并将最初在肌电描记术(EMG)中取得成功的最先进(SOTA)序列到序列深度学习架构应用于基于EEG的语音解码任务。此外,我们适应这个架构的词分类任务。使用Brennan数据集,其中包含受试者听叙述语音的EEG记录,我们对数据进行预处理,并评估EEG到单词/句子任务的分类和序列到序列模型。我们的实验表明,VAE有潜力重建人工EEG数据增强。与此同时,我们的序列到序列模型在生成句子方面比我们的分类模型表现得更有希望,尽管两者都仍然是具有挑战性的任务。这些发现为未来的EEG语音感知解码研究奠定了基础,并可能扩展到语音产生任务,如无声或想象语音。
摘要:Decoding speech from non-invasive brain signals, such aselectroencephalography (EEG), has the potential to advance brain-computerinterfaces (BCIs), with applications in silent communication and assistivetechnologies for individuals with speech impairments. However, EEG-based speechdecoding faces major challenges, such as noisy data, limited datasets, and poorperformance on complex tasks like speech perception. This study attempts toaddress these challenges by employing variational autoencoders (VAEs) for EEGdata augmentation to improve data quality and applying a state-of-the-art(SOTA) sequence-to-sequence deep learning architecture, originally successfulin electromyography (EMG) tasks, to EEG-based speech decoding. Additionally, weadapt this architecture for word classification tasks. Using the Brennandataset, which contains EEG recordings of subjects listening to narratedspeech, we preprocess the data and evaluate both classification andsequence-to-sequence models for EEG-to-words/sentences tasks. Our experimentsshow that VAEs have the potential to reconstruct artificial EEG data foraugmentation. Meanwhile, our sequence-to-sequence model achieves more promisingperformance in generating sentences compared to our classification model,though both remain challenging tasks. These findings lay the groundwork forfuture research on EEG speech perception decoding, with possible extensions tospeech production tasks such as silent or imagined speech.
标题:基于DNN的音频处理闭环系统中的无伪影音质
链接:https://arxiv.org/abs/2501.04116
摘要:深度神经网络(DNN)的最新进展显着改善了各种音频处理应用,包括语音增强,合成和助听器算法。基于DNN的闭环系统由于其强大的性能和适应不同条件的能力而在这些应用中受到欢迎。尽管它们的有效性,当前基于DNN的闭环系统经常遭受由次优采样方法引入的伪影引起的声音质量下降。为了应对这一挑战,我们引入了dCoNNear,这是一种新型的DNN架构,旨在无缝集成到闭环框架中。该架构专门用于防止生成虚假伪像。我们通过闭环框架内的原理验证示例证明了dCoNNear的有效性,该框架采用正常和听力受损配置文件的听觉处理的生物物理学现实模型来设计个性化助听器算法。我们的研究结果表明,dCoNNear不仅准确地模拟了现有非DNN生物物理模型的所有处理阶段,而且还消除了可听伪影,从而提高了助听器算法的音质。这项研究提出了一种新颖的,无伪影的闭环框架,提高了音频处理系统的音质,为音频和听力技术中的高保真应用提供了一个有前途的解决方案。
摘要:Recent advances in deep neural networks (DNNs) have significantly improvedvarious audio processing applications, including speech enhancement, synthesis,and hearing aid algorithms. DNN-based closed-loop systems have gainedpopularity in these applications due to their robust performance and ability toadapt to diverse conditions. Despite their effectiveness, current DNN-basedclosed-loop systems often suffer from sound quality degradation caused byartifacts introduced by suboptimal sampling methods. To address this challenge,we introduce dCoNNear, a novel DNN architecture designed for seamlessintegration into closed-loop frameworks. This architecture specifically aims toprevent the generation of spurious artifacts. We demonstrate the effectivenessof dCoNNear through a proof-of-principle example within a closed-loop frameworkthat employs biophysically realistic models of auditory processing for bothnormal and hearing-impaired profiles to design personalized hearing aidalgorithms. Our results show that dCoNNear not only accurately simulates allprocessing stages of existing non-DNN biophysical models but also eliminatesaudible artifacts, thereby enhancing the sound quality of the resulting hearingaid algorithms. This study presents a novel, artifact-free closed-loopframework that improves the sound quality of audio processing systems, offeringa promising solution for high-fidelity applications in audio and hearingtechnologies.
标题:符号音乐分析中音调表示的基于区间的标记化评估
链接:https://arxiv.org/abs/2501.04630
备注:Accepted at Artificial Intelligence for Music Workshop at AAAI 2025 (this https URL)
摘要:符号音乐分析任务通常由最初为自然语言处理开发的模型执行,例如Transformers。这种模型需要将输入数据表示为序列,这是通过标记化过程实现的。符号音乐的符号化策略通常依赖于绝对的音高值来表示音高信息。然而,音乐研究在很大程度上促进了更高层次的表现,如旋律轮廓和和声关系,其中音高间隔比绝对音高更具表现力。在这项工作中,我们介绍了一个通用的框架,用于构建基于区间的标记。通过对三个音乐分析任务进行评估,我们表明这种基于间隔的标记化提高了模型性能,并促进了其可解释性。
摘要:Symbolic music analysis tasks are often performed by models originallydeveloped for Natural Language Processing, such as Transformers. Such modelsrequire the input data to be represented as sequences, which is achievedthrough a process of tokenization. Tokenization strategies for symbolic musicoften rely on absolute MIDI values to represent pitch information. However,music research largely promotes the benefit of higher-level representationssuch as melodic contour and harmonic relations for which pitch intervals turnout to be more expressive than absolute pitches. In this work, we introduce ageneral framework for building interval-based tokenizations. By evaluatingthese tokenizations on three music analysis tasks, we show that suchinterval-based tokenizations improve model performances and facilitate theirexplainability.
标题:时间同步ASB模型的端到端训练中的正确标签上下文
链接:https://arxiv.org/abs/2501.04521
备注:Accepted for presentation at ICASSP 2025
摘要:当前的时间同步序列到序列自动语音识别(ASR)模型是通过使用对所有对齐求和的序列级交叉熵来训练的。由于区别性公式,将正确的标签上下文并入训练标准的梯度会导致归一化问题,并且在数学上没有明确定义。经典的混合神经网络隐马尔可夫模型(NN-HMM)具有其固有的生成公式,能够对正确的标签上下文进行调节。然而,由于HMM状态绑定,正确标签上下文的身份从未显式建模。在这项工作中,我们提出了一个因子损失与辅助左,右标签上下文,总结了所有的路线。我们表明,当训练数据资源有限时,包含正确的标签上下文特别有益。此外,我们还表明,可以完全依赖全和标准来构建分解混合隐马尔可夫模型系统。实验在Switchboard 300 h和LibriSpeech 960 h上进行。
摘要:Current time-synchronous sequence-to-sequence automatic speech recognition(ASR) models are trained by using sequence level cross-entropy that sums overall alignments. Due to the discriminative formulation, incorporating the rightlabel context into the training criterion's gradient causes normalizationproblems and is not mathematically well-defined. The classic hybrid neuralnetwork hidden Markov model (NN-HMM) with its inherent generative formulationenables conditioning on the right label context. However, due to the HMMstate-tying the identity of the right label context is never modeledexplicitly. In this work, we propose a factored loss with auxiliary left andright label contexts that sums over all alignments. We show that the inclusionof the right label context is particularly beneficial when training dataresources are limited. Moreover, we also show that it is possible to build afactored hybrid HMM system by relying exclusively on the full-sum criterion.Experiments were conducted on Switchboard 300h and LibriSpeech 960h.
标题:语音纯度引导的离散令牌用于发音障碍语音识别
链接:https://arxiv.org/abs/2501.04379
备注:ICASSP 2025
摘要:提取的离散标记提供了有效的和域自适应的语音特征。他们的应用程序,表现出清晰度不精确和大的不匹配对正常的声音混乱的讲话仍然是未知的。为了改善他们的语音歧视,削弱在无监督K-均值或矢量量化的连续特征,本文提出了新的音素纯度指导(PPG)离散令牌构音障碍语音识别。语音标签监督是用来正规化的最大似然和重建误差成本标准的K-均值和VAE-VQ为基础的离散令牌提取。在UASpeech语料库上进行的实验表明,从HuBERT中提取的PPG离散令牌特征始终优于混合TDNN和端到端(E2 E)Conformer系统,该系统使用基于非PPG的K-means或VAE-VQ令牌跨越不同码本大小,统计上显著的单词错误率(WER)降低高达0.99%和1.77%的绝对值在16名构音障碍者的UASpeech测试集上,两组间的差异分别为3.21%和4.82%。通过组合使用不同令牌特征的系统获得了最低WER 23.25%。手机纯度指标也得到了持续改善。T-SNE可视化进一步表明,在引入电话纯度指导后,K-means/VAE-VQ集群之间产生了更清晰的决策边界。
摘要:Discrete tokens extracted provide efficient and domain adaptable speechfeatures. Their application to disordered speech that exhibits articulationimprecision and large mismatch against normal voice remains unexplored. Toimprove their phonetic discrimination that is weakened during unsupervisedK-means or vector quantization of continuous features, this paper proposesnovel phone-purity guided (PPG) discrete tokens for dysarthric speechrecognition. Phonetic label supervision is used to regularize maximumlikelihood and reconstruction error costs used in standard K-means and VAE-VQbased discrete token extraction. Experiments conducted on the UASpeech corpussuggest that the proposed PPG discrete token features extracted from HuBERTconsistently outperform hybrid TDNN and End-to-End (E2E) Conformer systemsusing non-PPG based K-means or VAE-VQ tokens across varying codebook sizes bystatistically significant word error rate (WER) reductions up to 0.99\% and1.77\% absolute (3.21\% and 4.82\% relative) respectively on the UASpeech testset of 16 dysarthric speakers. The lowest WER of 23.25\% was obtained bycombining systems using different token features. Consistent improvements onthe phone purity metric were also achieved. T-SNE visualization furtherdemonstrates sharper decision boundaries were produced between K-means/VAE-VQclusters after introducing phone-purity guidance.
标题:MAD-紫外线:第一次通过超声发声挑战检测INTERSPEECT小鼠自闭症
链接:https://arxiv.org/abs/2501.04292
备注:5 pages, 1 figure and 2 tables. For MAD-UV Challenge 2025
摘要:小鼠自闭症检测通过超声发声(MAD-UV)挑战介绍了第一个INTERSPEECH挑战,重点是通过小鼠的发声检测自闭症谱系障碍(ASD)。参与者的任务是开发模型,根据高采样率的记录将小鼠自动分类为野生型或ASD模型。我们的基线系统采用了一个简单的基于CNN的分类,使用三种不同的频谱图特征。结果表明,自动ASD检测的可行性,与所考虑的听觉范围的功能实现最佳性能(UAR为0.600的段级和0.625的主题级分类)。这一挑战将语音技术和生物医学研究联系起来,为我们通过机器学习方法提高对ASD模型的理解提供了机会。这些发现为发声分析提出了有希望的方向,并突出了听觉和超声发声在ASD检测中的潜在价值。
摘要:The Mice Autism Detection via Ultrasound Vocalization (MAD-UV) Challengeintroduces the first INTERSPEECH challenge focused on detecting autism spectrumdisorder (ASD) in mice through their vocalizations. Participants are taskedwith developing models to automatically classify mice as either wild-type orASD models based on recordings with a high sampling rate. Our baseline systememploys a simple CNN-based classification using three different spectrogramfeatures. Results demonstrate the feasibility of automated ASD detection, withthe considered audible-range features achieving the best performance (UAR of0.600 for segment-level and 0.625 for subject-level classification). Thischallenge bridges speech technology and biomedical research, offeringopportunities to advance our understanding of ASD models through machinelearning approaches. The findings suggest promising directions for vocalizationanalysis and highlight the potential value of audible and ultrasoundvocalizations in ASD detection.
标题:DrawSpeech:使用韵律草图作为控制条件的表达性语音合成
链接:https://arxiv.org/abs/2501.04256
备注:Accepted by ICASSP 2025
摘要:None
摘要:Controlling text-to-speech (TTS) systems to synthesize speech with theprosodic characteristics expected by users has attracted much attention. Toachieve controllability, current studies focus on two main directions: (1)using reference speech as prosody prompt to guide speech synthesis, and (2)using natural language descriptions to control the generation process. However,finding reference speech that exactly contains the prosody that users want tosynthesize takes a lot of effort. Description-based guidance in TTS systems canonly determine the overall prosody, which has difficulty in achievingfine-grained prosody control over the synthesized speech. In this paper, wepropose DrawSpeech, a sketch-conditioned diffusion model capable of generatingspeech based on any prosody sketches drawn by users. Specifically, the prosodysketches are fed to DrawSpeech to provide a rough indication of the expectedprosody trends. DrawSpeech then recovers the detailed pitch and energy contoursbased on the coarse sketches and synthesizes the desired speech. Experimentalresults show that DrawSpeech can generate speech with a wide variety of prosodyand can precisely control the fine-grained prosody in a user-friendly manner.Our implementation and audio samples are publicly available.
标题:再次聆听和观看:视听语音识别的生成性错误纠正
链接:https://arxiv.org/abs/2501.04038
摘要:与传统的自动语音识别(ASR)不同,视听语音识别(AVSR)同时获取音频和视频信号来推断转录。最近的研究表明,大语言模型(LLM)可以有效地用于生成错误校正(GER)在ASR中通过预测最好的转录从ASR生成的N-最好的假设。然而,这些LLM缺乏同时理解音频和视频的能力,使得GER方法在AVSR中应用具有挑战性。在这项工作中,我们提出了一个新的GER范式AVSR,称为AVGER,遵循的概念“听和看了一遍”。具体来说,我们首先使用强大的AVSR系统读取音频和视频信号,以获得N-Best假设,然后使用基于Q-former的多模态同步编码器再次读取音频和视频信息,并将其分别转换为LLM可以理解的音频和视频压缩表示。之后,视听压缩表示和N-Best假设一起构成跨模态提示,以指导LLM产生最佳转录。此外,本文还提出了一种多级一致性约束训练准则,包括逻辑级、话语级和表示级,以提高校正精度,同时增强音视频压缩表示的可解释性。在LRS 3数据集上的实验结果表明,该方法优于当前主流的AVSR系统。与它们相比,拟议的AVGER可以将字错误率(WER)降低24%。代码和模型可以在https://github.com/CircleRedRain/AVGER上找到。
摘要:Unlike traditional Automatic Speech Recognition (ASR), Audio-Visual SpeechRecognition (AVSR) takes audio and visual signals simultaneously to infer thetranscription. Recent studies have shown that Large Language Models (LLMs) canbe effectively used for Generative Error Correction (GER) in ASR by predictingthe best transcription from ASR-generated N-best hypotheses. However, theseLLMs lack the ability to simultaneously understand audio and visual, making theGER approach challenging to apply in AVSR. In this work, we propose a novel GERparadigm for AVSR, termed AVGER, that follows the concept of ``listening andseeing again''. Specifically, we first use the powerful AVSR system to read theaudio and visual signals to get the N-Best hypotheses, and then use theQ-former-based Multimodal Synchronous Encoder to read the audio and visualinformation again and convert them into an audio and video compressionrepresentation respectively that can be understood by LLM. Afterward, theaudio-visual compression representation and the N-Best hypothesis togetherconstitute a Cross-modal Prompt to guide the LLM in producing the besttranscription. In addition, we also proposed a Multi-Level ConsistencyConstraint training criterion, including logits-level, utterance-level andrepresentations-level, to improve the correction accuracy while enhancing theinterpretability of audio and visual compression representations. Theexperimental results on the LRS3 dataset show that our method outperformscurrent mainstream AVSR systems. The proposed AVGER can reduce the Word ErrorRate (WER) by 24% compared to them. Code and models can be found at:https://github.com/CircleRedRain/AVGER.
