本文经arXiv每日学术速递授权转载
【1】Representation Purification for End-to-End Speech Translation
链接:https://arxiv.org/abs/2412.04266
备注:Accepted by COLING 2025
摘要:语音到文本翻译(ST)是一项跨模态任务,涉及将口语转换为不同语言的文本。以前的研究主要集中在通过促进机器翻译的知识转移来增强语音翻译,探索各种方法来弥合语音和文本模态之间的差距。尽管取得了实质性的进展,但语音中与翻译内容无关的因素,如音色和节奏,往往限制了知识转移的效率。在本文中,我们将语音表示概念化为内容不可知和内容相关因素的组合。我们通过初步实验研究内容不可知因素对翻译性能的影响,并观察到一个显着的性能恶化时,内容不可知的扰动引入语音信号。为了解决这个问题,我们提出了一个\textbf{S}语音\textbf{R}演示\textbf{P}纯化与\textbf{S}视觉\textbf{E}增强(SRPSE)框架,其排除语音表示内的内容不可知分量以减轻其对ST的负面影响。2个数据集表明,SRPSE在三种设置下显著提高了所有翻译方向的翻译性能,并在一个textit{transcript-free}设置。
摘要:Speech-to-text translation (ST) is a cross-modal task that involvesconverting spoken language into text in a different language. Previous researchprimarily focused on enhancing speech translation by facilitating knowledgetransfer from machine translation, exploring various methods to bridge the gapbetween speech and text modalities. Despite substantial progress made, factorsin speech that are not relevant to translation content, such as timbre andrhythm, often limit the efficiency of knowledge transfer. In this paper, weconceptualize speech representation as a combination of content-agnostic andcontent-relevant factors. We examine the impact of content-agnostic factors ontranslation performance through preliminary experiments and observe asignificant performance deterioration when content-agnostic perturbations areintroduced to speech signals. To address this issue, we propose a\textbf{S}peech \textbf{R}epresentation \textbf{P}urification with\textbf{S}upervision \textbf{E}nhancement (SRPSE) framework, which excludes thecontent-agnostic components within speech representations to mitigate theirnegative impact on ST. Experiments on MuST-C and CoVoST-2 datasets demonstratethat SRPSE significantly improves translation performance across alltranslation directions in three settings and achieves preeminent performanceunder a \textit{transcript-free} setting.
标题:抒情音乐中关键词与强音的关系
链接:https://arxiv.org/abs/2412.04202
备注:Accepted by IEEE BigData 2024
摘要:人工智能(AI)歌曲生成已经成为一个热门话题,但对探索特定歌词和节奏特征之间潜在相关性的关注仍然有限。相比之下,这项试点研究特别调查了关键词与节奏强调特征(例如歌曲中的强节拍)之间的关系。它侧重于几个关键要素:关键词或非关键词,重读或非重读音节,强或弱节拍,目的是揭示有见地的相关性。实验结果表明,平均而言,80.8%的关键字落在强节拍上,而62%的非关键字落在弱节拍上。重读音节与强节拍或弱节拍之间的关系较弱,这表明关键词与强节拍的关系最强。此外,歌词-节奏匹配分数(测量跨各种时间签名的强节拍上的关键词和弱节拍上的非关键词的关键匹配度量)为0.765,而音节类型的匹配分数为0.495。本研究表明,词的类型与其对应的节拍类型强烈一致,证明了不同的模式,而音节类型表现出弱得多的对齐。这种差异强调了词类型在捕捉音乐节奏结构方面的更大可靠性,突出了它们在有效的节奏匹配和分析中的关键作用。我们还得出结论,始终与强节拍保持一致的关键字是歌词节奏关联的更可靠指标,通过增强的结构分析为AI驱动的歌曲生成提供了有价值的见解。此外,我们开发的量身定制的歌词-节奏匹配(LRM)指标最大限度地提高了歌词与相应节拍重音的一致性,我们新颖的LRM文件格式可以捕获关键的歌词和节奏信息,而无需原始乐谱。
摘要:Artificial Intelligence (AI) song generation has emerged as a popular topic,yet the focus on exploring the latent correlations between specific lyrical andrhythmic features remains limited. In contrast, this pilot study particularlyinvestigates the relationships between keywords and rhythmically stressedfeatures such as strong beats in songs. It focuses on several key elements:keywords or non-keywords, stressed or unstressed syllables, and strong or weakbeats, with the aim of uncovering insightful correlations. Experimental resultsindicate that, on average, 80.8\% of keywords land on strong beats, whereas62\% of non-keywords fall on weak beats. The relationship between stressedsyllables and strong or weak beats is weak, revealing that keywords have thestrongest relationships with strong beats. Additionally, the lyrics-rhythmmatching score, a key matching metric measuring keywords on strong beats andnon-keywords on weak beats across various time signatures, is 0.765, while thematching score for syllable types is 0.495. This study demonstrates that wordtypes strongly align with their corresponding beat types, as evidenced by thedistinct patterns, whereas syllable types exhibit a much weaker alignment. Thisdisparity underscores the greater reliability of word types in capturingrhythmic structures in music, highlighting their crucial role in effectiverhythmic matching and analysis. We also conclude that keywords thatconsistently align with strong beats are more reliable indicators oflyrics-rhythm associations, providing valuable insights for AI-driven songgeneration through enhanced structural analysis. Furthermore, our developmentof tailored Lyrics-Rhythm Matching (LRM) metrics maximizes lyrical alignmentswith corresponding beat stresses, and our novel LRM file format capturescritical lyrical and rhythmic information without needing original sheet music.
链接:https://arxiv.org/abs/2412.04100
备注:Submitted to CACM, 12 pages, 2 figures
摘要:生成式人工智能的最新进展引发了人们对音乐生成的新兴趣,并扩大了音乐生成的可能性。然而,这些系统跨音乐流派的性能和通用性受到训练数据可用性的严重影响。我们对人工智能音乐生成研究中使用的超过100万小时的音频数据集进行了广泛的分析,并手动审查了来自11个着名的人工智能和音乐会议和组织的200多篇论文(AAAI、ACM、EUSIPCO、EURASIP、ICASSP、ICML、IJCAI、ISMIR、NeurIPS、NIME,SMC),以确定在人工智能研究中公平代表和纳入全球南方音乐流派方面的关键差距。我们的研究结果揭示了一种明显的不平衡:大约86%的总数据集时间和超过93%的研究人员主要关注来自全球北方的音乐。然而,这些数据集中约40%包括某种形式的非西方音乐,来自南半球的流派仅占数据的14.6%。此外,大约51%的调查论文集中在象征性音乐的生成上,这种方法往往无法捕捉到南亚、中东和非洲等地区音乐中固有的文化细微差别。随着人工智能越来越多地影响音乐的创作和传播,音乐类型在数据集和研究中的代表性严重不足,对全球音乐多样性构成了严重威胁。我们还提出了一些重要的措施来减轻这些风险,并为人工智能驱动的音乐创作创造一个更具包容性的未来。
摘要:Recent advances in generative AI have sparked renewed interest and expandedpossibilities for music generation. However, the performance and versatility ofthese systems across musical genres are heavily influenced by the availabilityof training data. We conducted an extensive analysis of over one million hoursof audio datasets used in AI music generation research and manually reviewedmore than 200 papers from eleven prominent AI and music conferences andorganizations (AAAI, ACM, EUSIPCO, EURASIP, ICASSP, ICML, IJCAI, ISMIR,NeurIPS, NIME, SMC) to identify a critical gap in the fair representation andinclusion of the musical genres of the Global South in AI research. Ourfindings reveal a stark imbalance: approximately 86% of the total dataset hoursand over 93% of researchers focus primarily on music from the Global North.However, around 40% of these datasets include some form of non-Western music,genres from the Global South account for only 14.6% of the data. Furthermore,approximately 51% of the papers surveyed concentrate on symbolic musicgeneration, a method that often fails to capture the cultural nuances inherentin music from regions such as South Asia, the Middle East, and Africa. As AIincreasingly shapes the creation and dissemination of music, the significantunderrepresentation of music genres in datasets and research presents a seriousthreat to global musical diversity. We also propose some important steps tomitigate these risks and foster a more inclusive future for AI-driven musicgeneration.
标题:基于语音识别的特征提取,用于增强型语音自动严重度分类
链接:https://arxiv.org/abs/2412.03784
备注:Accepted to SLT 2024
摘要:由于目前临床评估的主观性,出现了构音障碍语音严重程度自动评估的需求。DNN模型优于ML模型,但缺乏用户友好的可解释性。ML模型在特征级别上提供了可解释的结果,但它们的性能相对较低。当前的ML模型从原始波形中提取各种特征来预测严重程度。然而,现有的方法不包括临床评价中使用的所有构音障碍特征。为了解决这一差距,我们提出了一种特征提取方法,最大限度地减少信息损失。我们引入了ASR转录作为一种新的特征提取源。我们对构音障碍语音的ASR模型进行了微调,并利用该模型对构音障碍语音进行了转录和分词边界信息的提取。它可以捕捉更精细的发音和更广泛的韵律特征。这些功能证明了对现有功能的改善严重度预测性能:平衡准确度为83.72%。
摘要:Due to the subjective nature of current clinical evaluation, the need forautomatic severity evaluation in dysarthric speech has emerged. DNN modelsoutperform ML models but lack user-friendly explainability. ML models offerexplainable results at a feature level, but their performance is comparativelylower. Current ML models extract various features from raw waveforms to predictseverity. However, existing methods do not encompass all dysarthric featuresused in clinical evaluation. To address this gap, we propose a featureextraction method that minimizes information loss. We introduce an ASRtranscription as a novel feature extraction source. We finetune the ASR modelfor dysarthric speech, then use this model to transcribe dysarthric speech andextract word segment boundary information. It enables capturing finerpronunciation and broader prosodic features. These features demonstrated animproved severity prediction performance to existing features: balancedaccuracy of 83.72%.
标题:环境音频Zero-Shot学习的传播
链接:https://arxiv.org/abs/2412.03771
备注:This work has been submitted to the IEEE for possible publication
摘要:Zero-shot学习使模型能够通过利用语义信息来泛化到不可见的类,从而弥合训练集和具有非重叠类的测试集之间的差距。虽然许多研究都集中在计算机视觉中的zero-shot学习上,但这些方法在环境音频中的应用仍然没有得到充分的探索,在现有的研究中表现不佳。生成方法在计算机视觉中已经证明是成功的,但在环境音频zero-shot学习中明显缺乏,其中基于分类的方法占主导地位。 为了解决这个差距,这项工作研究生成方法的zero-shot学习环境音频。采用了两种成功的计算机视觉生成模型:交叉对齐和分布对齐的变分自动编码器(CADA-VAE)和利用不变侧生成对抗网络(LisGAN)。此外,一种新的扩散模型的条件下,类辅助数据。扩散模型为看不见的类生成合成数据,该合成数据与看不见的类数据相结合以训练分类器。 实验在ESC-50和FSC 22两个环境音频数据集上进行。结果表明,扩散模型显着优于所有的基线方法,实现超过25%的ESC-50测试分区的准确性。 这项工作建立了扩散模型作为一个有前途的生成方法的zero-shot学习,并介绍了第一个基准的生成方法的环境音频zero-shot学习,在该领域的未来研究提供了基础。 在https://github.com/ysims/ZeroDiffusion上提供了新颖的零扩散方法的代码。
摘要:Zero-shot learning enables models to generalize to unseen classes byleveraging semantic information, bridging the gap between training and testingsets with non-overlapping classes. While much research has focused on zero-shotlearning in computer vision, the application of these methods to environmentalaudio remains underexplored, with poor performance in existing studies.Generative methods, which have demonstrated success in computer vision, arenotably absent from environmental audio zero-shot learning, whereclassification-based approaches dominate. To address this gap, this work investigates generative methods for zero-shotlearning in environmental audio. Two successful generative models from computervision are adapted: a cross-aligned and distribution-aligned variationalautoencoder (CADA-VAE) and a leveraging invariant side generative adversarialnetwork (LisGAN). Additionally, a novel diffusion model conditioned on classauxiliary data is introduced. The diffusion model generates synthetic data forunseen classes, which is combined with seen-class data to train a classifier. Experiments are conducted on two environmental audio datasets, ESC-50 andFSC22. Results show that the diffusion model significantly outperforms allbaseline methods, achieving more than 25% higher accuracy on the ESC-50 testpartition. This work establishes the diffusion model as a promising generative approachfor zero-shot learning and introduces the first benchmark of generative methodsfor environmental audio zero-shot learning, providing a foundation for futureresearch in the field. Code is provided at https://github.com/ysims/ZeroDiffusion for the novelZeroDiffusion method.
标题:NBM:欧洲夜间候鸟声学监测的开放数据集
链接:https://arxiv.org/abs/2412.03633
摘要:对候鸟种群的持续威胁突出表明,迫切需要有效的监测技术,以协助保护它们。其中,被动声学监测是一个重要的工具,特别是对夜间迁徙的物种,否则很难跟踪。这项工作提出了夜间鸟类迁徙(NBM)数据集,收集了来自117个物种的13,359个注释的声音。该数据集包括精确的时间和频率注释,由法国各地的数十名鸟类爱好者收集,实现了新颖的下游声学分析。特别是,我们证明了一个两阶段的对象检测模型,为音频数据的处理量身定制,可以在我们的数据集上进行训练,以检索频谱图中每个感兴趣的信号周围的局部边界框坐标。这种对象检测方法在鸟类声音识别文献中很大程度上被忽视,通过潜在地区分音频窗口内的单个鸟类,允许重要的应用。此外,我们还表明,我们的识别模型在数据集的45个主要物种上的准确性与在更大的数据集上训练的最先进的系统相竞争。这突出了促进类似的开放科学倡议的兴趣,以获得昂贵但有价值的音频文件的细粒度注释。所有数据和代码都是公开的。
摘要:The persisting threats on migratory bird populations highlights the urgentneed for effective monitoring techniques that could assist in theirconservation. Among these, passive acoustic monitoring is an essential tool,particularly for nocturnal migratory species that are difficult to trackotherwise. This work presents the Nocturnal Bird Migration (NBM) dataset, acollection of 13,359 annotated vocalizations from 117 species of the WesternPalearctic. The dataset includes precise time and frequency annotations,gathered by dozens of bird enthusiasts across France, enabling novel downstreamacoustic analysis. In particular, we demonstrate that a two-stage objectdetection model, tailored for the processing of audio data, can be trained onour dataset to retrieve localized bounding box coordinates around each signalof interest in a spectrogram. This object detection approach, which is largelyoverlooked in the bird sound recognition literature, allows importantapplications by potentially differentiating individual birds within audiowindows. Further, we show that the accuracy of our recognition model on the 45main species of the dataset competes with state-of-the-art systems trained onmuch larger datasets. This highlights the interest of fostering similaropen-science initiatives to acquire costly but valuable fine-grainedannotations of audio files. All data and code are made openly available.
标题:CA-SSLR:用于广义语音处理的条件感知自监督学习表示
链接:https://arxiv.org/abs/2412.04425
备注:38th Conference on Neural Information Processing Systems (NeurIPS 2024)
摘要:我们介绍条件感知的自我监督学习表示(CA-SSLR),一个通用的条件反射模型,广泛适用于各种语音处理任务。与针对下游模型进行优化的标准微调方法相比,CA-SSLR集成了早期层的语言和说话者嵌入,使SSL模型能够感知当前语言和说话者上下文。这种方法减少了对输入音频特征的依赖,同时保持了基本SSLR的完整性。CA-SSLR提高了模型的能力,并以最小的任务特定调优展示了其在不可见任务上的通用性。我们的方法采用线性调制来动态调整内部表示,实现细粒度的适应性,而不会显着改变原始模型的行为。实验表明,CA-SSLR减少了可训练参数的数量,减轻了过拟合,并且在资源不足和看不见的任务中表现出色。具体而言,CA-SSLR实现了LID错误相对减少10%,在ML-SUPERB基准上ASR CER提高37%,在VoxCeleb-1上SV EER降低27%,证明了其有效性。
摘要:We introduce Condition-Aware Self-Supervised Learning Representation(CA-SSLR), a generalist conditioning model broadly applicable to variousspeech-processing tasks. Compared to standard fine-tuning methods that optimizefor downstream models, CA-SSLR integrates language and speaker embeddings fromearlier layers, making the SSL model aware of the current language and speakercontext. This approach reduces the reliance on input audio features whilepreserving the integrity of the base SSLR. CA-SSLR improves the model'scapabilities and demonstrates its generality on unseen tasks with minimaltask-specific tuning. Our method employs linear modulation to dynamicallyadjust internal representations, enabling fine-grained adaptability withoutsignificantly altering the original model behavior. Experiments show thatCA-SSLR reduces the number of trainable parameters, mitigates overfitting, andexcels in under-resourced and unseen tasks. Specifically, CA-SSLR achieves a10% relative reduction in LID errors, a 37% improvement in ASR CER on theML-SUPERB benchmark, and a 27% decrease in SV EER on VoxCeleb-1, demonstratingits effectiveness.
标题:CA-SSLR:用于广义语音处理的条件感知自监督学习表示
链接:https://arxiv.org/abs/2412.04425
备注:38th Conference on Neural Information Processing Systems (NeurIPS 2024)
摘要:我们介绍条件感知的自我监督学习表示(CA-SSLR),一个通用的条件反射模型,广泛适用于各种语音处理任务。与针对下游模型进行优化的标准微调方法相比,CA-SSLR集成了早期层的语言和说话者嵌入,使SSL模型能够感知当前语言和说话者上下文。这种方法减少了对输入音频特征的依赖,同时保持了基本SSLR的完整性。CA-SSLR提高了模型的能力,并以最小的任务特定调优展示了其在不可见任务上的通用性。我们的方法采用线性调制来动态调整内部表示,实现细粒度的适应性,而不会显着改变原始模型的行为。实验表明,CA-SSLR减少了可训练参数的数量,减轻了过拟合,并且在资源不足和看不见的任务中表现出色。具体而言,CA-SSLR实现了LID错误相对减少10%,在ML-SUPERB基准上ASR CER提高37%,在VoxCeleb-1上SV EER降低27%,证明了其有效性。
摘要:We introduce Condition-Aware Self-Supervised Learning Representation(CA-SSLR), a generalist conditioning model broadly applicable to variousspeech-processing tasks. Compared to standard fine-tuning methods that optimizefor downstream models, CA-SSLR integrates language and speaker embeddings fromearlier layers, making the SSL model aware of the current language and speakercontext. This approach reduces the reliance on input audio features whilepreserving the integrity of the base SSLR. CA-SSLR improves the model'scapabilities and demonstrates its generality on unseen tasks with minimaltask-specific tuning. Our method employs linear modulation to dynamicallyadjust internal representations, enabling fine-grained adaptability withoutsignificantly altering the original model behavior. Experiments show thatCA-SSLR reduces the number of trainable parameters, mitigates overfitting, andexcels in under-resourced and unseen tasks. Specifically, CA-SSLR achieves a10% relative reduction in LID errors, a 37% improvement in ASR CER on theML-SUPERB benchmark, and a 27% decrease in SV EER on VoxCeleb-1, demonstratingits effectiveness.
标题:用于组合声学回声消除和降噪的集成最小均方误差算法
链接:https://arxiv.org/abs/2412.04267
摘要:在许多语音记录应用中,噪声和声学回声破坏了期望的语音。因此,需要组合降噪(NR)和声学回声消除(AEC)。通常,遵循级联方法,即,AEC和NR通过选择单独的信号模型、制定单独的成本函数以及使用单独的解决方案策略而被隔离地设计。然后,AEC和NR一个接一个地级联,而不考虑它们的相互作用。然而,在本文中,提出了一种综合的方法来考虑这种相互作用,在一般的多麦克风/多扬声器设置。因此,选择通过堆叠麦克风和扬声器信号而获得的麦克风信号向量或扩展信号向量的单个信号模型,制定单个均方误差成本函数,并且使用共同的解决方案策略。利用该麦克风信号模型,推导出了多通道维纳滤波器(MWF)。使用扩展的信号模型,一个扩展的MWF(MWFext)的推导,并找到几个等价的表达式,但仍然可以解释为级联算法。具体地,MWFext被示出为等同于其中AEC先于NR(AEC NR)、NR先于AEC(NR-AEC)以及扩展NR(NRext)先于AEC和后滤波器(PF)(NRext-AECPF)的算法。在秩亏条件下的MWFext是非唯一的,这样,这种等价相当于具体的表达式,不一定是最小范数的解决方案,这个MWFext。然而,由于非平稳性和不完美的相关矩阵估计,实际性能不同,导致AEC-NR和Next-AEC-PF达到最佳的整体性能。
摘要:In many speech recording applications, noise and acoustic echo corrupt thedesired speech. Consequently, combined noise reduction (NR) and acoustic echocancellation (AEC) is required. Generally, a cascade approach is followed,i.e., the AEC and NR are designed in isolation by selecting a separate signalmodel, formulating a separate cost function, and using a separate solutionstrategy. The AEC and NR are then cascaded one after the other, not accountingfor their interaction. In this paper, however, an integrated approach isproposed to consider this interaction in a generalmulti-microphone/multi-loudspeaker setup. Therefore, a single signal model ofeither the microphone signal vector or the extended signal vector, obtained bystacking microphone and loudspeaker signals, is selected, a single mean squarederror cost function is formulated, and a common solution strategy is used.Using this microphone signal model, a multi channel Wiener filter (MWF) isderived. Using the extended signal model, an extended MWF (MWFext) is derived,and several equivalent expressions are found, which nevertheless areinterpretable as cascade algorithms. Specifically, the MWFext is shown to beequivalent to algorithms where the AEC precedes the NR (AEC NR), the NRprecedes the AEC (NR-AEC), and the extended NR (NRext) precedes the AEC andpost-filter (PF) (NRext-AECPF). Under rank-deficiency conditions the MWFext isnon-unique, such that this equivalence amounts to the expressions beingspecific, not necessarily minimum-norm solutions for this MWFext. The practicalperformances nonetheless differ due to non-stationarities and imperfectcorrelation matrix estimation, resulting in the AEC-NR and NRext-AEC-PFattaining best overall performance.
标题:具有集成专家模型和上下文理解的全面音频查询处理系统
链接:https://arxiv.org/abs/2412.03980
摘要:本文提出了一个全面的聊天机器人系统,旨在通过集成多个专门的音频处理模型来处理广泛的音频相关查询。所提出的系统使用在不同的音频查询数据集上训练的意图分类器,将有关音频内容的查询路由到专家模型,例如自动语音识别(ASR)、扬声器日记、音乐识别和文本到音频生成。然后,3.8 B LLM模型从音频上下文检测(ACD)模块获取输入,从音频中提取音频事件信息,并对来自专家模型的文本域输出进行后处理,以计算对用户的最终响应。我们评估了自定义音频任务和MMAU声音集基准的系统。自定义数据集的动机是行业基准中未涵盖的目标用例,包括ACD-时间戳-QA(问题推理)以及ACD-时间-QA数据集,分别用于评估时间戳和时间推理问题。首先,我们确定基于BERT的意图分类器在路由查询中优于LLM-fewshot意图分类器。实验进一步表明,与最先进的大型音频语言模型相比,我们的方法显着提高了一些自定义任务的准确性,并在MMAU基准的声音测试集上优于7 B参数大小范围内的模型,从而为设备部署提供了一个有吸引力的选择。
摘要:This paper presents a comprehensive chatbot system designed to handle a widerange of audio-related queries by integrating multiple specialized audioprocessing models. The proposed system uses an intent classifier, trained on adiverse audio query dataset, to route queries about audio content to expertmodels such as Automatic Speech Recognition (ASR), Speaker Diarization, MusicIdentification, and Text-to-Audio generation. A 3.8 B LLM model then takesinputs from an Audio Context Detection (ACD) module extracting audio eventinformation from the audio and post processes text domain outputs from theexpert models to compute the final response to the user. We evaluated thesystem on custom audio tasks and MMAU sound set benchmarks. The custom datasetswere motivated by target use cases not covered in industry benchmarks andincluded ACD-timestamp-QA (Question Answering) as well as ACD-temporal-QAdatasets to evaluate timestamp and temporal reasoning questions, respectively.First we determined that a BERT based Intent Classifier outperforms LLM-fewshotintent classifier in routing queries. Experiments further show that ourapproach significantly improves accuracy on some custom tasks compared tostate-of-the-art Large Audio Language Models and outperforms models in the 7Bparameter size range on the sound testset of the MMAU benchmark, therebyoffering an attractive option for on device deployment.
标题:端到端语音翻译的表示净化
链接:https://arxiv.org/abs/2412.04266
备注:Accepted by COLING 2025
摘要:语音到文本翻译(ST)是一项跨模态任务,涉及将口语转换为不同语言的文本。以前的研究主要集中在通过促进机器翻译的知识转移来增强语音翻译,探索各种方法来弥合语音和文本模态之间的差距。尽管取得了实质性的进展,但语音中与翻译内容无关的因素,如音色和节奏,往往限制了知识转移的效率。在本文中,我们将语音表示概念化为内容不可知和内容相关因素的组合。我们通过初步实验研究内容不可知因素对翻译性能的影响,并观察到一个显着的性能恶化时,内容不可知的扰动引入语音信号。为了解决这个问题,我们提出了一个\textbf{S}语音\textbf{R}演示\textbf{P}纯化与\textbf{S}视觉\textbf{E}增强(SRPSE)框架,其排除语音表示内的内容不可知分量以减轻其对ST的负面影响。2个数据集表明,SRPSE在三种设置下显著提高了所有翻译方向的翻译性能,并在一个textit{transcript-free}设置。
摘要:Speech-to-text translation (ST) is a cross-modal task that involvesconverting spoken language into text in a different language. Previous researchprimarily focused on enhancing speech translation by facilitating knowledgetransfer from machine translation, exploring various methods to bridge the gapbetween speech and text modalities. Despite substantial progress made, factorsin speech that are not relevant to translation content, such as timbre andrhythm, often limit the efficiency of knowledge transfer. In this paper, weconceptualize speech representation as a combination of content-agnostic andcontent-relevant factors. We examine the impact of content-agnostic factors ontranslation performance through preliminary experiments and observe asignificant performance deterioration when content-agnostic perturbations areintroduced to speech signals. To address this issue, we propose a\textbf{S}peech \textbf{R}epresentation \textbf{P}urification with\textbf{S}upervision \textbf{E}nhancement (SRPSE) framework, which excludes thecontent-agnostic components within speech representations to mitigate theirnegative impact on ST. Experiments on MuST-C and CoVoST-2 datasets demonstratethat SRPSE significantly improves translation performance across alltranslation directions in three settings and achieves preeminent performanceunder a \textit{transcript-free} setting.
标题:抒情音乐中关键词与强音的关系
链接:https://arxiv.org/abs/2412.04202
备注:Accepted by IEEE BigData 2024
摘要:人工智能(AI)歌曲生成已经成为一个热门话题,但对探索特定歌词和节奏特征之间潜在相关性的关注仍然有限。相比之下,这项试点研究特别调查关键字和节奏强调的功能,如歌曲中的强节拍之间的关系。它侧重于几个关键要素:关键词或非关键词,重读或非重读音节,强或弱节拍,目的是揭示有见地的相关性。实验结果表明,平均而言,80.8%的关键字落在强节拍上,而62%的非关键字落在弱节拍上。重读音节与强节拍或弱节拍之间的关系较弱,这表明关键词与强节拍的关系最强。此外,歌词-节奏匹配分数(测量跨各种时间签名的强节拍上的关键词和弱节拍上的非关键词的关键匹配度量)为0.765,而音节类型的匹配分数为0.495。这项研究表明,单词类型与其相应的节拍类型强烈一致,正如不同的模式所证明的那样,而音节类型的一致性要弱得多。这种差异强调了词类型在捕捉音乐节奏结构方面的更大可靠性,突出了它们在有效的节奏匹配和分析中的关键作用。我们还得出结论,始终与强节拍保持一致的关键字是歌词节奏关联的更可靠指标,通过增强的结构分析为AI驱动的歌曲生成提供了有价值的见解。此外,我们开发的量身定制的歌词-节奏匹配(LRM)指标最大限度地提高了歌词与相应节拍重音的一致性,我们新颖的LRM文件格式可以捕获关键的歌词和节奏信息,而无需原始乐谱。
摘要:Artificial Intelligence (AI) song generation has emerged as a popular topic,yet the focus on exploring the latent correlations between specific lyrical andrhythmic features remains limited. In contrast, this pilot study particularlyinvestigates the relationships between keywords and rhythmically stressedfeatures such as strong beats in songs. It focuses on several key elements:keywords or non-keywords, stressed or unstressed syllables, and strong or weakbeats, with the aim of uncovering insightful correlations. Experimental resultsindicate that, on average, 80.8\% of keywords land on strong beats, whereas62\% of non-keywords fall on weak beats. The relationship between stressedsyllables and strong or weak beats is weak, revealing that keywords have thestrongest relationships with strong beats. Additionally, the lyrics-rhythmmatching score, a key matching metric measuring keywords on strong beats andnon-keywords on weak beats across various time signatures, is 0.765, while thematching score for syllable types is 0.495. This study demonstrates that wordtypes strongly align with their corresponding beat types, as evidenced by thedistinct patterns, whereas syllable types exhibit a much weaker alignment. Thisdisparity underscores the greater reliability of word types in capturingrhythmic structures in music, highlighting their crucial role in effectiverhythmic matching and analysis. We also conclude that keywords thatconsistently align with strong beats are more reliable indicators oflyrics-rhythm associations, providing valuable insights for AI-driven songgeneration through enhanced structural analysis. Furthermore, our developmentof tailored Lyrics-Rhythm Matching (LRM) metrics maximizes lyrical alignmentswith corresponding beat stresses, and our novel LRM file format capturescritical lyrical and rhythmic information without needing original sheet music.
链接:https://arxiv.org/abs/2412.04100
备注:Submitted to CACM, 12 pages, 2 figures
摘要:生成式人工智能的最新进展引发了人们对音乐生成的新兴趣,并扩大了音乐生成的可能性。然而,这些系统跨音乐流派的性能和通用性受到训练数据可用性的严重影响。我们对人工智能音乐生成研究中使用的超过100万小时的音频数据集进行了广泛的分析,并手动审查了来自11个着名的人工智能和音乐会议和组织的200多篇论文(AAAI、ACM、EUSIPCO、EURASIP、ICASSP、ICML、IJCAI、ISMIR、NeurIPS、NIME,SMC),以确定在人工智能研究中公平代表和纳入全球南方音乐流派方面的关键差距。我们的研究结果揭示了一种明显的不平衡:大约86%的总数据集时间和超过93%的研究人员主要关注来自全球北方的音乐。然而,这些数据集中约40%包括某种形式的非西方音乐,来自南半球的流派仅占数据的14.6%。此外,大约51%的调查论文集中在象征性音乐的生成上,这种方法往往无法捕捉到南亚、中东和非洲等地区音乐中固有的文化细微差别。随着人工智能越来越多地影响音乐的创作和传播,音乐类型在数据集和研究中的代表性严重不足,对全球音乐多样性构成了严重威胁。我们还提出了一些重要的措施来减轻这些风险,并为人工智能驱动的音乐创作创造一个更具包容性的未来。
摘要:Recent advances in generative AI have sparked renewed interest and expandedpossibilities for music generation. However, the performance and versatility ofthese systems across musical genres are heavily influenced by the availabilityof training data. We conducted an extensive analysis of over one million hoursof audio datasets used in AI music generation research and manually reviewedmore than 200 papers from eleven prominent AI and music conferences andorganizations (AAAI, ACM, EUSIPCO, EURASIP, ICASSP, ICML, IJCAI, ISMIR,NeurIPS, NIME, SMC) to identify a critical gap in the fair representation andinclusion of the musical genres of the Global South in AI research. Ourfindings reveal a stark imbalance: approximately 86% of the total dataset hoursand over 93% of researchers focus primarily on music from the Global North.However, around 40% of these datasets include some form of non-Western music,genres from the Global South account for only 14.6% of the data. Furthermore,approximately 51% of the papers surveyed concentrate on symbolic musicgeneration, a method that often fails to capture the cultural nuances inherentin music from regions such as South Asia, the Middle East, and Africa. As AIincreasingly shapes the creation and dissemination of music, the significantunderrepresentation of music genres in datasets and research presents a seriousthreat to global musical diversity. We also propose some important steps tomitigate these risks and foster a more inclusive future for AI-driven musicgeneration.
标题:基于语音识别的特征提取,用于增强型语音自动严重度分类
链接:https://arxiv.org/abs/2412.03784
备注:Accepted to SLT 2024
摘要:由于目前临床评估的主观性,出现了构音障碍语音严重程度自动评估的需求。DNN模型优于ML模型,但缺乏用户友好的可解释性。ML模型在特征级别上提供了可解释的结果,但它们的性能相对较低。当前的ML模型从原始波形中提取各种特征来预测严重程度。然而,现有的方法不包括临床评价中使用的所有构音障碍特征。为了解决这一差距,我们提出了一种特征提取方法,最大限度地减少信息损失。我们引入了ASR转录作为一种新的特征提取源。我们对构音障碍语音的ASR模型进行了微调,并利用该模型对构音障碍语音进行了转录和分词边界信息的提取。它可以捕捉更精细的发音和更广泛的韵律特征。这些功能证明了对现有功能的改善严重度预测性能:平衡准确度为83.72%。
摘要:Due to the subjective nature of current clinical evaluation, the need forautomatic severity evaluation in dysarthric speech has emerged. DNN modelsoutperform ML models but lack user-friendly explainability. ML models offerexplainable results at a feature level, but their performance is comparativelylower. Current ML models extract various features from raw waveforms to predictseverity. However, existing methods do not encompass all dysarthric featuresused in clinical evaluation. To address this gap, we propose a featureextraction method that minimizes information loss. We introduce an ASRtranscription as a novel feature extraction source. We finetune the ASR modelfor dysarthric speech, then use this model to transcribe dysarthric speech andextract word segment boundary information. It enables capturing finerpronunciation and broader prosodic features. These features demonstrated animproved severity prediction performance to existing features: balancedaccuracy of 83.72%.
标题:环境音频Zero-Shot学习的传播
链接:https://arxiv.org/abs/2412.03771
备注:This work has been submitted to the IEEE for possible publication
摘要:Zero-shot学习使模型能够通过利用语义信息来泛化到不可见的类,从而弥合训练集和具有非重叠类的测试集之间的差距。虽然许多研究都集中在计算机视觉中的zero-shot学习上,但这些方法在环境音频中的应用仍然没有得到充分的探索,在现有的研究中表现不佳。生成方法在计算机视觉中已经证明是成功的,但在环境音频zero-shot学习中明显缺乏,其中基于分类的方法占主导地位。 为了解决这个差距,这项工作研究生成方法的zero-shot学习环境音频。采用了两种成功的计算机视觉生成模型:交叉对齐和分布对齐的变分自动编码器(CADA-VAE)和利用不变侧生成对抗网络(LisGAN)。此外,一种新的扩散模型的条件下,类辅助数据。扩散模型为看不见的类生成合成数据,该合成数据与看不见的类数据相结合以训练分类器。 在两个环境音频数据集ESC-50和FSC 22上进行了实验。结果表明,扩散模型显着优于所有的基线方法,实现了超过25%的ESC-50测试分区的准确性。 这项工作建立了扩散模型作为一个有前途的生成方法的zero-shot学习,并介绍了第一个基准的生成方法的环境音频zero-shot学习,在该领域的未来研究提供了基础。 在https://github.com/ysims/ZeroDiffusion上提供了新颖的零扩散方法的代码。
摘要:Zero-shot learning enables models to generalize to unseen classes byleveraging semantic information, bridging the gap between training and testingsets with non-overlapping classes. While much research has focused on zero-shotlearning in computer vision, the application of these methods to environmentalaudio remains underexplored, with poor performance in existing studies.Generative methods, which have demonstrated success in computer vision, arenotably absent from environmental audio zero-shot learning, whereclassification-based approaches dominate. To address this gap, this work investigates generative methods for zero-shotlearning in environmental audio. Two successful generative models from computervision are adapted: a cross-aligned and distribution-aligned variationalautoencoder (CADA-VAE) and a leveraging invariant side generative adversarialnetwork (LisGAN). Additionally, a novel diffusion model conditioned on classauxiliary data is introduced. The diffusion model generates synthetic data forunseen classes, which is combined with seen-class data to train a classifier. Experiments are conducted on two environmental audio datasets, ESC-50 andFSC22. Results show that the diffusion model significantly outperforms allbaseline methods, achieving more than 25% higher accuracy on the ESC-50 testpartition. This work establishes the diffusion model as a promising generative approachfor zero-shot learning and introduces the first benchmark of generative methodsfor environmental audio zero-shot learning, providing a foundation for futureresearch in the field. Code is provided at https://github.com/ysims/ZeroDiffusion for the novelZeroDiffusion method.
标题:NBM:欧洲夜间候鸟声学监测的开放数据集
链接:https://arxiv.org/abs/2412.03633
摘要:对候鸟种群的持续威胁突出表明,迫切需要有效的监测技术,以协助保护它们。其中,被动声学监测是一个重要的工具,特别是对夜间迁徙的物种,否则很难跟踪。这项工作提出了夜间鸟类迁徙(NBM)数据集,收集了来自117个物种的13,359个注释的声音。该数据集包括精确的时间和频率注释,由法国各地的数十名鸟类爱好者收集,实现了新颖的下游声学分析。特别是,我们证明了一个两阶段的对象检测模型,为音频数据的处理量身定制,可以在我们的数据集上进行训练,以检索频谱图中每个感兴趣的信号周围的局部边界框坐标。这种对象检测方法在鸟类声音识别文献中很大程度上被忽视,通过潜在地区分音频窗口内的单个鸟类,允许重要的应用。此外,我们还表明,我们的识别模型在数据集的45个主要物种上的准确性与在更大的数据集上训练的最先进的系统相竞争。这突出了促进类似的开放科学倡议的兴趣,以获得昂贵但有价值的音频文件的细粒度注释。所有数据和代码都是公开的。
摘要:The persisting threats on migratory bird populations highlights the urgentneed for effective monitoring techniques that could assist in theirconservation. Among these, passive acoustic monitoring is an essential tool,particularly for nocturnal migratory species that are difficult to trackotherwise. This work presents the Nocturnal Bird Migration (NBM) dataset, acollection of 13,359 annotated vocalizations from 117 species of the WesternPalearctic. The dataset includes precise time and frequency annotations,gathered by dozens of bird enthusiasts across France, enabling novel downstreamacoustic analysis. In particular, we demonstrate that a two-stage objectdetection model, tailored for the processing of audio data, can be trained onour dataset to retrieve localized bounding box coordinates around each signalof interest in a spectrogram. This object detection approach, which is largelyoverlooked in the bird sound recognition literature, allows importantapplications by potentially differentiating individual birds within audiowindows. Further, we show that the accuracy of our recognition model on the 45main species of the dataset competes with state-of-the-art systems trained onmuch larger datasets. This highlights the interest of fostering similaropen-science initiatives to acquire costly but valuable fine-grainedannotations of audio files. All data and code are made openly available.
