本文经arXiv每日学术速递授权转载
链接:https://arxiv.org/abs/2412.07406
备注:Published in the Proceedings of the International Symposium on Visual Computing, 2021 this https URL
摘要:我们提出了一种新的自我监督的方法来学习音频和视觉表示从未标记的视频,基于它们的对应关系。该方法使用注意力机制来学习以不同分辨率从音频和视频流中提取的卷积特征的相对重要性,并使用注意力特征基于它们的对应关系对音频和视频输入进行编码。我们评估了模型学习的表示,以分类视听相关性,并为视觉场景推荐声音效果。我们的研究结果表明,与基线相比,注意力模型生成的表示将相关准确性提高了18%,VGG-Sound的推荐准确性提高了10%,VGG-Sound是一个公共视频数据集。此外,通过跨模态对比学习训练注意力模型来学习的视听表示进一步提高了推荐性能,这是基于我们使用VGG-Sound和更具挑战性的游戏视频记录数据集进行的评估。
摘要:We propose a novel self-supervised approach for learning audio and visualrepresentations from unlabeled videos, based on their correspondence. Theapproach uses an attention mechanism to learn the relative importance ofconvolutional features extracted at different resolutions from the audio andvisual streams and uses the attention features to encode the audio and visualinput based on their correspondence. We evaluated the representations learnedby the model to classify audio-visual correlation as well as to recommend soundeffects for visual scenes. Our results show that the representations generatedby the attention model improves the correlation accuracy compared to thebaseline, by 18% and the recommendation accuracy by 10% for VGG-Sound, which isa public video dataset. Additionally, audio-visual representations learned bytraining the attention model with cross-modal contrastive learning furtherimproves the recommendation performance, based on our evaluation usingVGG-Sound and a more challenging dataset consisting of gameplay videorecordings.
标题:与历史对接:利用音频增强对象进行策展
链接:https://arxiv.org/abs/2412.07345
备注:32 pages, 5 figures
摘要:本文介绍并讨论了游客与包含音频增强对象的两个音频增强现实体验的交互结果;虚拟音频源已附加到的物理真实世界对象。然后,它继续讨论共同确定的主题所产生的观察游客的行为在这些经验和分析他们的口头和书面反馈。音频增强对象的策展潜力进行了讨论,并通过结论的方式,他们的功能作为接口的数字音频档案内容提出,随着他们的能力,重新框架,重新语境化和创造新的经验与现有的收藏沉默的博物馆展品。
摘要:This article presents and discusses the results from visitors' interactionswith two audio augmented reality experiences containing audio augmentedobjects; physical, real-world objects to which virtual audio sources have beenattached. It then proceeds to discusses the commonly identified themes arisingfrom the observation of visitors' behaviour within these experiences and theanalysis of their verbal and written feedback. The curatorial potential ofaudio augmented objects is discussed and, by way of conclusion, theirfunctionality as interfaces to digital audio archival content is proposed,along with their ability to reframe, re-contextualise and create renewedexperiences with existing collections of silenced museum exhibits.
标题:利用非自回归生成和预训练在直接语音到语音翻译中保留说话人信息
链接:https://arxiv.org/abs/2412.07316
摘要:语音到语音翻译是指将一种语言的语音转换成另一种语言的语义对等的语音,以促进不同语言使用者之间的交流。语音到离散单元转换(S2 UT)是端到端S2 ST的主流方法,可解决传统级联系统中经常遇到的跨模块错误传播和推理速度慢等挑战。然而,由于离散单元主要捕获内容信息,因此常规S2 UT方法无法保留来自源的说话者特定特性。我们以前的工作,SC-S2 UT,介绍了一个扬声器适配器和一个单元到梅尔结构,使扬声器信息的保存和非自回归语音生成。在此基础上,本研究提出了一种自我监督的预训练方法,以丰富由扬声器适配器和单元到梅尔结构提取的信息。此外,我们研究了不同的特征融合策略,以进一步提高说话人和内容特征的集成。在ES-EN和FR-EN任务的CVSS-T数据集上进行的实验表明,与SC-S2 UT相比,我们提出的方法实现了1.14的BLEU分数提高,同时MOS和说话人相似性也有显着增强。此外,我们的方法实现了翻译质量相媲美传统的S2 UT,只有一个最小的增加0.04秒,每个话语的推理时间,同时保持高的说话人相似性。这些结果验证了该方法的有效性。
摘要:Speech-to-Speech Translation (S2ST) refers to the conversion of speech in onelanguage into semantically equivalent speech in another language, facilitatingcommunication between speakers of different languages. Speech-to-Discrete UnitTranslation (S2UT), a mainstream approach for end-to-end S2ST, addresseschallenges such as error propagation across modules and slow inference speedoften encountered in traditional cascade systems. However, as discrete unitsprimarily capture content information, conventional S2UT methods fail to retainspeaker-specific characteristics from the source. Our previous work, SC-S2UT,introduced a speaker adapter and a unit-to-mel structure, enabling thepreservation of speaker information and non-autoregressive speech generation.Building on this foundation, this study proposes a self-supervised pretrainingmethod to enrich the information extracted by both the speaker adapter and theunit-to-mel structure. Additionally, we investigate different feature fusionstrategies to further improve the integration of speaker and content features.Experiments conducted on the CVSS-T dataset for ES-EN and FR-EN tasksdemonstrate that our proposed method achieves a BLEU score improvement of 1.14compared to SC-S2UT, along with significant enhancements in MOS and speakersimilarity. Furthermore, our approach achieves translation quality comparableto traditional S2UT, with only a minimal increase of 0.04s per utterance ininference time, while maintaining high speaker similarity. These resultsvalidate the effectiveness of the proposed method.
标题:通过软提示微调对基于LLM的ASB进行有效的文本自适应
链接:https://arxiv.org/abs/2412.06967
备注:accepted as SLT 2024 proceeding
摘要:大语言模型(LLM)的出现对自动语音识别(ASR)产生了巨大的影响.将LLM与音频嵌入结合以生成transmittance成为新的最先进的ASR。尽管LLM使用大量的文本语料库进行训练,但高质量的特定于领域的文本数据仍然可以显着提高领域适应任务的ASR性能。虽然基于LLM的ASR可以通过微调LLM解码器自然地并入更多的文本语料库,但是在没有配对提示的情况下对仅文本数据微调这样的ASR可能会降低特定于领域的知识的有效性。为了缓解这个问题,我们提出了一个两步软提示微调策略,提高特定领域的文本适应。实验结果表明,与基线ASR相比,我们提出的方法的文本自适应在目标域上实现了相对高达9%的字错误率(WER)降低和高达18%的实体错误率(EER)降低。将其与特定领域语言模型(LM)融合相结合可以进一步提高EER相对2-5%
摘要:The advent of Large Language Models (LLM) has reformed the Automatic SpeechRecognition (ASR). Prompting LLM with audio embeddings to generatetranscriptions becomes the new state-of-the-art ASR. Despite LLMs being trainedwith an extensive amount of text corpora, high-quality domain-specific textdata can still significantly enhance ASR performance on domain adaptationtasks. Although LLM-based ASR can naturally incorporate more text corpora byfine-tuning the LLM decoder, fine-tuning such ASR on text-only data withoutpaired prompts may diminish the effectiveness of domain-specific knowledge. Tomitigate this issue, we propose a two-step soft prompt fine-tuning strategythat enhances domain-specific text adaptation. Experimental results show thattext adaptation with our proposed method achieved a relative up to 9% WordError Rate (WER) reduction and up to 18% Entity Error Rate (EER) reduction onthe target domain compared to the baseline ASR. Combining this withdomain-specific Language Model (LM) fusion can further improve the EER by arelative 2-5%
标题:利用扩散和一致性模型改进源提取
链接:https://arxiv.org/abs/2412.06965
摘要:在这项工作中,我们展示了集成的分数匹配扩散模型到一个确定性的架构,时域音乐源提取,从而提高音频质量。为了解决扩散模型通常缓慢的迭代采样过程,我们应用一致性蒸馏,并将采样过程减少到一个步骤,实现与扩散模型相当的性能,并且具有两个或更多个步骤,甚至超过它们。在四种乐器(低音、鼓、吉他和钢琴)的Slakh2100数据集上进行训练后,我们的模型显示出与基线方法相比在客观指标上的显着改进。正确的示例可在https://consistency-separation.github.io/上找到。
摘要:In this work, we demonstrate the integration of a score-matching diffusionmodel into a deterministic architecture for time-domain musical sourceextraction, resulting in enhanced audio quality. To address the typically slowiterative sampling process of diffusion models, we apply consistencydistillation and reduce the sampling process to a single step, achievingperformance comparable to that of diffusion models, and with two or more steps,even surpassing them. Trained on the Slakh2100 dataset for four instruments(bass, drums, guitar, and piano), our model shows significant improvementsacross objective metrics compared to baseline methods. Sound examples areavailable at https://consistency-separation.github.io/.
标题:野外声景分析的时空潜在表示
链接:https://arxiv.org/abs/2412.07648
备注:9 pages, 6 figures
摘要:在声学场景分析领域,本文提出了一种新的方法,从野外音频数据中发现时空潜在表示。通过使用WE-LIVE,一个内部收集的数据集,包括不同现实环境中的音频记录以及稀疏的GPS坐标,自我注释的情感和情境标签,我们解决了将每个音频片段与其相应位置相关联的挑战性任务作为借口任务,最终目标是声学检测暴力(异常)上下文,作为进一步的工作。通过生成声学嵌入并使用自监督学习范式,我们的目标是使用模型生成的潜在空间来声学表征时空上下文。我们使用YAMNet,一个在AudioSet中训练的声学事件分类器,在WE-LIVE中暂时定位和识别声学事件。为了将离散的声学事件转换为嵌入,我们比较了基于信息检索的TF-IDF算法和Node 2 Vec作为自然语言处理技术的类比。然后训练VAE以提供进一步适配的潜在空间。通过测量余弦距离并通过t分布随机邻居嵌入可视化数据分布来进行分析,揭示了不同的声学场景。具体来说,我们辨别室内和地铁环境之间的变化。值得注意的是,这些区别出现在VAE的潜在空间中,与编码前数据点的随机分布形成鲜明对比。总之,我们的研究为从野外音频数据中提取时空潜在表征提供了一种开创性的方法。
摘要:In the field of acoustic scene analysis, this paper presents a novel approachto find spatio-temporal latent representations from in-the-wild audio data. Byusing WE-LIVE, an in-house collected dataset that includes audio recordings indiverse real-world environments together with sparse GPS coordinates,self-annotated emotional and situational labels, we tackle the challenging taskof associating each audio segment with its corresponding location as a pretexttask, with the final aim of acoustically detecting violent (anomalous)contexts, left as further work. By generating acoustic embeddings and using theself-supervised learning paradigm, we aim to use the model-generated latentspace to acoustically characterize the spatio-temporal context. We use YAMNet,an acoustic events classifier trained in AudioSet to temporally locate andidentify acoustic events in WE-LIVE. In order to transform the discreteacoustic events into embeddings, we compare the information-retrieval-basedTF-IDF algorithm and Node2Vec as an analogy to Natural Language Processingtechniques. A VAE is then trained to provide a further adapted latent space.The analysis was carried out by measuring the cosine distance and visualizingdata distribution via t-Distributed Stochastic Neighbor Embedding, revealingdistinct acoustic scenes. Specifically, we discern variations between indoorand subway environments. Notably, these distinctions emerge within the latentspace of the VAE, a stark contrast to the random distribution of data pointsbefore encoding. In summary, our research contributes a pioneering approach forextracting spatio-temporal latent representations from in-the-wild audio data.
标题:学习自我监督的视听表示以获得合理的建议
链接:https://arxiv.org/abs/2412.07406
备注:Published in the Proceedings of the International Symposium on Visual Computing, 2021 this https URL
摘要:我们提出了一种新的自我监督的方法来学习音频和视觉表示从未标记的视频,基于它们的对应关系。该方法使用注意力机制来学习以不同分辨率从音频和视频流中提取的卷积特征的相对重要性,并使用注意力特征基于它们的对应关系对音频和视频输入进行编码。我们评估了模型学习的表示,以分类视听相关性,并为视觉场景推荐声音效果。我们的研究结果表明,与基线相比,注意力模型生成的表示将相关准确性提高了18%,VGG-Sound的推荐准确性提高了10%,VGG-Sound是一个公共视频数据集。此外,通过跨模态对比学习训练注意力模型来学习的视听表示进一步提高了推荐性能,这是基于我们使用VGG-Sound和更具挑战性的游戏视频记录数据集进行的评估。
摘要:We propose a novel self-supervised approach for learning audio and visualrepresentations from unlabeled videos, based on their correspondence. Theapproach uses an attention mechanism to learn the relative importance ofconvolutional features extracted at different resolutions from the audio andvisual streams and uses the attention features to encode the audio and visualinput based on their correspondence. We evaluated the representations learnedby the model to classify audio-visual correlation as well as to recommend soundeffects for visual scenes. Our results show that the representations generatedby the attention model improves the correlation accuracy compared to thebaseline, by 18% and the recommendation accuracy by 10% for VGG-Sound, which isa public video dataset. Additionally, audio-visual representations learned bytraining the attention model with cross-modal contrastive learning furtherimproves the recommendation performance, based on our evaluation usingVGG-Sound and a more challenging dataset consisting of gameplay videorecordings.
标题:与历史对接:利用音频增强对象进行策展
链接:https://arxiv.org/abs/2412.07345
备注:32 pages, 5 figures
摘要:本文介绍并讨论了游客与包含音频增强对象的两个音频增强现实体验的交互结果;虚拟音频源已附加到的物理真实世界对象。然后,它继续讨论共同确定的主题所产生的观察游客的行为在这些经验和分析他们的口头和书面反馈。音频增强对象的策展潜力进行了讨论,并通过结论的方式,他们的功能作为接口的数字音频档案内容提出,随着他们的能力,重新框架,重新语境化和创造新的经验与现有的收藏沉默的博物馆展品。
摘要:This article presents and discusses the results from visitors' interactionswith two audio augmented reality experiences containing audio augmentedobjects; physical, real-world objects to which virtual audio sources have beenattached. It then proceeds to discusses the commonly identified themes arisingfrom the observation of visitors' behaviour within these experiences and theanalysis of their verbal and written feedback. The curatorial potential ofaudio augmented objects is discussed and, by way of conclusion, theirfunctionality as interfaces to digital audio archival content is proposed,along with their ability to reframe, re-contextualise and create renewedexperiences with existing collections of silenced museum exhibits.
标题:利用非自回归生成和预训练在直接语音到语音翻译中保留说话人信息
链接:https://arxiv.org/abs/2412.07316
摘要:语音到语音翻译是指将一种语言的语音转换成另一种语言的语义对等的语音,以促进不同语言使用者之间的交流。语音到离散单元转换(S2 UT)是端到端S2 ST的主流方法,可解决传统级联系统中经常遇到的跨模块错误传播和推理速度慢等挑战。然而,由于离散单元主要捕获内容信息,因此常规S2 UT方法无法保留来自源的说话者特定特性。我们以前的工作,SC-S2 UT,介绍了一个扬声器适配器和一个单元到梅尔结构,使扬声器信息的保存和非自回归语音生成。在此基础上,本研究提出了一种自监督预训练方法,以丰富扬声器适配器和单元到梅尔结构提取的信息。此外,我们研究了不同的特征融合策略,以进一步提高说话人和内容特征的集成。在ES-EN和FR-EN任务的CVSS-T数据集上进行的实验表明,与SC-S2 UT相比,我们提出的方法实现了1.14的BLEU分数提高,同时MOS和说话人相似性也有显着增强。此外,我们的方法实现了翻译质量相媲美传统的S2 UT,只有一个最小的增加0.04秒,每个话语的推理时间,同时保持高的说话人相似性。这些结果验证了该方法的有效性。
摘要:Speech-to-Speech Translation (S2ST) refers to the conversion of speech in onelanguage into semantically equivalent speech in another language, facilitatingcommunication between speakers of different languages. Speech-to-Discrete UnitTranslation (S2UT), a mainstream approach for end-to-end S2ST, addresseschallenges such as error propagation across modules and slow inference speedoften encountered in traditional cascade systems. However, as discrete unitsprimarily capture content information, conventional S2UT methods fail to retainspeaker-specific characteristics from the source. Our previous work, SC-S2UT,introduced a speaker adapter and a unit-to-mel structure, enabling thepreservation of speaker information and non-autoregressive speech generation.Building on this foundation, this study proposes a self-supervised pretrainingmethod to enrich the information extracted by both the speaker adapter and theunit-to-mel structure. Additionally, we investigate different feature fusionstrategies to further improve the integration of speaker and content features.Experiments conducted on the CVSS-T dataset for ES-EN and FR-EN tasksdemonstrate that our proposed method achieves a BLEU score improvement of 1.14compared to SC-S2UT, along with significant enhancements in MOS and speakersimilarity. Furthermore, our approach achieves translation quality comparableto traditional S2UT, with only a minimal increase of 0.04s per utterance ininference time, while maintaining high speaker similarity. These resultsvalidate the effectiveness of the proposed method.
标题:通过软提示微调对基于LLM的ASB进行有效的文本自适应
链接:https://arxiv.org/abs/2412.06967
备注:accepted as SLT 2024 proceeding
摘要:大语言模型(LLM)的出现对自动语音识别(ASR)产生了巨大的影响.将LLM与音频嵌入结合以生成transmittance成为新的最先进的ASR。尽管LLM使用大量的文本语料库进行训练,但高质量的特定于领域的文本数据仍然可以显着提高领域适应任务的ASR性能。虽然基于LLM的ASR可以通过微调LLM解码器自然地并入更多的文本语料库,但是在没有配对提示的情况下对仅文本数据微调这样的ASR可能会降低特定于领域的知识的有效性。为了缓解这个问题,我们提出了一个两步软提示微调策略,提高特定领域的文本适应。实验结果表明,与基线ASR相比,我们提出的方法的文本自适应在目标域上实现了相对高达9%的字错误率(WER)降低和高达18%的实体错误率(EER)降低。将其与特定领域语言模型(LM)融合相结合,可以进一步将EER提高相对2-5%
摘要:The advent of Large Language Models (LLM) has reformed the Automatic SpeechRecognition (ASR). Prompting LLM with audio embeddings to generatetranscriptions becomes the new state-of-the-art ASR. Despite LLMs being trainedwith an extensive amount of text corpora, high-quality domain-specific textdata can still significantly enhance ASR performance on domain adaptationtasks. Although LLM-based ASR can naturally incorporate more text corpora byfine-tuning the LLM decoder, fine-tuning such ASR on text-only data withoutpaired prompts may diminish the effectiveness of domain-specific knowledge. Tomitigate this issue, we propose a two-step soft prompt fine-tuning strategythat enhances domain-specific text adaptation. Experimental results show thattext adaptation with our proposed method achieved a relative up to 9% WordError Rate (WER) reduction and up to 18% Entity Error Rate (EER) reduction onthe target domain compared to the baseline ASR. Combining this withdomain-specific Language Model (LM) fusion can further improve the EER by arelative 2-5%
标题:利用扩散和一致性模型改进源提取
链接:https://arxiv.org/abs/2412.06965
摘要:在这项工作中,我们展示了集成的分数匹配扩散模型到一个确定性的架构,时域音乐源提取,从而提高音频质量。为了解决扩散模型通常缓慢的迭代采样过程,我们应用一致性蒸馏,并将采样过程减少到一个步骤,实现与扩散模型相当的性能,并且具有两个或更多步骤,甚至超过它们。在四种乐器(低音、鼓、吉他和钢琴)的Slakh2100数据集上进行训练后,我们的模型显示出与基线方法相比在客观指标上的显着改进。在https://consistency-separation.github.io/上可以找到正确的例子。
摘要:In this work, we demonstrate the integration of a score-matching diffusionmodel into a deterministic architecture for time-domain musical sourceextraction, resulting in enhanced audio quality. To address the typically slowiterative sampling process of diffusion models, we apply consistencydistillation and reduce the sampling process to a single step, achievingperformance comparable to that of diffusion models, and with two or more steps,even surpassing them. Trained on the Slakh2100 dataset for four instruments(bass, drums, guitar, and piano), our model shows significant improvementsacross objective metrics compared to baseline methods. Sound examples areavailable at https://consistency-separation.github.io/.
