【1】SongComposer: A Large LanguageModel for Lyric and Melody Composition in Song Generation标题:SongComposer:歌词和旋律创作的大型语言模型 链接:https://arxiv.org/abs/2402.17645作者:Shuangrui Ding,Zihan Liu,Xiaoyi Dong,Pan Zhang,Rui Qian,Conghui He,Dahua Lin,Jiaqi Wang备注:project page: this https URL code: this https URL摘要:我们提出SongComposer,一个创新的LLM专为歌曲创作。它可以理解和生成符号歌曲表示的旋律和歌词,通过利用LLM的能力。现有的音乐相关的LLM处理的音乐作为量化的音频信号,而这种隐式编码导致编码效率低,灵活性差。相比之下,我们诉诸于符号化的歌曲表示,这是人类为音乐设计的成熟而有效的方式,并使LLM能够像人类一样明确地创作歌曲。在实践中,我们设计了一种新的元组设计来格式化歌词和旋律中的三个音符属性(音高,持续时间和休止持续时间),这保证了LLM对音乐符号的正确理解,并实现了歌词和旋律之间的精确对齐。为了向LLM传授基本的音乐理解,我们仔细收集了SongCompose-PT,这是一个大规模的歌曲预训练数据集,包括中文或英文的歌词,旋律和成对的歌词旋律。经过充分的预训练后,10 K精心制作的QA对被用于为LLM提供预防跟踪能力并解决各种任务。通过大量的实验,SongComposer在歌词到旋律生成,旋律到歌词生成,歌曲延续和文本到歌曲创作方面表现出色,优于GPT-4等高级LLM。摘要:We present SongComposer, an innovative LLM designed for song composition. It could understand and generate melodies and lyrics in symbolic song representations, by leveraging the capability of LLM. Existing music-related LLM treated the music as quantized audio signals, while such implicit encoding leads to inefficient encoding and poor flexibility. In contrast, we resort to symbolic song representation, the mature and efficient way humans designed for music, and enable LLM to explicitly compose songs like humans. In practice, we design a novel tuple design to format lyric and three note attributes (pitch, duration, and rest duration) in the melody, which guarantees the correct LLM understanding of musical symbols and realizes precise alignment between lyrics and melody. To impart basic music understanding to LLM, we carefully collected SongCompose-PT, a large-scale song pretraining dataset that includes lyrics, melodies, and paired lyrics-melodies in either Chinese or English. After adequate pre-training, 10K carefully crafted QA pairs are used to empower the LLM with the instruction-following capability and solve diverse tasks. With extensive experiments, SongComposer demonstrates superior performance in lyric-to-melody generation, melody-to-lyric generation, song continuation, and text-to-song creation, outperforming advanced LLMs like GPT-4.
【2】 Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages标题:情绪语音信息(EMOVOME)数据库:自发语音信息中的情绪识别链接:https://arxiv.org/abs/2402.17496作者:Lucía Gómez Zaragozá,Rocío del Amor,Elena Parra Vargas,Valery Naranjo,Mariano Alcañiz Raya,Javier Marín-Morales备注:10 pages, 6 figures, submitted to Scientific Data摘要:情感语音信息(EMOVOME)是一个自发的语音数据集,包含来自100名西班牙语使用者的999条语音信息。在招募参与者之前,语音信息是在野外条件下产生的,避免了由于实验室环境而产生的任何有意识的偏见。由三名非专家和两名专家在效价和唤醒维度上对音频进行标记,然后将其组合以获得每个维度的最终标签。专家们还提供了一个额外的标签,对应于七种情绪类别。为了使用EMOVOME为未来的调查设定基线,我们使用语音和音频transmittance实现了情感识别模型。对于语音,我们使用标准的eGeMAPS特征集和支持向量机,分别获得49.27%和44.71%的未加权准确率效价和唤醒。对于文本,我们微调了多语言BERT模型,并分别实现了61.15%和47.43%的未加权准确率效价和唤醒。该数据库将为野外情绪识别研究做出重大贡献,同时也为西班牙语提供了一个独特的自然和免费访问的资源。摘要:Emotional Voice Messages (EMOVOME) is a spontaneous speech dataset containing 999 audio messages from real conversations on a messaging app from 100 Spanish speakers, gender balanced. Voice messages were produced in-the-wild conditions before participants were recruited, avoiding any conscious bias due to laboratory environment. Audios were labeled in valence and arousal dimensions by three non-experts and two experts, which were then combined to obtain a final label per dimension. The experts also provided an extra label corresponding to seven emotion categories. To set a baseline for future investigations using EMOVOME, we implemented emotion recognition models using both speech and audio transcriptions. For speech, we used the standard eGeMAPS feature set and support vector machines, obtaining 49.27% and 44.71% unweighted accuracy for valence and arousal respectively. For text, we fine-tuned a multilingual BERT model and achieved 61.15% and 47.43% unweighted accuracy for valence and arousal respectively. This database will significantly contribute to research on emotion recognition in the wild, while also providing a unique natural and freely accessible resource for Spanish.
【3】 Automated Classification of Phonetic Segments in Child Speech Using Raw Ultrasound Imaging标题:基于原始超声成像的儿童语音段自动分类链接:https://arxiv.org/abs/2402.17482作者:Saja Al Ani,Joanne Cleland,Ahmed Zoha备注:None摘要:言语声音障碍(speech sound disorder,SSD)是指言语发声的持续性障碍,导致言语清晰度降低和言语交流障碍。对患有SSD的儿童进行早期识别和干预,并及时转介给言语和语言治疗师(SLT)进行治疗至关重要。语音障碍的自动检测被认为是一种有效的方法,检查和筛选大规模的人群。这项研究的重点是通过提出一种将超声舌成像(UTI)与深度学习模型相结合的技术解决方案来推进儿童早期SSD的自动诊断。引入FusionNet模型结合UTI数据和提取的纹理特征对UTI进行分类。总体目标是提高UTI分析的准确性和效率,特别是对与SSD相关的语音进行分类。该研究将FusionNet方法与标准深度学习方法进行了比较,突出了FusionNet模型在UTI分类中的出色改进效果以及多学习在改善语音治疗诊所中UTI分类方面的潜力。摘要:Speech sound disorder (SSD) is defined as a persistent impairment in speech sound production leading to reduced speech intelligibility and hindered verbal communication. Early recognition and intervention of children with SSD and timely referral to speech and language therapists (SLTs) for treatment are crucial. Automated detection of speech impairment is regarded as an efficient method for examining and screening large populations. This study focuses on advancing the automatic diagnosis of SSD in early childhood by proposing a technical solution that integrates ultrasound tongue imaging (UTI) with deep-learning models. The introduced FusionNet model combines UTI data with the extracted texture features to classify UTI. The overarching aim is to elevate the accuracy and efficiency of UTI analysis, particularly for classifying speech sounds associated with SSD. This study compared the FusionNet approach with standard deep-learning methodologies, highlighting the excellent improvement results of the FusionNet model in UTI classification and the potential of multi-learning in improving UTI classification in speech therapy clinics.
【4】 Natural Language Processing Methods for Symbolic Music Generation and Information Retrieval: a Survey标题:用于符号音乐生成和信息检索的自然语言处理方法综述链接:https://arxiv.org/abs/2402.17467作者:Dinh-Viet-Toan Le,Louis Bigo,Mikaela Keller,Dorien Herremans备注:36 pages, 5 figures, 4 tables摘要:自从Transformers模型在自然语言处理(NLP)领域取得突破性进展以来,在各个领域都出现了对该模型的一些改进。这种趋势已经蔓延到音乐信息检索(MIR)领域,包括对音乐数据处理的研究。然而,利用NLP工具处理符号音乐数据的做法在MIR中并不新鲜。音乐经常被比作语言,因为它们有几个相似之处,包括文本和音乐的顺序表示。这些类比也反映在MIR和NLP中的类似任务中。本文从两个方面综述了自然语言处理方法在符号音乐生成和信息检索研究中的应用。首先,我们提出了一个概述的表征符号音乐改编自自然语言的顺序表示。这种表现是通过考虑象征性音乐的特殊性而设计的。然后这些表示由模型处理。这些模型可能最初是为文本开发的,并适用于象征性音乐,在各种任务上进行训练。我们通过不同的棱镜来描述这些模型,特别是深度学习模型,突出音乐专用机制。最后,我们提出了一个讨论围绕符号音乐数据的NLP工具的有效使用。这包括有关NLP方法的技术问题以及文本和音乐之间的根本差异,这可能为进一步研究更有效地使NLP工具适应符号MIR打开几扇门。摘要:Several adaptations of Transformers models have been developed in various domains since its breakthrough in Natural Language Processing (NLP). This trend has spread into the field of Music Information Retrieval (MIR), including studies processing music data. However, the practice of leveraging NLP tools for symbolic music data is not novel in MIR. Music has been frequently compared to language, as they share several similarities, including sequential representations of text and music. These analogies are also reflected through similar tasks in MIR and NLP. This survey reviews NLP methods applied to symbolic music generation and information retrieval studies following two axes. We first propose an overview of representations of symbolic music adapted from natural language sequential representations. Such representations are designed by considering the specificities of symbolic music. These representations are then processed by models. Such models, possibly originally developed for text and adapted for symbolic music, are trained on various tasks. We describe these models, in particular deep learning models, through different prisms, highlighting music-specialized mechanisms. We finally present a discussion surrounding the effective use of NLP tools for symbolic music data. This includes technical issues regarding NLP methods and fundamental differences between text and music, which may open several doors for further research into more effectively adapting NLP tools to symbolic MIR. 【5】 EDTC: enhance depth of text comprehension in automated audio captioning标题:EDTC:增强自动音频字幕的文本理解深度链接:https://arxiv.org/abs/2402.17259作者:Liwen Tan,Yin Cao,Yi Zhou摘要:模态差异一直是自动音频字幕(AAC)领域和所有多模态领域的重大挑战。促进模型理解文本信息在建立文本和音频两种模式之间的无缝连接方面起着关键作用。虽然最近的研究集中在通过对比学习缩小这两种模式之间的差距,但仅使用简单的对比损失来弥合这两种模式之间的差异是具有挑战性的。本文介绍了增强文本理解深度(EDTC),它从三个不同的角度增强了模型对文本信息的理解。首先,我们提出了一个新的融合模块,FUSER,它的目的是提取共享的语义信息,从不同的音频特征,通过特征融合。然后,我们介绍了TRANSLATOR,这是一种新颖的对齐模块,旨在沿着张量水平对齐音频特征和文本特征。最后,通过向孪生结构添加动量来更新权重,使得模型可以同时学习关于两种模态的信息。由此产生的方法在AudioCaps数据集上实现了最先进的性能,并展示了与Clotho数据集上的最先进性能相当的结果。摘要:Modality discrepancies have perpetually posed significant challenges within the realm of Automated Audio Captioning (AAC) and across all multi-modal domains. Facilitating models in comprehending text information plays a pivotal role in establishing a seamless connection between the two modalities of text and audio. While recent research has focused on closing the gap between these two modalities through contrastive learning, it is challenging to bridge the difference between both modalities using only simple contrastive loss. This paper introduces Enhance Depth of Text Comprehension (EDTC), which enhances the model's understanding of text information from three different perspectives. First, we propose a novel fusion module, FUSER, which aims to extract shared semantic information from different audio features through feature fusion. We then introduced TRANSLATOR, a novel alignment module designed to align audio features and text features along the tensor level. Finally, the weights are updated by adding momentum to the twin structure so that the model can learn information about both modalities at the same time. The resulting method achieves state-of-the-art performance on AudioCaps datasets and demonstrates results comparable to the state-of-the-art on Clotho datasets.
【6】 An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement标题:利用编码器解纠缠的有效混合专家码切换语音识别方法链接:https://arxiv.org/abs/2402.17189作者:Tzu-Ting Yang,Hsin-Wei Wang,Yi-Cheng Wang,Chi-Han Lin,Berlin Chen备注:ICASSP 2024摘要:随着端到端(E2E)神经网络的大规模发展,近年来自动语音识别(ASR)取得了前所未有的突破。然而,由于标注数据的缺乏和语言间的差异,语码转换现象仍然是阻碍ASR性能提高的主要障碍。在本文中,我们专注于改进E2E ASR的声学编码器,以解决由代码切换现象所带来的挑战。我们的主要贡献有三个方面:首先,我们引入了一种新的解纠缠损失,使编码器的较低层能够捕获语言间的声学信息,同时减轻编码器的较高层的语言混乱。其次,通过综合实验,我们验证了我们提出的方法优于现有技术的方法,使用预训练的双编码器,同时只访问代码转换语料库和消耗一半的参数化。第三,编码器输出特征的明显差异也证实了解纠缠损失和专家混合(MoE)架构之间的互补性。摘要:With the massive developments of end-to-end (E2E) neural networks, recent years have witnessed unprecedented breakthroughs in automatic speech recognition (ASR). However, the codeswitching phenomenon remains a major obstacle that hinders ASR from perfection, as the lack of labeled data and the variations between languages often lead to degradation of ASR performance. In this paper, we focus exclusively on improving the acoustic encoder of E2E ASR to tackle the challenge caused by the codeswitching phenomenon. Our main contributions are threefold: First, we introduce a novel disentanglement loss to enable the lower-layer of the encoder to capture inter-lingual acoustic information while mitigating linguistic confusion at the higher-layer of the encoder. Second, through comprehensive experiments, we verify that our proposed method outperforms the prior-art methods using pretrained dual-encoders, meanwhile having access only to the codeswitching corpus and consuming half of the parameterization. Third, the apparent differentiation of the encoders' output features also corroborates the complementarity between the disentanglement loss and the mixture-of-experts (MoE) architecture.
【7】 Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models标题:极限编码器输出帧速率降低:改善大型端到端模型的计算延迟链接:https://arxiv.org/abs/2402.17184作者:Rohit Prabhavalkar,Zhong Meng,Weiran Wang,Adam Stooke,Xingyu Cai,Yanzhang He,Arun Narayanan,Dongseong Hwang,Tara N. Sainath,Pedro J. Moreno备注:Accepted to 2024 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2024)摘要:端到端(E2E)自动语音识别(ASR)模型的准确性随着规模的扩大而不断提高,有些模型的参数已经达到数十亿个。然而,这些模型的广泛部署和采用需要用于解码的计算高效的策略。在目前的工作中,我们研究了这样一种策略:在编码器中应用多个帧缩减层,将编码器输出压缩成少量的输出帧。虽然类似的技术已经在以前的工作中进行了研究,我们实现了显着更多的减少比以前已经证明通过使用多个漏斗减少层。通过消融,我们研究了编码器中各种架构选择的影响,以确定最有效的策略。我们证明,我们可以生成一个编码器输出帧的每2.56秒的输入语音,而不会显着影响词的错误率在一个大规模的语音搜索任务,同时提高编码器和解码器的延迟分别为48%和92%,相对于一个强大的,但计算昂贵的基线。摘要:The accuracy of end-to-end (E2E) automatic speech recognition (ASR) models continues to improve as they are scaled to larger sizes, with some now reaching billions of parameters. Widespread deployment and adoption of these models, however, requires computationally efficient strategies for decoding. In the present work, we study one such strategy: applying multiple frame reduction layers in the encoder to compress encoder outputs into a small number of output frames. While similar techniques have been investigated in previous work, we achieve dramatically more reduction than has previously been demonstrated through the use of multiple funnel reduction layers. Through ablations, we study the impact of various architectural choices in the encoder to identify the most effective strategies. We demonstrate that we can generate one encoder output frame for every 2.56 sec of input speech, without significantly affecting word error rate on a large-scale voice search task, while improving encoder and decoder latencies by 48% and 92% respectively, relative to a strong but computationally expensive baseline. 【8】 Experimental Study: Enhancing Voice Spoofing Detection Models with wav2vec 2.0标题:Wav2vec 2.0增强语音欺骗检测模型的实验研究链接:https://arxiv.org/abs/2402.17127作者:Taein Kang,Soyul Han,Sunmook Choi,Jaejin Seo,Sanghyeok Chung,Seungeun Lee,Seungsang Oh,Il-Youp Kwak备注:5 pages摘要:传统的欺骗检测系统严重依赖于使用从语音数据导出的手工特征。然而,最近出现了一个显着的转变,直接利用原始语音波形,如SincNet滤波器等方法所示。这种转变强调了对更复杂的音频样本功能的需求。此外,深度学习模型的成功,特别是那些使用大型预训练wav2vec 2.0作为特征化前端的模型,突出了精细特征编码器的重要性。作为回应,这项研究评估了wav2vec 2.0作为音频特征提取器的表示能力,通过两个关键调整来修改其预训练的Transformer层的大小:(1)从最左边的一个开始选择一个层子集,(2)从最右边的一个开始微调所选层的一部分。我们用五个欺骗检测后端模型补充了这一分析,主要关注AASIST,使我们能够为选择和微调过程确定最佳配置。与传统的手工制作功能相比,我们的调查确定了几个欺骗检测系统,这些系统在ASVspoof 2019 LA数据集中实现了最先进的性能。这种全面的探索为特征选择策略提供了有价值的见解,推进了欺骗检测领域。摘要:Conventional spoofing detection systems have heavily relied on the use of handcrafted features derived from speech data. However, a notable shift has recently emerged towards the direct utilization of raw speech waveforms, as demonstrated by methods like SincNet filters. This shift underscores the demand for more sophisticated audio sample features. Moreover, the success of deep learning models, particularly those utilizing large pretrained wav2vec 2.0 as a featurization front-end, highlights the importance of refined feature encoders. In response, this research assessed the representational capability of wav2vec 2.0 as an audio feature extractor, modifying the size of its pretrained Transformer layers through two key adjustments: (1) selecting a subset of layers starting from the leftmost one and (2) fine-tuning a portion of the selected layers from the rightmost one. We complemented this analysis with five spoofing detection back-end models, with a primary focus on AASIST, enabling us to pinpoint the optimal configuration for the selection and fine-tuning process. In contrast to conventional handcrafted features, our investigation identified several spoofing detection systems that achieve state-of-the-art performance in the ASVspoof 2019 LA dataset. This comprehensive exploration offers valuable insights into feature selection strategies, advancing the field of spoofing detection.
【9】 What Do Language Models Hear? Probing for Auditory Representations in Language Models标题:语言模型听到的是什么?语言模型中听觉表征的探索链接:https://arxiv.org/abs/2402.16998作者:Jerry Ngo,Yoon Kim摘要:这项工作探讨了语言模型是否编码有意义的接地表示的声音的对象。我们学习了一个线性探测器,该探测器在给定与该对象相关的音频片段的情况下检索该对象的正确文本表示,其中声音表示由预训练的音频模型给出。这种探测器是通过对比损失来训练的,这种对比损失会使物体的语言表征和声音表征彼此接近。在训练之后,测试探测器泛化到训练期间未看到的对象的能力。在不同的语言模型和音频模型中,我们发现探测泛化在许多情况下都是偶然的,这表明尽管只在原始文本上进行训练,但语言模型对某些对象的声音知识进行了编码。摘要:This work explores whether language models encode meaningfully grounded representations of sounds of objects. We learn a linear probe that retrieves the correct text representation of an object given a snippet of audio related to that object, where the sound representation is given by a pretrained audio model. This probe is trained via a contrastive loss that pushes the language representations and sound representations of an object to be close to one another. After training, the probe is tested on its ability to generalize to objects that were not seen during training. Across different language models and audio models, we find that the probe generalization is above chance in many cases, indicating that despite being trained only on raw text, language models encode grounded knowledge of sounds for some objects. 【10】 Towards Decoding Brain Activity During Passive Listening of Speech标题:论被动听音过程中的脑活动解码链接:https://arxiv.org/abs/2402.16996作者:Milán András Fodor,Tamás Gábor Csapó,Frigyes Viktor Arthur备注:27 pages, 7 figures摘要:这项研究的目的是研究言语感知的复杂机制,并最终解码听言语时大脑中产生的电变化。我们尝试使用深度学习方法从颅内脑电图(iEEG)数据中解码听到的语音。我们的目标是帮助脑机接口(BCI)技术的语音合成的进步,并希望,提供一个额外的角度对语音感知的认知过程。这种方法偏离了传统的专注于语音产生,而是选择研究感知语音的神经表征。这个角度开辟了一个复杂的视角,可能使我们能够研究更复杂的神经模式。利用深度学习模型的力量,该研究旨在建立这些复杂的神经活动与相应语音之间的联系。尽管这种方法尚未取得突破性进展,但这项研究揭示了在语音感知过程中解码神经活动的潜力。我们目前的努力可以作为一个基础,我们对扩大和改进这项工作的潜力持乐观态度,以更接近更先进的BCI,更好地理解感知语音的过程及其与口语的关系。摘要:The aim of the study is to investigate the complex mechanisms of speech perception and ultimately decode the electrical changes in the brain accruing while listening to speech. We attempt to decode heard speech from intracranial electroencephalographic (iEEG) data using deep learning methods. The goal is to aid the advancement of brain-computer interface (BCI) technology for speech synthesis, and, hopefully, to provide an additional perspective on the cognitive processes of speech perception. This approach diverges from the conventional focus on speech production and instead chooses to investigate neural representations of perceived speech. This angle opened up a complex perspective, potentially allowing us to study more sophisticated neural patterns. Leveraging the power of deep learning models, the research aimed to establish a connection between these intricate neural activities and the corresponding speech sounds. Despite the approach not having achieved a breakthrough yet, the research sheds light on the potential of decoding neural activity during speech perception. Our current efforts can serve as a foundation, and we are optimistic about the potential of expanding and improving upon this work to move closer towards more advanced BCIs, better understanding of processes underlying perceived speech and its relation to spoken speech. 【11】 The ICASSP 2024 Audio Deep Packet Loss Concealment Challenge标题:ICASSP 2024音频深度丢包隐藏挑战赛链接:https://arxiv.org/abs/2402.16927作者:Lorenz Diener,Solomiya Branets,Ando Saabas,Ross Cutler摘要:音频丢包隐藏是指对网络丢包造成的VoIP音频流中的间隙进行隐藏。通过ICASSP 2024音频深度丢包隐蔽性大挑战赛,我们在INTERSPEECH 2022上举办的音频PLC挑战赛的成功基础上再接再厉。我们在整体较硬的数据集上评估模型,并使用新的ITU-T P.804评估程序来更密切地评估系统在PLC任务上的性能。我们评估了9个系统,其中8个满足严格的实时性能要求的挑战,使用P.804和字的准确性评估。摘要:Audio packet loss concealment is the hiding of gaps in VoIP audio streams caused by network packet loss. With the ICASSP 2024 Audio Deep Packet Loss Concealment Grand Challenge, we build on the success of the previous Audio PLC Challenge held at INTERSPEECH 2022. We evaluate models on an overall harder dataset, and use the new ITU-T P.804 evaluation procedure to more closely evaluate the performance of systems specifically on the PLC task. We evaluate a total of 9 systems, 8 of which satisfy the strict real-time performance requirements of the challenge, using both P.804 and Word Accuracy evaluations.
【12】 High-Fidelity Neural Phonetic Posteriorgrams标题:高保真神经语音后处理图链接:https://arxiv.org/abs/2402.17735作者:Cameron Churchwell,Max Morrison,Bryan Pardo备注:Accepted to ICASSP 2024 Workshop on Explainable Machine Learning for Speech and Audio摘要:语音后验图(PPG)是语音的声学单元上的时变分类分布(例如,音素)。PPG是语音生成中的流行表示,这是由于其能够将发音特征与说话者身份分离,从而允许发音的准确重建(例如,语音转换)和粗粒度发音编辑(例如,外国口音转换)。在本文中,我们证明了提高质量的PPG产生一个国家的最先进的可解释的PPG表示。我们训练一个现成的语音合成器使用我们的PPG表示,并表明,高品质的PPG产生独立的控制音高和发音。我们进一步展示了PPG的新用途,例如声学发音距离和细粒度发音控制。摘要:A phonetic posteriorgram (PPG) is a time-varying categorical distribution over acoustic units of speech (e.g., phonemes). PPGs are a popular representation in speech generation due to their ability to disentangle pronunciation features from speaker identity, allowing accurate reconstruction of pronunciation (e.g., voice conversion) and coarse-grained pronunciation editing (e.g., foreign accent conversion). In this paper, we demonstrably improve the quality of PPGs to produce a state-of-the-art interpretable PPG representation. We train an off-the-shelf speech synthesizer using our PPG representation and show that high-quality PPGs yield independent control over pitch and pronunciation. We further demonstrate novel uses of PPGs, such as an acoustic pronunciation distance and fine-grained pronunciation control.
【13】 Real-time Low-latency Music Source Separation using Hybrid Spectrogram-TasNet标题:基于混合谱图-TasNet的实时低延迟音乐源分离链接:https://arxiv.org/abs/2402.17701作者:Satvik Venkatesh,Arthur Benilov,Philip Coleman,Frederic Roskam备注:Accepted to ICASSP 2024摘要:近年来,用于音乐混音的深度学习取得了重大进展。然而,很少有人关注这些神经网络如何适应实时低延迟应用,这可能有助于助听器,混音音频流和现场表演。在本文中,我们研究了在文献中针对此用例调整当前分层模型所涉及的各种挑战。随后,受混合Demucs架构的启发,我们提出了混合频谱图时域音频分离网络HS-TasNet,它利用了频谱和波形域的优势。对于23 ms的延迟,HS-TasNet在MusDB测试集上获得了4.65的总体信号失真比(SDR),并在额外的训练数据下增加到5.55。这些结果证明了实时低延迟音乐应用的高效混音的潜力。摘要:There have been significant advances in deep learning for music demixing in recent years. However, there has been little attention given to how these neural networks can be adapted for real-time low-latency applications, which could be helpful for hearing aids, remixing audio streams and live shows. In this paper, we investigate the various challenges involved in adapting current demixing models in the literature for this use case. Subsequently, inspired by the Hybrid Demucs architecture, we propose the Hybrid Spectrogram Time-domain Audio Separation Network HS-TasNet, which utilises the advantages of spectral and waveform domains. For a latency of 23 ms, the HS-TasNet obtains an overall signal-to-distortion ratio (SDR) of 4.65 on the MusDB test set, and increases to 5.55 with additional training data. These results demonstrate the potential of efficient demixing for real-time low-latency music applications.
eess.AS音频处理【1】 High-Fidelity Neural Phonetic Posteriorgrams标题:高保真神经语音后处理图链接:https://arxiv.org/abs/2402.17735作者:Cameron Churchwell,Max Morrison,Bryan Pardo备注:Accepted to ICASSP 2024 Workshop on Explainable Machine Learning for Speech and Audio摘要:语音后验图(PPG)是语音的声学单元上的时变分类分布(例如,音素)。PPG是语音生成中的流行表示,这是由于其能够将发音特征与说话者身份分离,从而允许发音的准确重建(例如,语音转换)和粗粒度发音编辑(例如,外国口音转换)。在本文中,我们证明了提高质量的PPG产生一个国家的最先进的可解释的PPG表示。我们训练一个现成的语音合成器使用我们的PPG表示,并表明,高品质的PPG产生独立的控制音高和发音。我们进一步展示了PPG的新用途,例如声学发音距离和细粒度发音控制。摘要:A phonetic posteriorgram (PPG) is a time-varying categorical distribution over acoustic units of speech (e.g., phonemes). PPGs are a popular representation in speech generation due to their ability to disentangle pronunciation features from speaker identity, allowing accurate reconstruction of pronunciation (e.g., voice conversion) and coarse-grained pronunciation editing (e.g., foreign accent conversion). In this paper, we demonstrably improve the quality of PPGs to produce a state-of-the-art interpretable PPG representation. We train an off-the-shelf speech synthesizer using our PPG representation and show that high-quality PPGs yield independent control over pitch and pronunciation. We further demonstrate novel uses of PPGs, such as an acoustic pronunciation distance and fine-grained pronunciation control.
【2】 Real-time Low-latency Music Source Separation using Hybrid Spectrogram-TasNet标题:基于混合谱图-TasNet的实时低延迟音乐源分离链接:https://arxiv.org/abs/2402.17701作者:Satvik Venkatesh,Arthur Benilov,Philip Coleman,Frederic Roskam备注:Accepted to ICASSP 2024摘要:近年来,用于音乐混音的深度学习取得了重大进展。然而,很少有人关注这些神经网络如何适应实时低延迟应用,这可能有助于助听器,混音音频流和现场表演。在本文中,我们研究了在文献中针对此用例调整当前分层模型所涉及的各种挑战。随后,受混合Demucs架构的启发,我们提出了混合频谱图时域音频分离网络HS-TasNet,它利用了频谱和波形域的优势。对于23 ms的延迟,HS-TasNet在MusDB测试集上获得了4.65的总体信号失真比(SDR),并在额外的训练数据下增加到5.55。这些结果证明了实时低延迟音乐应用的高效混音的潜力。摘要:There have been significant advances in deep learning for music demixing in recent years. However, there has been little attention given to how these neural networks can be adapted for real-time low-latency applications, which could be helpful for hearing aids, remixing audio streams and live shows. In this paper, we investigate the various challenges involved in adapting current demixing models in the literature for this use case. Subsequently, inspired by the Hybrid Demucs architecture, we propose the Hybrid Spectrogram Time-domain Audio Separation Network HS-TasNet, which utilises the advantages of spectral and waveform domains. For a latency of 23 ms, the HS-TasNet obtains an overall signal-to-distortion ratio (SDR) of 4.65 on the MusDB test set, and increases to 5.55 with additional training data. These results demonstrate the potential of efficient demixing for real-time low-latency music applications.
【3】 CLAPSep: Leveraging Contrastive Pre-trained Models for Multi-Modal Query-Conditioned Target Sound Extraction标题:CLAPSep:利用对比预训练模型进行多模态查询条件下的目标声音提取链接:https://arxiv.org/abs/2402.17455作者:Hao Ma,Zhiyuan Peng,Mingjie Shao,Ju Liu,Xu Li,Xixin Wu摘要:通用声音分离(USS)旨在从真实世界的录音中提取任意类型的声音。目标声查询提取是实现USS的有效途径。这种系统由两个组件组成:一个查询网络,将用户查询转换为条件嵌入,以及一个分离网络,根据条件嵌入提取目标声音。现有方法主要存在两个问题:首先,它们需要从头开始训练随机初始化的模型,缺乏对预训练模型的利用,并且需要大量的数据和计算资源来确保模型收敛;其次,现有方法需要联合训练查询网络和分离网络,这往往导致过拟合。为了解决这些问题,我们建立了基于对比语言音频预训练模型(CLAP)的CLAPSep模型。我们通过使用CLAP的预训练文本编码器作为查询网络,并将CLAP的预训练音频编码器权重引入分离网络,以充分利用嵌入在预训练模型中的先验知识来辅助目标声音提取任务。大量的实验结果表明,该方法在保证模型性能和泛化能力的同时,节省了训练资源。此外,我们探索模型的能力,综合利用语言/音频多模态和积极/消极的多价用户查询,提高系统性能,同时提供多样化的应用模式。摘要:Universal sound separation (USS) aims to extract arbitrary types of sounds from real-world sound recordings. Language-queried target sound extraction (TSE) is an effective approach to achieving USS. Such systems consist of two components: a query network that converts user queries into conditional embeddings, and a separation network that extracts the target sound based on conditional embeddings. Existing methods mainly suffer from two issues: firstly, they require training a randomly initialized model from scratch, lacking the utilization of pre-trained models, and substantial data and computational resources are needed to ensure model convergence; secondly, existing methods need to jointly train a query network and a separation network, which tends to lead to overfitting. To address these issues, we build the CLAPSep model based on contrastive language-audio pre-trained model (CLAP). We achieve this by using a pre-trained text encoder of CLAP as the query network and introducing pre-trained audio encoder weights of CLAP into the separation network to fully utilize the prior knowledge embedded in the pre-trained model to assist in target sound extraction tasks. Extensive experimental results demonstrate that the proposed method saves training resources while ensuring the model's performance and generalizability. Additionally, we explore the model's ability to comprehensively utilize language/audio multi-modal and positive/negative multi-valent user queries, enhancing system performance while providing diversified application modes.
【4】 Ambisonics Encoding For Arbitrary Microphone Arrays Incorporating Residual Channels For Binaural Reproduction标题:用于双耳重放的包含剩余声道的任意麦克风阵列的双声编码链接:https://arxiv.org/abs/2402.17362作者:Yhonatan Gayer,Vladimir Tourbabin,Zamir Ben-Hur,Jacob Donley,Boaz Rafaely备注:Accepted for presentation at HSCMA 2024摘要:在快速发展的虚拟和增强现实领域,准确的空间音频捕获和再现至关重要。对于这些应用,高保真度立体声已成为标准格式。然而,用于对来自任意麦克风阵列的高保真度立体声响复制信号进行编码的现有方法面临挑战,诸如由于不规则的阵列配置而导致的误差以及由通常少量的麦克风导致的有限的空间分辨率。为了解决这些限制和挑战,提出了一种用于研究高保真度立体声响复制编码的数学框架,突出了结合全转向函数的重要性,并提供了一种用于预测仅从转向函数编码每个高保真度立体声响复制通道的准确性的新措施。此外,新的残留通道制定补充高保真度立体声响复制通道。几个阵列配置的仿真研究表明,这种方法的双耳误差减少。摘要:In the rapidly evolving fields of virtual and augmented reality, accurate spatial audio capture and reproduction are essential. For these applications, Ambisonics has emerged as a standard format. However, existing methods for encoding Ambisonics signals from arbitrary microphone arrays face challenges, such as errors due to the irregular array configurations and limited spatial resolution resulting from a typically small number of microphones. To address these limitations and challenges, a mathematical framework for studying Ambisonics encoding is presented, highlighting the importance of incorporating the full steering function, and providing a novel measure for predicting the accuracy of encoding each Ambisonics channel from the steering functions alone. Furthermore, novel residual channels are formulated supplementing the Ambisonics channels. A simulation study for several array configurations demonstrates a reduction in binaural error for this approach.
【5】 Target Speaker Extraction by Directly Exploiting Contextual Information in the Time-Frequency Domain标题:直接利用时频域上下文信息的目标说话人提取链接:https://arxiv.org/abs/2402.17146作者:Xue Yang,Changchun Bao,Jing Zhou,Xianhong Chen备注:Accepted by ICASSP 2024摘要:在目标说话人提取中,许多研究依赖于说话人嵌入,这是从目标说话人的登记获得的,并用作指导。然而,单独使用说话人嵌入可能无法充分利用注册中包含的上下文信息。在本文中,我们直接利用这种上下文信息的时间-频率(T-F)域。具体地,登记和混合信号的T-F表示被交互以通过注意机制来计算加权矩阵。这些加权矩阵反映了T-F表示的不同帧之间的相似性,并进一步用于获得注册的一致T-F表示。这些一致的表示作为指导,允许更好地利用上下文信息。此外,该方法在基准数据集上实现了最先进的性能,并在复杂场景中显示了其有效性。摘要:In target speaker extraction, many studies rely on the speaker embedding which is obtained from an enrollment of the target speaker and employed as the guidance. However, solely using speaker embedding may not fully utilize the contextual information contained in the enrollment. In this paper, we directly exploit this contextual information in the time-frequency (T-F) domain. Specifically, the T-F representations of the enrollment and the mixed signal are interacted to compute the weighting matrices through an attention mechanism. These weighting matrices reflect the similarity among different frames of the T-F representations and are further employed to obtain the consistent T-F representations of the enrollment. These consistent representations are served as the guidance, allowing for better exploitation of the contextual information. Furthermore, the proposed method achieves the state-of-the-art performance on the benchmark dataset and shows its effectiveness in the complex scenarios. 【6】 SongComposer: A Large Language Model for Lyric and Melody Composition in Song Generation标题:SongComposer:歌词和旋律创作的大型语言模型链接:https://arxiv.org/abs/2402.17645作者:Shuangrui Ding,Zihan Liu,Xiaoyi Dong,Pan Zhang,Rui Qian,Conghui He,Dahua Lin,Jiaqi Wang备注:project page: this https URL code: this https URL摘要:我们提出SongComposer,一个创新的LLM专为歌曲创作。它可以理解和生成符号歌曲表示的旋律和歌词,通过利用LLM的能力。现有的音乐相关的LLM处理的音乐作为量化的音频信号,而这种隐式编码导致编码效率低,灵活性差。相比之下,我们诉诸于符号化的歌曲表示,这是人类为音乐设计的成熟而有效的方式,并使LLM能够像人类一样明确地创作歌曲。在实践中,我们设计了一种新的元组设计来格式化歌词和旋律中的三个音符属性(音高,持续时间和休止持续时间),这保证了LLM对音乐符号的正确理解,并实现了歌词和旋律之间的精确对齐。为了向LLM传授基本的音乐理解,我们仔细收集了SongCompose-PT,这是一个大规模的歌曲预训练数据集,包括中文或英文的歌词,旋律和成对的歌词旋律。经过充分的预训练后,10 K精心制作的QA对被用于为LLM提供预防跟踪能力并解决各种任务。通过大量的实验,SongComposer在歌词到旋律生成,旋律到歌词生成,歌曲延续和文本到歌曲创作方面表现出色,优于GPT-4等高级LLM。摘要:We present SongComposer, an innovative LLM designed for song composition. It could understand and generate melodies and lyrics in symbolic song representations, by leveraging the capability of LLM. Existing music-related LLM treated the music as quantized audio signals, while such implicit encoding leads to inefficient encoding and poor flexibility. In contrast, we resort to symbolic song representation, the mature and efficient way humans designed for music, and enable LLM to explicitly compose songs like humans. In practice, we design a novel tuple design to format lyric and three note attributes (pitch, duration, and rest duration) in the melody, which guarantees the correct LLM understanding of musical symbols and realizes precise alignment between lyrics and melody. To impart basic music understanding to LLM, we carefully collected SongCompose-PT, a large-scale song pretraining dataset that includes lyrics, melodies, and paired lyrics-melodies in either Chinese or English. After adequate pre-training, 10K carefully crafted QA pairs are used to empower the LLM with the instruction-following capability and solve diverse tasks. With extensive experiments, SongComposer demonstrates superior performance in lyric-to-melody generation, melody-to-lyric generation, song continuation, and text-to-song creation, outperforming advanced LLMs like GPT-4.
【7】 Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages标题:情绪语音信息(EMOVOME)数据库:自发语音信息中的情绪识别链接:https://arxiv.org/abs/2402.17496作者:Lucía Gómez Zaragozá,Rocío del Amor,Elena Parra Vargas,Valery Naranjo,Mariano Alcañiz Raya,Javier Marín-Morales备注:10 pages, 6 figures, submitted to Scientific Data摘要:情感语音信息(EMOVOME)是一个自发的语音数据集,包含来自100名西班牙语使用者的999条语音信息。在招募参与者之前,语音信息是在野外条件下产生的,避免了由于实验室环境而产生的任何有意识的偏见。由三名非专家和两名专家在效价和唤醒维度上对音频进行标记,然后将其组合以获得每个维度的最终标签。专家们还提供了一个额外的标签,对应于七种情绪类别。为了使用EMOVOME为未来的调查设定基线,我们使用语音和音频transmittance实现了情感识别模型。对于语音,我们使用标准的eGeMAPS特征集和支持向量机,分别获得49.27%和44.71%的未加权准确率效价和唤醒。对于文本,我们微调了多语言BERT模型,并分别实现了61.15%和47.43%的未加权准确率效价和唤醒。该数据库将为野外情绪识别研究做出重大贡献,同时也为西班牙语提供了一个独特的自然和免费访问的资源。摘要:Emotional Voice Messages (EMOVOME) is a spontaneous speech dataset containing 999 audio messages from real conversations on a messaging app from 100 Spanish speakers, gender balanced. Voice messages were produced in-the-wild conditions before participants were recruited, avoiding any conscious bias due to laboratory environment. Audios were labeled in valence and arousal dimensions by three non-experts and two experts, which were then combined to obtain a final label per dimension. The experts also provided an extra label corresponding to seven emotion categories. To set a baseline for future investigations using EMOVOME, we implemented emotion recognition models using both speech and audio transcriptions. For speech, we used the standard eGeMAPS feature set and support vector machines, obtaining 49.27% and 44.71% unweighted accuracy for valence and arousal respectively. For text, we fine-tuned a multilingual BERT model and achieved 61.15% and 47.43% unweighted accuracy for valence and arousal respectively. This database will significantly contribute to research on emotion recognition in the wild, while also providing a unique natural and freely accessible resource for Spanish. 【8】 Automated Classification of Phonetic Segments in Child Speech Using Raw Ultrasound Imaging标题:基于原始超声成像的儿童语音段自动分类链接:https://arxiv.org/abs/2402.17482作者:Saja Al Ani,Joanne Cleland,Ahmed Zoha备注:None摘要:言语声音障碍(speech sound disorder,SSD)是指言语发声的持续性障碍,导致言语清晰度降低和言语交流障碍。对患有SSD的儿童进行早期识别和干预,并及时转介给言语和语言治疗师(SLT)进行治疗至关重要。语音障碍的自动检测被认为是一种有效的方法,检查和筛选大规模的人群。这项研究的重点是通过提出一种将超声舌成像(UTI)与深度学习模型相结合的技术解决方案来推进儿童早期SSD的自动诊断。引入FusionNet模型结合UTI数据和提取的纹理特征对UTI进行分类。总体目标是提高UTI分析的准确性和效率,特别是对与SSD相关的语音进行分类。该研究将FusionNet方法与标准深度学习方法进行了比较,突出了FusionNet模型在UTI分类中的出色改进效果以及多学习在改善语音治疗诊所中UTI分类方面的潜力。摘要:Speech sound disorder (SSD) is defined as a persistent impairment in speech sound production leading to reduced speech intelligibility and hindered verbal communication. Early recognition and intervention of children with SSD and timely referral to speech and language therapists (SLTs) for treatment are crucial. Automated detection of speech impairment is regarded as an efficient method for examining and screening large populations. This study focuses on advancing the automatic diagnosis of SSD in early childhood by proposing a technical solution that integrates ultrasound tongue imaging (UTI) with deep-learning models. The introduced FusionNet model combines UTI data with the extracted texture features to classify UTI. The overarching aim is to elevate the accuracy and efficiency of UTI analysis, particularly for classifying speech sounds associated with SSD. This study compared the FusionNet approach with standard deep-learning methodologies, highlighting the excellent improvement results of the FusionNet model in UTI classification and the potential of multi-learning in improving UTI classification in speech therapy clinics.
【9】 Natural Language Processing Methods for Symbolic Music Generation and Information Retrieval: a Survey标题:用于符号音乐生成和信息检索的自然语言处理方法综述链接:https://arxiv.org/abs/2402.17467作者:Dinh-Viet-Toan Le,Louis Bigo,Mikaela Keller,Dorien Herremans备注:36 pages, 5 figures, 4 tables摘要:自从Transformers模型在自然语言处理(NLP)领域取得突破性进展以来,在各个领域都出现了对该模型的一些改进。这种趋势已经蔓延到音乐信息检索(MIR)领域,包括对音乐数据处理的研究。然而,利用NLP工具处理符号音乐数据的做法在MIR中并不新鲜。音乐经常被比作语言,因为它们有几个相似之处,包括文本和音乐的顺序表示。这些类比也反映在MIR和NLP中的类似任务中。本文从两个方面综述了自然语言处理方法在符号音乐生成和信息检索研究中的应用。首先,我们提出了一个概述的表征符号音乐改编自自然语言的顺序表示。这种表现是通过考虑象征性音乐的特殊性而设计的。然后这些表示由模型处理。这些模型可能最初是为文本开发的,并适用于象征性音乐,在各种任务上进行训练。我们通过不同的棱镜来描述这些模型,特别是深度学习模型,突出音乐专用机制。最后,我们提出了一个讨论围绕符号音乐数据的NLP工具的有效使用。这包括有关NLP方法的技术问题以及文本和音乐之间的根本差异,这可能为进一步研究更有效地使NLP工具适应符号MIR打开几扇门。摘要:Several adaptations of Transformers models have been developed in various domains since its breakthrough in Natural Language Processing (NLP). This trend has spread into the field of Music Information Retrieval (MIR), including studies processing music data. However, the practice of leveraging NLP tools for symbolic music data is not novel in MIR. Music has been frequently compared to language, as they share several similarities, including sequential representations of text and music. These analogies are also reflected through similar tasks in MIR and NLP. This survey reviews NLP methods applied to symbolic music generation and information retrieval studies following two axes. We first propose an overview of representations of symbolic music adapted from natural language sequential representations. Such representations are designed by considering the specificities of symbolic music. These representations are then processed by models. Such models, possibly originally developed for text and adapted for symbolic music, are trained on various tasks. We describe these models, in particular deep learning models, through different prisms, highlighting music-specialized mechanisms. We finally present a discussion surrounding the effective use of NLP tools for symbolic music data. This includes technical issues regarding NLP methods and fundamental differences between text and music, which may open several doors for further research into more effectively adapting NLP tools to symbolic MIR.
【10】 EDTC: enhance depth of text comprehension in automated audio captioning标题:EDTC:在自动音频字幕中增强文本理解深度链接:https://arxiv.org/abs/2402.17259作者:Liwen Tan,Yin Cao,Yi Zhou摘要:模态差异一直是自动音频字幕(AAC)领域和所有多模态领域的重大挑战。促进模型理解文本信息在建立文本和音频两种模式之间的无缝连接方面起着关键作用。虽然最近的研究集中在通过对比学习缩小这两种模式之间的差距,但仅使用简单的对比损失来弥合这两种模式之间的差异是具有挑战性的。本文介绍了增强文本理解深度(EDTC),它从三个不同的角度增强了模型对文本信息的理解。首先,我们提出了一个新的融合模块,FUSER,它的目的是提取共享的语义信息,从不同的音频特征,通过特征融合。然后,我们介绍了TRANSLATOR,这是一种新颖的对齐模块,旨在沿着张量水平对齐音频特征和文本特征。最后,通过向孪生结构添加动量来更新权重,使得模型可以同时学习关于两种模态的信息。由此产生的方法在AudioCaps数据集上实现了最先进的性能,并展示了与Clotho数据集上的最先进性能相当的结果。摘要:Modality discrepancies have perpetually posed significant challenges within the realm of Automated Audio Captioning (AAC) and across all multi-modal domains. Facilitating models in comprehending text information plays a pivotal role in establishing a seamless connection between the two modalities of text and audio. While recent research has focused on closing the gap between these two modalities through contrastive learning, it is challenging to bridge the difference between both modalities using only simple contrastive loss. This paper introduces Enhance Depth of Text Comprehension (EDTC), which enhances the model's understanding of text information from three different perspectives. First, we propose a novel fusion module, FUSER, which aims to extract shared semantic information from different audio features through feature fusion. We then introduced TRANSLATOR, a novel alignment module designed to align audio features and text features along the tensor level. Finally, the weights are updated by adding momentum to the twin structure so that the model can learn information about both modalities at the same time. The resulting method achieves state-of-the-art performance on AudioCaps datasets and demonstrates results comparable to the state-of-the-art on Clotho datasets.
【11】 An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement标题:一种有效的编码解缠码切换码语音识别的混合专家方法链接:https://arxiv.org/abs/2402.17189作者:Tzu-Ting Yang,Hsin-Wei Wang,Yi-Cheng Wang,Chi-Han Lin,Berlin Chen备注:ICASSP 2024摘要:随着端到端(E2E)神经网络的大规模发展,近年来自动语音识别(ASR)取得了前所未有的突破。然而,由于标注数据的缺乏和语言间的差异,语码转换现象仍然是阻碍ASR性能提高的主要障碍。在本文中,我们专注于改进E2E ASR的声学编码器,以解决由代码切换现象所带来的挑战。我们的主要贡献有三个方面:首先,我们引入了一种新的解纠缠损失,使编码器的较低层能够捕获语言间的声学信息,同时减轻编码器的较高层的语言混乱。其次,通过综合实验,我们验证了我们提出的方法优于现有技术的方法,使用预训练的双编码器,同时只访问代码转换语料库和消耗一半的参数化。第三,编码器输出特征的明显差异也证实了解纠缠损失和专家混合(MoE)架构之间的互补性。摘要:With the massive developments of end-to-end (E2E) neural networks, recent years have witnessed unprecedented breakthroughs in automatic speech recognition (ASR). However, the codeswitching phenomenon remains a major obstacle that hinders ASR from perfection, as the lack of labeled data and the variations between languages often lead to degradation of ASR performance. In this paper, we focus exclusively on improving the acoustic encoder of E2E ASR to tackle the challenge caused by the codeswitching phenomenon. Our main contributions are threefold: First, we introduce a novel disentanglement loss to enable the lower-layer of the encoder to capture inter-lingual acoustic information while mitigating linguistic confusion at the higher-layer of the encoder. Second, through comprehensive experiments, we verify that our proposed method outperforms the prior-art methods using pretrained dual-encoders, meanwhile having access only to the codeswitching corpus and consuming half of the parameterization. Third, the apparent differentiation of the encoders' output features also corroborates the complementarity between the disentanglement loss and the mixture-of-experts (MoE) architecture.
【12】 Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models标题:极限编码器输出帧速率降低:改善大型端到端模型的计算延迟链接:https://arxiv.org/abs/2402.17184作者:Rohit Prabhavalkar,Zhong Meng,Weiran Wang,Adam Stooke,Xingyu Cai,Yanzhang He,Arun Narayanan,Dongseong Hwang,Tara N. Sainath,Pedro J. Moreno备注:Accepted to 2024 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2024)摘要:端到端(E2E)自动语音识别(ASR)模型的准确性随着规模的扩大而不断提高,有些模型的参数已经达到数十亿个。然而,这些模型的广泛部署和采用需要用于解码的计算高效的策略。在目前的工作中,我们研究了这样一种策略:在编码器中应用多个帧缩减层,将编码器输出压缩成少量的输出帧。虽然类似的技术已经在以前的工作中进行了研究,我们实现了显着更多的减少比以前已经证明通过使用多个漏斗减少层。通过消融,我们研究了编码器中各种架构选择的影响,以确定最有效的策略。我们证明,我们可以生成一个编码器输出帧的每2.56秒的输入语音,而不会显着影响词的错误率在一个大规模的语音搜索任务,同时提高编码器和解码器的延迟分别为48%和92%,相对于一个强大的,但计算昂贵的基线。摘要:The accuracy of end-to-end (E2E) automatic speech recognition (ASR) models continues to improve as they are scaled to larger sizes, with some now reaching billions of parameters. Widespread deployment and adoption of these models, however, requires computationally efficient strategies for decoding. In the present work, we study one such strategy: applying multiple frame reduction layers in the encoder to compress encoder outputs into a small number of output frames. While similar techniques have been investigated in previous work, we achieve dramatically more reduction than has previously been demonstrated through the use of multiple funnel reduction layers. Through ablations, we study the impact of various architectural choices in the encoder to identify the most effective strategies. We demonstrate that we can generate one encoder output frame for every 2.56 sec of input speech, without significantly affecting word error rate on a large-scale voice search task, while improving encoder and decoder latencies by 48% and 92% respectively, relative to a strong but computationally expensive baseline.
【13】 Experimental Study: Enhancing Voice Spoofing Detection Models with wav2vec 2.0标题:Wav2vec 2.0增强语音欺骗检测模型的实验研究链接:https://arxiv.org/abs/2402.17127作者:Taein Kang,Soyul Han,Sunmook Choi,Jaejin Seo,Sanghyeok Chung,Seungeun Lee,Seungsang Oh,Il-Youp Kwak备注:5 pages摘要:传统的欺骗检测系统严重依赖于使用从语音数据导出的手工特征。然而,最近出现了一个显着的转变,直接利用原始语音波形,如SincNet滤波器等方法所示。这种转变强调了对更复杂的音频样本功能的需求。此外,深度学习模型的成功,特别是那些使用大型预训练wav2vec 2.0作为特征化前端的模型,突出了精细特征编码器的重要性。作为回应,这项研究评估了wav2vec 2.0作为音频特征提取器的表示能力,通过两个关键调整来修改其预训练的Transformer层的大小:(1)从最左边的一个开始选择一个层子集,(2)从最右边的一个开始微调所选层的一部分。我们用五个欺骗检测后端模型补充了这一分析,主要关注AASIST,使我们能够为选择和微调过程确定最佳配置。与传统的手工制作功能相比,我们的调查确定了几个欺骗检测系统,这些系统在ASVspoof 2019 LA数据集中实现了最先进的性能。这种全面的探索为特征选择策略提供了有价值的见解,推进了欺骗检测领域。摘要:Conventional spoofing detection systems have heavily relied on the use of handcrafted features derived from speech data. However, a notable shift has recently emerged towards the direct utilization of raw speech waveforms, as demonstrated by methods like SincNet filters. This shift underscores the demand for more sophisticated audio sample features. Moreover, the success of deep learning models, particularly those utilizing large pretrained wav2vec 2.0 as a featurization front-end, highlights the importance of refined feature encoders. In response, this research assessed the representational capability of wav2vec 2.0 as an audio feature extractor, modifying the size of its pretrained Transformer layers through two key adjustments: (1) selecting a subset of layers starting from the leftmost one and (2) fine-tuning a portion of the selected layers from the rightmost one. We complemented this analysis with five spoofing detection back-end models, with a primary focus on AASIST, enabling us to pinpoint the optimal configuration for the selection and fine-tuning process. In contrast to conventional handcrafted features, our investigation identified several spoofing detection systems that achieve state-of-the-art performance in the ASVspoof 2019 LA dataset. This comprehensive exploration offers valuable insights into feature selection strategies, advancing the field of spoofing detection.
【14】 What Do Language Models Hear? Probing for Auditory Representations in Language Models标题:语言模型听到的是什么?语言模型中听觉表征的探索链接:https://arxiv.org/abs/2402.16998作者:Jerry Ngo,Yoon Kim摘要:这项工作探讨了语言模型是否编码有意义的接地表示的声音的对象。我们学习了一个线性探测器,该探测器在给定与该对象相关的音频片段的情况下检索该对象的正确文本表示,其中声音表示由预训练的音频模型给出。这种探测器是通过对比损失来训练的,这种对比损失会使物体的语言表征和声音表征彼此接近。在训练之后,测试探测器泛化到训练期间未看到的对象的能力。在不同的语言模型和音频模型中,我们发现探测泛化在许多情况下都是偶然的,这表明尽管只在原始文本上进行训练,但语言模型对某些对象的声音知识进行了编码。摘要:This work explores whether language models encode meaningfully grounded representations of sounds of objects. We learn a linear probe that retrieves the correct text representation of an object given a snippet of audio related to that object, where the sound representation is given by a pretrained audio model. This probe is trained via a contrastive loss that pushes the language representations and sound representations of an object to be close to one another. After training, the probe is tested on its ability to generalize to objects that were not seen during training. Across different language models and audio models, we find that the probe generalization is above chance in many cases, indicating that despite being trained only on raw text, language models encode grounded knowledge of sounds for some objects. 【15】 Towards Decoding Brain Activity During Passive Listening of Speech标题:论被动听音过程中的脑活动解码链接:https://arxiv.org/abs/2402.16996作者:Milán András Fodor,Tamás Gábor Csapó,Frigyes Viktor Arthur备注:27 pages, 7 figures摘要:这项研究的目的是研究言语感知的复杂机制,并最终解码听言语时大脑中产生的电变化。我们尝试使用深度学习方法从颅内脑电图(iEEG)数据中解码听到的语音。我们的目标是帮助脑机接口(BCI)技术的语音合成的进步,并希望,提供一个额外的角度对语音感知的认知过程。这种方法偏离了传统的专注于语音产生,而是选择研究感知语音的神经表征。这个角度开辟了一个复杂的视角,可能使我们能够研究更复杂的神经模式。利用深度学习模型的力量,该研究旨在建立这些复杂的神经活动与相应语音之间的联系。尽管这种方法尚未取得突破性进展,但这项研究揭示了在语音感知过程中解码神经活动的潜力。我们目前的努力可以作为一个基础,我们对扩大和改进这项工作的潜力持乐观态度,以更接近更先进的BCI,更好地理解感知语音的过程及其与口语的关系。摘要:The aim of the study is to investigate the complex mechanisms of speech perception and ultimately decode the electrical changes in the brain accruing while listening to speech. We attempt to decode heard speech from intracranial electroencephalographic (iEEG) data using deep learning methods. The goal is to aid the advancement of brain-computer interface (BCI) technology for speech synthesis, and, hopefully, to provide an additional perspective on the cognitive processes of speech perception. This approach diverges from the conventional focus on speech production and instead chooses to investigate neural representations of perceived speech. This angle opened up a complex perspective, potentially allowing us to study more sophisticated neural patterns. Leveraging the power of deep learning models, the research aimed to establish a connection between these intricate neural activities and the corresponding speech sounds. Despite the approach not having achieved a breakthrough yet, the research sheds light on the potential of decoding neural activity during speech perception. Our current efforts can serve as a foundation, and we are optimistic about the potential of expanding and improving upon this work to move closer towards more advanced BCIs, better understanding of processes underlying perceived speech and its relation to spoken speech.
【16】 The ICASSP 2024 Audio Deep Packet Loss Concealment Challenge标题:ICASSP 2024音频深度丢包隐藏挑战赛链接:https://arxiv.org/abs/2402.16927作者:Lorenz Diener,Solomiya Branets,Ando Saabas,Ross Cutler摘要:音频丢包隐藏是指对网络丢包造成的VoIP音频流中的间隙进行隐藏。通过ICASSP 2024音频深度丢包隐蔽性大挑战赛,我们在INTERSPEECH 2022上举办的音频PLC挑战赛的成功基础上再接再厉。我们在整体较硬的数据集上评估模型,并使用新的ITU-T P.804评估程序来更密切地评估系统在PLC任务上的性能。我们评估了9个系统,其中8个满足严格的实时性能要求的挑战,使用P.804和字的准确性评估。摘要:Audio packet loss concealment is the hiding of gaps in VoIP audio streams caused by network packet loss. With the ICASSP 2024 Audio Deep Packet Loss Concealment Grand Challenge, we build on the success of the previous Audio PLC Challenge held at INTERSPEECH 2022. We evaluate models on an overall harder dataset, and use the new ITU-T P.804 evaluation procedure to more closely evaluate the performance of systems specifically on the PLC task. We evaluate a total of 9 systems, 8 of which satisfy the strict real-time performance requirements of the challenge, using both P.804 and Word Accuracy evaluations. 机器翻译由腾讯交互翻译提供,仅供参考