本文经arXiv每日学术速递授权转载
【1】 ChordSync: Conformer-Based Alignment of Chord Annotations to Music Audio
标题: Chord同步:基于协调器的Chord注释与音乐音频的对齐
作者:Andrea Poltronieri,Valentina Presutti,Martín Rocamora
Journal-ref:Sound and Music Computing Conference (SMC2024)
链接:点击下载PDF文件
【2】 Towards Explainable and Interpretable Musical Difficulty Estimation: A Parameter-efficient Approach
标题: 迈向可解释和可解释的音乐难度估计:参数高效的方法
作者:Pedro Ramoneda,Vsevolod Eremenko,Alexandre D'Hooge,Emilia Parada-Cabaleiro,Xavier Serra
链接:点击下载PDF文件
【3】 DiM-Gesture: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2 framework
标题: DiM-Gesture:采用自适应层规范化Mamba-2框架的联合语音手势生成
作者:Fan Zhang,Naye Ji,Fuxing Gao,Bozuo Zhao,Jingmei Wu,Yanbing Jiang,Hui Du,Zhenqing Ye,Jiayang Zhu,WeiFan Zhong,Leyao Yan,Xiaomeng Ma
备注:10 pages,10 figures. arXiv admin note: text overlap with arXiv:2403.10805
链接:点击下载PDF文件
【4】 Interaural time difference loss for binaural target sound extraction
标题: 双耳目标声音提取的耳间时差损失
作者:Carlos Hernandez-Olivan,Marc Delcroix,Tsubasa Ochiai,Naohiro Tawara,Tomohiro Nakatani,Shoko Araki
备注:Accepted in the International Workshop on Acoustic Signal Enhancement (IWAENC 2024)
链接:点击下载PDF文件
【5】 Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
标题: 模糊语音情感识别的迭代原型细化
作者:Haoqin Sun,Shiwan Zhao,Xiangyu Kong,Xuechen Wang,Hui Wang,Jiaming Zhou,Yong Qin
链接:点击下载PDF文件
【6】 Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
标题: Bailing-TTC:面向类人自发表示的汉语方言语音合成
作者:Xinhan Di,Zihao Chen,Yunming Liang,Junjie Zheng,Yihua Wang,Chaofan Ding
备注:8 pages, 2 figures
链接:点击下载PDF文件
【7】 Combining audio control and style transfer using latent diffusion
标题: 使用潜在扩散将音频控制和风格转移结合起来
作者:Nils Demerlé,Philippe Esling,Guillaume Doras,David Genova
Journal-ref:Proceedings of the 25th Int. Society for Music Information Retrieval Conference, San Francisco, United States, 2024
链接:点击下载PDF文件
【8】 Towards a Universal Method for Meaningful Signal Detection
标题: 迈向有意义信号检测的通用方法
作者:Louis Mahon
链接:点击下载PDF文件
【9】 Framework for Curating Speech Datasets and Evaluating ASR Systems: A Case Study for Polish
标题: 策展语音数据集和评估ASB系统的框架:波兰语案例研究
作者:Michał Junczyk
备注:Submitted to NeurIPS 2024 Datasets and Benchmarks Track
链接:点击下载PDF文件
标题: 对使用到达时间差自定位的自本地化的担忧
作者:Faxian Cao
备注:2 pages
链接:点击下载PDF文件
【2】 SynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation via Language Model and Synthetic Data
标题: SynesLM:通过语言模型和合成数据进行视听语音识别和翻译的统一方法
作者:Yichen Lu,Jiaqi Song,Xuankai Chang,Hengwei Bian,Soumi Maiti,Shinji Watanabe
链接:点击下载PDF文件
【3】 Long-Term Conversation Analysis: Privacy-Utility Trade-off under Noise and Reverberation
标题: 长期对话分析:噪音和回响下的隐私与公用事业权衡
作者:Jule Pohlhausen,Francesco Nespoli,Joerg Bitzer
备注:Accepted for publication at IWAENC 2024
链接:点击下载PDF文件
【4】 Towards a Universal Method for Meaningful Signal Detection
标题: 迈向有意义信号检测的通用方法
作者:Louis Mahon
链接:点击下载PDF文件
【5】 Framework for Curating Speech Datasets and Evaluating ASR Systems: A Case Study for Polish
标题: 策展语音数据集和评估ASB系统的框架:波兰语案例研究
作者:Michał Junczyk
备注:Submitted to NeurIPS 2024 Datasets and Benchmarks Track
链接:点击下载PDF文件
【6】 Handling Numeric Expressions in Automatic Speech Recognition
标题: 自动语音识别中数字表达的处理
作者:Christian Huber,Alexander Waibel
链接:点击下载PDF文件
【7】 ChordSync: Conformer-Based Alignment of Chord Annotations to Music Audio
标题: Chord同步:基于协调器的Chord注释与音乐音频的对齐
作者:Andrea Poltronieri,Valentina Presutti,Martín Rocamora
Journal-ref:Sound and Music Computing Conference (SMC2024)
链接:点击下载PDF文件
【8】 Towards Explainable and Interpretable Musical Difficulty Estimation: A Parameter-efficient Approach
标题: 迈向可解释和可解释的音乐难度估计:参数高效的方法
作者:Pedro Ramoneda,Vsevolod Eremenko,Alexandre D'Hooge,Emilia Parada-Cabaleiro,Xavier Serra
链接:点击下载PDF文件
【9】 Interaural time difference loss for binaural target sound extraction
标题: 双耳目标声音提取的耳间时差损失
作者:Carlos Hernandez-Olivan,Marc Delcroix,Tsubasa Ochiai,Naohiro Tawara,Tomohiro Nakatani,Shoko Araki
备注:Accepted in the International Workshop on Acoustic Signal Enhancement (IWAENC 2024)
链接:点击下载PDF文件
【10】 Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
标题: 模糊语音情感识别的迭代原型细化
作者:Haoqin Sun,Shiwan Zhao,Xiangyu Kong,Xuechen Wang,Hui Wang,Jiaming Zhou,Yong Qin
链接:点击下载PDF文件
【11】 Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
标题: Bailing-TTC:面向类人自发表示的汉语方言语音合成
作者:Xinhan Di,Zihao Chen,Yunming Liang,Junjie Zheng,Yihua Wang,Chaofan Ding
备注:8 pages, 2 figures
链接:点击下载PDF文件
【12】 Sentence-wise Speech Summarization: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation
标题: 逐句语音摘要:任务、数据集和使用LM知识提炼的端到端建模
作者:Kohei Matsuura,Takanori Ashihara,Takafumi Moriya,Masato Mimura,Takatomo Kano,Atsunori Ogawa,Marc Delcroix
备注:Accepted to Interspeech2024. Dataset: this https URL
链接:点击下载PDF文件
【13】 Combining audio control and style transfer using latent diffusion
标题: 使用潜在扩散将音频控制和风格转移结合起来
作者:Nils Demerlé,Philippe Esling,Guillaume Doras,David Genova
Journal-ref:Proceedings of the 25th Int. Society for Music Information Retrieval Conference, San Francisco, United States, 2024
链接:点击下载PDF文件
标题: Chord同步:基于协调器的Chord注释与音乐音频的对齐
作者:Andrea Poltronieri,Valentina Presutti,Martín Rocamora
Journal-ref:Sound and Music Computing Conference (SMC2024)
链接:点击下载PDF文件
摘要:在西方音乐传统中,和弦是和声的主要组成部分,和声是音乐的基本维度。尽管它与几个音乐信息检索(MIR)任务相关,但和弦注释的音频数据集是有限的,需要更多的多样性。改进这些资源的一种方法是利用在线提供的大量和弦注释,但这需要将它们与音乐音频对齐。然而,现有的音频到乐谱对齐技术(通常依赖于动态时间规整(DTW))未能解决这一挑战,因为它们需要弱对齐的数据来进行精确同步。在本文中,我们介绍了ChordSync,一种新的基于一致性的模型,旨在无缝对齐和弦注释与音频,消除了弱对齐的需要。我们还提供了一个预先训练的模型和一个用户友好的库,使用户能够轻松地同步和弦注释与音轨。通过这种方式,ChordSync为利用众包和弦数据进行MIR创造了机会,特别是在音频和弦估计中,从而促进了新数据集的生成。此外,我们的系统将其实用性扩展到音乐教育,通过提供准确对齐的注释来增强音乐学习体验,从而使学习者能够参与同步的音乐实践。摘要:In the Western music tradition, chords are the main constituent components of harmony, a fundamental dimension of music. Despite its relevance for several Music Information Retrieval (MIR) tasks, chord-annotated audio datasets are limited and need more diversity. One way to improve those resources is to leverage the large number of chord annotations available online, but this requires aligning them with music audio. However, existing audio-to-score alignment techniques, which typically rely on Dynamic Time Warping (DTW), fail to address this challenge, as they require weakly aligned data for precise synchronisation. In this paper, we introduce ChordSync, a novel conformer-based model designed to seamlessly align chord annotations with audio, eliminating the need for weak alignment. We also provide a pre-trained model and a user-friendly library, enabling users to synchronise chord annotations with audio tracks effortlessly. In this way, ChordSync creates opportunities for harnessing crowd-sourced chord data for MIR, especially in audio chord estimation, thereby facilitating the generation of novel datasets. Additionally, our system extends its utility to music education, enhancing music learning experiences by providing accurately aligned annotations, thus enabling learners to engage in synchronised musical practices.
【2】 Towards Explainable and Interpretable Musical Difficulty Estimation: A Parameter-efficient Approach
标题: 迈向可解释和可解释的音乐难度估计:参数高效的方法
作者:Pedro Ramoneda,Vsevolod Eremenko,Alexandre D'Hooge,Emilia Parada-Cabaleiro,Xavier Serra
链接:点击下载PDF文件
摘要:估计音乐作品难度对于组织教育音乐收藏是重要的。这一过程可以部分自动化,以促进教育者的作用。然而,流行的深度学习模型所做的决定很难理解,这可能会影响音乐教育课程对这种技术的接受。我们的工作采用可解释的描述符,在符号音乐表示的难度估计。此外,通过一个新的参数高效的白盒模型,我们超越了以前的努力,同时提供可解释的结果。这些可理解的结果模仿了在音乐教育中广泛使用的一种工具--量规的功能。我们的方法,在钢琴曲目分为9个类进行评估,独立达到41.4%的准确率,均方误差(MSE)为1.7,显示精确的难度估计。通过我们的基线,我们说明了如何建立在过去的研究可以提供替代音乐难度评估是可解释和可解释的。有了这个,我们的目标是促进音乐信息检索(MIR)社区和音乐教育之间更有效的沟通。摘要:Estimating music piece difficulty is important for organizing educational music collections. This process could be partially automatized to facilitate the educator's role. Nevertheless, the decisions performed by prevalent deep-learning models are hardly understandable, which may impair the acceptance of such a technology in music education curricula. Our work employs explainable descriptors for difficulty estimation in symbolic music representations. Furthermore, through a novel parameter-efficient white-box model, we outperform previous efforts while delivering interpretable results. These comprehensible outcomes emulate the functionality of a rubric, a tool widely used in music education. Our approach, evaluated in piano repertoire categorized in 9 classes, achieved 41.4% accuracy independently, with a mean squared error (MSE) of 1.7, showing precise difficulty estimation. Through our baseline, we illustrate how building on top of past research can offer alternatives for music difficulty assessment which are explainable and interpretable. With this, we aim to promote a more effective communication between the Music Information Retrieval (MIR) community and the music education one.
【3】 DiM-Gesture: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2 framework
标题: DiM-Gesture:采用自适应层规范化Mamba-2框架的联合语音手势生成
作者:Fan Zhang,Naye Ji,Fuxing Gao,Bozuo Zhao,Jingmei Wu,Yanbing Jiang,Hui Du,Zhenqing Ye,Jiayang Zhu,WeiFan Zhong,Leyao Yan,Xiaomeng Ma
备注:10 pages,10 figures. arXiv admin note: text overlap with arXiv:2403.10805
链接:点击下载PDF文件
摘要:语音驱动的手势生成是虚拟人创建中的一个新兴领域,其中当前的方法主要利用基于transformer的架构,该架构需要大量的存储器,并且其特征在于推理速度慢。为了应对这些限制,我们提出了 textit{DiM-Gestures},这是一种新颖的端到端生成模型,可以仅从原始语音音频中创建高度个性化的3D全身手势,采用基于Mamba的架构。该模型将基于Mamba的模糊特征提取器与非自回归自适应层归一化(AdaLN)Mamba-2扩散架构相集成。该提取器利用Mamba框架和WavLM预训练模型,自主导出隐式连续模糊特征,然后将其统一为奇异潜在特征。此功能由AdaLN Mamba-2处理,它在所有令牌中实现了统一的条件机制,以鲁棒地对模糊特征和所得手势序列之间的相互作用进行建模。这种创新的方法保证了手势语音同步的高保真度,同时保持手势的自然性。采用扩散模型进行训练和推理,我们的框架在ZEGGS和BEAT数据集上进行了广泛的主观和客观评估。这些评估证实了我们的模型相对于当代最先进的方法的增强性能,展示了DiTs架构(Persona-Gestors)的竞争结果,同时优化了内存使用并加快了推理速度。摘要:Speech-driven gesture generation is an emerging domain within virtual human creation, where current methods predominantly utilize Transformer-based architectures that necessitate extensive memory and are characterized by slow inference speeds. In response to these limitations, we propose textit{DiM-Gestures}, a novel end-to-end generative model crafted to create highly personalized 3D full-body gestures solely from raw speech audio, employing Mamba-based architectures. This model integrates a Mamba-based fuzzy feature extractor with a non-autoregressive Adaptive Layer Normalization (AdaLN) Mamba-2 diffusion architecture. The extractor, leveraging a Mamba framework and a WavLM pre-trained model, autonomously derives implicit, continuous fuzzy features, which are then unified into a singular latent feature. This feature is processed by the AdaLN Mamba-2, which implements a uniform conditional mechanism across all tokens to robustly model the interplay between the fuzzy features and the resultant gesture sequence. This innovative approach guarantees high fidelity in gesture-speech synchronization while maintaining the naturalness of the gestures. Employing a diffusion model for training and inference, our framework has undergone extensive subjective and objective evaluations on the ZEGGS and BEAT datasets. These assessments substantiate our model's enhanced performance relative to contemporary state-of-the-art methods, demonstrating competitive outcomes with the DiTs architecture (Persona-Gestors) while optimizing memory usage and accelerating inference speed.
【4】 Interaural time difference loss for binaural target sound extraction
标题: 双耳目标声音提取的耳间时差损失
作者:Carlos Hernandez-Olivan,Marc Delcroix,Tsubasa Ochiai,Naohiro Tawara,Tomohiro Nakatani,Shoko Araki
备注:Accepted in the International Workshop on Acoustic Signal Enhancement (IWAENC 2024)
链接:点击下载PDF文件
摘要:双耳目标声提取(TSE)的目的是从任意声音的双耳混合中提取出所需的声音,同时保留所需声音的空间线索。实际上,对于许多应用,目标声音信号及其空间线索携带关于声源的重要信息。双耳TSE可以用神经网络来实现,该神经网络被训练为在给定双耳混合的情况下仅输出期望的声音,并且嵌入表征作为输入的期望的声音类别。传统的TSE系统是使用信号电平损失来训练的,信号电平损失测量左声道和右声道的提取信号与参考信号之间的差异。在本文中,我们建议增加明确的空间损失,以更好地保持目标声音的空间线索。特别是,我们探讨旨在保留耳间水平(ILD),相位(IPD)和时间差(ITD)的损失。我们的实验表明,增加这样的空间损失,特别是我们新提出的ITD损失,有助于保持更好的空间线索,同时保持信号水平的指标。摘要:Binaural target sound extraction (TSE) aims to extract a desired sound from a binaural mixture of arbitrary sounds while preserving the spatial cues of the desired sound. Indeed, for many applications, the target sound signal and its spatial cues carry important information about the sound source. Binaural TSE can be realized with a neural network trained to output only the desired sound given a binaural mixture and an embedding characterizing the desired sound class as inputs. Conventional TSE systems are trained using signal-level losses, which measure the difference between the extracted and reference signals for the left and right channels. In this paper, we propose adding explicit spatial losses to better preserve the spatial cues of the target sound. In particular, we explore losses aiming at preserving the interaural level (ILD), phase (IPD), and time differences (ITD). We show experimentally that adding such spatial losses, particularly our newly proposed ITD loss, helps preserve better spatial cues while maintaining the signal-level metrics.
【5】 Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
标题: 模糊语音情感识别的迭代原型细化
作者:Haoqin Sun,Shiwan Zhao,Xiangyu Kong,Xuechen Wang,Hui Wang,Jiaming Zhou,Yong Qin
链接:点击下载PDF文件
摘要:由于表情的微妙性和模糊性,从语音中识别情感是一项艰巨的任务。传统的语音情感识别(SER)系统,通常依赖于一个单一的,精确的情感标签,挣扎与这种复杂性。因此,对情感的内在模糊性进行建模是一个迫切的问题。在本文中,我们提出了一个迭代的原型精化框架(IPR)模糊SER。IPR包括两个相互关联的组件:对比学习和类原型。前者提供了一种有效的方法来获得高质量的表示模糊样本。后者基于模糊标签(模糊数据与所有原型的相似性)进行动态更新。这些细化的嵌入产生精确的伪标签,从而增强了表示质量。在IEMOCAP数据集上进行的实验评估验证了IPR相对于最先进方法的优越性能,从而证明了我们提出的方法的有效性。摘要:Recognizing emotions from speech is a daunting task due to the subtlety and ambiguity of expressions. Traditional speech emotion recognition (SER) systems, which typically rely on a singular, precise emotion label, struggle with this complexity. Therefore, modeling the inherent ambiguity of emotions is an urgent problem. In this paper, we propose an iterative prototype refinement framework (IPR) for ambiguous SER. IPR comprises two interlinked components: contrastive learning and class prototypes. The former provides an efficient way to obtain high-quality representations of ambiguous samples. The latter are dynamically updated based on ambiguous labels -- the similarity of the ambiguous data to all prototypes. These refined embeddings yield precise pseudo labels, thus reinforcing representation quality. Experimental evaluations conducted on the IEMOCAP dataset validate the superior performance of IPR over state-of-the-art methods, thus proving the effectiveness of our proposed method.
【6】 Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
标题: Bailing-TTC:面向类人自发表示的汉语方言语音合成
作者:Xinhan Di,Zihao Chen,Yunming Liang,Junjie Zheng,Yihua Wang,Chaofan Ding
备注:8 pages, 2 figures
链接:点击下载PDF文件
摘要:近年来,大规模文语转换模型取得了很大的进展,但在汉语方言语音生成方面仍存在不足。为了解决这一问题,我们提出了白灵TTS,一个家庭的大规模TTS模型,能够产生高质量的汉语方言语音。百灵TTS是汉语方言语音生成的基础模型。首先,提出了连续半监督学习,以促进文本标记和语音标记的对齐。其次,使用特定的Transformer架构和多阶段训练过程开发了汉语方言表征学习。基于本文提出的网络结构和相应的策略,百灵文语转换系统能够有效地从文本生成汉语方言语音。实验表明,百灵文语转换系统生成的汉语方言语音具有类人的自发表征。鼓励读者在 url{https: c9412600.github.io bltts_tech_report index.html}上收听演示。摘要:Large-scale text-to-speech (TTS) models have made significant progress recently.However, they still fall short in the generation of Chinese dialectal speech. Toaddress this, we propose Bailing-TTS, a family of large-scale TTS models capable of generating high-quality Chinese dialectal speech. Bailing-TTS serves as a foundation model for Chinese dialectal speech generation. First, continual semi-supervised learning is proposed to facilitate the alignment of text tokens and speech tokens. Second, the Chinese dialectal representation learning is developed using a specific transformer architecture and multi-stage training processes. With the proposed design of novel network architecture and corresponding strategy, Bailing-TTS is able to generate Chinese dialectal speech from text effectively and efficiently. Experiments demonstrate that Bailing-TTS generates Chinese dialectal speech towards human-like spontaneous representation. Readers are encouraged to listen to demos at url{https: c9412600.github.io bltts_tech_report index.html}.
【7】 Combining audio control and style transfer using latent diffusion
标题: 使用潜在扩散将音频控制和风格转移结合起来
作者:Nils Demerlé,Philippe Esling,Guillaume Doras,David Genova
Journal-ref:Proceedings of the 25th Int. Society for Music Information Retrieval Conference, San Francisco, United States, 2024
链接:点击下载PDF文件
摘要:深度生成模型现在能够合成高质量的音频信号,将其开发的关键方面从音频质量转移到控制能力。虽然文本到音乐的生成在很大程度上被大众所采用,但显式控制和基于示例的风格转换是更适合捕捉艺术家和音乐家意图的模式。 在本文中,我们的目标是统一的显式控制和风格转移在一个单一的模型分离的局部和全局信息,分别捕捉音乐结构和音色。为此,我们利用扩散自动编码器的功能来提取语义特征,以构建两个表示空间。我们使用对抗性标准和两阶段训练策略来强制这些空间之间的解开。我们得到的模型可以生成与音色目标匹配的音频,同时使用显式控件或通过另一个音频示例指定结构。我们评估了我们的模型在乐器录音的单次音色传输和MIDI到音频任务上的表现,并表明我们在音频质量和目标保真度方面优于现有的基线。此外,我们表明,我们的方法可以通过将节奏和旋律内容转移到不同类型的目标音频的风格来生成完整音乐作品的封面版本。摘要:Deep generative models are now able to synthesize high-quality audio signals, shifting the critical aspect in their development from audio quality to control capabilities. Although text-to-music generation is getting largely adopted by the general public, explicit control and example-based style transfer are more adequate modalities to capture the intents of artists and musicians. In this paper, we aim to unify explicit control and style transfer within a single model by separating local and global information to capture musical structure and timbre respectively. To do so, we leverage the capabilities of diffusion autoencoders to extract semantic features, in order to build two representation spaces. We enforce disentanglement between those spaces using an adversarial criterion and a two-stage training strategy. Our resulting model can generate audio matching a timbre target, while specifying structure either with explicit controls or through another audio example. We evaluate our model on one-shot timbre transfer and MIDI-to-audio tasks on instrumental recordings and show that we outperform existing baselines in terms of audio quality and target fidelity. Furthermore, we show that our method can generate cover versions of complete musical pieces by transferring rhythmic and melodic content to the style of a target audio in a different genre.
【8】 Towards a Universal Method for Meaningful Signal Detection
标题: 迈向有意义信号检测的通用方法
作者:Louis Mahon
链接:点击下载PDF文件
摘要:众所周知,人类的语言和某些动物的发声可以传达有意义的内容,因为我们可以破译给定话语所传达的内容。本文探讨了另一种方法来确定是否是有意义的信号,一个只分析信号本身,是独立的传达的意义可能是什么。我们设计了一种方法,以波形作为输入,并输出一个分数,表明其程度的“意义”。我们对输入的连续部分进行聚类,以最小化总描述长度,然后将分配的聚类标签的代码长度作为有意义分数。我们根据经验评估我们的方法,对几个基线,并表明它是唯一一个给高分的人类语音在各种语言和各种扬声器,一个中等的分数,从鸟类和虎鲸的动物发声,和一个低分数,从各种来源的环境噪声。摘要:It is known that human speech and certain animal vocalizations can convey meaningful content because we can decipher the content that a given utterance does convey. This paper explores an alternative approach to determining whether a signal is meaningful, one that analyzes only the signal itself and is independent of what the conveyed meaning might be. We devise a method that takes a waveform as input and outputs a score indicating its degree of meaningfulness . We cluster contiguous portions of the input to minimize the total description length, and then take the length of the code of the assigned cluster labels as meaningfulness score. We evaluate our method empirically, against several baselines, and show that it is the only one to give a high score to human speech in various languages and with various speakers, a moderate score to animal vocalizations from birds and orcas, and a low score to ambient noise from various sources.
【9】 Framework for Curating Speech Datasets and Evaluating ASR Systems: A Case Study for Polish
标题: 策展语音数据集和评估ASB系统的框架:波兰语案例研究
作者:Michał Junczyk
备注:Submitted to NeurIPS 2024 Datasets and Benchmarks Track
链接:点击下载PDF文件
摘要:公共领域中可用的语音数据集通常未得到充分利用,因为在可兼容性和互操作性方面存在挑战。已经设计了一个全面的框架来调查,编目和策划可用的语音数据集,这使得自动语音识别(ASR)系统的可复制评估。进行了一项以波兰语为重点的案例研究;该框架被应用于管理超过24个数据集,并评估了ASR系统和模型的25种组合。这项研究是迄今为止对波兰语的商业和免费ASR系统进行的最广泛的比较。它从600个系统模型测试集评估中汲取了见解,标志着规模和全面性的重大进步。调查和性能比较的结果可通过交互式仪表板(https: huggingface.co spaces amu-cai pl-asr-leaderboard)、精选数据集(https: huggingface.co taskets amu-cai pl-asr-bigos-v2、https: huggingface.co taskets pelcra pl-asr-pelcra-for-bigos)和公开挑战电话(https: poleval.pl tasks task3)获得。用于评价的工具是开源的(https: github.com goodmike31 pl-asr-bigos-tools),便于复制和改编为其他语文,并随着新的数据集和系统不断扩展。摘要:Speech datasets available in the public domain are often underutilized because of challenges in discoverability and interoperability. A comprehensive framework has been designed to survey, catalog, and curate available speech datasets, which allows replicable evaluation of automatic speech recognition (ASR) systems. A case study focused on the Polish language was conducted; the framework was applied to curate more than 24 datasets and evaluate 25 combinations of ASR systems and models. This research constitutes the most extensive comparison to date of both commercial and free ASR systems for the Polish language. It draws insights from 600 system-model-test set evaluations, marking a significant advancement in both scale and comprehensiveness. The results of surveys and performance comparisons are available as interactive dashboards (https: huggingface.co spaces amu-cai pl-asr-leaderboard) along with curated datasets (https: huggingface.co datasets amu-cai pl-asr-bigos-v2, https: huggingface.co datasets pelcra pl-asr-pelcra-for-bigos) and the open challenge call (https: poleval.pl tasks task3). Tools used for evaluation are open-sourced (https: github.com goodmike31 pl-asr-bigos-tools), facilitating replication and adaptation for other languages, as well as continuous expansion with new datasets and systems.
eess.AS音频处理
【1】 Concerns for Self-Localization of Ad-Hoc Arrays Using Time Difference of Arrivals标题: 对使用到达时间差自定位的自本地化的担忧
作者:Faxian Cao
备注:2 pages
链接:点击下载PDF文件
摘要:本文介绍了关于IEEE Transactions on Signal Processing(TSP)上发表的题为“Self-Localization of Ad-Hoc Arrays Using Time Difference of Arrivals”的论文的一些见解和观察。本着建设性反馈的精神,我想强调两个主要考虑领域。第一个方面涉及到的方法,实验结果,并在文件中所作的声明。第二部分解决了具体的公式 印刷错误。这项工作的目的是发起一个建设性的对话有关的某些方面发表在IEEE TSP的文件。我们的目的是提供反馈,有助于不断改进的文件的鲁棒性和清晰度。摘要:This document presents some insights and observations regarding the paper that was published in IEEE Transactions on Signal Processing (TSP), titled "Self-Localization of Ad-Hoc Arrays Using Time Difference of Arrivals". In the spirit of constructive feedback, I wish to highlight two key areas of consideration. The first pertains to aspects related to methodology, experimental results, and statements made in the paper. The second part addresses specific equation typographical errors. This work aims to initiate a constructive dialogue concerning certain aspects of the paper published in IEEE TSP. Our intention is to provide feedback that contributes to the ongoing improvement of the paper's robustness and clarity.
【2】 SynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation via Language Model and Synthetic Data
标题: SynesLM:通过语言模型和合成数据进行视听语音识别和翻译的统一方法
作者:Yichen Lu,Jiaqi Song,Xuankai Chang,Hengwei Bian,Soumi Maiti,Shinji Watanabe
链接:点击下载PDF文件
摘要:在这项工作中,我们提出了SynesLM,一个统一的模型,可以执行三个多模态语言理解任务:视听自动语音识别(AV-ASR)和视觉辅助语音 机器翻译(VST VMT)。与以前的研究不同,我们的研究集中在嘴唇运动作为语音信号的视觉线索,我们的工作探索了整个框架内的更一般的视觉信息,如物体和动作。此外,我们使用合成图像数据来增强图像和语音数据之间的相关性。我们将SynesLM与How2数据集进行基准测试,在保持我们的多任务框架的同时,展示了与专用于AV-ASR的最先进(SOTA)模型相当的性能。值得注意的是,对于zero-shot AV-ASR,SynesLM通过将VisSpeech数据集上的字错误率(WER)从43.4%降低到39.4%,实现了SOTA性能。此外,我们在VST和VMT中的结果优于以前的结果,将BLEU分数从VST的37.2提高到43.5,VMT从54.4提高到54.8。摘要:In this work, we present SynesLM, an unified model which can perform three multimodal language understanding tasks: audio-visual automatic speech recognition(AV-ASR) and visual-aided speech machine translation(VST VMT). Unlike previous research that focused on lip motion as visual cues for speech signals, our work explores more general visual information within entire frames, such as objects and actions. Additionally, we use synthetic image data to enhance the correlation between image and speech data. We benchmark SynesLM against the How2 dataset, demonstrating performance on par with state-of-the-art (SOTA) models dedicated to AV-ASR while maintaining our multitasking framework. Remarkably, for zero-shot AV-ASR, SynesLM achieved SOTA performance by lowering the Word Error Rate (WER) from 43.4% to 39.4% on the VisSpeech Dataset. Furthermore, our results in VST and VMT outperform the previous results, improving the BLEU score to 43.5 from 37.2 for VST, and to 54.8 from 54.4 for VMT.
【3】 Long-Term Conversation Analysis: Privacy-Utility Trade-off under Noise and Reverberation
标题: 长期对话分析:噪音和回响下的隐私与公用事业权衡
作者:Jule Pohlhausen,Francesco Nespoli,Joerg Bitzer
备注:Accepted for publication at IWAENC 2024
链接:点击下载PDF文件
摘要:日常生活中的录音需要对讲话内容和讲话者身份进行隐私保护。这篇文章探讨了噪声和混响对边缘计算可行的低成本隐私保护方法的隐私和效用之间的权衡的影响。这些方法折衷了频谱和时间平滑、使用McAdams系数的说话者匿名化、以非常低的采样率采样以及组合。隐私通过自动语音和说话者识别来评估,而我们的实用程序则考虑语音活动检测和说话者日记。总的来说,我们的评估表明,额外的噪声降低了所有模型的性能比混响。这种退化对应于增强的语音隐私,而对于某些方法,实用性更少地恶化。摘要:Recordings in everyday life require privacy preservation of the speech content and speaker identity. This contribution explores the influence of noise and reverberation on the trade-off between privacy and utility for low-cost privacy-preserving methods feasible for edge computing. These methods compromise spectral and temporal smoothing, speaker anonymization using the McAdams coefficient, sampling with a very low sampling rate, and combinations. Privacy is assessed by automatic speech and speaker recognition, while our utility considers voice activity detection and speaker diarization. Overall, our evaluation shows that additional noise degrades the performance of all models more than reverberation. This degradation corresponds to enhanced speech privacy, while utility is less deteriorated for some methods.
【4】 Towards a Universal Method for Meaningful Signal Detection
标题: 迈向有意义信号检测的通用方法
作者:Louis Mahon
链接:点击下载PDF文件
摘要:众所周知,人类的语言和某些动物的发声可以传达有意义的内容,因为我们可以破译给定话语所传达的内容。本文探讨了另一种方法来确定是否是有意义的信号,一个只分析信号本身,是独立的传达的意义可能是什么。我们设计了一种方法,以波形作为输入,并输出一个分数,表明其程度的“意义”。我们对输入的连续部分进行聚类,以最小化总描述长度,然后将分配的聚类标签的代码长度作为有意义分数。我们根据经验评估我们的方法,对几个基线,并表明它是唯一一个给高分的人类语音在各种语言和各种扬声器,一个中等的分数,从鸟类和虎鲸的动物发声,和一个低分数,从各种来源的环境噪声。摘要:It is known that human speech and certain animal vocalizations can convey meaningful content because we can decipher the content that a given utterance does convey. This paper explores an alternative approach to determining whether a signal is meaningful, one that analyzes only the signal itself and is independent of what the conveyed meaning might be. We devise a method that takes a waveform as input and outputs a score indicating its degree of meaningfulness . We cluster contiguous portions of the input to minimize the total description length, and then take the length of the code of the assigned cluster labels as meaningfulness score. We evaluate our method empirically, against several baselines, and show that it is the only one to give a high score to human speech in various languages and with various speakers, a moderate score to animal vocalizations from birds and orcas, and a low score to ambient noise from various sources.
【5】 Framework for Curating Speech Datasets and Evaluating ASR Systems: A Case Study for Polish
标题: 策展语音数据集和评估ASB系统的框架:波兰语案例研究
作者:Michał Junczyk
备注:Submitted to NeurIPS 2024 Datasets and Benchmarks Track
链接:点击下载PDF文件
摘要:公共领域中可用的语音数据集通常未得到充分利用,因为在可兼容性和互操作性方面存在挑战。已经设计了一个全面的框架来调查,编目和策划可用的语音数据集,这使得自动语音识别(ASR)系统的可复制评估。进行了一项以波兰语为重点的案例研究;该框架被应用于管理超过24个数据集,并评估了ASR系统和模型的25种组合。这项研究是迄今为止对波兰语的商业和免费ASR系统进行的最广泛的比较。它从600个系统模型测试集评估中汲取了见解,标志着规模和全面性的重大进步。调查和性能比较的结果可通过交互式仪表板(https: huggingface.co spaces amu-cai pl-asr-leaderboard)、精选数据集(https: huggingface.co taskets amu-cai pl-asr-bigos-v2、https: huggingface.co taskets pelcra pl-asr-pelcra-for-bigos)和公开挑战电话(https: poleval.pl tasks task3)获得。用于评价的工具是开源的(https: github.com goodmike31 pl-asr-bigos-tools),便于复制和改编为其他语文,并随着新的数据集和系统不断扩展。摘要:Speech datasets available in the public domain are often underutilized because of challenges in discoverability and interoperability. A comprehensive framework has been designed to survey, catalog, and curate available speech datasets, which allows replicable evaluation of automatic speech recognition (ASR) systems. A case study focused on the Polish language was conducted; the framework was applied to curate more than 24 datasets and evaluate 25 combinations of ASR systems and models. This research constitutes the most extensive comparison to date of both commercial and free ASR systems for the Polish language. It draws insights from 600 system-model-test set evaluations, marking a significant advancement in both scale and comprehensiveness. The results of surveys and performance comparisons are available as interactive dashboards (https: huggingface.co spaces amu-cai pl-asr-leaderboard) along with curated datasets (https: huggingface.co datasets amu-cai pl-asr-bigos-v2, https: huggingface.co datasets pelcra pl-asr-pelcra-for-bigos) and the open challenge call (https: poleval.pl tasks task3). Tools used for evaluation are open-sourced (https: github.com goodmike31 pl-asr-bigos-tools), facilitating replication and adaptation for other languages, as well as continuous expansion with new datasets and systems.
【6】 Handling Numeric Expressions in Automatic Speech Recognition
标题: 自动语音识别中数字表达的处理
作者:Christian Huber,Alexander Waibel
链接:点击下载PDF文件
摘要:本文讨论了自动语音识别(ASR)成绩单中正确格式化数字表达式的问题。这是具有挑战性的,因为预期的转录本格式取决于上下文,例如,1945年(年)与19:45(时间戳)。我们比较了级联和端到端的方法来识别和格式化数字表达式,如年份,时间戳,货币金额和数量。对于端到端的方法,我们采用了一个数据生成策略,使用一个大的语言模型(LLM)与文本到语音(TTS)模型来生成自适应数据。我们的测试数据集上的结果表明,虽然基于LLM的方法在识别格式化的数值表达式方面表现良好,但自适应的端到端模型具有较低的延迟和推理成本的优势,具有竞争力的性能。摘要:This paper addresses the problem of correctly formatting numeric expressions in automatic speech recognition (ASR) transcripts. This is challenging since the expected transcript format depends on the context, e.g., 1945 (year) vs. 19:45 (timestamp). We compare cascaded and end-to-end approaches to recognize and format numeric expression, such as years, timestamps, currency amounts, and quantities. For the end-to-end approach we employed a data generation strategy using a large language model (LLM) together with a text to speech (TTS) model to generate adaptation data. The results on our test dataset show that while approaches based on LLMs perform well on recognizing formatted numeric expressions, adapted end-to-end models offer competitive performance with the advantage of lower latency and inference cost.
【7】 ChordSync: Conformer-Based Alignment of Chord Annotations to Music Audio
标题: Chord同步:基于协调器的Chord注释与音乐音频的对齐
作者:Andrea Poltronieri,Valentina Presutti,Martín Rocamora
Journal-ref:Sound and Music Computing Conference (SMC2024)
链接:点击下载PDF文件
摘要:在西方音乐传统中,和弦是和声的主要组成部分,和声是音乐的基本维度。尽管它与几个音乐信息检索(MIR)任务相关,但和弦注释的音频数据集是有限的,需要更多的多样性。改进这些资源的一种方法是利用在线提供的大量和弦注释,但这需要将它们与音乐音频对齐。然而,现有的音频到乐谱对齐技术(通常依赖于动态时间规整(DTW))未能解决这一挑战,因为它们需要弱对齐的数据来进行精确同步。在本文中,我们介绍了ChordSync,一种新的基于一致性的模型,旨在无缝对齐和弦注释与音频,消除了弱对齐的需要。我们还提供了一个预先训练的模型和一个用户友好的库,使用户能够轻松地同步和弦注释与音轨。通过这种方式,ChordSync为利用众包和弦数据进行MIR创造了机会,特别是在音频和弦估计中,从而促进了新数据集的生成。此外,我们的系统将其实用性扩展到音乐教育,通过提供准确对齐的注释来增强音乐学习体验,从而使学习者能够参与同步的音乐实践。摘要:In the Western music tradition, chords are the main constituent components of harmony, a fundamental dimension of music. Despite its relevance for several Music Information Retrieval (MIR) tasks, chord-annotated audio datasets are limited and need more diversity. One way to improve those resources is to leverage the large number of chord annotations available online, but this requires aligning them with music audio. However, existing audio-to-score alignment techniques, which typically rely on Dynamic Time Warping (DTW), fail to address this challenge, as they require weakly aligned data for precise synchronisation. In this paper, we introduce ChordSync, a novel conformer-based model designed to seamlessly align chord annotations with audio, eliminating the need for weak alignment. We also provide a pre-trained model and a user-friendly library, enabling users to synchronise chord annotations with audio tracks effortlessly. In this way, ChordSync creates opportunities for harnessing crowd-sourced chord data for MIR, especially in audio chord estimation, thereby facilitating the generation of novel datasets. Additionally, our system extends its utility to music education, enhancing music learning experiences by providing accurately aligned annotations, thus enabling learners to engage in synchronised musical practices.
【8】 Towards Explainable and Interpretable Musical Difficulty Estimation: A Parameter-efficient Approach
标题: 迈向可解释和可解释的音乐难度估计:参数高效的方法
作者:Pedro Ramoneda,Vsevolod Eremenko,Alexandre D'Hooge,Emilia Parada-Cabaleiro,Xavier Serra
链接:点击下载PDF文件
摘要:估计音乐作品难度对于组织教育音乐收藏是重要的。这一过程可以部分自动化,以促进教育者的作用。然而,流行的深度学习模型所做的决定很难理解,这可能会影响音乐教育课程对这种技术的接受。我们的工作采用可解释的描述符,在符号音乐表示的难度估计。此外,通过一个新的参数高效的白盒模型,我们超越了以前的努力,同时提供可解释的结果。这些可理解的结果模仿了在音乐教育中广泛使用的一种工具--量规的功能。我们的方法,在钢琴曲目分为9个类进行评估,独立达到41.4%的准确率,均方误差(MSE)为1.7,显示精确的难度估计。通过我们的基线,我们说明了如何建立在过去的研究可以提供替代音乐难度评估是可解释和可解释的。有了这个,我们的目标是促进音乐信息检索(MIR)社区和音乐教育之间更有效的沟通。摘要:Estimating music piece difficulty is important for organizing educational music collections. This process could be partially automatized to facilitate the educator's role. Nevertheless, the decisions performed by prevalent deep-learning models are hardly understandable, which may impair the acceptance of such a technology in music education curricula. Our work employs explainable descriptors for difficulty estimation in symbolic music representations. Furthermore, through a novel parameter-efficient white-box model, we outperform previous efforts while delivering interpretable results. These comprehensible outcomes emulate the functionality of a rubric, a tool widely used in music education. Our approach, evaluated in piano repertoire categorized in 9 classes, achieved 41.4% accuracy independently, with a mean squared error (MSE) of 1.7, showing precise difficulty estimation. Through our baseline, we illustrate how building on top of past research can offer alternatives for music difficulty assessment which are explainable and interpretable. With this, we aim to promote a more effective communication between the Music Information Retrieval (MIR) community and the music education one.
【9】 Interaural time difference loss for binaural target sound extraction
标题: 双耳目标声音提取的耳间时差损失
作者:Carlos Hernandez-Olivan,Marc Delcroix,Tsubasa Ochiai,Naohiro Tawara,Tomohiro Nakatani,Shoko Araki
备注:Accepted in the International Workshop on Acoustic Signal Enhancement (IWAENC 2024)
链接:点击下载PDF文件
摘要:双耳目标声提取(TSE)的目的是从任意声音的双耳混合中提取出所需的声音,同时保留所需声音的空间线索。实际上,对于许多应用,目标声音信号及其空间线索携带关于声源的重要信息。双耳TSE可以用神经网络来实现,该神经网络被训练为在给定双耳混合的情况下仅输出期望的声音,并且嵌入表征作为输入的期望的声音类别。传统的TSE系统是使用信号电平损失来训练的,信号电平损失测量左声道和右声道的提取信号与参考信号之间的差异。在本文中,我们建议增加明确的空间损失,以更好地保持目标声音的空间线索。特别是,我们探讨旨在保留耳间水平(ILD),相位(IPD)和时间差(ITD)的损失。我们的实验表明,增加这样的空间损失,特别是我们新提出的ITD损失,有助于保持更好的空间线索,同时保持信号水平的指标。摘要:Binaural target sound extraction (TSE) aims to extract a desired sound from a binaural mixture of arbitrary sounds while preserving the spatial cues of the desired sound. Indeed, for many applications, the target sound signal and its spatial cues carry important information about the sound source. Binaural TSE can be realized with a neural network trained to output only the desired sound given a binaural mixture and an embedding characterizing the desired sound class as inputs. Conventional TSE systems are trained using signal-level losses, which measure the difference between the extracted and reference signals for the left and right channels. In this paper, we propose adding explicit spatial losses to better preserve the spatial cues of the target sound. In particular, we explore losses aiming at preserving the interaural level (ILD), phase (IPD), and time differences (ITD). We show experimentally that adding such spatial losses, particularly our newly proposed ITD loss, helps preserve better spatial cues while maintaining the signal-level metrics.
【10】 Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
标题: 模糊语音情感识别的迭代原型细化
作者:Haoqin Sun,Shiwan Zhao,Xiangyu Kong,Xuechen Wang,Hui Wang,Jiaming Zhou,Yong Qin
链接:点击下载PDF文件
摘要:由于表情的微妙性和模糊性,从语音中识别情感是一项艰巨的任务。传统的语音情感识别(SER)系统,通常依赖于一个单一的,精确的情感标签,挣扎与这种复杂性。因此,对情感的内在模糊性进行建模是一个迫切的问题。在本文中,我们提出了一个迭代的原型精化框架(IPR)模糊SER。IPR包括两个相互关联的组件:对比学习和类原型。前者提供了一种有效的方法来获得高质量的表示模糊样本。后者基于模糊标签(模糊数据与所有原型的相似性)进行动态更新。这些细化的嵌入产生精确的伪标签,从而增强了表示质量。在IEMOCAP数据集上进行的实验评估验证了IPR的优越性能,从而证明了我们所提出的方法的有效性。摘要:Recognizing emotions from speech is a daunting task due to the subtlety and ambiguity of expressions. Traditional speech emotion recognition (SER) systems, which typically rely on a singular, precise emotion label, struggle with this complexity. Therefore, modeling the inherent ambiguity of emotions is an urgent problem. In this paper, we propose an iterative prototype refinement framework (IPR) for ambiguous SER. IPR comprises two interlinked components: contrastive learning and class prototypes. The former provides an efficient way to obtain high-quality representations of ambiguous samples. The latter are dynamically updated based on ambiguous labels -- the similarity of the ambiguous data to all prototypes. These refined embeddings yield precise pseudo labels, thus reinforcing representation quality. Experimental evaluations conducted on the IEMOCAP dataset validate the superior performance of IPR over state-of-the-art methods, thus proving the effectiveness of our proposed method.
【11】 Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
标题: Bailing-TTC:面向类人自发表示的汉语方言语音合成
作者:Xinhan Di,Zihao Chen,Yunming Liang,Junjie Zheng,Yihua Wang,Chaofan Ding
备注:8 pages, 2 figures
链接:点击下载PDF文件
摘要:近年来,大规模文语转换模型取得了很大的进展,但在汉语方言语音生成方面仍存在不足。为了解决这一问题,我们提出了白灵TTS,一个家庭的大规模TTS模型,能够产生高质量的汉语方言语音。百灵TTS是汉语方言语音生成的基础模型。首先,提出了连续半监督学习,以促进文本标记和语音标记的对齐。其次,使用特定的Transformer架构和多阶段训练过程开发了汉语方言表征学习。基于本文提出的网络结构和相应的策略,百灵文语转换系统能够有效地从文本中生成汉语方言语音。实验表明,百灵文语转换系统生成的汉语方言语音具有类人的自发表征。鼓励读者在 url{https: c9412600.github.io bltts_tech_report index.html}上收听演示。摘要:Large-scale text-to-speech (TTS) models have made significant progress recently.However, they still fall short in the generation of Chinese dialectal speech. Toaddress this, we propose Bailing-TTS, a family of large-scale TTS models capable of generating high-quality Chinese dialectal speech. Bailing-TTS serves as a foundation model for Chinese dialectal speech generation. First, continual semi-supervised learning is proposed to facilitate the alignment of text tokens and speech tokens. Second, the Chinese dialectal representation learning is developed using a specific transformer architecture and multi-stage training processes. With the proposed design of novel network architecture and corresponding strategy, Bailing-TTS is able to generate Chinese dialectal speech from text effectively and efficiently. Experiments demonstrate that Bailing-TTS generates Chinese dialectal speech towards human-like spontaneous representation. Readers are encouraged to listen to demos at url{https: c9412600.github.io bltts_tech_report index.html}.
【12】 Sentence-wise Speech Summarization: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation
标题: 逐句语音摘要:任务、数据集和使用LM知识提炼的端到端建模
作者:Kohei Matsuura,Takanori Ashihara,Takafumi Moriya,Masato Mimura,Takatomo Kano,Atsunori Ogawa,Marc Delcroix
备注:Accepted to Interspeech2024. Dataset: this https URL
链接:点击下载PDF文件
摘要:本文介绍了一种新的方法,称为逐句语音摘要(Sen-SSum),它生成文本摘要从一个口语文档中的逐句的方式。Sen-SSUM将自动语音识别(ASR)的实时处理与语音摘要的简洁性相结合。为了探索这种方法,我们为Sen-SSum提供了两个数据集:Mega-SSum和CSJ-SSum。使用这些数据集,我们的研究评估了两种基于Transformer的模型:1)结合ASR和强文本摘要模型的级联模型,以及2)直接将语音转换为文本摘要的端到端(E2 E)模型。虽然E2 E模型对开发计算效率高的模型很有吸引力,但它们的性能比级联模型差。因此,我们提出了知识蒸馏E2 E模型使用的级联模型产生的伪摘要。我们的实验表明,这种提出的知识蒸馏有效地提高了E2 E模型在两个数据集上的性能。摘要:This paper introduces a novel approach called sentence-wise speech summarization (Sen-SSum), which generates text summaries from a spoken document in a sentence-by-sentence manner. Sen-SSum combines the real-time processing of automatic speech recognition (ASR) with the conciseness of speech summarization. To explore this approach, we present two datasets for Sen-SSum: Mega-SSum and CSJ-SSum. Using these datasets, our study evaluates two types of Transformer-based models: 1) cascade models that combine ASR and strong text summarization models, and 2) end-to-end (E2E) models that directly convert speech into a text summary. While E2E models are appealing to develop compute-efficient models, they perform worse than cascade models. Therefore, we propose knowledge distillation for E2E models using pseudo-summaries generated by the cascade models. Our experiments show that this proposed knowledge distillation effectively improves the performance of the E2E model on both datasets.
【13】 Combining audio control and style transfer using latent diffusion
标题: 使用潜在扩散将音频控制和风格转移结合起来
作者:Nils Demerlé,Philippe Esling,Guillaume Doras,David Genova
Journal-ref:Proceedings of the 25th Int. Society for Music Information Retrieval Conference, San Francisco, United States, 2024
链接:点击下载PDF文件
摘要:深度生成模型现在能够合成高质量的音频信号,将其开发的关键方面从音频质量转移到控制能力。虽然文本到音乐的生成在很大程度上被大众所采用,但显式控制和基于示例的风格转换是更适合捕捉艺术家和音乐家意图的模式。 在本文中,我们的目标是统一的显式控制和风格转移在一个单一的模型分离的局部和全局信息,分别捕捉音乐结构和音色。为此,我们利用扩散自动编码器的功能来提取语义特征,以构建两个表示空间。我们使用对抗性标准和两阶段训练策略来强制这些空间之间的解开。我们得到的模型可以生成与音色目标匹配的音频,同时使用显式控件或通过另一个音频示例指定结构。我们评估了我们的模型上的单次音色转移和MIDI到音频任务的乐器录音,并表明我们优于现有的基线在音频质量和目标保真度。此外,我们表明,我们的方法可以通过将节奏和旋律内容转移到不同类型的目标音频的风格来生成完整音乐作品的封面版本。摘要:Deep generative models are now able to synthesize high-quality audio signals, shifting the critical aspect in their development from audio quality to control capabilities. Although text-to-music generation is getting largely adopted by the general public, explicit control and example-based style transfer are more adequate modalities to capture the intents of artists and musicians. In this paper, we aim to unify explicit control and style transfer within a single model by separating local and global information to capture musical structure and timbre respectively. To do so, we leverage the capabilities of diffusion autoencoders to extract semantic features, in order to build two representation spaces. We enforce disentanglement between those spaces using an adversarial criterion and a two-stage training strategy. Our resulting model can generate audio matching a timbre target, while specifying structure either with explicit controls or through another audio example. We evaluate our model on one-shot timbre transfer and MIDI-to-audio tasks on instrumental recordings and show that we outperform existing baselines in terms of audio quality and target fidelity. Furthermore, we show that our method can generate cover versions of complete musical pieces by transferring rhythmic and melodic content to the style of a target audio in a different genre.
机器翻译,仅供参考
