今日论文合集:cs.SD语音8篇,eess.AS音频处理9篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

cs.SD语音
【1】 Benchmarking Open-ended Audio Dialogue Understanding for Large  Audio-Language Models
标题:大型音频语言模型的开放式音频对话理解基准
链接:https://arxiv.org/abs/2412.05167
作者:Kuofeng Gao,  Shu-Tao Xia,  Ke Xu,  Philip Torr,  Jindong Gu
摘要:大型音频语言模型(LALM)具有非锁定的音频对话功能,其中音频对话是LALM和人类之间的口头语言的直接交换。最近的进展,如GPT-4 o,使LALM能够与人类进行来回的音频对话。这一进展不仅强调了LALM的潜力,而且还扩大了其在音频对话支持的各种实际场景中的适用性。然而,鉴于这些进步,目前仍然缺乏一个全面的基准来评估LALM在开放式音频对话理解中的性能。为了解决这个问题,我们提出了一个音频对话理解基准(ADU-Bench),它由4个基准数据集组成。他们评估了LALM在3种一般场景、12种技能、9种多语言和4种歧义处理类别中的开放式音频对话能力。值得注意的是,我们首先提出了在音频对话中表达不同意图的歧义处理的评估,这些意图超出了句子的相同字面意义,例如,“真的?“用不同的语调总之,ADU-Bench包括20,000多个用于LALM评估的开放式音频对话。通过对13个LALM进行的大量实验,我们的分析表明,现有LALM的音频对话理解能力仍有相当大的改进空间。特别是,他们与数学符号和公式斗争,理解人类行为,如角色扮演,理解多种语言,并处理来自不同语音元素的音频对话歧义,如语调,停顿位置和同音异义词。
摘要:Large Audio-Language Models (LALMs) have unclocked audio dialoguecapabilities, where audio dialogues are a direct exchange of spoken languagebetween LALMs and humans. Recent advances, such as GPT-4o, have enabled LALMsin back-and-forth audio dialogues with humans. This progression not onlyunderscores the potential of LALMs but also broadens their applicability acrossa wide range of practical scenarios supported by audio dialogues. However,given these advancements, a comprehensive benchmark to evaluate the performanceof LALMs in the open-ended audio dialogue understanding remains absentcurrently. To address this gap, we propose an Audio Dialogue UnderstandingBenchmark (ADU-Bench), which consists of 4 benchmark datasets. They assess theopen-ended audio dialogue ability for LALMs in 3 general scenarios, 12 skills,9 multilingual languages, and 4 categories of ambiguity handling. Notably, wefirstly propose the evaluation of ambiguity handling in audio dialogues thatexpresses different intentions beyond the same literal meaning of sentences,e.g., "Really!?" with different intonations. In summary, ADU-Bench includesover 20,000 open-ended audio dialogues for the assessment of LALMs. Throughextensive experiments conducted on 13 LALMs, our analysis reveals that there isstill considerable room for improvement in the audio dialogue understandingabilities of existing LALMs. In particular, they struggle with mathematicalsymbols and formulas, understanding human behavior such as roleplay,comprehending multiple languages, and handling audio dialogue ambiguities fromdifferent phonetic elements, such as intonations, pause positions, andhomophones.

【2】 Applying Automatic Differentiation to Optimize Differential Microphone  Array Designs
标题:应用自动差异优化差异麦克风阵列设计
链接:https://arxiv.org/abs/2412.05123
作者:Siminfar Samakoush Galougah,  Ramani Duraiswami
备注:6 pages, 9 figures
摘要:本文介绍了一种新的方法,利用可微规划设计高效,约束自适应非均匀线性差分麦克风阵列(LDMAs),降低了实施成本。利用自动微分框架,我们提出了一种可微凸的方法,使滤波器的自适应设计与所需的声音方向的无失真约束,同时也施加限制麦克风定位,以确保一致的性能。该方法在宽的频率范围上实现了期望的方向性因子(DF),并且有助于以较低的实现成本有效地恢复宽带语音信号。
摘要:This paper introduces a novel methodology leveraging differentiableprogramming to design efficient, constrained adaptive non-uniform LinearDifferential Microphone Arrays (LDMAs) with reduced implementation costs.Utilizing an automatic differentiation framework, we propose a differentiableconvex approach that enables the adaptive design of a filter with adistortionless constraint in the desired sound direction, while also imposingconstraints on microphone positioning to ensure consistent performance. Thisapproach achieves the desired Directivity Factor (DF) over a wide frequencyrange and facilitates effective recovery of wide-band speech signals at lowerimplementation costs.

【3】 Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
标题:连续语音令牌使LLM成为强大的多模式学习者
链接:https://arxiv.org/abs/2412.04917
作者:Ze Yuan,  Yanqing Liu,  Shujie Liu,  Sheng Zhao
摘要:GPT-40类多模态模型的最新进展已经证明了直接语音到语音对话的显着进步,具有实时语音交互体验和强大的语音理解能力。然而,目前的研究集中在离散语音令牌上,以与离散文本令牌对齐用于语言建模,这取决于具有剩余连接或独立组令牌的音频编解码器,这样的编解码器通常利用大规模和多样化的数据集训练来确保离散语音代码对各种域,噪声,风格的数据重建以及一个精心设计的编解码器量化器和编码器-解码器架构的离散令牌语言建模。本文介绍了Flow-Omni,一个基于连续语音令牌的GPT-4 o模型,能够实时语音交互和低流延迟。具体来说,首先,我们将流匹配损失与预训练的自回归LLM和小型MLP网络相结合,而不是仅使用交叉熵损失来预测语音提示中连续值语音令牌的概率分布。第二,我们将连续语音令牌合并到Flow-Omni多模态训练中,从而实现了离散文本令牌和连续语音令牌一起的鲁棒语音到语音性能。实验表明,与离散文本和语音多模态训练及其变体相比,连续语音令牌通过避免LLM的离散语音代码的表示损失的固有缺陷来减轻鲁棒性问题。
摘要:Recent advances in GPT-4o like multi-modality models have demonstratedremarkable progress for direct speech-to-speech conversation, with real-timespeech interaction experience and strong speech understanding ability. However,current research focuses on discrete speech tokens to align with discrete texttokens for language modelling, which depends on an audio codec with residualconnections or independent group tokens, such a codec usually leverages largescale and diverse datasets training to ensure that the discrete speech codeshave good representation for varied domain, noise, style data reconstruction aswell as a well-designed codec quantizer and encoder-decoder architecture fordiscrete token language modelling. This paper introduces Flow-Omni, acontinuous speech token based GPT-4o like model, capable of real-time speechinteraction and low streaming latency. Specifically, first, instead ofcross-entropy loss only, we combine flow matching loss with a pretrainedautoregressive LLM and a small MLP network to predict the probabilitydistribution of the continuous-valued speech tokens from speech prompt. second,we incorporated the continuous speech tokens to Flow-Omni multi-modalitytraining, thereby achieving robust speech-to-speech performance with discretetext tokens and continuous speech tokens together. Experiments demonstratethat, compared to discrete text and speech multi-modality training and itsvariants, the continuous speech tokens mitigate robustness issues by avoidingthe inherent flaws of discrete speech code's representation loss for LLM.

【4】 Diff4Steer: Steerable Diffusion Prior for Generative Music Retrieval  with Semantic Guidance
标题:迪夫4Steer:具有语义指导的生成音乐检索的可操纵扩散优先级
链接:https://arxiv.org/abs/2412.04746
作者:Xuchan Bao,  Judith Yue Li,  Zhong Yi Wan,  Kun Su,  Timo Denk,  Joonseok Lee,  Dima Kuzmin,  Fei Sha
备注:NeurIPS 2024 Creative AI Track
摘要:现代音乐检索系统通常依赖于用户偏好的固定表示,限制了它们捕捉用户多样且不确定的检索需求的能力。为了解决这一限制,我们介绍了Diff4Steer,一种新的生成检索框架,采用轻量级的扩散模型来合成不同的种子嵌入用户查询,代表音乐探索的潜在方向。与将用户查询映射到嵌入空间中的单个点的确定性方法不同,Diff 4Steer提供了用于检索的目标模态(音频)的统计先验,有效地捕获了用户偏好的不确定性和多方面性质。此外,Diff4Steer可以通过图像或文本输入进行控制,从而实现更灵活和可控的音乐发现与最近邻搜索相结合。我们的框架在检索和排名指标方面优于确定性回归方法和基于LLM的生成检索基线,证明了其在捕获用户偏好方面的有效性,从而提供更多样化和相关的建议。听力示例可在tinyurl.com/diff4steer上找到。
摘要:Modern music retrieval systems often rely on fixed representations of userpreferences, limiting their ability to capture users' diverse and uncertainretrieval needs. To address this limitation, we introduce Diff4Steer, a novelgenerative retrieval framework that employs lightweight diffusion models tosynthesize diverse seed embeddings from user queries that represent potentialdirections for music exploration. Unlike deterministic methods that map userquery to a single point in embedding space, Diff4Steer provides a statisticalprior on the target modality (audio) for retrieval, effectively capturing theuncertainty and multi-faceted nature of user preferences. Furthermore,Diff4Steer can be steered by image or text inputs, enabling more flexible andcontrollable music discovery combined with nearest neighbor search. Ourframework outperforms deterministic regression methods and LLM-based generativeretrieval baseline in terms of retrieval and ranking metrics, demonstrating itseffectiveness in capturing user preferences, leading to more diverse andrelevant recommendations. Listening examples are available attinyurl.com/diff4steer.

【5】 Exploring Transformer-Based Music Overpainting for Jazz Piano Variations
标题:探索爵士钢琴变奏曲中基于变形者的音乐过度绘制
链接:https://arxiv.org/abs/2412.04610
作者:Eleanor Row,  Ivan Shanin,  György Fazekas
备注:Accepted and presented as a Late-Breaking Demo at the 25th International Society for Music Information Retrieval (ISMIR) in San Francisco, US, 2024
摘要:本文探讨了基于transformer的音乐覆盖模型,重点是爵士钢琴变奏曲。音乐覆盖产生新的变化,同时保留输入的旋律和和声结构。现有的方法受到小数据集的限制,限制了可扩展性和多样性。我们引入了VAR 4000,这是一个更大的爵士钢琴演奏数据集的子集,由4,352个训练对组成。使用半自动流水线,我们评估两个Transformer配置的VAR 4000,比较其性能与较小的JAZZVAR数据集。初步结果显示,在更大的数据集配置的泛化和性能有希望的改进,突出了Transformer模型的潜力,以有效地扩展更大,更多样化的数据集上的音乐覆盖。
摘要:This paper explores transformer-based models for music overpainting, focusingon jazz piano variations. Music overpainting generates new variations whilepreserving the melodic and harmonic structure of the input. Existing approachesare limited by small datasets, restricting scalability and diversity. Weintroduce VAR4000, a subset of a larger dataset for jazz piano performances,consisting of 4,352 training pairs. Using a semi-automatic pipeline, weevaluate two transformer configurations on VAR4000, comparing their performancewith the smaller JAZZVAR dataset. Preliminary results show promisingimprovements in generalisation and performance with the larger datasetconfiguration, highlighting the potential of transformer models to scaleeffectively for music overpainting on larger and more diverse datasets.

【6】 MoD-ART: Modal Decomposition of Acoustic Radiance Transfer
标题:MoD-ART:声辐射传递的模式分解
链接:https://arxiv.org/abs/2412.04534
作者:Matteo Scerbo,  Sebastian J. Schlecht,  Randall Ali,  Lauri Savioja,  Enzo De Sena
摘要:当多个声源和听众存在于同一环境中时,以交互速度对后期混响进行建模是一项具有挑战性的任务。当环境在几何上复杂和/或具有不均匀的能量吸收(例如,耦合体积)时,这是特别成问题的,因为在这种情况下,后期混响取决于声源和收听者的位置,因此必须实时地适应他们的运动。我们提出了一种新的方法来完成这项任务,称为声辐射传输模态分解(MoD-ART),它可以高效地处理高度复杂的场景。该方法基于声辐射传递的几何声学方法,从中提取一组能量衰减模式及其与源和收听者的位置关系。在本文中,我们描述了物理和数学意义的MoD-ART,突出其优势和适用于不同的场景。通过分析该方法的计算复杂度,我们表明,它比较非常有利的射线追踪。我们还提出了模拟结果表明,MoD-ART可以捕获多个衰减斜率和颤振回波。
摘要:Modeling late reverberation at interactive speeds is a challenging task whenmultiple sound sources and listeners are present in the same environment. Thisis especially problematic when the environment is geometrically complex and/orfeatures uneven energy absorption (e.g. coupled volumes), because in such casesthe late reverberation is dependent on the sound sources' and listeners'positions, and therefore must be adapted to their movements in real time. Wepresent a novel approach to the task, named modal decomposition of AcousticRadiance Transfer (MoD-ART), which can handle highly complex scenarios withefficiency. The approach is based on the geometrical acoustics method ofAcoustic Radiance Transfer, from which we extract a set of energy decay modesand their positional relationships with sources and listeners. In this paper,we describe the physical and mathematical meaningfulness of MoD-ART,highlighting its advantages and applicability to different scenarios. Throughan analysis of the method's computational complexity, we show that it comparesvery favourably with ray-tracing. We also present simulation results showingthat MoD-ART can capture multiple decay slopes and flutter echoes.

【7】 Perceptually Transparent Binaural Auralization of Simulated Sound Fields
标题:模拟音场的感知透明的双耳可听化
链接:https://arxiv.org/abs/2412.05015
作者:Jens Ahrens
摘要:与空间信息以有形形式可用的基于几何声学的模拟相反,将基于波的模拟可听化并不简单。已经提出了各种方法,其通过对模拟声场的声压或粒子速度(或两者)进行采样来计算具有已知头部相关传递函数的虚拟收听者的耳朵信号。这些方法的可用感知评估结果并不全面,因此不清楚实现感知透明可听化所需的采样点的数量和布置,即用于实现在感知上与地面实况不可区分的可听化。本文提出了一种最常见的双耳可听化方法的感知评估,有和没有中间立体混响表示的体积采样声压或声压和粒子速度采样的球形或立方体表面。我们的结果证实,如果球面网格上的289个采样点处有声压和粒子速度,则感知透明的可听化是可能的。其他栅格几何体需要更多的点。所有测试的方法都可以在本文附带的Chalmers Auralization Toolbox中开源。
摘要:Contrary to geometric acoustics-based simulations where the spatialinformation is available in a tangible form, it is not straightforward toauralize wave-based simulations. A variety of methods have been proposed thatcompute the ear signals of a virtual listener with known head-related transferfunctions from sampling either the sound pressure or the particle velocity (orboth) of the simulated sound field. The available perceptual evaluation resultsof such methods are not comprehensive so that it is unclear what number andarrangement of sampling points is required for achieving perceptuallytransparent auralization, i.e.~for achieving an auralization that isperceptually indistinguishable from the ground truth. This article presents aperceptual evaluation of the most common binaural auralization methods with andwithout intermediate ambisonic representation of volumetrically sampled soundpressure or sound pressure and particle velocity sampled on spherical orcubical surfaces. Our results confirm that perceptually transparentauralization is possible if sound pressure and particle velocity are availableat 289 sampling points on a spherical surface grid. Other grid geometriesrequire considerably more points. All tested methods are available open sourcein the Chalmers Auralization Toolbox that accompanies this article.

【8】 StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional  Flow Matching
标题:StableVC:具有条件流匹配的风格可控Zero-Shot语音转换
链接:https://arxiv.org/abs/2412.04724
作者:Jixun Yao,  Yuguang Yan,  Yu Pan,  Ziqian Ning,  Jiaohao Ye,  Hongbin Zhou,  Lei Xie
摘要:Zero-shot语音转换(VC)的目的是将源说话人的音色转换到一个任意的不可见说话人,同时保留原始的语言内容。尽管最近在使用基于语言模型或基于扩散的方法的zero-shot VC中取得了进步,但仍然存在几个挑战:1)当前方法主要集中于适应来自不可见说话者的音色,并且不能独立地将风格和音色转移到不同的不可见说话者; 2)由于自回归建模方法或需要大量采样步骤,这些方法通常遭受较慢的推断速度;(3)转换后的样本质量和相似度仍不完全令人满意。为了解决这些问题,我们提出了一种风格可控的zero-shot VC方法命名为StableVC,其目的是将音色和风格从源语音转移到不同的看不见的目标说话人。具体来说,我们将语音分解为语言内容,音色和风格,然后采用一个条件流匹配模块来重建高质量的梅尔频谱图的基础上,这些分解的功能。为了以zero-shot的方式有效地捕获音色和风格,我们引入了一种新的具有自适应门的双重注意机制,而不是使用传统的特征级联。通过这种非自回归设计,StableVC可以有效地捕捉来自不同的未见过的扬声器的复杂音色和风格,并以比实时更快的速度生成高质量的语音。实验表明,我们提出的StableVC优于国家的最先进的基线系统在zero-shot VC和实现灵活的控制音色和风格,从不同的看不见的发言人。此外,与自回归和基于扩散的基线相比,StableVC提供的采样速度约为25倍和1.65倍。
摘要:Zero-shot voice conversion (VC) aims to transfer the timbre from the sourcespeaker to an arbitrary unseen speaker while preserving the original linguisticcontent. Despite recent advancements in zero-shot VC using language model-basedor diffusion-based approaches, several challenges remain: 1) current approachesprimarily focus on adapting timbre from unseen speakers and are unable totransfer style and timbre to different unseen speakers independently; 2) theseapproaches often suffer from slower inference speeds due to the autoregressivemodeling methods or the need for numerous sampling steps; 3) the quality andsimilarity of the converted samples are still not fully satisfactory. Toaddress these challenges, we propose a style controllable zero-shot VC approachnamed StableVC, which aims to transfer timbre and style from source speech todifferent unseen target speakers. Specifically, we decompose speech intolinguistic content, timbre, and style, and then employ a conditional flowmatching module to reconstruct the high-quality mel-spectrogram based on thesedecomposed features. To effectively capture timbre and style in a zero-shotmanner, we introduce a novel dual attention mechanism with an adaptive gate,rather than using conventional feature concatenation. With thisnon-autoregressive design, StableVC can efficiently capture the intricatetimbre and style from different unseen speakers and generate high-qualityspeech significantly faster than real-time. Experiments demonstrate that ourproposed StableVC outperforms state-of-the-art baseline systems in zero-shot VCand achieves flexible control over timbre and style from different unseenspeakers. Moreover, StableVC offers approximately 25x and 1.65x faster samplingcompared to autoregressive and diffusion-based baselines.

eess.AS音频处理

【1】 Perceptually Transparent Binaural Auralization of Simulated Sound Fields
标题:模拟音场的感知透明的双耳可听化
链接:https://arxiv.org/abs/2412.05015
作者:Jens Ahrens
摘要:与空间信息以有形形式可用的基于几何声学的模拟相反,将基于波的模拟可听化并不简单。已经提出了各种方法,其通过对模拟声场的声压或粒子速度(或两者)进行采样来计算具有已知头部相关传递函数的虚拟收听者的耳朵信号。这些方法的可用感知评估结果并不全面,因此不清楚实现感知透明可听化所需的采样点的数量和布置,即用于实现在感知上与地面实况不可区分的可听化。本文提出了一种最常见的双耳可听化方法的感知评估,有和没有中间立体混响表示的体积采样声压或声压和粒子速度采样的球形或立方体表面。我们的研究结果证实,感知透明的可听化是可能的,如果声压和粒子速度可在289个采样点上的球面网格。其他栅格几何体需要更多的点。所有测试的方法都可以在本文附带的Chalmers Auralization Toolbox中开源。
摘要:Contrary to geometric acoustics-based simulations where the spatialinformation is available in a tangible form, it is not straightforward toauralize wave-based simulations. A variety of methods have been proposed thatcompute the ear signals of a virtual listener with known head-related transferfunctions from sampling either the sound pressure or the particle velocity (orboth) of the simulated sound field. The available perceptual evaluation resultsof such methods are not comprehensive so that it is unclear what number andarrangement of sampling points is required for achieving perceptuallytransparent auralization, i.e.~for achieving an auralization that isperceptually indistinguishable from the ground truth. This article presents aperceptual evaluation of the most common binaural auralization methods with andwithout intermediate ambisonic representation of volumetrically sampled soundpressure or sound pressure and particle velocity sampled on spherical orcubical surfaces. Our results confirm that perceptually transparentauralization is possible if sound pressure and particle velocity are availableat 289 sampling points on a spherical surface grid. Other grid geometriesrequire considerably more points. All tested methods are available open sourcein the Chalmers Auralization Toolbox that accompanies this article.

【2】 StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional  Flow Matching
标题:StableVC:具有条件流匹配的风格可控Zero-Shot语音转换
链接:https://arxiv.org/abs/2412.04724
作者:Jixun Yao,  Yuguang Yan,  Yu Pan,  Ziqian Ning,  Jiaohao Ye,  Hongbin Zhou,  Lei Xie
摘要:Zero-shot语音转换(VC)的目的是将源说话人的音色转换到一个任意的不可见说话人,同时保留原始的语言内容。尽管最近在使用基于语言模型或基于扩散的方法的zero-shot VC中取得了进步,但仍然存在几个挑战:1)当前方法主要集中于适应来自不可见说话者的音色,并且不能独立地将风格和音色转移到不同的不可见说话者; 2)由于自回归建模方法或需要大量采样步骤,这些方法通常遭受较慢的推断速度;(3)转换后的样本质量和相似度仍不完全令人满意。为了解决这些问题,我们提出了一种风格可控的zero-shot VC方法命名为StableVC,其目的是将音色和风格从源语音转移到不同的看不见的目标说话人。具体来说,我们将语音分解为语言内容,音色和风格,然后采用一个条件流匹配模块来重建高质量的梅尔频谱图的基础上,这些分解的功能。为了以zero-shot的方式有效地捕获音色和风格,我们引入了一种新的具有自适应门的双重注意机制,而不是使用传统的特征级联。通过这种非自回归设计,StableVC可以有效地捕捉来自不同的未见过的扬声器的复杂音色和风格,并以比实时更快的速度生成高质量的语音。实验表明,我们提出的StableVC优于国家的最先进的基线系统在zero-shot VC和实现灵活的控制音色和风格,从不同的看不见的发言人。此外,与自回归和基于扩散的基线相比,StableVC提供了大约25倍和1.65倍的采样速度。
摘要:Zero-shot voice conversion (VC) aims to transfer the timbre from the sourcespeaker to an arbitrary unseen speaker while preserving the original linguisticcontent. Despite recent advancements in zero-shot VC using language model-basedor diffusion-based approaches, several challenges remain: 1) current approachesprimarily focus on adapting timbre from unseen speakers and are unable totransfer style and timbre to different unseen speakers independently; 2) theseapproaches often suffer from slower inference speeds due to the autoregressivemodeling methods or the need for numerous sampling steps; 3) the quality andsimilarity of the converted samples are still not fully satisfactory. Toaddress these challenges, we propose a style controllable zero-shot VC approachnamed StableVC, which aims to transfer timbre and style from source speech todifferent unseen target speakers. Specifically, we decompose speech intolinguistic content, timbre, and style, and then employ a conditional flowmatching module to reconstruct the high-quality mel-spectrogram based on thesedecomposed features. To effectively capture timbre and style in a zero-shotmanner, we introduce a novel dual attention mechanism with an adaptive gate,rather than using conventional feature concatenation. With thisnon-autoregressive design, StableVC can efficiently capture the intricatetimbre and style from different unseen speakers and generate high-qualityspeech significantly faster than real-time. Experiments demonstrate that ourproposed StableVC outperforms state-of-the-art baseline systems in zero-shot VCand achieves flexible control over timbre and style from different unseenspeakers. Moreover, StableVC offers approximately 25x and 1.65x faster samplingcompared to autoregressive and diffusion-based baselines.

【3】 Benchmarking Open-ended Audio Dialogue Understanding for Large  Audio-Language Models
标题:大型音频语言模型的开放式音频对话理解基准
链接:https://arxiv.org/abs/2412.05167
作者:Kuofeng Gao,  Shu-Tao Xia,  Ke Xu,  Philip Torr,  Jindong Gu
摘要:大型音频语言模型(LALM)具有非锁定的音频对话功能,其中音频对话是LALM和人类之间的口头语言的直接交换。最近的进展,如GPT-4 o,使LALM能够与人类进行来回的音频对话。这一进展不仅强调了LALM的潜力,而且还扩大了其在音频对话支持的各种实际场景中的适用性。然而,鉴于这些进步,目前仍然缺乏一个全面的基准来评估LALM在开放式音频对话理解中的性能。为了解决这个问题,我们提出了一个音频对话理解基准(ADU-Bench),它由4个基准数据集组成。他们评估了LALM在3种一般场景、12种技能、9种多语言和4种歧义处理类别中的开放式音频对话能力。值得注意的是,我们首先提出了在音频对话中表达不同意图的歧义处理的评估,这些意图超出了句子的相同字面意义,例如,“真的?“用不同的语调总之,ADU-Bench包括20,000多个用于LALM评估的开放式音频对话。通过对13个LALM进行的大量实验,我们的分析表明,现有LALM的音频对话理解能力仍有相当大的改进空间。特别是,他们与数学符号和公式斗争,理解人类行为,如角色扮演,理解多种语言,并处理来自不同语音元素的音频对话歧义,如语调,停顿位置和同音异义词。
摘要:Large Audio-Language Models (LALMs) have unclocked audio dialoguecapabilities, where audio dialogues are a direct exchange of spoken languagebetween LALMs and humans. Recent advances, such as GPT-4o, have enabled LALMsin back-and-forth audio dialogues with humans. This progression not onlyunderscores the potential of LALMs but also broadens their applicability acrossa wide range of practical scenarios supported by audio dialogues. However,given these advancements, a comprehensive benchmark to evaluate the performanceof LALMs in the open-ended audio dialogue understanding remains absentcurrently. To address this gap, we propose an Audio Dialogue UnderstandingBenchmark (ADU-Bench), which consists of 4 benchmark datasets. They assess theopen-ended audio dialogue ability for LALMs in 3 general scenarios, 12 skills,9 multilingual languages, and 4 categories of ambiguity handling. Notably, wefirstly propose the evaluation of ambiguity handling in audio dialogues thatexpresses different intentions beyond the same literal meaning of sentences,e.g., "Really!?" with different intonations. In summary, ADU-Bench includesover 20,000 open-ended audio dialogues for the assessment of LALMs. Throughextensive experiments conducted on 13 LALMs, our analysis reveals that there isstill considerable room for improvement in the audio dialogue understandingabilities of existing LALMs. In particular, they struggle with mathematicalsymbols and formulas, understanding human behavior such as roleplay,comprehending multiple languages, and handling audio dialogue ambiguities fromdifferent phonetic elements, such as intonations, pause positions, andhomophones.

【4】 Applying Automatic Differentiation to Optimize Differential Microphone  Array Designs
标题:应用自动差异优化差异麦克风阵列设计
链接:https://arxiv.org/abs/2412.05123
作者:Siminfar Samakoush Galougah,  Ramani Duraiswami
备注:6 pages, 9 figures
摘要:本文介绍了一种新的方法,利用可微规划设计高效,约束自适应非均匀线性差分麦克风阵列(LDMAs),降低了实施成本。利用自动微分框架,我们提出了一种可微凸的方法,使滤波器的自适应设计与所需的声音方向的无失真约束,同时也施加限制麦克风定位,以确保一致的性能。该方法在宽的频率范围上实现了期望的方向性因子(DF),并且有助于以较低的实现成本有效地恢复宽带语音信号。
摘要:This paper introduces a novel methodology leveraging differentiableprogramming to design efficient, constrained adaptive non-uniform LinearDifferential Microphone Arrays (LDMAs) with reduced implementation costs.Utilizing an automatic differentiation framework, we propose a differentiableconvex approach that enables the adaptive design of a filter with adistortionless constraint in the desired sound direction, while also imposingconstraints on microphone positioning to ensure consistent performance. Thisapproach achieves the desired Directivity Factor (DF) over a wide frequencyrange and facilitates effective recovery of wide-band speech signals at lowerimplementation costs.

【5】 Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
标题:连续语音令牌使LLM成为强大的多模式学习者
链接:https://arxiv.org/abs/2412.04917
作者:Ze Yuan,  Yanqing Liu,  Shujie Liu,  Sheng Zhao
摘要:GPT-40类多模态模型的最新进展已经证明了直接语音到语音对话的显着进步,具有实时语音交互体验和强大的语音理解能力。然而,目前的研究集中在离散语音令牌上,以与离散文本令牌对齐用于语言建模,这取决于具有剩余连接或独立组令牌的音频编解码器,这样的编解码器通常利用大规模和多样化的数据集训练来确保离散语音代码对各种域,噪声,风格的数据重建以及一个精心设计的编解码器量化器和编码器-解码器架构的离散令牌语言建模。本文介绍了Flow-Omni,一个基于连续语音令牌的GPT-4 o模型,能够实时语音交互和低流延迟。具体来说,首先,我们将流匹配损失与预训练的自回归LLM和小型MLP网络相结合,而不是仅使用交叉熵损失来预测语音提示中连续值语音令牌的概率分布。第二,我们将连续语音令牌合并到Flow-Omni多模态训练中,从而实现了离散文本令牌和连续语音令牌一起的鲁棒语音到语音性能。实验表明,与离散文本和语音多模态训练及其变体相比,连续语音令牌通过避免LLM的离散语音代码的表示损失的固有缺陷来减轻鲁棒性问题。
摘要:Recent advances in GPT-4o like multi-modality models have demonstratedremarkable progress for direct speech-to-speech conversation, with real-timespeech interaction experience and strong speech understanding ability. However,current research focuses on discrete speech tokens to align with discrete texttokens for language modelling, which depends on an audio codec with residualconnections or independent group tokens, such a codec usually leverages largescale and diverse datasets training to ensure that the discrete speech codeshave good representation for varied domain, noise, style data reconstruction aswell as a well-designed codec quantizer and encoder-decoder architecture fordiscrete token language modelling. This paper introduces Flow-Omni, acontinuous speech token based GPT-4o like model, capable of real-time speechinteraction and low streaming latency. Specifically, first, instead ofcross-entropy loss only, we combine flow matching loss with a pretrainedautoregressive LLM and a small MLP network to predict the probabilitydistribution of the continuous-valued speech tokens from speech prompt. second,we incorporated the continuous speech tokens to Flow-Omni multi-modalitytraining, thereby achieving robust speech-to-speech performance with discretetext tokens and continuous speech tokens together. Experiments demonstratethat, compared to discrete text and speech multi-modality training and itsvariants, the continuous speech tokens mitigate robustness issues by avoidingthe inherent flaws of discrete speech code's representation loss for LLM.

【6】 Adaptive Dropout for Pruning Conformers
标题:修剪Conformers的自适应退出
链接:https://arxiv.org/abs/2412.04836
作者:Yotaro Kubo,  Xingyu Cai,  Michiel Bacchiani
摘要:本文提出了一种有效地执行联合训练和修剪的方法,基于自适应丢弃层与单位保留概率。所提出的方法是基于一个单元的保留概率在一个辍学层的估计。估计保留概率较小的单元可被视为不可打印。使用反向传播和Gumbel-Softmax技术估计单元的保留概率。这种修剪方法应用于Conformers中的几个应用点,从而可以显着减少有效的参数数量。具体而言,自适应丢弃层被引入每个Conformer块中的三个位置:(a)前馈网络组件的隐藏层,(b)自注意组件的查询向量和值向量,以及(c)LConv组件的输入向量。通过在LibriSpeech任务上进行语音识别实验来评估所提出的方法。结果表明,这种方法可以同时实现参数减少和精度提高。单词错误率提高了约1%,同时参数数量减少了54%。
摘要:This paper proposes a method to effectively perform jointtraining-and-pruning based on adaptive dropout layers with unit-wise retentionprobabilities. The proposed method is based on the estimation of a unit-wiseretention probability in a dropout layer. A unit that is estimated to have asmall retention probability can be considered to be prunable. The retentionprobability of the unit is estimated using back-propagation and theGumbel-Softmax technique. This pruning method is applied at several applicationpoints in Conformers such that the effective number of parameters can besignificantly reduced. Specifically, adaptive dropout layers are introduced inthree locations in each Conformer block: (a) the hidden layer of thefeed-forward-net component, (b) the query vectors and the value vectors of theself-attention component, and (c) the input vectors of the LConv component. Theproposed method is evaluated by conducting a speech recognition experiment onthe LibriSpeech task. It was shown that this approach could simultaneouslyachieve a parameter reduction and accuracy improvement. The word error ratesimproved by approx 1% while reducing the number of parameters by 54%.

【7】 Diff4Steer: Steerable Diffusion Prior for Generative Music Retrieval  with Semantic Guidance
标题:迪夫4Steer:具有语义指导的生成音乐检索的可操纵扩散优先级
链接:https://arxiv.org/abs/2412.04746
作者:Xuchan Bao,  Judith Yue Li,  Zhong Yi Wan,  Kun Su,  Timo Denk,  Joonseok Lee,  Dima Kuzmin,  Fei Sha
备注:NeurIPS 2024 Creative AI Track
摘要:现代音乐检索系统通常依赖于用户偏好的固定表示,限制了它们捕捉用户多样且不确定的检索需求的能力。为了解决这一限制,我们介绍了Diff4Steer,一种新的生成检索框架,采用轻量级的扩散模型来合成不同的种子嵌入用户查询,代表音乐探索的潜在方向。与将用户查询映射到嵌入空间中的单个点的确定性方法不同,Diff4Steer提供了用于检索的目标模态(音频)的统计先验,有效地捕获了用户偏好的不确定性和多方面性质。此外,Diff4Steer可以通过图像或文本输入进行控制,从而实现更灵活和可控的音乐发现与最近邻搜索相结合。我们的框架在检索和排名指标方面优于确定性回归方法和基于LLM的生成检索基线,证明了其在捕获用户偏好方面的有效性,从而提供更多样化和相关的建议。听力示例可在tinyurl.com/diff4steer上找到。
摘要:Modern music retrieval systems often rely on fixed representations of userpreferences, limiting their ability to capture users' diverse and uncertainretrieval needs. To address this limitation, we introduce Diff4Steer, a novelgenerative retrieval framework that employs lightweight diffusion models tosynthesize diverse seed embeddings from user queries that represent potentialdirections for music exploration. Unlike deterministic methods that map userquery to a single point in embedding space, Diff4Steer provides a statisticalprior on the target modality (audio) for retrieval, effectively capturing theuncertainty and multi-faceted nature of user preferences. Furthermore,Diff4Steer can be steered by image or text inputs, enabling more flexible andcontrollable music discovery combined with nearest neighbor search. Ourframework outperforms deterministic regression methods and LLM-based generativeretrieval baseline in terms of retrieval and ranking metrics, demonstrating itseffectiveness in capturing user preferences, leading to more diverse andrelevant recommendations. Listening examples are available attinyurl.com/diff4steer.

【8】 Exploring Transformer-Based Music Overpainting for Jazz Piano Variations
标题:探索爵士钢琴变奏曲中基于变形者的音乐过度绘制
链接:https://arxiv.org/abs/2412.04610
作者:Eleanor Row,  Ivan Shanin,  György Fazekas
备注:Accepted and presented as a Late-Breaking Demo at the 25th International Society for Music Information Retrieval (ISMIR) in San Francisco, US, 2024
摘要:本文探讨了基于transformer的音乐覆盖模型,重点是爵士钢琴变奏曲。音乐覆盖产生新的变化,同时保留输入的旋律和和声结构。现有的方法受到小数据集的限制,限制了可扩展性和多样性。我们引入了VAR 4000,这是一个更大的爵士钢琴演奏数据集的子集,由4,352个训练对组成。使用半自动流水线,我们评估两个Transformer配置的VAR 4000,比较其性能与较小的JAZZVAR数据集。初步结果显示,在更大的数据集配置的泛化和性能有希望的改进,突出了Transformer模型的潜力,以有效地扩展更大,更多样化的数据集上的音乐覆盖。
摘要:This paper explores transformer-based models for music overpainting, focusingon jazz piano variations. Music overpainting generates new variations whilepreserving the melodic and harmonic structure of the input. Existing approachesare limited by small datasets, restricting scalability and diversity. Weintroduce VAR4000, a subset of a larger dataset for jazz piano performances,consisting of 4,352 training pairs. Using a semi-automatic pipeline, weevaluate two transformer configurations on VAR4000, comparing their performancewith the smaller JAZZVAR dataset. Preliminary results show promisingimprovements in generalisation and performance with the larger datasetconfiguration, highlighting the potential of transformer models to scaleeffectively for music overpainting on larger and more diverse datasets.

【9】 MoD-ART: Modal Decomposition of Acoustic Radiance Transfer
标题:MoD-ART:声辐射传递的模式分解
链接:https://arxiv.org/abs/2412.04534
作者:Matteo Scerbo,  Sebastian J. Schlecht,  Randall Ali,  Lauri Savioja,  Enzo De Sena
摘要:当多个声源和听众存在于同一环境中时,以交互速度对后期混响进行建模是一项具有挑战性的任务。当环境在几何上复杂和/或具有不均匀的能量吸收(例如,耦合体积)时,这是特别成问题的,因为在这种情况下,后期混响取决于声源和收听者的位置,因此必须实时地适应他们的运动。我们提出了一种新的方法,命名为模态分解的声辐射传递(MoD-ART),它可以处理高度复杂的情况下,效率。该方法基于声辐射传递的几何声学方法,从中提取一组能量衰减模式及其与源和收听者的位置关系。在本文中,我们描述了物理和数学意义的MoD-ART,突出其优势和适用于不同的场景。通过分析该方法的计算复杂度,我们表明,它比较非常有利的射线追踪。我们还提出了模拟结果表明,MoD-ART可以捕获多个衰减斜率和颤振回波。
摘要:Modeling late reverberation at interactive speeds is a challenging task whenmultiple sound sources and listeners are present in the same environment. Thisis especially problematic when the environment is geometrically complex and/orfeatures uneven energy absorption (e.g. coupled volumes), because in such casesthe late reverberation is dependent on the sound sources' and listeners'positions, and therefore must be adapted to their movements in real time. Wepresent a novel approach to the task, named modal decomposition of AcousticRadiance Transfer (MoD-ART), which can handle highly complex scenarios withefficiency. The approach is based on the geometrical acoustics method ofAcoustic Radiance Transfer, from which we extract a set of energy decay modesand their positional relationships with sources and listeners. In this paper,we describe the physical and mathematical meaningfulness of MoD-ART,highlighting its advantages and applicability to different scenarios. Throughan analysis of the method's computational complexity, we show that it comparesvery favourably with ray-tracing. We also present simulation results showingthat MoD-ART can capture multiple decay slopes and flutter echoes.

机器翻译由腾讯交互翻译提供,仅供参考