今日论文合集:cs.SD语音11篇,eess.AS音频处理13篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】 Are Deep Speech Denoising Models Robust to Adversarial Noise?

标题: 深度语音去噪模型对对抗噪音是否稳健?
链接:https://arxiv.org/abs/2503.11627
作者: Will Schwarzer,  Philip S. Thomas,  Andrea Fanelli,  Xiaoyu Liu
备注:13 pages, 5 figures
摘要:深度噪声抑制(DNS)模型在各种高风险语音应用中得到广泛使用。然而,在本文中,我们证明了四个最近的DNS模型都可以通过添加不可察觉的对抗性噪声而被简化为输出无法理解的胡言乱语。此外,我们的研究结果显示了有针对性的攻击,这可能会导致模型输出任意的话语,和空中攻击的短期可扩展性。虽然这些攻击的成功因模型和设置而异,但当特定于模型(即,白盒和不可转让),我们的研究结果强调了迫切需要在DNS系统的实际对策。
摘要:Deep noise suppression (DNS) models enjoy widespread use throughout a varietyof high-stakes speech applications. However, in this paper, we show that fourrecent DNS models can each be reduced to outputting unintelligible gibberishthrough the addition of imperceptible adversarial noise. Furthermore, ourresults show the near-term plausibility of targeted attacks, which could inducemodels to output arbitrary utterances, and over-the-air attacks. While thesuccess of these attacks varies by model and setting, and attacks appear to bestrongest when model-specific (i.e., white-box and non-transferable), ourresults highlight a pressing need for practical countermeasures in DNS systems.

【2】 Designing Neural Synthesizers for Low Latency Interaction
标题: 设计低延迟交互的神经合成器
链接:https://arxiv.org/abs/2503.11562
作者: Franco Caspe,  Jordie Shier,  Mark Sandler,  Charalampos Saitis,  Andrew McPherson
备注:See website at fcaspe.github.io/brave - 13 pages, 5 figures, accepted to the Journal of the Audio Engineering Society
摘要:神经音频合成(NAS)模型提供了对高质量、富有表现力的音频生成器的交互式音乐控制。虽然这些模型可以实时操作,但它们通常具有高延迟,使其不适合亲密的音乐交互。深度学习模型中的架构选择对音频延迟的影响在NAS文献中仍然很大程度上未被探索。在这项工作中,我们调查的延迟和抖动的来源,通常发现在交互式NAS模型。然后,我们将此分析应用于使用RAVE的音色传输任务,RAVE是Caillon等人于2021年推出的音频波形卷积变分自动编码器。最后,我们提出了一个迭代的设计方法来优化延迟。我们称之为BRAVE(Bravely Realtime Audio Variational autoEncoder,勇敢实时音频变分自动编码器)的模型达到了高潮,该模型具有低延迟,具有更好的音高和响度复制能力,同时显示出类似于RAVE的音色修改功能。我们实现它在一个专门的推理框架低延迟,实时推理,并提出了一个概念验证的音频插件兼容的音频信号从乐器。我们希望本文中描述的挑战和指导方针能够支持NAS研究人员从头开始设计低延迟推理模型,丰富音乐家的可能性。
摘要:Neural Audio Synthesis (NAS) models offer interactive musical control overhigh-quality, expressive audio generators. While these models can operate inreal-time, they often suffer from high latency, making them unsuitable forintimate musical interaction. The impact of architectural choices in deeplearning models on audio latency remains largely unexplored in the NASliterature. In this work, we investigate the sources of latency and jittertypically found in interactive NAS models. We then apply this analysis to thetask of timbre transfer using RAVE, a convolutional variational autoencoder foraudio waveforms introduced by Caillon et al. in 2021. Finally, we present aniterative design approach for optimizing latency. This culminates with a modelwe call BRAVE (Bravely Realtime Audio Variational autoEncoder), which islow-latency and exhibits better pitch and loudness replication while showingtimbre modification capabilities similar to RAVE. We implement it in aspecialized inference framework for low-latency, real-time inference andpresent a proof-of-concept audio plugin compatible with audio signals frommusical instruments. We expect the challenges and guidelines described in thisdocument to support NAS researchers in designing models for low-latencyinference from the ground up, enriching the landscape of possibilities formusicians.

【3】 Exploring Performance-Complexity Trade-Offs in Sound Event Detection
标题: 探索声音事件检测中的性能复杂性权衡
链接:https://arxiv.org/abs/2503.11373
作者: Tobias Morocutti,  Florian Schmid,  Jonathan Greif,  Francesco Foscarin,  Gerhard Widmer
摘要:我们的目标是为声音事件检测任务开发新的低复杂度网络的问题。我们的目标是仔细分析性能复杂性权衡,旨在与大型国家的最先进的模型,在一小部分的计算要求的竞争。我们发现,以前提出的用于音频标记的低复杂度卷积模型可以通过调整卷积步长,删除全局池,以及重要的是,在(现在的逐帧)分类头之前添加序列模型,有效地适应事件检测(需要逐帧预测)。系统的实验表明,序列模型类型的最佳选择取决于哪个复杂性度量对给定的应用程序最重要。我们还调查了知识蒸馏等强化培训策略的影响。最后,我们表明,结合优化的训练策略,我们可以达到与最先进的Transformers相当的事件检测性能,同时只需要大约5%的参数。我们在https://github.com/theMoro/EfficientSED上发布了所有预训练的模型和用于重现这项工作的代码,以支持未来在低复杂度声音事件检测方面的研究。
摘要:We target the problem of developing new low-complexity networks for the soundevent detection task. Our goal is to meticulously analyze theperformance-complexity trade-off, aiming to be competitive with the largestate-of-the-art models, at a fraction of the computational requirements. Wefind that low-complexity convolutional models previously proposed for audiotagging can be effectively adapted for event detection (which requiresframe-wise prediction) by adjusting convolutional strides, removing the globalpooling, and, importantly, adding a sequence model before the (now frame-wise)classification heads. Systematic experiments reveal that the best choice forthe sequence model type depends on which complexity metric is most importantfor the given application. We also investigate the impact of enhanced trainingstrategies such as knowledge distillation. In the end, we show that combinedwith an optimized training strategy, we can reach event detection performancecomparable to state-of-the-art transformers while requiring only around 5% ofthe parameters. We release all our pre-trained models and the code forreproducing this work to support future research in low-complexity sound eventdetection at https://github.com/theMoro/EfficientSED.

【4】 Creating a Good Teacher for Knowledge Distillation in Acoustic Scene  Classification
标题: 打造声学场景分类知识提炼的好老师
链接:https://arxiv.org/abs/2503.11363
作者: Tobias Morocutti,  Florian Schmid,  Khaled Koutini,  Gerhard Widmer
摘要:知识蒸馏(KD)是一种广泛应用的技术,用于将大型模型的知识压缩成更紧凑和有效的模型。KD已被证明在构建性能良好的低复杂度声学场景分类(ASC)系统方面非常有效,并在过去三年中用于年度DCASE挑战的所有顶级提交任务。有广泛的研究可用于建立KD过程,设计有效的学生模型,并形成良好的教师合奏。然而,较少的研究已经进行了调查,教师模型的属性是有益的低复杂度的学生。在这项工作中,我们试图通过研究使用不同的教师网络架构时对学生表现的影响来缩小这一差距,改变教师模型的大小,用不同的设备泛化方法训练它们,并应用不同的集成策略。结果表明,教师模型的大小,设备泛化方法,集成策略和集成规模是一个良好的学生网络的关键因素。
摘要:Knowledge Distillation (KD) is a widespread technique for compressing theknowledge of large models into more compact and efficient models. KD has provedto be highly effective in building well-performing low-complexity AcousticScene Classification (ASC) systems and was used in all the top-rankedsubmissions to this task of the annual DCASE challenge in the past three years.There is extensive research available on establishing the KD process, designingefficient student models, and forming well-performing teacher ensembles.However, less research has been conducted on investigating which teacher modelattributes are beneficial for low-complexity students. In this work, we try toclose this gap by studying the effects on the student's performance when usingdifferent teacher network architectures, varying the teacher model size,training them with different device generalization methods, and applyingdifferent ensembling strategies. The results show that teacher model sizes,device generalization methods, the ensembling strategy and the ensemble sizeare key factors for a well-performing student network.

【5】 MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with  Minimal Multimodal Speech Tokens
标题: MMS-LLaMA:具有最小多模式语音令牌的高效基于LLM的视听语音识别
链接:https://arxiv.org/abs/2503.11315
作者: Jeong Hun Yeo,  Hyeongseop Rha,  Se Jin Park,  Yong Man Ro
备注:The code and models are available this https URL
摘要:视听语音识别(AVSR)通过结合听觉和视觉信息实现噪声环境下的鲁棒语音识别。然而,最近基于大语言模型(LLM)的AVSR系统由于LLM处理的视听语音的高时间分辨率而招致高计算成本。在这项工作中,我们引入了一个有效的多模态语音LLM框架,最大限度地减少令牌长度,同时保留必要的语言内容。我们的方法采用了一个早期的AV融合模块,简化功能集成,视听语音Q-Former,动态分配令牌的基础上输入持续时间,和一个完善的查询分配策略与语音速率预测器,以调整令牌分配根据每个音频样本的说话速度。在LRS 3数据集上进行的大量实验表明,我们的方法实现了最先进的性能,WER为0.74%,而每秒仅使用3.5个令牌。此外,与以前的多模态语音LLM框架相比,我们的方法不仅减少了86%的令牌使用,而且通过将FLOPs减少35.7%来提高计算效率。
摘要:Audio-Visual Speech Recognition (AVSR) achieves robust speech recognition innoisy environments by combining auditory and visual information. However,recent Large Language Model (LLM) based AVSR systems incur high computationalcosts due to the high temporal resolution of audio-visual speech processed byLLMs. In this work, we introduce an efficient multimodal speech LLM frameworkthat minimizes token length while preserving essential linguistic content. Ourapproach employs an early av-fusion module for streamlined feature integration,an audio-visual speech Q-Former that dynamically allocates tokens based oninput duration, and a refined query allocation strategy with a speech ratepredictor to adjust token allocation according to speaking speed of each audiosample. Extensive experiments on the LRS3 dataset show that our method achievesstate-of-the-art performance with a WER of 0.74% while using only 3.5 tokensper second. Moreover, our approach not only reduces token usage by 86% comparedto the previous multimodal speech LLM framework, but also improvescomputational efficiency by reducing FLOPs by 35.7%.

【6】 Exploring the Potential of Large Multimodal Models as Effective  Alternatives for Pronunciation Assessment
标题: 探索大型多模式模型作为发音评估有效替代方案的潜力
链接:https://arxiv.org/abs/2503.11229
作者: Ke Wang,  Lei He,  Kun Liu,  Yan Deng,  Wenning Wei,  Sheng Zhao
备注:7 pages
摘要:大型多模态模型(LLM)在广泛的领域中表现出卓越的性能。本文探讨了它们在发音评估任务中的潜力,特别关注评估生成预训练Transformer(GPT)模型,特别是GPT-4 o的能力。我们的研究调查了它处理语音和音频的能力,以进行多层次的粒度和维度的发音评估,重点是反馈生成和评分。在我们的实验中,我们使用公开的Speechocean 762数据集。评估集中在两个关键方面:多层次评分和生成的反馈的实用性。评分结果与Speechocean 762数据集中提供的手动评分进行比较,而反馈质量则使用大型语言模型(LLM)进行评估。研究结果强调了将Lemp与传统的发音评估方法相结合的有效性,提供了对模型优势的见解,并确定了进一步改进的领域。
摘要:Large Multimodal Models (LMMs) have demonstrated exceptional performanceacross a wide range of domains. This paper explores their potential inpronunciation assessment tasks, with a particular focus on evaluating thecapabilities of the Generative Pre-trained Transformer (GPT) model,specifically GPT-4o. Our study investigates its ability to process speech andaudio for pronunciation assessment across multiple levels of granularity anddimensions, with an emphasis on feedback generation and scoring. For ourexperiments, we use the publicly available Speechocean762 dataset. Theevaluation focuses on two key aspects: multi-level scoring and the practicalityof the generated feedback. Scoring results are compared against the manualscores provided in the Speechocean762 dataset, while feedback quality isassessed using Large Language Models (LLMs). The findings highlight theeffectiveness of integrating LMMs with traditional methods for pronunciationassessment, offering insights into the model's strengths and identifying areasfor further improvement.

【7】 Comparative Study of Spike Encoding Methods for Environmental Sound  Classification
标题: 环境声分类中尖峰编码方法的比较研究
链接:https://arxiv.org/abs/2503.11206
作者: Andres Larroza,  Javier Naranjo-Alcazar,  Vicent Ortiz Castelló,  Pedro Zuccarello
备注:Under review EUSIPCO 2025
摘要:尖峰神经网络(SNN)提供了一种很有前途的方法来降低能耗和计算需求,使其特别有利于边缘应用中的嵌入式机器学习。然而,来自传统数字传感器的数据必须首先转换成尖峰序列,以使用神经形态计算技术进行处理。由于频率、背景噪声和重叠声学事件的高度可变性,环境声音的分类提出了独特的挑战。尽管存在这些挑战,但大多数关于基于尖峰的音频编码的研究都集中在语音处理上,使得非语音环境声音的研究不足。在这项工作中,我们对广泛使用的尖峰编码技术进行了全面的比较,评估了它们在ESC-10数据集上的有效性。通过了解编码选择对环境声音处理的影响,研究人员和从业人员可以选择最适合实际应用的方法,如智能监控,环境监测和工业声学分析。本研究为环境声音分类中的尖峰信号编码提供了一个基准,为今后神经形态音频处理的研究提供了基础性参考。
摘要:Spiking Neural Networks (SNNs) offer a promising approach to reduce energyconsumption and computational demands, making them particularly beneficial forembedded machine learning in edge applications. However, data from conventionaldigital sensors must first be converted into spike trains to be processed usingneuromorphic computing technologies. The classification of environmental soundspresents unique challenges due to the high variability of frequencies,background noise, and overlapping acoustic events. Despite these challenges,most studies on spike-based audio encoding focus on speech processing, leavingnon-speech environmental sounds underexplored. In this work, we conduct acomprehensive comparison of widely used spike encoding techniques, evaluatingtheir effectiveness on the ESC-10 dataset. By understanding the impact ofencoding choices on environmental sound processing, researchers andpractitioners can select the most suitable approach for real-world applicationssuch as smart surveillance, environmental monitoring, and industrial acousticanalysis. This study serves as a benchmark for spike encoding in environmentalsound classification, providing a foundational reference for future research inneuromorphic audio processing.

【8】 Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study  on Audio Question Answering
标题: 强化学习优于监督微调:音频问题回答的案例研究
链接:https://arxiv.org/abs/2503.11197
作者: Gang Li,  Jizhong Liu,  Heinrich Dinkel,  Yadong Niu,  Junbo Zhang,  Jian Luan
摘要:最近,强化学习(RL)已被证明可以大大提高大型语言模型(LLM)的推理能力,并且基于RL的方法已逐步应用于视觉多模态任务。然而,在这些发展中,音频模式在很大程度上被忽视了。因此,我们在音频理解和推理方面进行了一系列RL探索,特别关注音频问答(AQA)任务。我们利用组相对策略优化(GRPO)算法来Qwen 2-Audio-7 B-Instruct,我们的实验在MMAU Test-mini基准测试中展示了最先进的性能,达到了64.5%的准确率。该技术报告的主要发现如下:1)GRPO算法可以有效地应用于大型音频语言模型(LALM),即使该模型只有8.2B参数; 2)只有38 k后训练样本,RL的性能显着优于监督微调(SFT),表明基于RL的方法可以在没有大型数据集的情况下有效; 3)显式推理过程对AQA任务没有显着的好处,如何有效地利用深度思维仍然是一个有待进一步研究的问题; 4)LALM仍然远远落后于人类的语义语言推理,这表明基于RL的方法值得进一步探索。我们的项目可以在https://github.com/xiaomi/r1-aqa和https://huggingface.co/mispeech/r1-aqa上找到。
摘要:Recently, reinforcement learning (RL) has been shown to greatly enhance thereasoning capabilities of large language models (LLMs), and RL-based approacheshave been progressively applied to visual multimodal tasks. However, the audiomodality has largely been overlooked in these developments. Thus, we conduct aseries of RL explorations in audio understanding and reasoning, specificallyfocusing on the audio question answering (AQA) task. We leverage the grouprelative policy optimization (GRPO) algorithm to Qwen2-Audio-7B-Instruct, andour experiments demonstrated state-of-the-art performance on the MMAU Test-minibenchmark, achieving an accuracy rate of 64.5%. The main findings in thistechnical report are as follows: 1) The GRPO algorithm can be effectivelyapplied to large audio language models (LALMs), even when the model has only8.2B parameters; 2) With only 38k post-training samples, RL significantlyoutperforms supervised fine-tuning (SFT), indicating that RL-based approachescan be effective without large datasets; 3) The explicit reasoning process hasnot shown significant benefits for AQA tasks, and how to efficiently utilizedeep thinking remains an open question for further research; 4) LALMs still lagfar behind humans auditory-language reasoning, suggesting that the RL-basedapproaches warrant further exploration. Our project is available athttps://github.com/xiaomi/r1-aqa and https://huggingface.co/mispeech/r1-aqa.

【9】 Cross-Modal Learning for Music-to-Music-Video Description Generation
标题: 用于音乐到音乐视频描述生成的跨模式学习
链接:https://arxiv.org/abs/2503.11190
作者: Zhuoyuan Mao,  Mengjie Zhao,  Qiyu Wu,  Zhi Zhong,  Wei-Hsiang Liao,  Hiromi Wakaki,  Yuki Mitsufuji
备注:Accepted by RepL4NLP 2025 @ NAACL 2025
摘要:由于音乐和视频模态之间的内在差异,音乐到音乐视频生成是一项具有挑战性的任务。强大的文本到视频扩散模型的出现为音乐视频(MV)生成开辟了一条有前途的途径,首先解决音乐到MV描述任务,随后利用这些模型进行视频生成。在这项研究中,我们专注于MV描述生成任务,并提出了一个全面的管道,包括训练数据构建和多模态模型微调。我们在基于Music4All数据集的新构建的音乐到MV描述数据集上微调了现有的预训练多模态模型,该数据集集成了音乐和视觉信息。我们的实验结果表明,音乐表示可以有效地映射到文本域,使有意义的MV描述直接从音乐输入的生成。我们还确定了数据集构建管道中的关键组件,这些组件对MV描述的质量产生了关键影响,并突出了特定的音乐属性,这些属性保证了对改进MV描述生成的更大关注。
摘要:Music-to-music-video generation is a challenging task due to the intrinsicdifferences between the music and video modalities. The advent of powerfultext-to-video diffusion models has opened a promising pathway for music-video(MV) generation by first addressing the music-to-MV description task andsubsequently leveraging these models for video generation. In this study, wefocus on the MV description generation task and propose a comprehensivepipeline encompassing training data construction and multimodal modelfine-tuning. We fine-tune existing pre-trained multimodal models on our newlyconstructed music-to-MV description dataset based on the Music4All dataset,which integrates both musical and visual information. Our experimental resultsdemonstrate that music representations can be effectively mapped to textualdomains, enabling the generation of meaningful MV description directly frommusic inputs. We also identify key components in the dataset constructionpipeline that critically impact the quality of MV description and highlightspecific musical attributes that warrant greater focus for improved MVdescription generation.

【10】 Joint Training And Decoding for Multilingual End-to-End Simultaneous  Speech Translation
标题: 多语言端到端同步语音翻译的联合训练和解码
链接:https://arxiv.org/abs/2503.11080
作者: Wuwei Huang,  Renren Jin,  Wen Zhang,  Jian Luan,  Bin Wang,  Deyi Xiong
备注:ICASSP 2023
摘要:端到端语音翻译(ST)的最新研究促进了多语种端到端ST和端到端同声ST的探索。在本文中,我们研究了一对多语种环境中的端到端同声翻译,这更接近于实际场景中的应用。我们探讨了一个单独的解码器架构和一个统一的架构,在这种情况下,联合同步训练。为了进一步探索跨语言的知识转移,我们提出了一个异步训练策略,建议统一的解码器架构。我们策划了一个多路对齐的多语言端到端ST数据集,作为评估我们方法的基准测试平台。实验结果表明,我们的模型在收集的数据集上的有效性。我们的代码和数据可在https://github.com/XiaoMi/TED-MMST上获得。
摘要:Recent studies on end-to-end speech translation(ST) have facilitated theexploration of multilingual end-to-end ST and end-to-end simultaneous ST. Inthis paper, we investigate end-to-end simultaneous speech translation in aone-to-many multilingual setting which is closer to applications in realscenarios. We explore a separate decoder architecture and a unifiedarchitecture for joint synchronous training in this scenario. To furtherexplore knowledge transfer across languages, we propose an asynchronoustraining strategy on the proposed unified decoder architecture. A multi-wayaligned multilingual end-to-end ST dataset was curated as a benchmark testbedto evaluate our methods. Experimental results demonstrate the effectiveness ofour models on the collected dataset. Our codes and data are available at:https://github.com/XiaoMi/TED-MMST.

【11】 A Data-Driven Exploration of Elevation Cues in HRTFs: An Explainable AI  Perspective Across Multiple Datasets
标题: HRTF中海拔线索的数据驱动探索:跨多个数据集中的可解释人工智能视角
链接:https://arxiv.org/abs/2503.11312
作者: Juan Antonio De Rus,  Mario Montagud,  Jesus Lopez-Ballester,  Francesc J. Ferri,  Maximo Cobos
备注:14 pages, 9 figures
摘要:双耳音频中的精确仰角感知仍然是一个挑战,尽管对头部相关传递函数(HRTF)和频谱线索进行了广泛的研究。虽然先前的研究已经推进了我们对声音定位线索的理解,但频谱特征和仰角感知之间的相互作用仍然没有完全理解。本文对来自11个不同的公共HRTF数据集的600多个主题进行了全面分析,采用卷积神经网络(CNN)模型结合可解释人工智能(XAI)技术来研究海拔线索。除了测试各种HRTF预处理方法外,我们还专注于数据集内和数据集间的泛化和可解释性,评估模型在受试者和测量设置引起的不同HRTF变化中的鲁棒性。通过利用类激活映射(CAM)显着性图,我们确定了可能有助于海拔感知的关键频段,从而更深入地了解驱动特定海拔分类的频谱特征。这项研究通过分析不同的数据集和预处理技术,为HRTF建模和海拔感知提供了新的视角,扩大了我们对各种条件下这些线索的理解。
摘要:Precise elevation perception in binaural audio remains a challenge, despiteextensive research on head-related transfer functions (HRTFs) and spectralcues. While prior studies have advanced our understanding of sound localizationcues, the interplay between spectral features and elevation perception is stillnot fully understood. This paper presents a comprehensive analysis of over 600subjects from 11 diverse public HRTF datasets, employing a convolutional neuralnetwork (CNN) model combined with explainable artificial intelligence (XAI)techniques to investigate elevation cues. In addition to testing various HRTFpre-processing methods, we focus on both within-dataset and inter-datasetgeneralization and explainability, assessing the model's robustness acrossdifferent HRTF variations stemming from subjects and measurement setups. Byleveraging class activation mapping (CAM) saliency maps, we identify keyfrequency bands that may contribute to elevation perception, providing deeperinsights into the spectral features that drive elevation-specificclassification. This study offers new perspectives on HRTF modeling andelevation perception by analyzing diverse datasets and pre-processingtechniques, expanding our understanding of these cues across a wide range ofconditions.

eess.AS音频处理

【1】 A Data-Driven Exploration of Elevation Cues in HRTFs: An Explainable AI  Perspective Across Multiple Datasets
标题: HRTF中海拔线索的数据驱动探索:跨多个数据集中的可解释人工智能视角
链接:https://arxiv.org/abs/2503.11312
作者: Juan Antonio De Rus,  Mario Montagud,  Jesus Lopez-Ballester,  Francesc J. Ferri,  Maximo Cobos
备注:14 pages, 9 figures
摘要:双耳音频中的精确仰角感知仍然是一个挑战,尽管对头部相关传递函数(HRTF)和频谱线索进行了广泛的研究。虽然先前的研究已经推进了我们对声音定位线索的理解,但频谱特征和仰角感知之间的相互作用仍然没有完全理解。本文对来自11个不同的公共HRTF数据集的600多个主题进行了全面分析,采用卷积神经网络(CNN)模型结合可解释人工智能(XAI)技术来研究海拔线索。除了测试各种HRTF预处理方法外,我们还专注于数据集内和数据集间的泛化和可解释性,评估模型在受试者和测量设置引起的不同HRTF变化中的鲁棒性。通过利用类激活映射(CAM)显着性图,我们确定了可能有助于海拔感知的关键频段,从而更深入地了解驱动特定海拔分类的频谱特征。这项研究通过分析不同的数据集和预处理技术,为HRTF建模和海拔感知提供了新的视角,扩大了我们对各种条件下这些线索的理解。
摘要:Precise elevation perception in binaural audio remains a challenge, despiteextensive research on head-related transfer functions (HRTFs) and spectralcues. While prior studies have advanced our understanding of sound localizationcues, the interplay between spectral features and elevation perception is stillnot fully understood. This paper presents a comprehensive analysis of over 600subjects from 11 diverse public HRTF datasets, employing a convolutional neuralnetwork (CNN) model combined with explainable artificial intelligence (XAI)techniques to investigate elevation cues. In addition to testing various HRTFpre-processing methods, we focus on both within-dataset and inter-datasetgeneralization and explainability, assessing the model's robustness acrossdifferent HRTF variations stemming from subjects and measurement setups. Byleveraging class activation mapping (CAM) saliency maps, we identify keyfrequency bands that may contribute to elevation perception, providing deeperinsights into the spectral features that drive elevation-specificclassification. This study offers new perspectives on HRTF modeling andelevation perception by analyzing diverse datasets and pre-processingtechniques, expanding our understanding of these cues across a wide range ofconditions.

【2】 MAVFlow: Preserving Paralinguistic Elements with Conditional Flow  Matching for Zero-Shot AV2AV Multilingual Translation
标题: MAVFlow:通过条件流匹配保留副语言元素,实现Zero-ShotAV 2AV多语言翻译
链接:https://arxiv.org/abs/2503.11026
作者: Sungwoo Cho,  Jeongsoo Choi,  Sungnyun Kim,  Se-Young Yun
备注:Preliminary work
摘要:尽管文本到语音(TTS)模型取得了最新进展,但视听到视听(AV2AV)翻译仍然面临着一个关键挑战:保持原始和翻译的声音和面部特征之间的说话人一致性。为了解决这个问题,我们提出了一个条件流匹配(CFM)的zero-shot视听渲染器,利用强大的双重指导,从音频和视觉模态。通过利用CFM的多模态引导,我们的模型鲁棒地保留了特定于说话者的特征,并显著增强了zero-shot AV2AV翻译能力。对于音频模态,我们通过将强大的扬声器嵌入与x向量相结合来增强CFM过程,这有助于增强扬声器的一致性。此外,我们传达情感的细微差别,面部渲染模块。音频和视觉提示提供的指导仍然独立于语义或语言内容,使我们的渲染器能够有效地处理不同语言的单语者的zero-shot翻译任务。我们的经验表明,包括高质量的梅尔频谱条件下的面部信息,不仅提高了合成语音的质量,而且还积极影响面部生成,导致整体性能的改善。
摘要:Despite recent advances in text-to-speech (TTS) models, audio-visual toaudio-visual (AV2AV) translation still faces a critical challenge: maintainingspeaker consistency between the original and translated vocal and facialfeatures. To address this issue, we propose a conditional flow matching (CFM)zero-shot audio-visual renderer that utilizes strong dual guidance from bothaudio and visual modalities. By leveraging multi-modal guidance with CFM, ourmodel robustly preserves speaker-specific characteristics and significantlyenhances zero-shot AV2AV translation abilities. For the audio modality, weenhance the CFM process by integrating robust speaker embeddings withx-vectors, which serve to bolster speaker consistency. Additionally, we conveyemotional nuances to the face rendering module. The guidance provided by bothaudio and visual cues remains independent of semantic or linguistic content,allowing our renderer to effectively handle zero-shot translation tasks formonolingual speakers in different languages. We empirically demonstrate thatthe inclusion of high-quality mel-spectrograms conditioned on facialinformation not only enhances the quality of the synthesized speech but alsopositively influences facial generation, leading to overall performanceimprovements.

【3】 EEG-Based Decoding of Sound Location: Comparing Free-Field to  Headphone-Based Non-Individual HRTFs
标题: 基于脑电波的声音位置解码:比较自由场与基于耳机的非个体HRTI
链接:https://arxiv.org/abs/2503.10783
作者: Nils Marggraf-Turley,  Martha Shiell,  Niels Pontoppidan,  Drew Cappotto,  Lorenzo Picinali
备注:40 pages, 6 figures, submitted to JASA
摘要:声源定位依赖于空间线索,如耳间时间差(ITD),耳间电平差(ILD)和单耳频谱线索。单独测量的头部相关传递函数(HRTF)有助于精确的空间听力,但测量是不切实际的,需要非个人的HRTF,这可能会损害定位精度和外部化。为了进一步研究这一现象,自由场和非个人HRTF听力的神经生理学差异进行了探讨,从EEG派生的事件相关电位(ERP)解码声音的位置。22名参与者在两种条件下定位刺激,记录EEG反应,并训练逻辑回归分类器以区分声源位置。   与自由场相比,KEMAR观察到较低的皮层反应幅度,特别是在前中央和枕顶叶区域。方差分析表明,听觉条件(F(1,21)= 34.56,p <0.0001)和位置(F(3,63)= 18.17,p <0.0001)对解码准确性(DA)有显著的主效应,在自由场和耳间线索主导的位置,DA更高。DA与前后混淆率呈负相关(r =-0.57,p <0.01),将神经DA与知觉混淆联系起来。   这些研究结果表明,耳机为基础的非个人HRTF引起较低幅度的皮质反应静态,方位角变化的位置比自由场条件。基于EEG的DA和前后混淆之间的相关性强调了神经生理学标记物评估空间听觉辨别的潜力。
摘要:Sound source localization relies on spatial cues such as interaural timedifferences (ITD), interaural level differences (ILD), and monaural spectralcues. Individually measured Head-Related Transfer Functions (HRTFs) facilitateprecise spatial hearing but are impractical to measure, necessitatingnon-individual HRTFs, which may compromise localization accuracy andexternalization. To further investigate this phenomenon, the neurophysiologicaldifferences between free-field and non-individual HRTF listening are exploredby decoding sound locations from EEG-derived Event-Related Potentials (ERPs).Twenty-two participants localized stimuli under both conditions with EEGresponses recorded and logistic regression classifiers trained to distinguishsound source locations. Lower cortical response amplitudes were observed for KEMAR compared tofree-field, especially in front-central and occipital-parietal regions. ANOVAidentified significant main effects of auralization condition (F(1, 21) =34.56, p < 0.0001) and location (F(3, 63) = 18.17, p < 0.0001) on decodingaccuracy (DA), which was higher in free-field and interaural-cue-dominatedlocations. DA negatively correlated with front-back confusion rates (r = -0.57,p < 0.01), linking neural DA to perceptual confusion. These findings demonstrate that headphone-based non-individual HRTFs elicitlower amplitude cortical responses to static, azimuthally-varying locationsthan free-field conditions. The correlation between EEG-based DA and front-backconfusion underscores neurophysiological markers' potential for assessingspatial auditory discrimination.

【4】 Are Deep Speech Denoising Models Robust to Adversarial Noise?
标题: 深度语音去噪模型对对抗噪音是否稳健?
链接:https://arxiv.org/abs/2503.11627
作者: Will Schwarzer,  Philip S. Thomas,  Andrea Fanelli,  Xiaoyu Liu
备注:13 pages, 5 figures
摘要:深度噪声抑制(DNS)模型在各种高风险语音应用中得到广泛使用。然而,在本文中,我们表明,四个最近的DNS模型可以通过添加不可感知的对抗性噪声,每个都可以减少到输出难以理解的胡言乱语。此外,我们的研究结果显示了有针对性的攻击,这可能会导致模型输出任意的话语,和空中攻击的短期可扩展性。虽然这些攻击的成功因模型和设置而异,但当特定于模型(即,白盒和不可转让),我们的研究结果强调了迫切需要在DNS系统的实际对策。
摘要:Deep noise suppression (DNS) models enjoy widespread use throughout a varietyof high-stakes speech applications. However, in this paper, we show that fourrecent DNS models can each be reduced to outputting unintelligible gibberishthrough the addition of imperceptible adversarial noise. Furthermore, ourresults show the near-term plausibility of targeted attacks, which could inducemodels to output arbitrary utterances, and over-the-air attacks. While thesuccess of these attacks varies by model and setting, and attacks appear to bestrongest when model-specific (i.e., white-box and non-transferable), ourresults highlight a pressing need for practical countermeasures in DNS systems.

【5】 Designing Neural Synthesizers for Low Latency Interaction
标题: 设计低延迟交互的神经合成器
链接:https://arxiv.org/abs/2503.11562
作者: Franco Caspe,  Jordie Shier,  Mark Sandler,  Charalampos Saitis,  Andrew McPherson
备注:See website at fcaspe.github.io/brave - 13 pages, 5 figures, accepted to the Journal of the Audio Engineering Society
摘要:神经音频合成(NAS)模型提供了对高质量、富有表现力的音频生成器的交互式音乐控制。虽然这些模型可以实时操作,但它们通常具有高延迟,使其不适合亲密的音乐交互。深度学习模型中的架构选择对音频延迟的影响在NAS文献中仍然很大程度上未被探索。在这项工作中,我们调查的延迟和抖动的来源,通常发现在交互式NAS模型。然后,我们将此分析应用于使用RAVE的音色传输任务,RAVE是Caillon等人于2021年推出的音频波形卷积变分自动编码器。最后,我们提出了一个迭代的设计方法来优化延迟。我们称之为BRAVE(Bravely Realtime Audio Variational autoEncoder,勇敢实时音频变分自动编码器)的模型达到了高潮,该模型具有低延迟,具有更好的音高和响度复制能力,同时显示出类似于RAVE的音色修改功能。我们实现它在一个专门的推理框架低延迟,实时推理,并提出了一个概念验证的音频插件兼容的音频信号从乐器。我们希望本文中描述的挑战和指导方针能够支持NAS研究人员从头开始设计低延迟推理模型,丰富音乐家的可能性。
摘要:Neural Audio Synthesis (NAS) models offer interactive musical control overhigh-quality, expressive audio generators. While these models can operate inreal-time, they often suffer from high latency, making them unsuitable forintimate musical interaction. The impact of architectural choices in deeplearning models on audio latency remains largely unexplored in the NASliterature. In this work, we investigate the sources of latency and jittertypically found in interactive NAS models. We then apply this analysis to thetask of timbre transfer using RAVE, a convolutional variational autoencoder foraudio waveforms introduced by Caillon et al. in 2021. Finally, we present aniterative design approach for optimizing latency. This culminates with a modelwe call BRAVE (Bravely Realtime Audio Variational autoEncoder), which islow-latency and exhibits better pitch and loudness replication while showingtimbre modification capabilities similar to RAVE. We implement it in aspecialized inference framework for low-latency, real-time inference andpresent a proof-of-concept audio plugin compatible with audio signals frommusical instruments. We expect the challenges and guidelines described in thisdocument to support NAS researchers in designing models for low-latencyinference from the ground up, enriching the landscape of possibilities formusicians.

【6】 Exploring Performance-Complexity Trade-Offs in Sound Event Detection
标题: 探索声音事件检测中的性能复杂性权衡
链接:https://arxiv.org/abs/2503.11373
作者: Tobias Morocutti,  Florian Schmid,  Jonathan Greif,  Francesco Foscarin,  Gerhard Widmer
摘要:我们的目标是为声音事件检测任务开发新的低复杂度网络的问题。我们的目标是仔细分析性能复杂性权衡,旨在与大型国家的最先进的模型,在一小部分的计算要求的竞争。我们发现,以前提出的用于音频标记的低复杂度卷积模型可以通过调整卷积步长,删除全局池,以及重要的是,在(现在的逐帧)分类头之前添加序列模型,有效地适应事件检测(需要逐帧预测)。系统的实验表明,序列模型类型的最佳选择取决于哪个复杂性度量对给定的应用程序最重要。我们还调查了知识蒸馏等强化培训策略的影响。最后,我们表明,结合优化的训练策略,我们可以达到与最先进的Transformers相当的事件检测性能,同时只需要大约5%的参数。我们在https://github.com/theMoro/EfficientSED上发布了所有预训练的模型和用于重现这项工作的代码,以支持未来在低复杂度声音事件检测方面的研究。
摘要:We target the problem of developing new low-complexity networks for the soundevent detection task. Our goal is to meticulously analyze theperformance-complexity trade-off, aiming to be competitive with the largestate-of-the-art models, at a fraction of the computational requirements. Wefind that low-complexity convolutional models previously proposed for audiotagging can be effectively adapted for event detection (which requiresframe-wise prediction) by adjusting convolutional strides, removing the globalpooling, and, importantly, adding a sequence model before the (now frame-wise)classification heads. Systematic experiments reveal that the best choice forthe sequence model type depends on which complexity metric is most importantfor the given application. We also investigate the impact of enhanced trainingstrategies such as knowledge distillation. In the end, we show that combinedwith an optimized training strategy, we can reach event detection performancecomparable to state-of-the-art transformers while requiring only around 5% ofthe parameters. We release all our pre-trained models and the code forreproducing this work to support future research in low-complexity sound eventdetection at https://github.com/theMoro/EfficientSED.

【7】 Creating a Good Teacher for Knowledge Distillation in Acoustic Scene  Classification
标题: 打造声学场景分类知识提炼的好老师
链接:https://arxiv.org/abs/2503.11363
作者: Tobias Morocutti,  Florian Schmid,  Khaled Koutini,  Gerhard Widmer
摘要:知识蒸馏(KD)是一种广泛应用的技术,用于将大型模型的知识压缩成更紧凑和有效的模型。KD已被证明在构建性能良好的低复杂度声学场景分类(ASC)系统方面非常有效,并在过去三年中用于年度DCASE挑战的所有顶级提交任务。有广泛的研究可用于建立KD过程,设计有效的学生模型,并形成良好的教师合奏。然而,较少的研究已经进行了调查,教师模型的属性是有益的低复杂度的学生。在这项工作中,我们试图通过研究使用不同的教师网络架构时对学生表现的影响来缩小这一差距,改变教师模型的大小,用不同的设备泛化方法训练它们,并应用不同的集成策略。结果表明,教师模型的大小,设备泛化方法,集成策略和集成规模是一个良好的学生网络的关键因素。
摘要:Knowledge Distillation (KD) is a widespread technique for compressing theknowledge of large models into more compact and efficient models. KD has provedto be highly effective in building well-performing low-complexity AcousticScene Classification (ASC) systems and was used in all the top-rankedsubmissions to this task of the annual DCASE challenge in the past three years.There is extensive research available on establishing the KD process, designingefficient student models, and forming well-performing teacher ensembles.However, less research has been conducted on investigating which teacher modelattributes are beneficial for low-complexity students. In this work, we try toclose this gap by studying the effects on the student's performance when usingdifferent teacher network architectures, varying the teacher model size,training them with different device generalization methods, and applyingdifferent ensembling strategies. The results show that teacher model sizes,device generalization methods, the ensembling strategy and the ensemble sizeare key factors for a well-performing student network.

【8】 MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with  Minimal Multimodal Speech Tokens
标题: MMS-LLaMA:具有最小多模式语音令牌的高效基于LLM的视听语音识别
链接:https://arxiv.org/abs/2503.11315
作者: Jeong Hun Yeo,  Hyeongseop Rha,  Se Jin Park,  Yong Man Ro
备注:The code and models are available this https URL
摘要:视听语音识别(AVSR)通过结合听觉和视觉信息实现噪声环境下的鲁棒语音识别。然而,最近基于大语言模型(LLM)的AVSR系统由于LLM处理的视听语音的高时间分辨率而招致高计算成本。在这项工作中,我们引入了一个有效的多模态语音LLM框架,最大限度地减少令牌长度,同时保留必要的语言内容。我们的方法采用了一个早期的AV融合模块,简化功能集成,视听语音Q-Former,动态分配令牌的基础上输入持续时间,和一个完善的查询分配策略与语音速率预测器,以调整令牌分配根据每个音频样本的说话速度。在LRS 3数据集上进行的大量实验表明,我们的方法实现了最先进的性能,WER为0.74%,而每秒仅使用3.5个令牌。此外,与以前的多模态语音LLM框架相比,我们的方法不仅减少了86%的令牌使用,而且通过将FLOPs减少35.7%来提高计算效率。
摘要:Audio-Visual Speech Recognition (AVSR) achieves robust speech recognition innoisy environments by combining auditory and visual information. However,recent Large Language Model (LLM) based AVSR systems incur high computationalcosts due to the high temporal resolution of audio-visual speech processed byLLMs. In this work, we introduce an efficient multimodal speech LLM frameworkthat minimizes token length while preserving essential linguistic content. Ourapproach employs an early av-fusion module for streamlined feature integration,an audio-visual speech Q-Former that dynamically allocates tokens based oninput duration, and a refined query allocation strategy with a speech ratepredictor to adjust token allocation according to speaking speed of each audiosample. Extensive experiments on the LRS3 dataset show that our method achievesstate-of-the-art performance with a WER of 0.74% while using only 3.5 tokensper second. Moreover, our approach not only reduces token usage by 86% comparedto the previous multimodal speech LLM framework, but also improvescomputational efficiency by reducing FLOPs by 35.7%.

【9】 Exploring the Potential of Large Multimodal Models as Effective  Alternatives for Pronunciation Assessment
标题: 探索大型多模式模型作为发音评估有效替代方案的潜力
链接:https://arxiv.org/abs/2503.11229
作者: Ke Wang,  Lei He,  Kun Liu,  Yan Deng,  Wenning Wei,  Sheng Zhao
备注:7 pages
摘要:大型多模态模型(LLM)在广泛的领域中表现出卓越的性能。本文探讨了它们在发音评估任务中的潜力,特别关注评估生成预训练Transformer(GPT)模型,特别是GPT-4 o的能力。我们的研究调查了它处理语音和音频的能力,以进行多层次的粒度和维度的发音评估,重点是反馈生成和评分。在我们的实验中,我们使用公开的Speechocean 762数据集。评估集中在两个关键方面:多层次评分和生成的反馈的实用性。评分结果与Speechocean 762数据集中提供的手动评分进行比较,而反馈质量则使用大型语言模型(LLM)进行评估。研究结果强调了将Lemp与传统的发音评估方法相结合的有效性,提供了对模型优势的见解,并确定了进一步改进的领域。
摘要:Large Multimodal Models (LMMs) have demonstrated exceptional performanceacross a wide range of domains. This paper explores their potential inpronunciation assessment tasks, with a particular focus on evaluating thecapabilities of the Generative Pre-trained Transformer (GPT) model,specifically GPT-4o. Our study investigates its ability to process speech andaudio for pronunciation assessment across multiple levels of granularity anddimensions, with an emphasis on feedback generation and scoring. For ourexperiments, we use the publicly available Speechocean762 dataset. Theevaluation focuses on two key aspects: multi-level scoring and the practicalityof the generated feedback. Scoring results are compared against the manualscores provided in the Speechocean762 dataset, while feedback quality isassessed using Large Language Models (LLMs). The findings highlight theeffectiveness of integrating LMMs with traditional methods for pronunciationassessment, offering insights into the model's strengths and identifying areasfor further improvement.

【10】 Comparative Study of Spike Encoding Methods for Environmental Sound  Classification
标题: 环境声分类中尖峰编码方法的比较研究
链接:https://arxiv.org/abs/2503.11206
作者: Andres Larroza,  Javier Naranjo-Alcazar,  Vicent Ortiz Castelló,  Pedro Zuccarello
备注:Under review EUSIPCO 2025
摘要:尖峰神经网络(SNN)提供了一种很有前途的方法来降低能耗和计算需求,使其特别有利于边缘应用中的嵌入式机器学习。然而,来自传统数字传感器的数据必须首先转换成尖峰序列,以使用神经形态计算技术进行处理。由于频率、背景噪声和重叠声学事件的高度可变性,环境声音的分类提出了独特的挑战。尽管存在这些挑战,但大多数关于基于尖峰的音频编码的研究都集中在语音处理上,使得非语音环境声音的研究不足。在这项工作中,我们对广泛使用的尖峰编码技术进行了全面的比较,评估了它们在ESC-10数据集上的有效性。通过了解编码选择对环境声音处理的影响,研究人员和从业人员可以选择最适合实际应用的方法,如智能监控,环境监测和工业声学分析。本研究为环境声音分类中的尖峰信号编码提供了一个基准,为今后神经形态音频处理的研究提供了基础性参考。
摘要:Spiking Neural Networks (SNNs) offer a promising approach to reduce energyconsumption and computational demands, making them particularly beneficial forembedded machine learning in edge applications. However, data from conventionaldigital sensors must first be converted into spike trains to be processed usingneuromorphic computing technologies. The classification of environmental soundspresents unique challenges due to the high variability of frequencies,background noise, and overlapping acoustic events. Despite these challenges,most studies on spike-based audio encoding focus on speech processing, leavingnon-speech environmental sounds underexplored. In this work, we conduct acomprehensive comparison of widely used spike encoding techniques, evaluatingtheir effectiveness on the ESC-10 dataset. By understanding the impact ofencoding choices on environmental sound processing, researchers andpractitioners can select the most suitable approach for real-world applicationssuch as smart surveillance, environmental monitoring, and industrial acousticanalysis. This study serves as a benchmark for spike encoding in environmentalsound classification, providing a foundational reference for future research inneuromorphic audio processing.

【11】 Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study  on Audio Question Answering
标题: 强化学习优于监督微调:音频问题回答的案例研究
链接:https://arxiv.org/abs/2503.11197
作者: Gang Li,  Jizhong Liu,  Heinrich Dinkel,  Yadong Niu,  Junbo Zhang,  Jian Luan
摘要:最近,强化学习(RL)已被证明可以大大提高大型语言模型(LLM)的推理能力,并且基于RL的方法已逐步应用于视觉多模态任务。然而,在这些发展中,音频模式在很大程度上被忽视了。因此,我们在音频理解和推理方面进行了一系列RL探索,特别关注音频问答(AQA)任务。我们利用组相对策略优化(GRPO)算法来Qwen 2-Audio-7 B-Instruct,我们的实验在MMAU Test-mini基准测试中展示了最先进的性能,达到了64.5%的准确率。该技术报告的主要发现如下:1)GRPO算法可以有效地应用于大型音频语言模型(LALM),即使该模型只有8.2B参数; 2)只有38 k后训练样本,RL的性能显着优于监督微调(SFT),表明基于RL的方法可以在没有大型数据集的情况下有效; 3)显式推理过程对AQA任务没有显着的好处,如何有效地利用深度思维仍然是一个有待进一步研究的问题; 4)LALM仍然远远落后于人类的语义语言推理,这表明基于RL的方法值得进一步探索。我们的项目可以在https://github.com/xiaomi/r1-aqa和https://huggingface.co/mispeech/r1-aqa上找到。
摘要:Recently, reinforcement learning (RL) has been shown to greatly enhance thereasoning capabilities of large language models (LLMs), and RL-based approacheshave been progressively applied to visual multimodal tasks. However, the audiomodality has largely been overlooked in these developments. Thus, we conduct aseries of RL explorations in audio understanding and reasoning, specificallyfocusing on the audio question answering (AQA) task. We leverage the grouprelative policy optimization (GRPO) algorithm to Qwen2-Audio-7B-Instruct, andour experiments demonstrated state-of-the-art performance on the MMAU Test-minibenchmark, achieving an accuracy rate of 64.5%. The main findings in thistechnical report are as follows: 1) The GRPO algorithm can be effectivelyapplied to large audio language models (LALMs), even when the model has only8.2B parameters; 2) With only 38k post-training samples, RL significantlyoutperforms supervised fine-tuning (SFT), indicating that RL-based approachescan be effective without large datasets; 3) The explicit reasoning process hasnot shown significant benefits for AQA tasks, and how to efficiently utilizedeep thinking remains an open question for further research; 4) LALMs still lagfar behind humans auditory-language reasoning, suggesting that the RL-basedapproaches warrant further exploration. Our project is available athttps://github.com/xiaomi/r1-aqa and https://huggingface.co/mispeech/r1-aqa.

【12】 Cross-Modal Learning for Music-to-Music-Video Description Generation
标题: 用于音乐到音乐视频描述生成的跨模式学习
链接:https://arxiv.org/abs/2503.11190
作者: Zhuoyuan Mao,  Mengjie Zhao,  Qiyu Wu,  Zhi Zhong,  Wei-Hsiang Liao,  Hiromi Wakaki,  Yuki Mitsufuji
备注:Accepted by RepL4NLP 2025 @ NAACL 2025
摘要:由于音乐和视频模态之间的内在差异,音乐到音乐视频生成是一项具有挑战性的任务。强大的文本到视频扩散模型的出现为音乐视频(MV)生成开辟了一条有前途的途径,首先解决音乐到MV描述任务,随后利用这些模型进行视频生成。在本研究中,我们重点关注MV描述生成任务,并提出了一个包含训练数据构建和多模式模型微调的全面管道。我们在基于Music4All数据集的新构建的音乐到MV描述数据集上微调了现有的预训练多模态模型,该数据集集成了音乐和视觉信息。我们的实验结果表明,音乐表示可以有效地映射到文本域,使有意义的MV描述直接从音乐输入的生成。我们还确定了数据集构建管道中的关键组件,这些组件对MV描述的质量产生了关键影响,并突出了特定的音乐属性,这些属性保证了对改进MV描述生成的更大关注。
摘要:Music-to-music-video generation is a challenging task due to the intrinsicdifferences between the music and video modalities. The advent of powerfultext-to-video diffusion models has opened a promising pathway for music-video(MV) generation by first addressing the music-to-MV description task andsubsequently leveraging these models for video generation. In this study, wefocus on the MV description generation task and propose a comprehensivepipeline encompassing training data construction and multimodal modelfine-tuning. We fine-tune existing pre-trained multimodal models on our newlyconstructed music-to-MV description dataset based on the Music4All dataset,which integrates both musical and visual information. Our experimental resultsdemonstrate that music representations can be effectively mapped to textualdomains, enabling the generation of meaningful MV description directly frommusic inputs. We also identify key components in the dataset constructionpipeline that critically impact the quality of MV description and highlightspecific musical attributes that warrant greater focus for improved MVdescription generation.

【13】 Joint Training And Decoding for Multilingual End-to-End Simultaneous  Speech Translation
标题: 多语言端到端同步语音翻译的联合训练和解码
链接:https://arxiv.org/abs/2503.11080
作者: Wuwei Huang,  Renren Jin,  Wen Zhang,  Jian Luan,  Bin Wang,  Deyi Xiong
备注:ICASSP 2023
摘要:端到端语音翻译(ST)的最新研究促进了多语种端到端ST和端到端同声ST的探索。在本文中,我们研究了一对多语种环境中的端到端同声翻译,这更接近于实际场景中的应用。我们探讨了一个单独的解码器架构和一个统一的架构,在这种情况下,联合同步训练。为了进一步探索跨语言的知识转移,我们提出了一个异步训练策略的建议统一解码器架构。我们策划了一个多路对齐的多语言端到端ST数据集,作为评估我们方法的基准测试平台。实验结果表明,我们的模型在收集的数据集上的有效性。我们的代码和数据可在https://github.com/XiaoMi/TED-MMST上获得。
摘要:Recent studies on end-to-end speech translation(ST) have facilitated theexploration of multilingual end-to-end ST and end-to-end simultaneous ST. Inthis paper, we investigate end-to-end simultaneous speech translation in aone-to-many multilingual setting which is closer to applications in realscenarios. We explore a separate decoder architecture and a unifiedarchitecture for joint synchronous training in this scenario. To furtherexplore knowledge transfer across languages, we propose an asynchronoustraining strategy on the proposed unified decoder architecture. A multi-wayaligned multilingual end-to-end ST dataset was curated as a benchmark testbedto evaluate our methods. Experimental results demonstrate the effectiveness ofour models on the collected dataset. Our codes and data are available at:https://github.com/XiaoMi/TED-MMST.

机器翻译由腾讯交互翻译提供,仅供参考