今日论文合集:cs.SD语音12篇,eess.AS音频处理3篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song  Generation

标题:SongGen:文本到歌曲生成的单级自回归Transformer
链接:https://arxiv.org/abs/2502.13128
作者:Zihan Liu,  Shuangrui Ding,  Zhixiong Zhang,  Xiaoyi Dong,  Pan Zhang,  Yuhang Zang,  Yuhang Cao,  Dahua Lin,  Jiaqi Wang
摘要:文本到歌曲生成,从文本输入创建人声和伴奏的任务,由于领域的复杂性和数据的稀缺性,提出了重大的挑战。现有的方法通常采用多阶段生成过程,导致繁琐的训练和推理管道。在本文中,我们提出了松根,一个完全开源的,单级自回归Transformer设计的可控歌曲生成。拟议的模型有助于对不同的音乐属性进行细粒度控制,包括歌词和乐器、流派、情绪和音色的文本描述,同时还提供可选的三秒参考剪辑用于语音克隆。在一个统一的自回归框架内,SongGen支持两种输出模式:混合模式,直接生成人声和伴奏的混合,以及双轨模式,分别合成它们,以便在下游应用中获得更大的灵活性。我们为每种模式探索了不同的令牌模式策略,从而带来了显着的改进和有价值的见解。此外,我们设计了一个自动化的数据预处理管道,有效的质量控制。为了促进社区参与和未来的研究,我们将发布我们的模型权重,训练代码,注释数据和预处理管道。生成的示例展示在我们的项目页面https://liuzh-19.github.io/SongGen/上,代码将在https://github.com/LiuZH-19/SongGen上提供。
摘要:Text-to-song generation, the task of creating vocals and accompaniment fromtextual inputs, poses significant challenges due to domain complexity and datascarcity. Existing approaches often employ multi-stage generation procedures,resulting in cumbersome training and inference pipelines. In this paper, wepropose SongGen, a fully open-source, single-stage auto-regressive transformerdesigned for controllable song generation. The proposed model facilitatesfine-grained control over diverse musical attributes, including lyrics andtextual descriptions of instrumentation, genre, mood, and timbre, while alsooffering an optional three-second reference clip for voice cloning. Within aunified auto-regressive framework, SongGen supports two output modes: mixedmode, which generates a mixture of vocals and accompaniment directly, anddual-track mode, which synthesizes them separately for greater flexibility indownstream applications. We explore diverse token pattern strategies for eachmode, leading to notable improvements and valuable insights. Furthermore, wedesign an automated data preprocessing pipeline with effective quality control.To foster community engagement and future research, we will release our modelweights, training code, annotated data, and preprocessing pipeline. Thegenerated samples are showcased on our project page athttps://liuzh-19.github.io/SongGen/ , and the code will be available athttps://github.com/LiuZH-19/SongGen .

【2】 A Dual-Stage Time-Context Network for Speech-Based Alzheimer's Disease  Detection
标题:基于言语的阿尔茨海默病检测的双阶段时间上下文网络
链接:https://arxiv.org/abs/2502.13064
作者:Yifan Gao,  Long Guo,  Hong Liu
摘要:阿尔茨海默病(Alzheimer's disease,AD)是一种进行性神经退行性疾病,可导致记忆和交流能力的不可逆性认知下降。通过语音分析早期发现AD对于延缓疾病进展至关重要。然而,现有的方法主要使用预先训练的声学模型进行特征提取,但在长持续时间的语音中对局部和全局模式进行建模的能力有限。在这封信中,我们介绍了一个双阶段时间上下文网络(DSTC-Net)用于基于语音的AD检测,在长时间录音中将局部声学特征与全局会话上下文相结合。我们首先将每个长时间录音划分为固定长度的片段,以减少计算开销并保留局部时间细节。接下来,我们将这些片段馈送到片段内时间注意力(ISTA)模块,其中,具有帧级注意力的双向长短期记忆(BiLSTM)网络提取增强的局部特征。随后,跨段上下文注意力(CSCA)模块应用基于卷积的上下文建模和自适应注意力来统一所有段的全局模式。在ADReSSo数据集上的大量实验表明,我们的DSTC-Net优于最先进的模型,准确率为83.10%,F1为83.15%。
摘要:Alzheimer's disease (AD) is a progressive neurodegenerative disorder thatleads to irreversible cognitive decline in memory and communication. Earlydetection of AD through speech analysis is crucial for delaying diseaseprogression. However, existing methods mainly use pre-trained acoustic modelsfor feature extraction but have limited ability to model both local and globalpatterns in long-duration speech. In this letter, we introduce a Dual-StageTime-Context Network (DSTC-Net) for speech-based AD detection, integratinglocal acoustic features with global conversational context in long-durationrecordings.We first partition each long-duration recording into fixed-lengthsegments to reduce computational overhead and preserve local temporaldetails.Next, we feed these segments into an Intra-Segment Temporal Attention(ISTA) module, where a bidirectional Long Short-Term Memory (BiLSTM) networkwith frame-level attention extracts enhanced local features.Subsequently, aCross-Segment Context Attention (CSCA) module applies convolution-based contextmodeling and adaptive attention to unify global patterns across allsegments.Extensive experiments on the ADReSSo dataset show that our DSTC-Netoutperforms state-of-the-art models, reaching 83.10% accuracy and 83.15% F1.

【3】 Skip That Beat: Augmenting Meter Tracking Models for Underrepresented  Time Signatures
标题:跳过节拍:增强代表性不足的时间签名的电表跟踪模型
链接:https://arxiv.org/abs/2502.12972
作者:Giovana Morais,  Brian McFee,  Magdalena Fuentes
备注:4 pages + references, 3 figures, 1st Latin American Music Information Retrieval (LAMIR) workshop
摘要:节拍和强拍跟踪模型主要使用具有4/4米音乐的数据集开发,这降低了它们对其他时间签名的曲目的泛化,例如2/4的巴西桑巴舞。在这项工作中,我们提出了一个简单的增强技术,以增加表示的时间签名超过4/4,即2/4和3/4。我们的增强过程通过从4/4注释轨道中删除节拍间隔来工作。我们表明,增强的数据有助于改善代表性不足的米的强拍跟踪,同时保持在两个不同的模型中的节拍跟踪的整体性能。我们还表明,这种技术有助于提高在一个看不见的桑巴舞数据集的强拍跟踪。
摘要:Beat and downbeat tracking models are predominantly developed using datasetswith music in 4/4 meter, which decreases their generalization to repertories inother time signatures, such as Brazilian samba which is in 2/4. In this work,we propose a simple augmentation technique to increase the representation oftime signatures beyond 4/4, namely 2/4 and 3/4. Our augmentation procedureworks by removing beat intervals from 4/4 annotated tracks. We show that theaugmented data helps to improve downbeat tracking for underrepresented meterswhile preserving the overall performance of beat tracking in two differentmodels. We also show that this technique helps improve downbeat tracking in anunseen samba dataset.

【4】 Keep what you need : extracting efficient subnetworks from large audio  representation models
标题:保留您所需的内容:从大型音频表示模型中提取高效子网络
链接:https://arxiv.org/abs/2502.12925
作者:David Genova,  Philippe Esling,  Tom Hurlin
摘要:最近,音频基础模型的研究取得了显着的进展,如复杂下游任务的不断改进的结果所示。随后,这些预先训练的网络很快被用于各种音频应用。然而,这些改进导致这些模型的尺寸和复杂性都有相当大的增加。随着这个问题引起的环境问题,这阻止了在消费级设备上部署这种网络,并排除了它们用于实时应用。此外,这似乎与使用这些模型的任务的特异性相矛盾,与从任何类型的音频数据中提取丰富的多用途表示相比,这些模型通常更简单。在本文中,我们解决了这个问题,一个简单而有效的方法来提取轻量级的专家子网络从大型基础模型。具体来说,我们在预训练表示模型的层之间引入可学习的二进制掩码。当在下游任务上训练端到端模型时,我们向总体目标添加了一个稀疏诱导损失,从而学习了一个专门用于单个任务的紧凑子网络。重要的是,基础模型的权重保持冻结,从而降低了额外的训练成本。一旦经过训练,被屏蔽的计算单元就可以从网络中移除,这意味着显著的性能提升。我们评估我们的方法在三个广泛的音频基础模型,每个模型基于不同的骨干架构,并说明其对常见的音频表示评估任务的有效性,以及其在语音,音乐和一般音频的多功能性。复制结果的代码和支持网页可在https://github.com/gnvIRCAM/Audio-representation-trimming上获得
摘要:Recently, research on audio foundation models has witnessed notable advances,as illustrated by the ever improving results on complex downstream tasks.Subsequently, those pretrained networks have quickly been used for variousaudio applications. These improvements have however resulted in a considerableincrease both in size and complexity of these models. Along the environmentalconcerns this issue raises, this prevents the deployment of such networks onconsumer-level devices, and precludes their use for real-time applications.Moreover, this appears contradictory with the specificity of the tasks forwhich these models are used, which are often simpler compared to extracting arich, multi-purpose representation from any type of audio data. In this paper,we address this issue with a simple, yet effective method to extractlightweight specialist subnetworks from large foundation models. Specifically,we introduce learnable binary masks in-between the layers of a pretrainedrepresentation model. When training the end-to-end model on a downstream task,we add a sparsity-inducing loss to the overall objective, hence learning acompact subnetwork specialized on a single task. Importantly, the weights ofthe foundation model are kept frozen, resulting into low additional trainingcosts. Once trained, the masked computational units can then be removed fromthe network, implying significant performance gains. We assess our method onthree widespread audio foundation models, each based on a different backbonearchitecture, and illustrate its effectiveness on common audio representationevaluation tasks, as well as its versatility on both speech, music, and generalaudio. Code for reproducing the results and supporting webpage are available athttps://github.com/gnvIRCAM/Audio-representation-trimming

【5】 Soundwave: Less is More for Speech-Text Alignment in LLMs
标题:Soundwave:LLM中语音文本对齐的少即是多
链接:https://arxiv.org/abs/2502.12900
作者:Yuhao Zhang,  Zhiheng Liu,  Fan Bu,  Ruiyu Zhang,  Benyou Wang,  Haizhou Li
摘要:现有的端到端语音大语言模型(LLM)通常依赖于大规模的注释数据进行训练,而数据有效的训练尚未深入讨论。我们专注于语音和文本之间的两个基本问题:表示空间差距和序列长度不一致。我们提出了Soundwave,它利用高效的训练策略和新颖的架构来解决这些问题。结果表明,声波在语音翻译和AIR-Bench语音任务中的表现优于高级Qwen 2-Audio,仅使用五十分之一的训练数据。进一步的分析表明,声波在对话过程中仍然保持其智能。该项目可在https://github.com/FreedomIntelligence/Soundwave上查阅。
摘要:Existing end-to-end speech large language models (LLMs) usually rely onlarge-scale annotated data for training, while data-efficient training has notbeen discussed in depth. We focus on two fundamental problems between speechand text: the representation space gap and sequence length inconsistency. Wepropose Soundwave, which utilizes an efficient training strategy and a novelarchitecture to address these issues. Results show that Soundwave outperformsthe advanced Qwen2-Audio in speech translation and AIR-Bench speech tasks,using only one-fiftieth of the training data. Further analysis shows thatSoundwave still retains its intelligence during conversation. The project isavailable at https://github.com/FreedomIntelligence/Soundwave.

【6】 High-Fidelity Music Vocoder using Neural Audio Codecs
标题:使用神经音频编解码器的高保真音乐声码器
链接:https://arxiv.org/abs/2502.12759
作者:Luca A. Lanzendörfer,  Florian Grötschla,  Michael Ungersböck,  Roger Wattenhofer
备注:Accepted at ICASSP 2025
摘要:虽然神经声码器在高保真语音合成方面取得了重大进展,但它们在复调音乐中的应用仍然有待探索。在这项工作中,我们提出了DisCoder,一种神经声码器,它利用神经音频编解码器通知的生成对抗编码器-解码器架构,从mel频谱图重建高保真44.1 kHz音频。我们的方法首先将梅尔频谱图转换成与描述音频编解码器(DAC)潜在空间对齐的低维表示,然后使用微调的DAC解码器将其重建为音频信号。DisCoder在多个客观指标和MUSHRA听力研究中实现了音乐合成的最佳性能。我们的方法在语音合成方面也表现出了竞争力,突出了它作为通用声码器的潜力。
摘要:While neural vocoders have made significant progress in high-fidelity speechsynthesis, their application on polyphonic music has remained underexplored. Inthis work, we propose DisCoder, a neural vocoder that leverages a generativeadversarial encoder-decoder architecture informed by a neural audio codec toreconstruct high-fidelity 44.1 kHz audio from mel spectrograms. Our approachfirst transforms the mel spectrogram into a lower-dimensional representationaligned with the Descript Audio Codec (DAC) latent space before reconstructingit to an audio signal using a fine-tuned DAC decoder. DisCoder achievesstate-of-the-art performance in music synthesis on several objective metricsand in a MUSHRA listening study. Our approach also shows competitiveperformance in speech synthesis, highlighting its potential as a universalvocoder.

【7】 Playing with Voices: Tabletop Role-Playing Game Recordings as a  Diarization Challenge
标题:玩声音:桌面角色扮演游戏录音作为日记化挑战
链接:https://arxiv.org/abs/2502.12714
作者:Lian Remme,  Kevin Tang
备注:15 pages, 14 figures, published in NAACL Findings 2025
摘要:本文提供了一个概念证明,桌面角色扮演游戏(TTRPG)的音频可以作为一个挑战,日记系统。TTRPG主要通过对话进行。参与者经常改变他们的声音,以表明他们正在作为虚构的角色说话。无论有没有技术帮助,音频处理系统都容易进行语音转换。TTRPG呈现了一种对话现象,其中语音转换是沉浸式游戏体验的固有特性。这可能会使diarizers更具挑战性,以选择真正的发言者,并确定模仿只是。我们提出了一个小TTRPG音频数据集的创建,并将其与AMI和ICSI语料库进行比较。评价了两种diarizer(pyannote.audio和wespeaker)的性能。我们观察到TTRPG的属性导致两种diarizer的混淆率更高。此外,wespeaker严重低估了TTRPG音频文件中扬声器的数量。我们提出TTRPG音频作为一个有前途的挑战日记系统。
摘要:This paper provides a proof of concept that audio of tabletop role-playinggames (TTRPG) could serve as a challenge for diarization systems. TTRPGs arecarried out mostly by conversation. Participants often alter their voices toindicate that they are talking as a fictional character. Audio processingsystems are susceptible to voice conversion with or without technologicalassistance. TTRPG present a conversational phenomenon in which voice conversionis an inherent characteristic for an immersive gaming experience. This couldmake it more challenging for diarizers to pick the real speaker and determinethat impersonating is just that. We present the creation of a small TTRPG audiodataset and compare it against the AMI and the ICSI corpus. The performance oftwo diarizers, pyannote.audio and wespeaker, were evaluated. We observed thatTTRPGs' properties result in a higher confusion rate for both diarizers.Additionally, wespeaker strongly underestimates the number of speakers in theTTRPG audio files. We propose TTRPG audio as a promising challenge fordiarization systems.

【8】 DeepResonance: Enhancing Multimodal Music Understanding via  Music-centric Multi-way Instruction Tuning
标题:DeepResonance:通过以音乐为中心的多路指令调音增强多模式音乐理解
链接:https://arxiv.org/abs/2502.12623
作者:Zhuoyuan Mao,  Mengjie Zhao,  Qiyu Wu,  Hiromi Wakaki,  Yuki Mitsufuji
摘要:音乐大语言模型(LLM)的最新进展显着改善了音乐理解任务,其中涉及模型分析和解释各种音乐元素的能力。这些改进主要集中在集成音乐和文本输入。然而,结合额外的形式,如图像,视频和文本音乐功能,以提高音乐理解的潜力仍然没有开发。为了弥合这一差距,我们提出了DeepResonance,这是一种多模式音乐理解LLM,通过多路指令调整进行微调,具有多路对齐的音乐,文本,图像和视频数据。为此,我们构建了Music 4 way-MI2 T、Music 4 way-MV 2 T和Music 4 way-Any 2 T三个4路训练和评估数据集,旨在使DeepResonance能够整合视觉和文本音乐特征内容。我们还引入了多采样ImageBind嵌入和预对齐Transformer,以在输入到文本LLM之前增强模态融合,为多路指令调整定制DeepResonance。我们的模型在六个音乐理解任务中实现了最先进的性能,突出了辅助模态的优势和DeepResonance的结构优势。我们计划开源模型和新构建的数据集。
摘要:Recent advancements in music large language models (LLMs) have significantlyimproved music understanding tasks, which involve the model's ability toanalyze and interpret various musical elements. These improvements primarilyfocused on integrating both music and text inputs. However, the potential ofincorporating additional modalities such as images, videos and textual musicfeatures to enhance music understanding remains unexplored. To bridge this gap,we propose DeepResonance, a multimodal music understanding LLM fine-tuned viamulti-way instruction tuning with multi-way aligned music, text, image, andvideo data. To this end, we construct Music4way-MI2T, Music4way-MV2T, andMusic4way-Any2T, three 4-way training and evaluation datasets designed toenable DeepResonance to integrate both visual and textual music featurecontent. We also introduce multi-sampled ImageBind embeddings and apre-alignment Transformer to enhance modality fusion prior to input into textLLMs, tailoring DeepResonance for multi-way instruction tuning. Our modelachieves state-of-the-art performances across six music understanding tasks,highlighting the benefits of the auxiliary modalities and the structuralsuperiority of DeepResonance. We plan to open-source the models and the newlyconstructed datasets.

【9】 TechSinger: Technique Controllable Multilingual Singing Voice Synthesis  via Flow Matching
标题:TechSinger:通过流匹配技术可控多语言歌唱语音合成
链接:https://arxiv.org/abs/2502.12572
作者:Wenxiang Guo,  Yu Zhang,  Changhao Pan,  Rongjie Huang,  Li Tang,  Ruiqi Li,  Zhiqing Hong,  Yongqi Wang,  Zhou Zhao
备注:Accepted by AAAI 2025
摘要:歌唱声音合成在生成自然、高质量的声音方面取得了显著的进展。然而,现有的方法很少提供精确的控制声乐技术,如强度,混合声,假声,泡沫,和呼吸音,从而限制了合成语音的表达潜力。我们介绍TechSinger,这是一个先进的可控歌唱声音合成系统,支持五种语言和七种声乐技巧。TechSinger利用基于流匹配的生成模型来产生歌声,并增强对各种技术的表达控制。为了提高训练数据的多样性,我们开发了一个技术检测模型,自动注释数据集与音素级的技术标签。此外,我们基于人工智能的技术预测模型使用户能够通过自然语言指定所需的声乐属性,对合成的歌唱进行细粒度控制。实验结果表明,TechSinger显著增强了合成歌声的表现力和真实感,在音频质量和特定技术控制方面优于现有方法。音频样本可以在https://tech-singer.github.io上找到。
摘要:Singing voice synthesis has made remarkable progress in generating naturaland high-quality voices. However, existing methods rarely provide precisecontrol over vocal techniques such as intensity, mixed voice, falsetto, bubble,and breathy tones, thus limiting the expressive potential of synthetic voices.We introduce TechSinger, an advanced system for controllable singing voicesynthesis that supports five languages and seven vocal techniques. TechSingerleverages a flow-matching-based generative model to produce singing voices withenhanced expressive control over various techniques. To enhance the diversityof training data, we develop a technique detection model that automaticallyannotates datasets with phoneme-level technique labels. Additionally, ourprompt-based technique prediction model enables users to specify desired vocalattributes through natural language, offering fine-grained control over thesynthesized singing. Experimental results demonstrate that TechSingersignificantly enhances the expressiveness and realism of synthetic singingvoices, outperforming existing methods in terms of audio quality andtechnique-specific control. Audio samples can be found athttps://tech-singer.github.io.

【10】 Myna: Masking-Based Contrastive Learning of Musical Representations
标题:八哥:基于掩蔽的音乐表现对比学习
链接:https://arxiv.org/abs/2502.12511
作者:Ori Yonay,  Tracy Hammond,  Tianbao Yang
备注:Submitted to ICML 2025
摘要:我们提出了Myna,一种简单而有效的自我监督音乐表征学习方法。基于对比学习框架,Myna引入了两项关键创新:(1)使用Vision Transformer(ViT)作为mel频谱图的骨干,以及(2)一种新颖的数据增强策略,令牌掩蔽,掩蔽了90%的频谱图令牌。这些创新提供了有效性和效率:(i)令牌掩蔽使每GPU批量大小从先前方法(CLMR,MULE)的48或120显著增加到4096。(ii)通过避免传统的增强,Myna保留了音高敏感性,提高了关键检测等任务的性能。(iii)垂直补丁的使用允许模型更好地捕获关键特征以进行关键检测。我们的混合模型Myna-22 M-Hybrid可处理16 x16和128 x2的贴片,实现了最先进的结果。在单个GPU上训练,它的平均性能优于MULE(62 M),并与分别在16和64个GPU上训练的MERT-95 M相媲美。此外,它超越了MERT-95 M-public,成为在公开数据上训练的性能最好的模型。我们发布我们的代码和模型,以促进可重复性和促进未来的研究。
摘要:We present Myna, a simple yet effective approach for self-supervised musicalrepresentation learning. Built on a contrastive learning framework, Mynaintroduces two key innovations: (1) the use of a Vision Transformer (ViT) onmel-spectrograms as the backbone and (2) a novel data augmentation strategy,token masking, that masks 90 percent of spectrogram tokens. These innovationsdeliver both effectiveness and efficiency: (i) Token masking enables asignificant increase in per-GPU batch size, from 48 or 120 in prior methods(CLMR, MULE) to 4096. (ii) By avoiding traditional augmentations, Myna retainspitch sensitivity, enhancing performance in tasks like key detection. (iii) Theuse of vertical patches allows the model to better capture critical featuresfor key detection. Our hybrid model, Myna-22M-Hybrid, processes both 16x16 and128x2 patches, achieving state-of-the-art results. Trained on a single GPU, itoutperforms MULE (62M) on average and rivals MERT-95M, which was trained on 16and 64 GPUs, respectively. Additionally, it surpasses MERT-95M-public,establishing itself as the best-performing model trained on publicly availabledata. We release our code and models to promote reproducibility and facilitatefuture research.

【11】 Note-Level Singing Melody Transcription for Time-Aligned Musical Score  Generation
标题:音符级歌唱旋律转录,用于时间一致的乐谱生成
链接:https://arxiv.org/abs/2502.12438
作者:Leekyung Kim,  Sungwook Jeon,  Wan Heo,  Jonghun Park
备注:Accepted by IEEE Transactions on Audio, Speech and Language Processing(TASLP)
摘要:自动音乐转录将音频记录转换为符号表示,便于音乐分析,检索和生成。音符的特征在于音频域中的音高、起始点和偏移,而其在乐谱域中根据音高和音符值来定义。从定时信息以及音高和音符值导出的时间对准的乐谱允许将乐谱的一部分与音乐音频的对应部分进行匹配,从而实现各种应用。在本文中,我们考虑了传统的音符级转录任务的扩展版本,该任务识别起始,偏移和音高,通过包括提取额外的音符值来从音频输入生成时间对齐的分数。为了解决这一新的挑战,我们提出了一个端到端的框架,集成了音符值,音高和时间信息的识别。这种方法避免了多阶段方法中固有的误差积累,并通过相互加强来提高精度。我们的框架采用专门针对这一任务的标记化表示,通过纳入注意值信息。此外,我们引入了一个伪标签技术,以解决注释的音符值数据的稀缺性问题。该技术从现有的数据集为传统的音符级转录产生近似的音符值标签。实验结果表明,该模型在笔记级转录任务相比,现有的国家的最先进的方法具有优越的性能。我们还引入了新的评估指标,评估时间和注意值方面,以证明模型的鲁棒性。此外,通过可视化的乐谱定性评估证实了我们的模型在捕捉音符值的有效性。
摘要:Automatic music transcription converts audio recordings into symbolicrepresentations, facilitating music analysis, retrieval, and generation. Amusical note is characterized by pitch, onset, and offset in an audio domain,whereas it is defined in terms of pitch and note value in a musical scoredomain. A time-aligned score, derived from timing information along with pitchand note value, allows matching a part of the score with the corresponding partof the music audio, enabling various applications. In this paper, we consideran extended version of the traditional note-level transcription task thatrecognizes onset, offset, and pitch, through including extraction of additionalnote value to generate a time-aligned score from an audio input. To addressthis new challenge, we propose an end-to-end framework that integratesrecognition of the note value, pitch, and temporal information. This approachavoids error accumulation inherent in multi-stage methods and enhances accuracythrough mutual reinforcement. Our framework employs tokenized representationsspecifically targeted for this task, through incorporating note valueinformation. Furthermore, we introduce a pseudo-labeling technique to address ascarcity problem of annotated note value data. This technique producesapproximate note value labels from existing datasets for the traditionalnote-level transcription. Experimental results demonstrate the superiorperformance of the proposed model in note-level transcription tasks whencompared to existing state-of-the-art approaches. We also introduce newevaluation metrics that assess both temporal and note value aspects todemonstrate the robustness of the model. Moreover, qualitative assessments viavisualized musical scores confirmed the effectiveness of our model in capturingthe note values.

【12】 An Attention-Assisted AI Model for Real-Time Underwater Sound Speed  Estimation Leveraging Remote Sensing Sea Surface Temperature Data
标题:利用遥感海面温度数据实时水下音速估计的注意力辅助人工智能模型
链接:https://arxiv.org/abs/2502.12817
作者:Pengfei Wu,  Wei Huang,  Yujie Shi,  Hao Zhang
摘要:由于声速的变化会影响信号传输的路径,因此,对水下声速分布的估计是促进有效水下通信和精确定位的关键基础。直接测量声速的常规技术以及涉及利用声场数据反演声速的方法需要现场数据收集。这一要求不仅对设备部署提出了很高的要求,而且对实现声速分布的实时估计提出了挑战。为了构建一个实时的声速场,并消除水下现场数据测量操作的需要,我们提出了一种自注意嵌入式多模态数据融合卷积神经网络(SA-CNN-CNN)的实时水下声速剖面(SSP)估计。该模型旨在阐明遥感海表温度(SST)数据,历史SSP的主要组成部分的特点,和他们的空间坐标之间的内在关系。这是通过使用CNN和注意力机制分别从输入数据中提取局部和全局相关性来实现的。最终目标是便于在指定的任务区域内的声速分布的快速和精确的估计。实验结果表明,本文提出的方法具有较低的均方根误差(RMSE)和较强的鲁棒性比其他国家的最先进的方法。
摘要:The estimation of underwater sound velocity distribution serves as a criticalbasis for facilitating effective underwater communication and precisepositioning, given that variations in sound velocity influence the path ofsignal transmission. Conventional techniques for the direct measurement ofsound velocity, as well as methods that involve the inversion of sound velocityutilizing acoustic field data, necessitate on--site data collection. Thisrequirement not only places high demands on device deployment, but alsopresents challenges in achieving real-time estimation of sound velocitydistribution. In order to construct a real-time sound velocity field andeliminate the need for underwater onsite data measurement operations, wepropose a self-attention embedded multimodal data fusion convolutional neuralnetwork (SA-MDF-CNN) for real-time underwater sound speed profile (SSP)estimation. The proposed model seeks to elucidate the inherent relationshipbetween remote sensing sea surface temperature (SST) data, the primarycomponent characteristics of historical SSPs, and their spatial coordinates.This is achieved by employing CNNs and attention mechanisms to extract localand global correlations from the input data, respectively. The ultimateobjective is to facilitate a rapid and precise estimation of sound velocitydistribution within a specified task area. Experimental results show that themethod proposed in this paper has lower root mean square error (RMSE) andstronger robustness than other state-of-the-art methods.

eess.AS音频处理

【1】 A Comprehensive Survey on Generative AI for Video-to-Music Generation
标题:视频到音乐生成的生成人工智能全面调查
链接:https://arxiv.org/abs/2502.12489
作者:Shulei Ji,  Songruoyao Wu,  Zihao Wang,  Shuyu Li,  Kejun Zhang
摘要:视频到音乐生成的蓬勃发展可以归因于多模态生成模型的优势。然而,缺乏全面梳理这一领域工作的文献。为了填补这一空白,本文全面回顾了使用深度生成AI技术生成视频到音乐的方法,重点关注三个关键部分:视觉特征提取、音乐生成框架和调节机制。我们对现有的方法进行了分类,根据它们对每个组件的设计,阐明了不同策略的作用。在此之前,我们提供了视频和音乐形态的细粒度分类,说明了不同的类别如何影响生成管道中组件的设计。此外,我们总结了可用的多模态数据集和评估指标,同时强调了该领域正在面临的挑战。
摘要:The burgeoning growth of video-to-music generation can be attributed to theascendancy of multimodal generative models. However, there is a lack ofliterature that comprehensively combs through the work in this field. To fillthis gap, this paper presents a comprehensive review of video-to-musicgeneration using deep generative AI techniques, focusing on three keycomponents: visual feature extraction, music generation frameworks, andconditioning mechanisms. We categorize existing approaches based on theirdesigns for each component, clarifying the roles of different strategies.Preceding this, we provide a fine-grained classification of video and musicmodalities, illustrating how different categories influence the design ofcomponents within the generation pipelines. Furthermore, we summarize availablemultimodal datasets and evaluation metrics while highlighting ongoingchallenges in the field.

【2】 DeepResonance: Enhancing Multimodal Music Understanding via  Music-centric Multi-way Instruction Tuning
标题:DeepResonance:通过以音乐为中心的多路指令调音增强多模式音乐理解
链接:https://arxiv.org/abs/2502.12623
作者:Zhuoyuan Mao,  Mengjie Zhao,  Qiyu Wu,  Hiromi Wakaki,  Yuki Mitsufuji
摘要:音乐大语言模型(LLM)的最新进展显着改善了音乐理解任务,其中涉及模型分析和解释各种音乐元素的能力。这些改进主要集中在集成音乐和文本输入。然而,结合额外的形式,如图像,视频和文本音乐功能,以提高音乐理解的潜力仍然没有开发。为了弥合这一差距,我们提出了DeepResonance,这是一种多模式音乐理解LLM,通过多路指令调整进行微调,具有多路对齐的音乐,文本,图像和视频数据。为此,我们构建了Music 4 way-MI2 T、Music 4 way-MV 2 T和Music 4 way-Any 2 T三个4路训练和评估数据集,旨在使DeepResonance能够整合视觉和文本音乐特征内容。我们还引入了多采样ImageBind嵌入和预对齐Transformer,以在输入到文本LLM之前增强模态融合,为多路指令调整定制DeepResonance。我们的模型在六项音乐理解任务中实现了最先进的表现,突出了辅助模态的优势和DeepResonance的结构优势。我们计划开源模型和新构建的数据集。
摘要:Recent advancements in music large language models (LLMs) have significantlyimproved music understanding tasks, which involve the model's ability toanalyze and interpret various musical elements. These improvements primarilyfocused on integrating both music and text inputs. However, the potential ofincorporating additional modalities such as images, videos and textual musicfeatures to enhance music understanding remains unexplored. To bridge this gap,we propose DeepResonance, a multimodal music understanding LLM fine-tuned viamulti-way instruction tuning with multi-way aligned music, text, image, andvideo data. To this end, we construct Music4way-MI2T, Music4way-MV2T, andMusic4way-Any2T, three 4-way training and evaluation datasets designed toenable DeepResonance to integrate both visual and textual music featurecontent. We also introduce multi-sampled ImageBind embeddings and apre-alignment Transformer to enhance modality fusion prior to input into textLLMs, tailoring DeepResonance for multi-way instruction tuning. Our modelachieves state-of-the-art performances across six music understanding tasks,highlighting the benefits of the auxiliary modalities and the structuralsuperiority of DeepResonance. We plan to open-source the models and the newlyconstructed datasets.

【3】 Note-Level Singing Melody Transcription for Time-Aligned Musical Score  Generation
标题:音符级歌唱旋律转录,用于时间一致的乐谱生成
链接:https://arxiv.org/abs/2502.12438
作者:Leekyung Kim,  Sungwook Jeon,  Wan Heo,  Jonghun Park
备注:Accepted by IEEE Transactions on Audio, Speech and Language Processing(TASLP)
摘要:自动音乐转录将音频记录转换为符号表示,便于音乐分析,检索和生成。音符的特征在于音频域中的音高、起始点和偏移,而其在乐谱域中根据音高和音符值来定义。从定时信息以及音高和音符值导出的时间对准的乐谱允许将乐谱的一部分与音乐音频的对应部分进行匹配,从而实现各种应用。在本文中,我们考虑了传统的音符级转录任务的扩展版本,该任务识别起始,偏移和音高,通过包括提取额外的音符值来从音频输入生成时间对齐的分数。为了解决这一新的挑战,我们提出了一个端到端的框架,集成了音符值,音高和时间信息的识别。这种方法避免了多阶段方法中固有的误差积累,并通过相互加强来提高精度。我们的框架采用专门针对这一任务的标记化表示,通过纳入注意值信息。此外,我们引入了一个伪标签技术,以解决注释的音符值数据的稀缺性问题。该技术从现有的数据集为传统的音符级转录产生近似的音符值标签。实验结果表明,该模型在笔记级转录任务相比,现有的国家的最先进的方法具有优越的性能。我们还引入了新的评估指标,评估时间和注意值方面,以证明模型的鲁棒性。此外,通过可视化的乐谱定性评估证实了我们的模型在捕捉音符值的有效性。
摘要:Automatic music transcription converts audio recordings into symbolicrepresentations, facilitating music analysis, retrieval, and generation. Amusical note is characterized by pitch, onset, and offset in an audio domain,whereas it is defined in terms of pitch and note value in a musical scoredomain. A time-aligned score, derived from timing information along with pitchand note value, allows matching a part of the score with the corresponding partof the music audio, enabling various applications. In this paper, we consideran extended version of the traditional note-level transcription task thatrecognizes onset, offset, and pitch, through including extraction of additionalnote value to generate a time-aligned score from an audio input. To addressthis new challenge, we propose an end-to-end framework that integratesrecognition of the note value, pitch, and temporal information. This approachavoids error accumulation inherent in multi-stage methods and enhances accuracythrough mutual reinforcement. Our framework employs tokenized representationsspecifically targeted for this task, through incorporating note valueinformation. Furthermore, we introduce a pseudo-labeling technique to address ascarcity problem of annotated note value data. This technique producesapproximate note value labels from existing datasets for the traditionalnote-level transcription. Experimental results demonstrate the superiorperformance of the proposed model in note-level transcription tasks whencompared to existing state-of-the-art approaches. We also introduce newevaluation metrics that assess both temporal and note value aspects todemonstrate the robustness of the model. Moreover, qualitative assessments viavisualized musical scores confirmed the effectiveness of our model in capturingthe note values.

机器翻译由腾讯交互翻译提供,仅供参考