本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
链接:https://arxiv.org/abs/2502.18434
备注:ISCA/ITG Workshop on Diversity in Large Speech and Language Models
摘要:这项研究调查了影响自动语音识别(ASR)系统的公平性和跨性别的性能的因素,超出了传统的人口统计学检查。使用LibriSpeech数据集和Whisper小模型,我们分析了训练数据中不同性别表示的性能差异。我们的研究结果表明,训练数据中的性别比例与ASR性能之间存在复杂的相互作用。最佳公平发生在特定的性别分布,而不是简单的50-50分裂。此外,我们的研究结果表明,音高变化等因素可以显着影响ASR的准确性。这项研究有助于更深入地了解ASR系统中的偏见,突出了精心策划的培训数据在减轻性别偏见方面的重要性。
摘要:This study investigates factors influencing Automatic Speech Recognition(ASR) systems' fairness and performance across genders, beyond the conventionalexamination of demographics. Using the LibriSpeech dataset and the Whispersmall model, we analyze how performance varies across different genderrepresentations in training data. Our findings suggest a complex interplaybetween the gender ratio in training data and ASR performance. Optimal fairnessoccurs at specific gender distributions rather than a simple 50-50 split.Furthermore, our findings suggest that factors like pitch variability cansignificantly affect ASR accuracy. This research contributes to a deeperunderstanding of biases in ASR systems, highlighting the importance ofcarefully curated training data in mitigating gender bias.
标题:从视觉到声音:利用基于视觉的算法推进音频异常检测
链接:https://arxiv.org/abs/2502.18328
摘要:视觉异常检测(VAD)的最新进展已经引入了利用预训练特征提取器生成的嵌入的复杂算法。受这些发展的启发,我们研究这些算法的适应音频域,以解决音频异常检测(AAD)的问题。与大多数现有的AAD方法(主要对异常样本进行分类)不同,我们的方法在频谱图中引入了异常的细粒度时频定位,显着提高了可解释性。此功能可以更精确地了解异常发生的位置和时间,使最终用户更容易操作结果。我们评估我们的方法对工业和环境的基准,证明VAD技术在检测音频信号中的异常的有效性。此外,它们通过启用局部异常识别来提高可解释性,使音频异常检测系统更具可解释性和实用性。
摘要:Recent advances in Visual Anomaly Detection (VAD) have introducedsophisticated algorithms leveraging embeddings generated by pre-trained featureextractors. Inspired by these developments, we investigate the adaptation ofsuch algorithms to the audio domain to address the problem of Audio AnomalyDetection (AAD). Unlike most existing AAD methods, which primarily classifyanomalous samples, our approach introduces fine-grained temporal-frequencylocalization of anomalies within the spectrogram, significantly improvingexplainability. This capability enables a more precise understanding of whereand when anomalies occur, making the results more actionable for end users. Weevaluate our approach on industrial and environmental benchmarks, demonstratingthe effectiveness of VAD techniques in detecting anomalies in audio signals.Moreover, they improve explainability by enabling localized anomalyidentification, making audio anomaly detection systems more interpretable andpractical.
标题:GCDance:由音乐驱动的流派控制的3D全身舞一代
链接:https://arxiv.org/abs/2502.18309
摘要:从音乐中生成高质量的全身舞蹈序列是一项具有挑战性的任务,因为它需要严格遵守特定类型的编舞。此外,生成的序列必须既物理上真实,又与音乐的节拍和节奏精确同步。为了克服这些挑战,我们提出了GCDance,一个无分类器的扩散框架,用于生成以音乐和文本提示为条件的特定类型的舞蹈动作。具体来说,我们的方法通过将高级预训练的音乐基础模型特征与手工制作的特征相结合来提取音乐特征,以进行多粒度特征融合。为了实现体裁可控性,我们利用CLIP在我们的舞蹈生成管道中的每个时间步有效地嵌入基于体裁的文本提示表示。我们的GCDance框架可以从同一首音乐中生成不同的舞蹈风格,同时确保与音乐的节奏和旋律的一致性。在FineDance数据集上获得的大量实验结果表明,GCDance的性能明显优于现有的最先进的方法,这些方法在AIST++数据集上也取得了有竞争力的结果。我们的消融和推理时间分析表明,GCDance提供了一个有效的解决方案,高品质的音乐驱动的舞蹈生成。
摘要:Generating high-quality full-body dance sequences from music is a challengingtask as it requires strict adherence to genre-specific choreography. Moreover,the generated sequences must be both physically realistic and preciselysynchronized with the beats and rhythm of the music. To overcome thesechallenges, we propose GCDance, a classifier-free diffusion framework forgenerating genre-specific dance motions conditioned on both music and textualprompts. Specifically, our approach extracts music features by combininghigh-level pre-trained music foundation model features with hand-craftedfeatures for multi-granularity feature fusion. To achieve genrecontrollability, we leverage CLIP to efficiently embed genre-based textualprompt representations at each time step within our dance generation pipeline.Our GCDance framework can generate diverse dance styles from the same piece ofmusic while ensuring coherence with the rhythm and melody of the music.Extensive experimental results obtained on the FineDance dataset demonstratethat GCDance significantly outperforms the existing state-of-the-artapproaches, which also achieve competitive results on the AIST++ dataset. Ourablation and inference time analysis demonstrate that GCDance provides aneffective solution for high-quality music-driven dance generation.
标题:通过上下文感知和思维链引导语言模型实现稳定的语音情感识别
链接:https://arxiv.org/abs/2502.18186
摘要:大规模音频语言模型(ALMs),如Qwen 2-Audio,能够理解不同的音频信号,执行音频分析和生成文本响应。然而,在语音情感识别(SER),ALMs经常遭受幻觉,导致错误分类或不相关的输出。为了解决这些挑战,我们提出了C$^2$SER,一种新的ALM,旨在通过上下文感知和思维链(CoT)来提高SER的稳定性和准确性。C$^2$SER集成了用于语义感知的Whisper编码器和用于声学感知的Proption 2 Vec-S,其中Proption 2 Vec-S通过半监督学习扩展了Proption 2 Vec,以增强情感识别。此外,C$^2$SER采用CoT方法,以逐步的方式处理SER,同时利用语音内容和说话风格来提高识别率。为了进一步增强稳定性,C$^2$SER引入了从显式CoT到隐式CoT的自蒸馏,减少了错误积累并提高了识别准确性。大量的实验表明,C$^2$SER优于现有流行的ALM,如Qwen 2-Audio和SECap,提供更稳定和精确的情感识别。我们发布了训练代码、检查点和测试集,以方便进一步的研究。
摘要:Large-scale audio language models (ALMs), such as Qwen2-Audio, are capable ofcomprehending diverse audio signal, performing audio analysis and generatingtextual responses. However, in speech emotion recognition (SER), ALMs oftensuffer from hallucinations, resulting in misclassifications or irrelevantoutputs. To address these challenges, we propose C$^2$SER, a novel ALM designedto enhance the stability and accuracy of SER through Contextual perception andChain of Thought (CoT). C$^2$SER integrates the Whisper encoder for semanticperception and Emotion2Vec-S for acoustic perception, where Emotion2Vec-Sextends Emotion2Vec with semi-supervised learning to enhance emotionaldiscrimination. Additionally, C$^2$SER employs a CoT approach, processing SERin a step-by-step manner while leveraging speech content and speaking styles toimprove recognition. To further enhance stability, C$^2$SER introducesself-distillation from explicit CoT to implicit CoT, mitigating erroraccumulation and boosting recognition accuracy. Extensive experiments show thatC$^2$SER outperforms existing popular ALMs, such as Qwen2-Audio and SECap,delivering more stable and precise emotion recognition. We release the trainingcode, checkpoints, and test sets to facilitate further research.
标题:基于Sinkhorn分歧的源功率最优分配确定盲源分离
链接:https://arxiv.org/abs/2502.18182
备注:Manuscript under review in TASLP
摘要:盲源分离(BSS)是指从传感器阵列记录的观测值中恢复多个源信号的过程。BSS的常用方法,包括独立向量分析(IVA)和独立低秩矩阵分析(ILRMA),通常依赖于二阶模型来捕获源信号的统计独立性以进行分离。然而,这些方法通常不考虑跨频带的隐式结构信息,这可能导致假设的源分布与从观测的混合物估计的分离的源信号的分布之间的模型失配。为了解决这些局限性,本文表明,传统的方法,如IVA和ILRMA可以很容易地利用Sinkhorn分歧,结合最佳运输(OT)框架,自适应校正源方差估计。这允许恢复源分布,同时对带间信号依赖性进行建模并跨带重新分配源功率。因此,这些算法的增强版本的开发,集成到他们的标准实现的Sinkhorn迭代计划。大量的仿真结果表明,所提出的方法一致提高BSS性能。
摘要:Blind source separation (BSS) refers to the process of recovering multiplesource signals from observations recorded by an array of sensors. Commonapproaches to BSS, including independent vector analysis (IVA), and independentlow-rank matrix analysis (ILRMA), typically rely on second-order models tocapture the statistical independence of source signals for separation. However,these methods generally do not account for the implicit structural informationacross frequency bands, which may lead to model mismatches between the assumedsource distributions and the distributions of the separated source signalsestimated from the observed mixtures. To tackle these limitations, this papershows that conventional approaches such as IVA and ILRMA can easily beleveraged by the Sinkhorn divergence, incorporating an optimal transport (OT)framework to adaptively correct source variance estimates. This allows for therecovery of the source distribution while modeling the inter-band signaldependence and reallocating source power across bands. As a result, enhancedversions of these algorithms are developed, integrating a Sinkhorn iterativescheme into their standard implementations. Extensive simulations demonstratethat the proposed methods consistently enhance BSS performance.
标题:NotaGen:利用大型语言模型训练范式提升象征性音乐生成中的音乐性
链接:https://arxiv.org/abs/2502.18008
摘要:我们介绍了NotaGen,一个符号音乐生成模型,旨在探索生产高品质的古典乐谱的潜力。受大型语言模型(LLM)成功的启发,NotaGen采用了预训练、微调和强化学习范式(以下简称LLM训练范式)。它预先训练了160万首音乐,然后根据“时期作曲家乐器”提示对大约9 K的高质量古典作品进行微调。对于强化学习,我们提出了CLaMP-DPO方法,它进一步提高了生成质量和可控性,而不需要人工注释或预定义的奖励。我们的实验证明了CLaMP-DPO在具有不同架构和编码方案的符号音乐生成模型中的有效性。此外,主观A/B测试表明,NotaGen在人类作品中的表现优于基线模型,极大地推进了符号音乐生成中的音乐美学。https://electricalexis.github.io/notagen-demo
摘要:We introduce NotaGen, a symbolic music generation model aiming to explore thepotential of producing high-quality classical sheet music. Inspired by thesuccess of Large Language Models (LLMs), NotaGen adopts pre-training,fine-tuning, and reinforcement learning paradigms (henceforth referred to asthe LLM training paradigms). It is pre-trained on 1.6M pieces of music, andthen fine-tuned on approximately 9K high-quality classical compositionsconditioned on "period-composer-instrumentation" prompts. For reinforcementlearning, we propose the CLaMP-DPO method, which further enhances generationquality and controllability without requiring human annotations or predefinedrewards. Our experiments demonstrate the efficacy of CLaMP-DPO in symbolicmusic generation models with different architectures and encoding schemes.Furthermore, subjective A/B tests show that NotaGen outperforms baseline modelsagainst human compositions, greatly advancing musical aesthetics in symbolicmusic generation.The project homepage ishttps://electricalexis.github.io/notagen-demo.
标题:通过集成BGRU和Transformer架构提高语音质量
链接:https://arxiv.org/abs/2502.17911
摘要:语音增强是提高噪声环境下语音信号质量的重要手段。本文研究了集成双向门控递归单元(BGRU)和Transformer模型用于语音增强任务的有效性。通过一个全面的实验评估,我们的研究表明,这种混合架构的优越性,传统的方法和独立的模型。组合的BGRU-Transformer框架在捕获时间依赖性和学习复杂信号模式方面表现出色,从而增强降噪和提高语音质量。结果表明,显着的性能增益相比,现有的方法,突出了这种集成模型在现实世界中的应用潜力。BGRU和Transformer架构的无缝集成不仅增强了系统的鲁棒性,而且为先进的语音处理技术开辟了道路。该研究有助于语音增强技术的持续努力,并为未来优化模型架构,探索许多应用场景以及推进噪声环境下的语音处理领域奠定了坚实的基础。
摘要:Speech enhancement plays an essential role in improving the quality of speechsignals in noisy environments. This paper investigates the efficacy ofintegrating Bidirectional Gated Recurrent Units (BGRU) and Transformer modelsfor speech enhancement tasks. Through a comprehensive experimental evaluation,our study demonstrates the superiority of this hybrid architecture overtraditional methods and standalone models. The combined BGRU-Transformerframework excels in capturing temporal dependencies and learning complex signalpatterns, leading to enhanced noise reduction and improved speech quality.Results show significant performance gains compared to existing approaches,highlighting the potential of this integrated model in real-world applications.The seamless integration of BGRU and Transformer architectures not onlyenhances system robustness but also opens the road for advanced speechprocessing techniques. This research contributes to the ongoing efforts inspeech enhancement technology and sets a solid foundation for futureinvestigations into optimizing model architectures, exploring many applicationscenarios, and advancing the field of speech processing in noisy environments.
标题:使用共形器和CSC算法的六轴加速度计识别无声语音句子
链接:https://arxiv.org/abs/2502.17829
摘要:无声语音接口(SSI)正在积极开发,以帮助长期遭受日常困难和生活质量下降的沟通障碍的个人。然而,由于省略和连接,无声句子很难分割和识别。提出了一种新的无声语音句子识别方法,将六轴加速度传感器采集的人脸运动信号转换为转录的单词和句子。一个基于Conformer的神经网络与Connectionist-Temporal-Classification算法被用来获得上下文的理解和翻译的非声学信号成单词序列,只要求在数据库中的组成词。测试结果表明,该方法实现了97.17%的准确率在句子识别中,超过了现有的无声语音识别方法的典型准确率为85%-95%,并展示了潜在的加速度计作为一个可用的SSI模态的高精度无声语音句子识别。
摘要:Silent speech interfaces (SSI) are being actively developed to assistindividuals with communication impairments who have long suffered from dailyhardships and a reduced quality of life. However, silent sentences aredifficult to segment and recognize due to elision and linking. A novel silentspeech sentence recognition method is proposed to convert the facial motionsignals collected by six-axis accelerometers into transcribed words andsentences. A Conformer-based neural network with theConnectionist-Temporal-Classification algorithm is used to gain contextualunderstanding and translate the non-acoustic signals into words sequences,solely requesting the constituent words in the database. Test results show thatthe proposed method achieves a 97.17% accuracy in sentence recognition,surpassing the existing silent speech recognition methods with a typicalaccuracy of 85%-95%, and demonstrating the potential of accelerometers as anavailable SSI modality for high-accuracy silent speech sentence recognition.
标题:具有表现性音乐表现检测功能的Gigailon数据集
链接:https://arxiv.org/abs/2502.17726
备注:Published at Transactions of the International Society for Music Information Retrieval (TISMIR), 8(1), 1-19
摘要:1983年推出的乐器数字接口(Musical Instrument Digital Interface,简称MIDI)通过允许计算机和乐器有效地通信,彻底改变了音乐制作。音乐文件编码音乐指令compatible,促进方便的音乐共享。它们有利于音乐信息检索(MIR),有助于音乐理解,计算音乐学和生成音乐的研究。GigaNotes数据集包含超过140万个独特的日志文件,包括18亿个日志事件和超过530万个日志跟踪。GigaMusic是目前最大的符号音乐收藏,以供公平交易下的研究目的。区分非表达性和表达性的XML文件是一个挑战,因为XML文件本身并不进行这种区分。为了解决这个问题,我们介绍了一套创新的检测表现力的音乐表现。其中包括独特音符速度比(DNVR)启发式算法,用于分析MIDI音符速度;独特音符开始偏差比(DNODR)启发式算法,用于检查音符开始时间的偏差;以及音符开始中值度量水平(NOMML)启发式算法,用于评估相对于度量水平的开始位置。我们的评估表明,这些算法有效地区分了非表达性和表达性的轨迹。此外,经过评估,我们创建了最丰富的表现力的数据集,采用我们的启发式,NOMML。GigaExclusive的这次策划迭代包含了NOMML检测到的表现力很强的乐器曲目,包含了所有的General Exclusive乐器,占GigaExclusive数据集的31%,共计1,655,649个曲目。
摘要:The Musical Instrument Digital Interface (MIDI), introduced in 1983,revolutionized music production by allowing computers and instruments tocommunicate efficiently. MIDI files encode musical instructions compactly,facilitating convenient music sharing. They benefit Music Information Retrieval(MIR), aiding in research on music understanding, computational musicology, andgenerative music. The GigaMIDI dataset contains over 1.4 million unique MIDIfiles, encompassing 1.8 billion MIDI note events and over 5.3 million MIDItracks. GigaMIDI is currently the largest collection of symbolic music in MIDIformat available for research purposes under fair dealing. Distinguishingbetween non-expressive and expressive MIDI tracks is challenging, as MIDI filesdo not inherently make this distinction. To address this issue, we introduce aset of innovative heuristics for detecting expressive music performance. Theseinclude the Distinctive Note Velocity Ratio (DNVR) heuristic, which analyzesMIDI note velocity; the Distinctive Note Onset Deviation Ratio (DNODR)heuristic, which examines deviations in note onset times; and the Note OnsetMedian Metric Level (NOMML) heuristic, which evaluates onset positions relativeto metric levels. Our evaluation demonstrates these heuristics effectivelydifferentiate between non-expressive and expressive MIDI tracks. Furthermore,after evaluation, we create the most substantial expressive MIDI dataset,employing our heuristic, NOMML. This curated iteration of GigaMIDI encompassesexpressively-performed instrument tracks detected by NOMML, containing allGeneral MIDI instruments, constituting 31% of the GigaMIDI dataset, totalling1,655,649 tracks.
标题:VANPY:语音分析框架
链接:https://arxiv.org/abs/2502.17579
摘要:语音数据越来越多地用于现代数字通信,但仍然缺乏用于自动语音分析和表征的综合工具。为此,我们开发了VANPY(Python中的语音分析)框架,用于语音数据的自动预处理,特征提取和分类。VANPY是一个开源端到端综合框架,旨在根据语音数据进行说话者特征描述。该框架在设计时考虑了可扩展性,允许轻松集成新组件并适应各种语音分析应用程序。它目前集成了超过15个语音分析组件-包括音乐/语音分离,语音活动检测,说话人嵌入,声音特征提取和各种分类模型。 VANPY的四个组件是内部开发的,并集成到框架中,以扩展其说话人特征描述功能:性别分类,情感分类,年龄回归和身高回归。这些模型在各种数据集上表现出强大的性能,尽管没有超过最先进的性能。 作为概念证明,我们在分析电影《低俗小说》中角色声音的用例挑战中展示了该框架提取说话者特征的能力。“结果说明了该框架提取多个说话者特征的能力,包括性别、年龄、身高、情绪类型和情绪强度,这些特征是在三个维度上测量的:唤醒、支配和效价。
摘要:Voice data is increasingly being used in modern digital communications, yetthere is still a lack of comprehensive tools for automated voice analysis andcharacterization. To this end, we developed the VANPY (Voice Analysis inPython) framework for automated pre-processing, feature extraction, andclassification of voice data. The VANPY is an open-source end-to-endcomprehensive framework that was developed for the purpose of speakercharacterization from voice data. The framework is designed with extensibilityin mind, allowing for easy integration of new components and adaptation tovarious voice analysis applications. It currently incorporates over fifteenvoice analysis components - including music/speech separation, voice activitydetection, speaker embedding, vocal feature extraction, and variousclassification models. Four of the VANPY's components were developed in-house and integrated intothe framework to extend its speaker characterization capabilities: genderclassification, emotion classification, age regression, and height regression.The models demonstrate robust performance across various datasets, although notsurpassing state-of-the-art performance. As a proof of concept, we demonstrate the framework's ability to extractspeaker characteristics on a use-case challenge of analyzing character voicesfrom the movie "Pulp Fiction." The results illustrate the framework'scapability to extract multiple speaker characteristics, including gender, age,height, emotion type, and emotion intensity measured across three dimensions:arousal, dominance, and valence.
标题:通过深频谱包封整形用音乐进行感知噪音掩蔽
链接:https://arxiv.org/abs/2502.17527
备注:None
摘要:人们经常在嘈杂的环境中听音乐,试图将自己与周围的声音隔离开来。实际上,由于同时掩蔽的效果,音乐信号可以掩蔽噪声的一些频率分量。在这篇文章中,我们提出了一种基于心理声学掩蔽模型的神经网络,旨在通过预测滤波器频率响应重塑其频谱包络来增强音乐掩蔽环境噪声的能力。该模型使用感知损失函数进行训练,该函数平衡了两个约束:有效地掩蔽噪声,同时保留原始音乐混音和用户选择的收听水平。我们评估我们的模拟数据的方法复制用户的体验听音乐的耳机在嘈杂的环境中。基于定义的客观指标的结果表明,我们的系统提高了最先进的水平。
摘要:People often listen to music in noisy environments, seeking to isolatethemselves from ambient sounds. Indeed, a music signal can mask some of thenoise's frequency components due to the effect of simultaneous masking. In thisarticle, we propose a neural network based on a psychoacoustic masking model,designed to enhance the music's ability to mask ambient noise by reshaping itsspectral envelope with predicted filter frequency responses. The model istrained with a perceptual loss function that balances two constraints:effectively masking the noise while preserving the original music mix and theuser's chosen listening level. We evaluate our approach on simulated datareplicating a user's experience of listening to music with headphones in anoisy environment. The results, based on defined objective metrics, demonstratethat our system improves the state of the art.
标题:探索自动语音识别技术中的性别差异
链接:https://arxiv.org/abs/2502.18434
备注:ISCA/ITG Workshop on Diversity in Large Speech and Language Models
摘要:这项研究调查了影响自动语音识别(ASR)系统的公平性和跨性别的性能的因素,超出了传统的人口统计学检查。使用LibriSpeech数据集和Whisper小模型,我们分析了训练数据中不同性别表示的性能差异。我们的研究结果表明,训练数据中的性别比例与ASR性能之间存在复杂的相互作用。最佳公平发生在特定的性别分布,而不是简单的50-50分裂。此外,我们的研究结果表明,音高变化等因素可以显着影响ASR的准确性。这项研究有助于更深入地了解ASR系统中的偏见,突出了精心策划的培训数据在减轻性别偏见方面的重要性。
摘要:This study investigates factors influencing Automatic Speech Recognition(ASR) systems' fairness and performance across genders, beyond the conventionalexamination of demographics. Using the LibriSpeech dataset and the Whispersmall model, we analyze how performance varies across different genderrepresentations in training data. Our findings suggest a complex interplaybetween the gender ratio in training data and ASR performance. Optimal fairnessoccurs at specific gender distributions rather than a simple 50-50 split.Furthermore, our findings suggest that factors like pitch variability cansignificantly affect ASR accuracy. This research contributes to a deeperunderstanding of biases in ASR systems, highlighting the importance ofcarefully curated training data in mitigating gender bias.
标题:从视觉到声音:利用基于视觉的算法推进音频异常检测
链接:https://arxiv.org/abs/2502.18328
摘要:视觉异常检测(VAD)的最新进展已经引入了利用预训练特征提取器生成的嵌入的复杂算法。受这些发展的启发,我们研究这些算法的适应音频域,以解决音频异常检测(AAD)的问题。与大多数现有的AAD方法(主要对异常样本进行分类)不同,我们的方法在频谱图中引入了异常的细粒度时频定位,显着提高了可解释性。此功能可以更精确地了解异常发生的位置和时间,使最终用户更容易操作结果。我们评估我们的方法对工业和环境的基准,证明VAD技术在检测音频信号中的异常的有效性。此外,它们通过启用局部异常识别来提高可解释性,使音频异常检测系统更具可解释性和实用性。
摘要:Recent advances in Visual Anomaly Detection (VAD) have introducedsophisticated algorithms leveraging embeddings generated by pre-trained featureextractors. Inspired by these developments, we investigate the adaptation ofsuch algorithms to the audio domain to address the problem of Audio AnomalyDetection (AAD). Unlike most existing AAD methods, which primarily classifyanomalous samples, our approach introduces fine-grained temporal-frequencylocalization of anomalies within the spectrogram, significantly improvingexplainability. This capability enables a more precise understanding of whereand when anomalies occur, making the results more actionable for end users. Weevaluate our approach on industrial and environmental benchmarks, demonstratingthe effectiveness of VAD techniques in detecting anomalies in audio signals.Moreover, they improve explainability by enabling localized anomalyidentification, making audio anomaly detection systems more interpretable andpractical.
标题:GCDance:由音乐驱动的流派控制的3D全身舞一代
链接:https://arxiv.org/abs/2502.18309
摘要:从音乐中生成高质量的全身舞蹈序列是一项具有挑战性的任务,因为它需要严格遵守特定类型的编舞。此外,生成的序列必须既物理上真实,又与音乐的节拍和节奏精确同步。为了克服这些挑战,我们提出了GCDance,一个无分类器的扩散框架,用于生成以音乐和文本提示为条件的特定类型的舞蹈动作。具体来说,我们的方法通过将高级预训练的音乐基础模型特征与手工制作的特征相结合来提取音乐特征,以进行多粒度特征融合。为了实现体裁可控性,我们利用CLIP在我们的舞蹈生成管道中的每个时间步有效地嵌入基于体裁的文本提示表示。我们的GCDance框架可以从同一首音乐中生成不同的舞蹈风格,同时确保与音乐的节奏和旋律的一致性。在FineDance数据集上获得的大量实验结果表明,GCDance的性能明显优于现有的最先进的方法,这些方法在AIST++数据集上也取得了有竞争力的结果。我们的消融和推理时间分析表明,GCDance提供了一个有效的解决方案,高品质的音乐驱动的舞蹈生成。
摘要:Generating high-quality full-body dance sequences from music is a challengingtask as it requires strict adherence to genre-specific choreography. Moreover,the generated sequences must be both physically realistic and preciselysynchronized with the beats and rhythm of the music. To overcome thesechallenges, we propose GCDance, a classifier-free diffusion framework forgenerating genre-specific dance motions conditioned on both music and textualprompts. Specifically, our approach extracts music features by combininghigh-level pre-trained music foundation model features with hand-craftedfeatures for multi-granularity feature fusion. To achieve genrecontrollability, we leverage CLIP to efficiently embed genre-based textualprompt representations at each time step within our dance generation pipeline.Our GCDance framework can generate diverse dance styles from the same piece ofmusic while ensuring coherence with the rhythm and melody of the music.Extensive experimental results obtained on the FineDance dataset demonstratethat GCDance significantly outperforms the existing state-of-the-artapproaches, which also achieve competitive results on the AIST++ dataset. Ourablation and inference time analysis demonstrate that GCDance provides aneffective solution for high-quality music-driven dance generation.
标题:通过上下文感知和思维链引导语言模型实现稳定的语音情感识别
链接:https://arxiv.org/abs/2502.18186
摘要:大规模音频语言模型(ALMs),如Qwen 2-Audio,能够理解不同的音频信号,执行音频分析和生成文本响应。然而,在语音情感识别(SER),ALMs经常遭受幻觉,导致错误分类或不相关的输出。为了解决这些挑战,我们提出了C$^2$SER,一种新的ALM,旨在通过上下文感知和思维链(CoT)来提高SER的稳定性和准确性。C$^2$SER集成了用于语义感知的Whisper编码器和用于声学感知的Proption 2 Vec-S,其中Proption 2 Vec-S通过半监督学习扩展了Proption 2 Vec,以增强情感识别。此外,C$^2$SER采用CoT方法,以逐步的方式处理SER,同时利用语音内容和说话风格来提高识别率。为了进一步增强稳定性,C$^2$SER引入了从显式CoT到隐式CoT的自蒸馏,减少了错误积累并提高了识别准确性。大量的实验表明,C$^2$SER优于现有流行的ALM,如Qwen 2-Audio和SECap,提供更稳定和精确的情感识别。我们发布了训练代码、检查点和测试集,以方便进一步的研究。
摘要:Large-scale audio language models (ALMs), such as Qwen2-Audio, are capable ofcomprehending diverse audio signal, performing audio analysis and generatingtextual responses. However, in speech emotion recognition (SER), ALMs oftensuffer from hallucinations, resulting in misclassifications or irrelevantoutputs. To address these challenges, we propose C$^2$SER, a novel ALM designedto enhance the stability and accuracy of SER through Contextual perception andChain of Thought (CoT). C$^2$SER integrates the Whisper encoder for semanticperception and Emotion2Vec-S for acoustic perception, where Emotion2Vec-Sextends Emotion2Vec with semi-supervised learning to enhance emotionaldiscrimination. Additionally, C$^2$SER employs a CoT approach, processing SERin a step-by-step manner while leveraging speech content and speaking styles toimprove recognition. To further enhance stability, C$^2$SER introducesself-distillation from explicit CoT to implicit CoT, mitigating erroraccumulation and boosting recognition accuracy. Extensive experiments show thatC$^2$SER outperforms existing popular ALMs, such as Qwen2-Audio and SECap,delivering more stable and precise emotion recognition. We release the trainingcode, checkpoints, and test sets to facilitate further research.
标题:基于Sinkhorn分歧的源功率最优分配确定盲源分离
链接:https://arxiv.org/abs/2502.18182
备注:Manuscript under review in TASLP
摘要:盲源分离(BSS)是指从传感器阵列记录的观测值中恢复多个源信号的过程。BSS的常用方法,包括独立向量分析(IVA)和独立低秩矩阵分析(ILRMA),通常依赖于二阶模型来捕获源信号的统计独立性以进行分离。然而,这些方法通常不考虑跨频带的隐式结构信息,这可能导致假设的源分布和从观测的混合物估计的分离的源信号的分布之间的模型失配。为了解决这些局限性,本文表明,传统的方法,如IVA和ILRMA可以很容易地利用Sinkhorn分歧,结合最佳运输(OT)框架,自适应校正源方差估计。这允许恢复源分布,同时对带间信号依赖性进行建模并跨带重新分配源功率。因此,这些算法的增强版本的开发,集成到他们的标准实现的Sinkhorn迭代计划。大量的仿真结果表明,所提出的方法一致提高BSS性能。
摘要:Blind source separation (BSS) refers to the process of recovering multiplesource signals from observations recorded by an array of sensors. Commonapproaches to BSS, including independent vector analysis (IVA), and independentlow-rank matrix analysis (ILRMA), typically rely on second-order models tocapture the statistical independence of source signals for separation. However,these methods generally do not account for the implicit structural informationacross frequency bands, which may lead to model mismatches between the assumedsource distributions and the distributions of the separated source signalsestimated from the observed mixtures. To tackle these limitations, this papershows that conventional approaches such as IVA and ILRMA can easily beleveraged by the Sinkhorn divergence, incorporating an optimal transport (OT)framework to adaptively correct source variance estimates. This allows for therecovery of the source distribution while modeling the inter-band signaldependence and reallocating source power across bands. As a result, enhancedversions of these algorithms are developed, integrating a Sinkhorn iterativescheme into their standard implementations. Extensive simulations demonstratethat the proposed methods consistently enhance BSS performance.
标题:NotaGen:利用大型语言模型训练范式提升象征性音乐生成中的音乐性
链接:https://arxiv.org/abs/2502.18008
摘要:我们介绍了NotaGen,一个符号音乐生成模型,旨在探索生产高品质的古典乐谱的潜力。受大型语言模型(LLM)成功的启发,NotaGen采用了预训练、微调和强化学习范式(以下简称LLM训练范式)。它预先训练了160万首音乐,然后根据“时期作曲家乐器”提示对大约9 K的高质量古典作品进行微调。对于强化学习,我们提出了CLaMP-DPO方法,它进一步提高了生成质量和可控性,而不需要人工注释或预定义的奖励。我们的实验证明了CLaMP-DPO在具有不同架构和编码方案的符号音乐生成模型中的有效性。此外,主观A/B测试表明,NotaGen在人类作品中的表现优于基线模型,极大地推进了符号音乐生成中的音乐美学。https://electricalexis.github.io/notagen-demo
摘要:We introduce NotaGen, a symbolic music generation model aiming to explore thepotential of producing high-quality classical sheet music. Inspired by thesuccess of Large Language Models (LLMs), NotaGen adopts pre-training,fine-tuning, and reinforcement learning paradigms (henceforth referred to asthe LLM training paradigms). It is pre-trained on 1.6M pieces of music, andthen fine-tuned on approximately 9K high-quality classical compositionsconditioned on "period-composer-instrumentation" prompts. For reinforcementlearning, we propose the CLaMP-DPO method, which further enhances generationquality and controllability without requiring human annotations or predefinedrewards. Our experiments demonstrate the efficacy of CLaMP-DPO in symbolicmusic generation models with different architectures and encoding schemes.Furthermore, subjective A/B tests show that NotaGen outperforms baseline modelsagainst human compositions, greatly advancing musical aesthetics in symbolicmusic generation.The project homepage ishttps://electricalexis.github.io/notagen-demo.
标题:通过集成BGRU和Transformer架构提高语音质量
链接:https://arxiv.org/abs/2502.17911
摘要:语音增强是提高噪声环境下语音信号质量的重要手段。本文研究了集成双向门控递归单元(BGRU)和Transformer模型用于语音增强任务的有效性。通过一个全面的实验评估,我们的研究表明,这种混合架构的优越性,传统的方法和独立的模型。组合的BGRU-Transformer框架在捕获时间依赖性和学习复杂信号模式方面表现出色,从而增强降噪和提高语音质量。结果表明,显着的性能增益相比,现有的方法,突出了这种集成模型在现实世界中的应用潜力。BGRU和Transformer架构的无缝集成不仅增强了系统的鲁棒性,而且为先进的语音处理技术开辟了道路。该研究有助于语音增强技术的持续努力,并为未来优化模型架构,探索许多应用场景以及推进噪声环境下的语音处理领域奠定了坚实的基础。
摘要:Speech enhancement plays an essential role in improving the quality of speechsignals in noisy environments. This paper investigates the efficacy ofintegrating Bidirectional Gated Recurrent Units (BGRU) and Transformer modelsfor speech enhancement tasks. Through a comprehensive experimental evaluation,our study demonstrates the superiority of this hybrid architecture overtraditional methods and standalone models. The combined BGRU-Transformerframework excels in capturing temporal dependencies and learning complex signalpatterns, leading to enhanced noise reduction and improved speech quality.Results show significant performance gains compared to existing approaches,highlighting the potential of this integrated model in real-world applications.The seamless integration of BGRU and Transformer architectures not onlyenhances system robustness but also opens the road for advanced speechprocessing techniques. This research contributes to the ongoing efforts inspeech enhancement technology and sets a solid foundation for futureinvestigations into optimizing model architectures, exploring many applicationscenarios, and advancing the field of speech processing in noisy environments.
标题:使用共形器和CSC算法的六轴加速度计识别无声语音句子
链接:https://arxiv.org/abs/2502.17829
摘要:无声语音接口(SSI)正在积极开发,以帮助长期遭受日常困难和生活质量下降的沟通障碍的个人。然而,由于省略和连接,无声句子很难分割和识别。提出了一种新的无声语音句子识别方法,将六轴加速度传感器采集的人脸运动信号转换为转录的单词和句子。一个基于Conformer的神经网络与Connectionist-Temporal-Classification算法被用来获得上下文的理解和翻译的非声学信号成单词序列,只要求在数据库中的组成词。测试结果表明,该方法实现了97.17%的准确率在句子识别中,超过了现有的无声语音识别方法的典型准确率为85%-95%,并展示了潜在的加速度计作为一个可用的SSI模态的高精度无声语音句子识别。
摘要:Silent speech interfaces (SSI) are being actively developed to assistindividuals with communication impairments who have long suffered from dailyhardships and a reduced quality of life. However, silent sentences aredifficult to segment and recognize due to elision and linking. A novel silentspeech sentence recognition method is proposed to convert the facial motionsignals collected by six-axis accelerometers into transcribed words andsentences. A Conformer-based neural network with theConnectionist-Temporal-Classification algorithm is used to gain contextualunderstanding and translate the non-acoustic signals into words sequences,solely requesting the constituent words in the database. Test results show thatthe proposed method achieves a 97.17% accuracy in sentence recognition,surpassing the existing silent speech recognition methods with a typicalaccuracy of 85%-95%, and demonstrating the potential of accelerometers as anavailable SSI modality for high-accuracy silent speech sentence recognition.
标题:URO长凳:端到端口语对话模型的综合基准
链接:https://arxiv.org/abs/2502.17810
摘要:近年来,随着大型语言模型(LLM)的发展,端到端口语对话模型(SDM)取得了重大进展。与基于文本的LLM相比,SDM的评估需要考虑语音相关的方面,例如语言信息和语音质量。然而,在语音到语音(S2 S)场景中仍然缺乏对SDM的全面评估。为了解决这一差距,我们提出了URO-Bench,一个广泛的SDMs基准。值得注意的是,URO-Bench是第一个涵盖多语言、多轮对话和多语言评估的S2 S基准。我们的基准分为两个难度级别:基本轨道和专业轨道,分别由16和20个数据集组成,评估模型在理解,推理和口语对话方面的能力。对我们提出的基准测试的评估表明,目前的开源SDM在日常QA任务中表现相当不错,但在预防跟踪能力方面落后于其骨干LLM,并且还遭受灾难性遗忘。他们的表现在高级评估的语言信息和音频理解仍然低于标准,强调需要在这一方向进一步研究。我们希望URO-Bench能够通过对现有模式进行多方面评估并帮助跟踪这一领域的进展,有效促进口语对话模式的发展。
摘要:In recent years, with advances in large language models (LLMs), end-to-endspoken dialogue models (SDMs) have made significant strides. Compared totext-based LLMs, the evaluation of SDMs needs to take speech-related aspectsinto account, such as paralinguistic information and speech quality. However,there is still a lack of comprehensive evaluations for SDMs in speech-to-speech(S2S) scenarios. To address this gap, we propose URO-Bench, an extensivebenchmark for SDMs. Notably, URO-Bench is the first S2S benchmark that coversevaluations about multilingualism, multi-round dialogues, and paralinguistics.Our benchmark is divided into two difficulty levels: basic track and pro track,consisting of 16 and 20 datasets respectively, evaluating the model's abilitiesin Understanding, Reasoning, and Oral conversation. Evaluations on our proposedbenchmark reveal that current open-source SDMs perform rather well in daily QAtasks, but lag behind their backbone LLMs in terms of instruction-followingability and also suffer from catastrophic forgetting. Their performance inadvanced evaluations of paralinguistic information and audio understandingremains subpar, highlighting the need for further research in this direction.We hope that URO-Bench can effectively facilitate the development of spokendialogue models by providing a multifaceted evaluation of existing models andhelping to track progress in this area.
标题:具有表现性音乐表现检测功能的Gigailon数据集
链接:https://arxiv.org/abs/2502.17726
备注:Published at Transactions of the International Society for Music Information Retrieval (TISMIR), 8(1), 1-19
摘要:1983年推出的乐器数字接口(Musical Instrument Digital Interface,简称MIDI)通过允许计算机和乐器有效地通信,彻底改变了音乐制作。音乐文件编码音乐指令compatible,促进方便的音乐共享。它们有利于音乐信息检索(MIR),有助于音乐理解,计算音乐学和生成音乐的研究。GigaNotes数据集包含超过140万个独特的日志文件,包括18亿个日志事件和超过530万个日志跟踪。GigaMusic是目前最大的符号音乐收藏,以供公平交易下的研究目的。区分非表达性和表达性的XML文件是一个挑战,因为XML文件本身并不进行这种区分。为了解决这个问题,我们介绍了一套创新的检测表现力的音乐表现。这些包括独特音符速度比(DNVR)启发式算法,它分析音符的速度;独特音符开始偏差比(DNODR)启发式算法,它检查音符开始时间的偏差;和音符开始中值度量水平(NOMML)启发式算法,它评估相对于度量水平的开始位置。我们的评估表明,这些算法有效地区分了非表达性和表达性的轨迹。此外,经过评估,我们创建了最丰富的表现力的数据集,采用我们的启发式,NOMML。GigaExclusive的这次策划迭代包含了NOMML检测到的表现力很强的乐器曲目,包含了所有的General Exclusive乐器,占GigaExclusive数据集的31%,共计1,655,649个曲目。
摘要:The Musical Instrument Digital Interface (MIDI), introduced in 1983,revolutionized music production by allowing computers and instruments tocommunicate efficiently. MIDI files encode musical instructions compactly,facilitating convenient music sharing. They benefit Music Information Retrieval(MIR), aiding in research on music understanding, computational musicology, andgenerative music. The GigaMIDI dataset contains over 1.4 million unique MIDIfiles, encompassing 1.8 billion MIDI note events and over 5.3 million MIDItracks. GigaMIDI is currently the largest collection of symbolic music in MIDIformat available for research purposes under fair dealing. Distinguishingbetween non-expressive and expressive MIDI tracks is challenging, as MIDI filesdo not inherently make this distinction. To address this issue, we introduce aset of innovative heuristics for detecting expressive music performance. Theseinclude the Distinctive Note Velocity Ratio (DNVR) heuristic, which analyzesMIDI note velocity; the Distinctive Note Onset Deviation Ratio (DNODR)heuristic, which examines deviations in note onset times; and the Note OnsetMedian Metric Level (NOMML) heuristic, which evaluates onset positions relativeto metric levels. Our evaluation demonstrates these heuristics effectivelydifferentiate between non-expressive and expressive MIDI tracks. Furthermore,after evaluation, we create the most substantial expressive MIDI dataset,employing our heuristic, NOMML. This curated iteration of GigaMIDI encompassesexpressively-performed instrument tracks detected by NOMML, containing allGeneral MIDI instruments, constituting 31% of the GigaMIDI dataset, totalling1,655,649 tracks.
标题:VANPY:语音分析框架
链接:https://arxiv.org/abs/2502.17579
摘要:语音数据越来越多地用于现代数字通信,但仍然缺乏用于自动语音分析和表征的综合工具。为此,我们开发了VANPY(Python中的语音分析)框架,用于语音数据的自动预处理,特征提取和分类。VANPY是一个开源端到端综合框架,旨在根据语音数据进行说话者特征描述。该框架在设计时考虑了可扩展性,允许轻松集成新组件并适应各种语音分析应用程序。它目前集成了超过15个语音分析组件-包括音乐/语音分离,语音活动检测,说话人嵌入,声音特征提取和各种分类模型。 VANPY的四个组件是内部开发的,并集成到框架中,以扩展其说话人特征描述功能:性别分类,情感分类,年龄回归和身高回归。这些模型在各种数据集上表现出强大的性能,尽管没有超过最先进的性能。 作为一个概念证明,我们展示了框架的能力,提取扬声器特征的用例的挑战,分析人物的声音,从电影“低俗小说。“结果说明了该框架提取多个说话者特征的能力,包括性别、年龄、身高、情绪类型和情绪强度,这些特征是在三个维度上测量的:唤醒、支配和效价。
摘要:Voice data is increasingly being used in modern digital communications, yetthere is still a lack of comprehensive tools for automated voice analysis andcharacterization. To this end, we developed the VANPY (Voice Analysis inPython) framework for automated pre-processing, feature extraction, andclassification of voice data. The VANPY is an open-source end-to-endcomprehensive framework that was developed for the purpose of speakercharacterization from voice data. The framework is designed with extensibilityin mind, allowing for easy integration of new components and adaptation tovarious voice analysis applications. It currently incorporates over fifteenvoice analysis components - including music/speech separation, voice activitydetection, speaker embedding, vocal feature extraction, and variousclassification models. Four of the VANPY's components were developed in-house and integrated intothe framework to extend its speaker characterization capabilities: genderclassification, emotion classification, age regression, and height regression.The models demonstrate robust performance across various datasets, although notsurpassing state-of-the-art performance. As a proof of concept, we demonstrate the framework's ability to extractspeaker characteristics on a use-case challenge of analyzing character voicesfrom the movie "Pulp Fiction." The results illustrate the framework'scapability to extract multiple speaker characteristics, including gender, age,height, emotion type, and emotion intensity measured across three dimensions:arousal, dominance, and valence.
标题:通过深频谱包封整形用音乐进行感知噪音掩蔽
链接:https://arxiv.org/abs/2502.17527
备注:None
摘要:人们经常在嘈杂的环境中听音乐,试图将自己与周围的声音隔离开来。实际上,由于同时掩蔽的效果,音乐信号可以掩蔽噪声的一些频率分量。在这篇文章中,我们提出了一种基于心理声学掩蔽模型的神经网络,旨在通过预测滤波器频率响应重塑其频谱包络来增强音乐掩蔽环境噪声的能力。该模型使用感知损失函数进行训练,该函数平衡了两个约束:有效地掩蔽噪声,同时保留原始音乐混音和用户选择的收听水平。我们评估我们的模拟数据的方法复制用户的体验听音乐的耳机在嘈杂的环境中。基于定义的客观指标的结果表明,我们的系统提高了最先进的水平。
摘要:People often listen to music in noisy environments, seeking to isolatethemselves from ambient sounds. Indeed, a music signal can mask some of thenoise's frequency components due to the effect of simultaneous masking. In thisarticle, we propose a neural network based on a psychoacoustic masking model,designed to enhance the music's ability to mask ambient noise by reshaping itsspectral envelope with predicted filter frequency responses. The model istrained with a perceptual loss function that balances two constraints:effectively masking the noise while preserving the original music mix and theuser's chosen listening level. We evaluate our approach on simulated datareplicating a user's experience of listening to music with headphones in anoisy environment. The results, based on defined objective metrics, demonstratethat our system improves the state of the art.
