本文经arXiv每日学术速递授权转载
链接:https://arxiv.org/abs/2502.08191
摘要:目标说话人提取的核心是利用注册信息从多个说话人的环境中提取目标语音信号。现有的方法主要依赖于从注册中获得的说话人嵌入,潜在地忽略了上下文信息以及混合和注册之间的内部交互。在本文中,我们提出了一种新的双流上下文融合网络(DualStream Contextual Fusion Network,DCF-Net)在时间-频率(T-F)域。具体而言,DualStream Fusion Block(DSFB)被引入以获得上下文信息并捕获跨空间和通道维度的上下文化注册和混合表示之间的交互,然后利用丰富且一致的表示来指导提取网络以实现更好的提取。实验结果表明,DCF-Net优于最先进的(SOTA)方法,在基准数据集上实现了21.6 dB的尺度不变信号失真比改善(SI-SDRi),并在噪声和混响场景中表现出其鲁棒性和有效性。此外,我们的模型的错误提取结果,称为目标混淆问题,减少到0.4%,这突出了DCF-Net的实际应用的潜力。
摘要:Target speaker extraction focuses on extracting a target speech signal froman environment with multiple speakers by leveraging an enrollment. Existingmethods predominantly rely on speaker embeddings obtained from the enrollment,potentially disregarding the contextual information and the internalinteractions between the mixture and enrollment. In this paper, we propose anovel DualStream Contextual Fusion Network (DCF-Net) in the time-frequency(T-F) domain. Specifically, DualStream Fusion Block (DSFB) is introduced toobtain contextual information and capture the interactions betweencontextualized enrollment and mixture representation across both spatial andchannel dimensions, and then rich and consistent representations are utilizedto guide the extraction network for better extraction. Experimental resultsdemonstrate that DCF-Net outperforms state-of-the-art (SOTA) methods, achievinga scale-invariant signal-to-distortion ratio improvement (SI-SDRi) of 21.6 dBon the benchmark dataset, and exhibits its robustness and effectiveness in bothnoise and reverberation scenarios. In addition, the wrong extraction results ofour model, called target confusion problem, reduce to 0.4%, which highlightsthe potential of DCF-Net for practical applications.
标题:当代流行音乐中的音调分析方法:强调Primaal商业作品中的音调不确定性
链接:https://arxiv.org/abs/2502.08131
备注:44 pages, 11 figures
摘要:我们确定的特点是如何音高是由超音乐,一个主流的商业音乐公司,专门从事广告音乐的全球性公司的表达目的操纵。研究表明,该公司的“Primaal”品牌中音高的使用和组织不同于西方古典音乐。通过对制作人的采访和对他们作品的深入分析,我们发现他们的方法集中在一个有意识的目标上,即构建一个基于音高不确定性的音乐话语,与西方古典传统中明确定义的音高的清晰传递形成对比。据Primaal制片人说,他们承认坎耶·韦斯特和傻朋克等艺术家的影响,并使用广泛可用的技术,音高的不确定性吸引了听众的注意力。我们提供了音乐摘录的分析,展示了他们的方法,以及为实现其表达目标所采用的工具和方法的描述。这些目标和方法被置于更广泛的历史背景中,与西方音乐中音高组织的基本原则形成对比。Hyper Music用来引入和控制音高不确定性的技术包括提高上分音,表现性地使用不和谐,围绕与特定“模式”相关的“极点”的连续音高分布,以及不断发展的音高。我们从心理声学的角度来研究这些技术,并进行听力测试证实了一些意见。本研究的最终目的是介绍一套适合于当代流行音乐音高分析的方法。
摘要:We identify characteristic features of how pitch is manipulated forexpressive purposes by Hyper Music, a mainstream commercial music companyspecialising in advertisement music for global corporations. The study showsthat the use and organisation of pitch in the company's `Primaal' brand differsfrom Western classical music. Through interviews with producers and in-depthanalysis of their work, we reveal that their methods centre on a conscious aimto construct a musical discourse based on pitch uncertainty, contrasting withthe clear transmission of well-defined pitches in Western classical traditions.According to the Primaal producers, who acknowledge the influence of artistssuch as Kanye West and Daft Punk and use widely available technology, pitchuncertainty captures the listener's attention. We provide analyses of musicalexcerpts demonstrating their approach, alongside descriptions of the tools andmethods employed to achieve their expressive goals. These goals and methods areplaced in a broader historical context, contrasting with fundamental principlesof pitch organisation in Western music. Techniques used by Hyper Music tointroduce and control pitch uncertainty include boosting upper partials,expressive use of inharmonicity, continuous pitch distributions around 'poles'tied to specific 'modes', and continuously evolving pitch. We examine thesetechniques from a psychoacoustic perspective, and conduct listening testscorroborating some of the observations. The ultimate goal of the study is tointroduce a set of methods suited to the analysis of pitch in contemporarypopular music.
标题:Hookpad Aria:词曲作者的副驾驶
链接:https://arxiv.org/abs/2502.08122
备注:Extended abstract presented in the Late-Breaking Demo Session at ISMIR 2024 (ISMIR LBD 2024)
摘要:Hookpad Aria是一个生成式AI系统,旨在帮助音乐家创作西方流行歌曲。我们的系统是无缝集成到Hookpad,一个基于网络的编辑器设计的组成铅表:象征性的乐谱,描述旋律和和声。Hookpad Aria具有许多生成功能,旨在帮助用户进行非顺序的合成工作流程,包括:(1)生成现有材料的从左到右的延续,(2)填充现有材料中间缺失的跨度,以及(3)从旋律生成和声,反之亦然。Hookpad Aria也是音乐共同创作的可扩展数据飞轮-自2024年3月发布以来,Aria已经为3 k用户生成了318 k建议,这些用户已经接受了74 k的歌曲。 有关Hookpad Aria的更多信息,请访问https://www.hooktheory.com/hookpad/aria
摘要:We present Hookpad Aria, a generative AI system designed to assist musiciansin writing Western pop songs. Our system is seamlessly integrated into Hookpad,a web-based editor designed for the composition of lead sheets: symbolic musicscores that describe melody and harmony. Hookpad Aria has numerous generationcapabilities designed to assist users in non-sequential composition workflows,including: (1) generating left-to-right continuations of existing material, (2)filling in missing spans in the middle of existing material, and (3) generatingharmony from melody and vice versa. Hookpad Aria is also a scalable dataflywheel for music co-creation -- since its release in March 2024, Aria hasgenerated 318k suggestions for 3k users who have accepted 74k into their songs. More information about Hookpad Aria is available athttps://www.hooktheory.com/hookpad/aria
标题:使用boostlet的稀疏波场重建和去噪
链接:https://arxiv.org/abs/2502.08230
备注:5 pages, 4 figures
摘要:Boostlets是时空函数,将非色散波场分解为由膨胀、双曲线旋转和平移参数化的局部波形集合。我们研究的稀疏性的boostlets,并发现所得到的分解是显着稀疏比其他国家的最先进的表示系统,如小波和剪切波。这转化为在对boostlet系数进行硬阈值化时改进的去噪性能。结果表明,boostlets提供了一个自然的框架稀疏分解波场在统一的时空。
摘要:Boostlets are spatiotemporal functions that decompose nondispersivewavefields into a collection of localized waveforms parametrized by dilations,hyperbolic rotations, and translations. We study the sparsity properties ofboostlets and find that the resulting decompositions are significantly sparserthan those of other state-of-the-art representation systems, such as waveletsand shearlets. This translates into improved denoising performance whenhard-thresholding the boostlet coefficients. The results suggest that boostletsoffer a natural framework for sparsely decomposing wavefields in unifiedspace-time.
标题:儿童ASB错误的原因分析:量化生理、认知和外在因素的影响
链接:https://arxiv.org/abs/2502.08587
备注:Submitted to Computer Speech & Language
摘要:近年来,儿童自动语音识别(ASR)系统的使用越来越多,这促使人们努力提高为儿童语音设计的模型的准确性。目前的方法要么直接利用开源语音基础模型(SFM),要么用儿童的语音数据对其进行微调。这些SFM,无论是开源的还是针对儿童的微调,与成人语音相比,通常表现出更高的单词错误率(WER)。然而,有一个缺乏系统的分析,这种性能下降的可持续森林管理系统的原因。理解和解决这种表现差异背后的原因对于提高儿童语音SFM的准确性至关重要。我们的研究通过调查准确性下降的原因和儿童言语中WER的主要贡献者来解决这一差距。在第一部分的研究中,我们进行了全面的基准研究两个自我监督的SFM(Wav2Vec2. 0和Hubert)和两个弱监督的SFM(耳语和MMS)在不同年龄组的两个儿童的语音语料库,建立原始数据的因果推理分析的第二部分。在研究的第二部分中,我们分析了生理因素(年龄,性别),认知因素(发音能力),和外部因素(词汇难度,背景噪音,字数)的影响,在儿童的言语中使用因果推理的SFM准确性。结果表明,生理(年龄)和特定的外部因素(音频中的字数)对准确率的影响最大,其次是背景噪音和发音能力。微调儿童语音的SFM降低了对生理和认知因素的敏感性,而对音频中单词数量的敏感性仍然存在。 关键词:儿童ASR,言语基础模型,因果推理,生理学,认知,发音
摘要:The increasing use of children's automatic speech recognition (ASR) systemshas spurred research efforts to improve the accuracy of models designed forchildren's speech in recent years. The current approach utilizes eitheropen-source speech foundation models (SFMs) directly or fine-tuning them withchildren's speech data. These SFMs, whether open-source or fine-tuned forchildren, often exhibit higher word error rates (WERs) compared to adultspeech. However, there is a lack of systemic analysis of the cause of thisdegraded performance of SFMs. Understanding and addressing the reasons behindthis performance disparity is crucial for improving the accuracy of SFMs forchildren's speech. Our study addresses this gap by investigating the causes ofaccuracy degradation and the primary contributors to WER in children's speech.In the first part of the study, we conduct a comprehensive benchmarking studyon two self-supervised SFMs (Wav2Vec2.0 and Hubert) and two weakly supervisedSFMs (Whisper and MMS) across various age groups on two children speechcorpora, establishing the raw data for the causal inference analysis in thesecond part. In the second part of the study, we analyze the impact ofphysiological factors (age, gender), cognitive factors (pronunciation ability),and external factors (vocabulary difficulty, background noise, and word count)on SFM accuracy in children's speech using causal inference. The resultsindicate that physiology (age) and particular external factor (number of wordsin audio) have the highest impact on accuracy, followed by background noise andpronunciation ability. Fine-tuning SFMs on children's speech reducessensitivity to physiological and cognitive factors, while sensitivity to thenumber of words in audio persists. Keywords: Children's ASR, Speech Foundational Models, Causal Inference,Physiology, Cognition, Pronunciation
标题:使用boostlet的稀疏波场重建和去噪
链接:https://arxiv.org/abs/2502.08230
备注:5 pages, 4 figures
摘要:Boostlets是时空函数,将非色散波场分解为由膨胀、双曲线旋转和平移参数化的局部波形集合。我们研究的稀疏性的boostlets,并发现所得到的分解是显着稀疏比其他国家的最先进的表示系统,如小波和剪切波。这转化为在对boostlet系数进行硬阈值化时改进的去噪性能。结果表明,boostlets提供了一个自然的框架稀疏分解波场在统一的时空。
摘要:Boostlets are spatiotemporal functions that decompose nondispersivewavefields into a collection of localized waveforms parametrized by dilations,hyperbolic rotations, and translations. We study the sparsity properties ofboostlets and find that the resulting decompositions are significantly sparserthan those of other state-of-the-art representation systems, such as waveletsand shearlets. This translates into improved denoising performance whenhard-thresholding the boostlet coefficients. The results suggest that boostletsoffer a natural framework for sparsely decomposing wavefields in unifiedspace-time.
标题:DualStream上下文融合网络:通过利用混合和注册交互来高效提取目标说话者
链接:https://arxiv.org/abs/2502.08191
摘要:目标说话人提取的核心是利用注册信息从多个说话人的环境中提取目标语音信号。现有的方法主要依赖于从注册中获得的说话人嵌入,潜在地忽略了上下文信息以及混合和注册之间的内部交互。在本文中,我们提出了一种新的双流上下文融合网络(DualStream Contextual Fusion Network,DCF-Net)在时间-频率(T-F)域。具体而言,DualStream Fusion Block(DSFB)被引入以获得上下文信息并捕获跨空间和通道维度的上下文化注册和混合表示之间的交互,然后利用丰富且一致的表示来指导提取网络以实现更好的提取。实验结果表明,DCF-Net优于最先进的(SOTA)方法,在基准数据集上实现了21.6 dB的尺度不变信号失真比改善(SI-SDRi),并在噪声和混响场景中表现出其鲁棒性和有效性。此外,我们的模型的错误提取结果,称为目标混淆问题,减少到0.4%,这突出了DCF-Net的实际应用的潜力。
摘要:Target speaker extraction focuses on extracting a target speech signal froman environment with multiple speakers by leveraging an enrollment. Existingmethods predominantly rely on speaker embeddings obtained from the enrollment,potentially disregarding the contextual information and the internalinteractions between the mixture and enrollment. In this paper, we propose anovel DualStream Contextual Fusion Network (DCF-Net) in the time-frequency(T-F) domain. Specifically, DualStream Fusion Block (DSFB) is introduced toobtain contextual information and capture the interactions betweencontextualized enrollment and mixture representation across both spatial andchannel dimensions, and then rich and consistent representations are utilizedto guide the extraction network for better extraction. Experimental resultsdemonstrate that DCF-Net outperforms state-of-the-art (SOTA) methods, achievinga scale-invariant signal-to-distortion ratio improvement (SI-SDRi) of 21.6 dBon the benchmark dataset, and exhibits its robustness and effectiveness in bothnoise and reverberation scenarios. In addition, the wrong extraction results ofour model, called target confusion problem, reduce to 0.4%, which highlightsthe potential of DCF-Net for practical applications.
标题:neuro2 panel:从神经活动中解码发声
链接:https://arxiv.org/abs/2502.07800
备注:Master Thesis
摘要:由于神经尖峰的固有稀疏性和长度以及大脑回路的复杂性,准确解码神经尖峰序列并将其与运动输出相关联是一项具有挑战性的任务。这个硕士项目研究解码斑胸草雀运动输出(在离散音节和连续频谱图)的实验方法,从神经像素获得的侵入性神经记录。 主要成果有三:(1)XGBoost结合SHAP分析揭示了神经元间的相互作用模式,这对音节分类至关重要。(2)新方法(用GPT 2标记神经数据)和架构(Mamba 2)证明了使用尖峰信号解码音节的潜力。(3)一个组合的对比学习-VAE框架成功地从分箱的神经数据中生成了谱图。 这项工作为复杂运动输出的神经解码奠定了良好的基础,并为处理稀疏神经数据提供了几种新的方法。
摘要:Accurate decoding of neural spike trains and relating them to motor output isa challenging task due to the inherent sparsity and length in neural spikes andthe complexity of brain circuits. This master project investigates experimentalmethods for decoding zebra finch motor outputs (in both discrete syllables andcontinuous spectrograms), from invasive neural recordings obtained fromNeuropixels. There are three major achievements: (1) XGBoost with SHAP analysis trained onspike rates revealed neuronal interaction patterns crucial for syllableclassification. (2) Novel method (tokenizing neural data with GPT2) andarchitecture (Mamba2) demonstrated potential for decoding of syllables usingspikes. (3) A combined contrastive learning-VAE framework successfullygenerated spectrograms from binned neural data. This work establishes a promising foundation for neural decoding of complexmotor outputs and offers several novel methodological approaches for processingsparse neural data.
