今日论文合集:cs.SD语音10篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Which Prosodic Features Matter Most for Pragmatics?
标题: 哪些韵律特征对修辞学最重要?
作者:Nigel G. Ward,Divette Marco,Olac Fuentes
备注:Submitted to ICASSP 2025. Audio illustrations available at this https URL
链接:点击下载PDF文件
摘要:我们调查的韵律功能最重要的传达韵律功能。我们使用的问题,预测人类感知的话语对之间的语用相似性,以评估不同类型的韵律特征的效用。例如,我们发现,与时长相关的特征比与音高相关的特征更重要,并且话语初始特征比话语最终特征更重要。此外,故障分析表明,使用音高功能的建模往往无法处理重要的语用功能,并建议几个通常被忽视的声学和韵律功能是语用意义重大,包括鼻音和颤音。这些发现可以指导未来的韵律基础研究,并建议如何提高语音合成评估,以及其他应用。摘要:We investigate which prosodic features matter most in conveying prosodic functions. We use the problem of predicting human perceptions of pragmatic similarity among utterance pairs to evaluate the utility of prosodic features of different types. We find, for example, that duration-related features are more important than pitch-related features, and that utterance-initial features are more important than utterance-final features. Further, failure analysis indicates that modeling using pitch features only often fails to handle important pragmatic functions, and suggests that several generally-neglected acoustic and prosodic features are pragmatically significant, including nasality and vibrato. These findings can guide future basic research in prosody, and suggest how to improve speech synthesis evaluation, among other applications.

【2】 EAViT: External Attention Vision Transformer for Audio Classification
标题: EAViT:音频分类的外部注意力视觉Transformer
作者:Aquib Iqbal,Abid Hasan Zim,Md Asaduzzaman Tonmoy,Limengnan Zhou,Asad Malik,Minoru Kuribayashi
链接:点击下载PDF文件
摘要:本文提出了外部注意力Vision Transformer(EAViT)模型,一种新的方法,旨在提高音频分类精度。随着数字音频资源的激增,对精确和高效的音频分类系统的需求已经加剧,这是由各种应用(包括音乐流媒体平台和环境声音识别)中对改进的推荐系统和用户个性化的需求所驱动的。准确的音频分类对于将庞大的音频库组织成连贯的类别至关重要,使用户能够更有效地找到他们喜欢的音频内容并与之交互。在这项研究中,我们利用GTZAN数据集,其中包括1,000个音乐摘录,跨越10个不同的流派。每个30秒的音频片段被分割成3秒的摘录,以增强数据集的鲁棒性并减轻过拟合风险,从而允许更细粒度的特征分析。EAViT模型将多头外部注意力(MEA)机制集成到Vision Transformer(ViT)框架中,有效地捕获样本之间的长期依赖性和潜在相关性。这种外部注意力(EA)机制采用可学习的记忆单元,增强了网络有效处理复杂音频特征的能力。研究表明,EAViT的总体准确度高达93.99%,超越了最先进的模型。摘要:This paper presents the External Attention Vision Transformer (EAViT) model, a novel approach designed to enhance audio classification accuracy. As digital audio resources proliferate, the demand for precise and efficient audio classification systems has intensified, driven by the need for improved recommendation systems and user personalization in various applications, including music streaming platforms and environmental sound recognition. Accurate audio classification is crucial for organizing vast audio libraries into coherent categories, enabling users to find and interact with their preferred audio content more effectively. In this study, we utilize the GTZAN dataset, which comprises 1,000 music excerpts spanning ten diverse genres. Each 30-second audio clip is segmented into 3-second excerpts to enhance dataset robustness and mitigate overfitting risks, allowing for more granular feature analysis. The EAViT model integrates multi-head external attention (MEA) mechanisms into the Vision Transformer (ViT) framework, effectively capturing long-range dependencies and potential correlations between samples. This external attention (EA) mechanism employs learnable memory units that enhance the network's capacity to process complex audio features efficiently. The study demonstrates that EAViT achieves a remarkable overall accuracy of 93.99%, surpassing state-of-the-art models.

【3】 Coarse-to-fine Alignment Makes Better Speech-image Retrieval
标题: 粗到细对齐可实现更好的语音图像检索
作者:Lifeng Zhou,Yuke Li
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个新的语音图像检索框架。我们利用语音-图像对比(SIC)学习任务在粗层次上对齐语音和图像表示,并利用语音-图像匹配(SIM)学习任务进一步细化细粒度的跨模态对齐。SIC和SIM学习任务以统一的方式联合训练。为了优化学习过程,我们利用嵌入队列来促进在SIC学习期间对高质量和多样化的负面表示进行高效采样。此外,它通过有效地挖掘基于SIC任务中计算的对比相似性的硬否定来增强SIM任务的学习。为了进一步优化噪声监督下的学习,我们将动量蒸馏纳入训练过程。实验结果表明,我们的框架优于国家的最先进的方法在R@1 超过4%的语音图像检索任务的两个基准数据集。此外,在zero-shot实验中观察到,我们的框架表现出出色的泛化能力。摘要:In this paper, we propose a novel framework for speech-image retrieval. We utilize speech-image contrastive (SIC) learning tasks to align speech and image representations at a coarse level and speech-image matching (SIM) learning tasks to further refine the fine-grained cross-modal alignment. SIC and SIM learning tasks are jointly trained in a unified manner. To optimize the learning process, we utilize an embedding queue that facilitates efficient sampling of high-quality and diverse negative representations during SIC learning. Additionally, it enhances the learning of SIM tasks by effectively mining hard negatives based on contrastive similarities calculated in SIC tasks. To further optimize learning under noisy supervision, we incorporate momentum distillation into the training process. Experimental results show that our framework outperforms the state-of-the-art method by more than 4% in R@1 on two benchmark datasets for the speech-image retrieval tasks. Moreover, as observed in zero-shot experiments, our framework demonstrates excellent generalization capabilities.

【4】 NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
标题: NEST:自我监督的快速Conformer作为语音处理任务的通用调味剂
作者:He Huang,Taejin Park,Kunal Dhawan,Ivan Medennikov,Krishna C. Puvvada,Nithin Rao Koluguri,Weiqing Wang,Jagadeesh Balam,Boris Ginsburg
链接:点击下载PDF文件
摘要:自监督学习已被证明有利于广泛的语音处理任务,如语音识别 翻译,说话人确认和日记等,但是,这些方法大多是计算密集型的,由于使用Transformer编码器和缺乏子采样。在本文中,我们提出了一个新的自监督学习模型称为神经编码器的自监督训练(NEST)。具体来说,我们采用FastConformer架构,该架构具有8倍子采样率,比Transformer或Conformer架构更快。而不是基于聚类的令牌生成,我们诉诸于固定的随机投影,其简单性和有效性。我们还提出了一个广义的嘈杂的语音增强,教模型解开噪声或其他扬声器的主要扬声器。实验表明,所提出的NEST模型改善了现有的自监督模型在各种语音处理任务。代码和检查点将通过NVIDIA NeMo工具包公开提供。摘要:Self-supervised learning has been proved to benefit a wide range of speech processing tasks, such as speech recognition translation, speaker verification and diarization, etc. However, most of these approaches are computationally intensive due to using transformer encoder and lack of sub-sampling. In this paper, we propose a new self-supervised learning model termed as Neural Encoder for Self-supervised Training (NEST). Specifically, we adopt the FastConformer architecture, which has an 8x sub-sampling rate and is faster than Transformer or Conformer architectures. Instead of clustering-based token generation, we resort to fixed random projection for its simplicity and effectiveness. We also propose a generalized noisy speech augmentation that teaches the model to disentangle the main speaker from noise or other speakers. Experiments show that the proposed NEST model improves over existing self-supervised models on a variety of speech processing tasks. Code and checkpoints will be publicly available via NVIDIA NeMo toolkit.

【5】 On Class Separability Pitfalls In Audio-Text Contrastive Zero-Shot Learning
标题: 音频文本对比Zero-Shot学习中的班级可分离性陷阱
作者:Tiago Tavares,Fabio Ayres,Zhepei Wang,Paris Smaragdis
链接:点击下载PDF文件
摘要:音频-文本跨通道对比学习的最新进展显示了其向zero-shot学习方向发展的潜力。一种可能性是将来自预先训练的骨干神经网络的项目嵌入投影到跨模态空间中,在该跨模态空间中,可以在任一域中计算项目相似性。这个过程依赖于骨干网络的强大的单峰预训练,以及投影仪的数据密集型训练任务。这两个过程可能会因无意的数据泄漏而产生偏差,这可能是由于在预训练中使用监督学习或使用来自zero-shot学习评估的标签无意中训练跨模态投影而引起的。在这项研究中,我们表明,测量的zero-shot学习准确性的一个重要部分是由于从音频和文本骨干继承的优势,也就是说,他们不是在跨模态域学习,并没有从一个模态转移到另一个。摘要:Recent advances in audio-text cross-modal contrastive learning have shown its potential towards zero-shot learning. One possibility for this is by projecting item embeddings from pre-trained backbone neural networks into a cross-modal space in which item similarity can be calculated in either domain. This process relies on a strong unimodal pre-training of the backbone networks, and on a data-intensive training task for the projectors. These two processes can be biased by unintentional data leakage, which can arise from using supervised learning in pre-training or from inadvertently training the cross-modal projection using labels from the zero-shot learning evaluation. In this study, we show that a significant part of the measured zero-shot learning accuracy is due to strengths inherited from the audio and text backbones, that is, they are not learned in the cross-modal domain and are not transferred from one modality to another.

【6】 Uncertainty-Aware Mean Opinion Score Prediction
标题: 不确定性意识的平均意见分数预测
作者:Hui Wang,Shiwan Zhao,Jiaming Zhou,Xiguang Zheng,Haoqin Sun,Xuechen Wang,Yong Qin
备注:Accepted by Interspeech 2024, oral
链接:点击下载PDF文件
摘要:平均意见得分(MOS)预测在特定领域取得了重大进展。然而,MOS预测模型在不同样本中的不稳定性能在这些系统的实际应用中提出了持续的挑战。在本文中,我们指出,不确定性建模的缺乏是一个显着的限制,阻碍MOS预测系统应用到现实和开放的世界。我们分析了MOS预测任务中不确定性的来源,并提出建立一个具有不确定性意识的MOS预测系统,该系统分别采用异方差回归和蒙特卡罗退出来建模偶然不确定性和认知不确定性。实验结果表明,该系统能够很好地捕捉不确定性,并能够进行选择性预测和域外检测。这种能力大大增强了MOS系统在各种真实和开放世界环境中的实际效用。摘要:Mean Opinion Score (MOS) prediction has made significant progress in specific domains. However, the unstable performance of MOS prediction models across diverse samples presents ongoing challenges in the practical application of these systems. In this paper, we point out that the absence of uncertainty modeling is a significant limitation hindering MOS prediction systems from applying to the real and open world. We analyze the sources of uncertainty in the MOS prediction task and propose to establish an uncertainty-aware MOS prediction system that models aleatory uncertainty and epistemic uncertainty by heteroscedastic regression and Monte Carlo dropout separately. The experimental results show that the system captures uncertainty well and is capable of performing selective prediction and out-of-domain detection. Such capabilities significantly enhance the practical utility of MOS systems in diverse real and open-world environments.

【7】 Towards measuring fairness in speech recognition: Fair-Speech dataset
标题: 衡量语音识别中的公平性:公平语音数据集
作者:Irina-Elena Veliche,Zhuangqun Huang,Vineeth Ayyat Kochaniyan,Fuchun Peng,Ozlem Kalinli,Michael L. Seltzer
链接:点击下载PDF文件
摘要:当前的公共语音识别数据集(ASR)往往不会专门关注公平性方面,例如不同人口群体的性能。本文介绍了一个新的数据集,公平演讲,一个公开发布的语料库,以帮助研究人员评估他们的ASR模型的准确性在一组不同的自我报告的人口统计信息,如年龄,性别,种族,地理变异和参与者是否认为自己的母语是英语。我们的数据集包括美国593人录制的大约26.5K的话语,这些人付费录制并提交自己说语音命令的音频。我们还提供ASR基线,包括在转录和未转录的社交媒体视频和开源模型上训练的模型。摘要:The current public datasets for speech recognition (ASR) tend not to focus specifically on the fairness aspect, such as performance across different demographic groups. This paper introduces a novel dataset, Fair-Speech, a publicly released corpus to help researchers evaluate their ASR models for accuracy across a diverse set of self-reported demographic information, such as age, gender, ethnicity, geographic variation and whether the participants consider themselves native English speakers. Our dataset includes approximately 26.5K utterances in recorded speech by 593 people in the United States, who were paid to record and submit audios of themselves saying voice commands. We also provide ASR baselines, including on models trained on transcribed and untranscribed social media videos and open source models.

【8】 Hierarchical Generative Modeling of Melodic Vocal Contours in Hindustani Classical Music
标题: 印度斯坦古典音乐旋律声乐轮廓的分层生成建模
作者:Nithya Shikarpur,Krishna Maneesha Dendukur,Yusong Wu,Antoine Caillon,Cheng-Zhi Anna Huang
备注:Accepted at International Society for Music Information Retrieval (ISMIR) 2024
链接:点击下载PDF文件
摘要:印度斯坦音乐是一种表演驱动的口头传统,展示了丰富的旋律模式。在本文中,我们专注于生成建模的歌手的声乐旋律提取的录音,因为声音是音乐突出的传统。在印度斯坦音乐中,先前的生成工作将旋律建模为粗糙的离散符号,无法捕捉歌唱中丰富的表达性旋律的复杂性。因此,我们建议使用精细量化的音高轮廓,作为分层音频建模的中间表示。我们提出了GaMaDHaNi,一个模块化的两级层次结构,包括一个生成模型的音高轮廓,和一个音高轮廓音频合成模型。我们比较我们的方法,非分层音频模型和分层模型,使用自监督的中间表示,通过收听测试和定性分析。我们还评估了音频模型的能力,忠实地代表音高轮廓输入使用皮尔逊相关系数。通过使用音高轮廓作为中间表示,我们表明,我们的模型可以通过突出两个潜在的交互用例(1)启动生成和(2)粗音高调节,更好地在人类-AI协作环境中倾听和响应音乐家。摘要:Hindustani music is a performance-driven oral tradition that exhibits the rendition of rich melodic patterns. In this paper, we focus on generative modeling of singers' vocal melodies extracted from audio recordings, as the voice is musically prominent within the tradition. Prior generative work in Hindustani music models melodies as coarse discrete symbols which fails to capture the rich expressive melodic intricacies of singing. Thus, we propose to use a finely quantized pitch contour, as an intermediate representation for hierarchical audio modeling. We propose GaMaDHaNi, a modular two-level hierarchy, consisting of a generative model on pitch contours, and a pitch contour to audio synthesis model. We compare our approach to non-hierarchical audio models and hierarchical models that use a self-supervised intermediate representation, through a listening test and qualitative analysis. We also evaluate audio model's ability to faithfully represent the pitch contour input using Pearson correlation coefficient. By using pitch contours as an intermediate representation, we show that our model may be better equipped to listen and respond to musicians in a human-AI collaborative setting by highlighting two potential interaction use cases (1) primed generation, and (2) coarse pitch conditioning.

【9】 Information and motor constraints shape melodic diversity across cultures
标题: 信息和运动限制塑造了跨文化的旋律多样性
作者:John M McBride,Nahie Kim,Yuri Nishikawa,Mekhmed Saadakeev,Marcus T Pearce,Tsvi Tlusty
链接:点击下载PDF文件
摘要:可能的旋律的数量是相当大的,然而,尽管旋律变化的潜力几乎是无限的,来自不同社会的旋律可以惊人地相似。运动约束假说解释了某些相似性,如标量运动和轮廓形状,但不解释其他主要的共同特征,如重复,歌曲长度和音阶大小。在这里,我们调查的信息限制所产生的人类记忆的局限性,在塑造这些标志性的旋律的作用。我们测量了跨越几大洲的62个民间旋律语料库的信息率的决定因素,发现多个权衡,所有的行为,以限制跨社会的信息率。相比之下,来自欧洲(包括土耳其)的39个艺术音乐语料库显示更长,更复杂的旋律,随着时间的推移,复杂性增加,这表明艺术和民间音乐中不同的文化进化选择压力,可能是由于使用书面与口头传播。我们的无参数模型预测的经验尺度度分布使用标量运动,旋律长度,最重要的是,信息率的信息约束。这提供了强有力的证据表明,在音乐的文化传播过程中的信息约束限制了音阶中音符的数量,并提出了对中间旋律复杂性的偏好是对旋律文化演变的根本约束。摘要:The number of possible melodies is unfathomably large, yet despite this virtually unlimited potential for melodic variation, melodies from different societies can be surprisingly similar. The motor constraint hypothesis accounts for certain similarities, such as scalar motion and contour shape, but not for other major common features, such as repetition, song length, and scale size. Here we investigate the role of information constraints arising from limitations on human memory in shaping these hallmarks of melodies. We measure determinants of information rate in 62 corpora of Folk melodies spanning several continents, finding multiple trade-offs that all act to constrain the information rate across societies. By contrast, 39 corpora of Art music from Europe (including Turkey) show longer, more complex melodies, and increased complexity over time, suggesting different cultural-evolutionary selection pressures in Art and Folk music, possibly due to the use of written versus oral transmission. Our parameter-free model predicts the empirical scale degree distribution using information constraints on scalar motion, melody length, and, most importantly, information rate. This provides strong evidence that information constraints during cultural transmission of music limit the number of notes in a scale, and proposes that preference for intermediate melodic complexity is a fundamental constraint on the cultural evolution of melody.

【10】 Melody predominates over harmony in the evolution of musical scales across 96 countries
标题: 在96个国家的音阶演变中,旋律优先于和声
作者:John M McBride,Elizabeth Phillips,Patrick E Savage,Steven Brown,Tsvi Tlusty
链接:点击下载PDF文件
摘要:自古以来,音阶的标准理论是基于和声,而不是旋律。最近的一些分析支持这两种观点,我们缺乏跨文化数据的比较测试。我们通过对主要理论与来自96个国家的1,314个量表进行严格的计算比较来解决这个长期存在的问题。旋律理论几乎得到了普遍的支持,它预测1-3个音阶的步长。和声解释了一些简单整数比音程的流行,特别是欧亚社会的音乐理论音阶,这可能解释了它们在西方学者中的统治地位。然而,和谐预测从人种学记录,特别是欧亚大陆以外的规模。总的来说,我们表明,历史上强调和谐是误导,旋律是世界音乐音阶的主要决定因素。摘要:The standard theory of musical scales since antiquity has been based on harmony, rather than melody. Some recent analyses support either view, and we lack a comparative test on cross-cultural data. We address this longstanding problem through a rigorous, computational comparison of the main theories against 1,314 scales from 96 countries. There is near-universal support for melodic theories, which predict step-sizes of 1-3 semitones. Harmony accounts for the prevalence of some simple-integer-ratio intervals, particularly for music-theoretic scales from Eurasian societies, which may explain their dominance amongst Western scholars. However, harmony poorly predicts scales measured from ethnographic recordings, particularly outside of Eurasia. Overall, we show that the historical emphasis on harmony is misguided and that melody is the primary determinant of the world's musical scales.


eess.AS音频处理
【1】 SpeechPrompt: Prompting Speech Language Models for Speech Processing Tasks
标题: SpeechPromise:为语音处理任务准备语音语言模型
作者:Kai-Wei Chang,Haibin Wu,Yu-Kai Wang,Yuan-Kuei Wu,Hua Shen,Wei-Cheng Tseng,Iu-thing Kang,Shang-Wen Li,Hung-yi Lee
Journal-ref:in IEEEACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 3730-3744, 2024
链接:点击下载PDF文件
摘要:翻译已经成为利用预训练语言模型(LM)的一种实用方法。这种方法有几个优点。它允许LM以最少的训练和参数更新来适应新任务,从而实现存储和计算的效率。此外,提示只修改LM的输入,并利用语言模型的生成能力,以统一的方式解决各种下游任务。这大大减少了在设计特定任务模型时对人力的需求。随着LM服务的任务数量的增加,这些优势变得更加明显。基于提示的优势,我们首次在语音处理领域探索了提示语音LM的潜力。近年来,将语音转换为离散单元进行语言建模的兴趣越来越大。我们的先驱研究表明,这些量化的语音单元是高度通用的,在我们的统一提示框架。它们不仅可以作为类别标签,而且还包含丰富的语音信息,可以重新合成为语音信号,用于语音生成任务。具体来说,我们重新制定语音处理任务到单元生成任务。因此,我们可以将语音分类、序列生成和语音生成等任务无缝集成到一个统一的提示框架中。实验结果表明,与基于自监督学习模型的强微调方法相比,在可训练参数数量相近的情况下,该方法具有较好的性能。提示方法在Few-Shot设置中也显示出有希望的结果。此外,随着高级言语模型的出现,所提出的激励框架具有很大的潜力。摘要:Prompting has become a practical method for utilizing pre-trained language models (LMs). This approach offers several advantages. It allows an LM to adapt to new tasks with minimal training and parameter updates, thus achieving efficiency in both storage and computation. Additionally, prompting modifies only the LM's inputs and harnesses the generative capabilities of language models to address various downstream tasks in a unified manner. This significantly reduces the need for human labor in designing task-specific models. These advantages become even more evident as the number of tasks served by the LM scales up. Motivated by the strengths of prompting, we are the first to explore the potential of prompting speech LMs in the domain of speech processing. Recently, there has been a growing interest in converting speech into discrete units for language modeling. Our pioneer research demonstrates that these quantized speech units are highly versatile within our unified prompting framework. Not only can they serve as class labels, but they also contain rich phonetic information that can be re-synthesized back into speech signals for speech generation tasks. Specifically, we reformulate speech processing tasks into speech-to-unit generation tasks. As a result, we can seamlessly integrate tasks such as speech classification, sequence generation, and speech generation within a single, unified prompting framework. The experiment results show that the prompting method can achieve competitive performance compared to the strong fine-tuning method based on self-supervised learning models with a similar number of trainable parameters. The prompting method also shows promising results in the few-shot setting. Moreover, with the advanced speech LMs coming into the stage, the proposed prompting framework attains great potential.

【2】 Inference-Adaptive Neural Steering for Real-Time Area-Based Sound Source Separation
标题: 基于实时区域的光源分离的推理自适应神经引导
作者:Martin Strauss,Wolfgang Mack,María Luis Valero,Okan Köpüklü
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:我们提出了一种新的神经转向技术,该技术在推理过程中适应空间感知多麦克风声源分离算法的目标区域,而无需重新训练深度神经网络(DNN)。为了实现这一点,我们首先训练DNN,旨在将语音保留在由角度跨度定义的目标区域内,同时抑制来自其他方向的声源。之后,将相移应用于麦克风信号,使我们能够在推理过程中移动目标区域的中心,而计算复杂性的额外成本可以忽略不计。此外,我们表明,所提出的方法在各种各样的声学场景中表现良好,包括目标区域内外的几个扬声器和额外的噪声。更确切地说,所提出的方法在DNSMOS和SI-SDR方面与针对转向目标区域明确训练的DNN表现相当。摘要:We propose a novel Neural Steering technique that adapts the target area of a spatial-aware multi-microphone sound source separation algorithm during inference without the necessity of retraining the deep neural network (DNN). To achieve this, we first train a DNN aiming to retain speech within a target region, defined by an angular span, while suppressing sound sources stemming from other directions. Afterward, a phase shift is applied to the microphone signals, allowing us to shift the center of the target area during inference at negligible additional cost in computational complexity. Further, we show that the proposed approach performs well in a wide variety of acoustic scenarios, including several speakers inside and outside the target area and additional noise. More precisely, the proposed approach performs on par with DNNs trained explicitly for the steered target area in terms of DNSMOS and SI-SDR.

【3】 Which Prosodic Features Matter Most for Pragmatics?
标题: 哪些韵律特征对修辞学最重要?
作者:Nigel G. Ward,Divette Marco,Olac Fuentes
备注:Submitted to ICASSP 2025. Audio illustrations available at this https URL
链接:点击下载PDF文件
摘要:我们调查的韵律功能最重要的传达韵律功能。我们使用的问题,预测人类感知的话语对之间的语用相似性,以评估不同类型的韵律特征的效用。例如,我们发现,与时长相关的特征比与音高相关的特征更重要,并且话语初始特征比话语最终特征更重要。此外,故障分析表明,使用音高功能的建模往往无法处理重要的语用功能,并建议几个通常被忽视的声学和韵律功能是语用意义重大,包括鼻音和颤音。这些发现可以指导未来的韵律基础研究,并建议如何提高语音合成评估,以及其他应用。摘要:We investigate which prosodic features matter most in conveying prosodic functions. We use the problem of predicting human perceptions of pragmatic similarity among utterance pairs to evaluate the utility of prosodic features of different types. We find, for example, that duration-related features are more important than pitch-related features, and that utterance-initial features are more important than utterance-final features. Further, failure analysis indicates that modeling using pitch features only often fails to handle important pragmatic functions, and suggests that several generally-neglected acoustic and prosodic features are pragmatically significant, including nasality and vibrato. These findings can guide future basic research in prosody, and suggest how to improve speech synthesis evaluation, among other applications.

【4】 EAViT: External Attention Vision Transformer for Audio Classification
标题: EAViT:音频分类的外部注意力视觉Transformer
作者:Aquib Iqbal,Abid Hasan Zim,Md Asaduzzaman Tonmoy,Limengnan Zhou,Asad Malik,Minoru Kuribayashi
链接:点击下载PDF文件
摘要:本文提出了外部注意力Vision Transformer(EAViT)模型,一种新的方法,旨在提高音频分类精度。随着数字音频资源的激增,对精确和高效的音频分类系统的需求已经加剧,这是由各种应用(包括音乐流媒体平台和环境声音识别)中对改进的推荐系统和用户个性化的需求所驱动的。准确的音频分类对于将庞大的音频库组织成连贯的类别至关重要,使用户能够更有效地找到他们喜欢的音频内容并与之交互。在这项研究中,我们利用GTZAN数据集,其中包括1,000个音乐摘录,跨越10个不同的流派。每个30秒的音频片段被分割成3秒的摘录,以增强数据集的鲁棒性并减轻过拟合风险,从而允许更细粒度的特征分析。EAViT模型将多头外部注意力(MEA)机制集成到Vision Transformer(ViT)框架中,有效地捕获样本之间的长期依赖性和潜在相关性。这种外部注意力(EA)机制采用可学习的记忆单元,增强了网络有效处理复杂音频特征的能力。该研究表明,EAViT实现了93.99%的显着整体准确性,超过了最先进的模型。摘要:This paper presents the External Attention Vision Transformer (EAViT) model, a novel approach designed to enhance audio classification accuracy. As digital audio resources proliferate, the demand for precise and efficient audio classification systems has intensified, driven by the need for improved recommendation systems and user personalization in various applications, including music streaming platforms and environmental sound recognition. Accurate audio classification is crucial for organizing vast audio libraries into coherent categories, enabling users to find and interact with their preferred audio content more effectively. In this study, we utilize the GTZAN dataset, which comprises 1,000 music excerpts spanning ten diverse genres. Each 30-second audio clip is segmented into 3-second excerpts to enhance dataset robustness and mitigate overfitting risks, allowing for more granular feature analysis. The EAViT model integrates multi-head external attention (MEA) mechanisms into the Vision Transformer (ViT) framework, effectively capturing long-range dependencies and potential correlations between samples. This external attention (EA) mechanism employs learnable memory units that enhance the network's capacity to process complex audio features efficiently. The study demonstrates that EAViT achieves a remarkable overall accuracy of 93.99%, surpassing state-of-the-art models.

【5】 Coarse-to-fine Alignment Makes Better Speech-image Retrieval
标题: 粗到细对齐可实现更好的语音图像检索
作者:Lifeng Zhou,Yuke Li
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个新的语音图像检索框架。我们利用语音-图像对比(SIC)学习任务在粗层次上对齐语音和图像表示,并利用语音-图像匹配(SIM)学习任务进一步细化细粒度的跨模态对齐。SIC和SIM学习任务以统一的方式联合训练。为了优化学习过程,我们利用嵌入队列,便于在SIC学习期间对高质量和多样化的否定表示进行有效采样。此外,它通过有效地挖掘基于SIC任务中计算的对比相似性的硬否定来增强SIM任务的学习。为了进一步优化噪声监督下的学习,我们将动量蒸馏纳入训练过程。实验结果表明,我们的框架优于国家的最先进的方法在R@1超过4%的语音图像检索任务的两个基准数据集。此外,在zero-shot实验中观察到,我们的框架表现出出色的泛化能力。摘要:In this paper, we propose a novel framework for speech-image retrieval. We utilize speech-image contrastive (SIC) learning tasks to align speech and image representations at a coarse level and speech-image matching (SIM) learning tasks to further refine the fine-grained cross-modal alignment. SIC and SIM learning tasks are jointly trained in a unified manner. To optimize the learning process, we utilize an embedding queue that facilitates efficient sampling of high-quality and diverse negative representations during SIC learning. Additionally, it enhances the learning of SIM tasks by effectively mining hard negatives based on contrastive similarities calculated in SIC tasks. To further optimize learning under noisy supervision, we incorporate momentum distillation into the training process. Experimental results show that our framework outperforms the state-of-the-art method by more than 4% in R@1 on two benchmark datasets for the speech-image retrieval tasks. Moreover, as observed in zero-shot experiments, our framework demonstrates excellent generalization capabilities.

【6】 NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
标题: NEST:自我监督的快速Conformer作为语音处理任务的通用调味剂
作者:He Huang,Taejin Park,Kunal Dhawan,Ivan Medennikov,Krishna C. Puvvada,Nithin Rao Koluguri,Weiqing Wang,Jagadeesh Balam,Boris Ginsburg
链接:点击下载PDF文件
摘要:自监督学习已被证明有利于广泛的语音处理任务,如语音识别 翻译,说话人确认和日记等,但是,这些方法大多是计算密集型的,由于使用Transformer编码器和缺乏子采样。在本文中,我们提出了一个新的自监督学习模型称为神经编码器的自监督训练(NEST)。具体来说,我们采用FastConformer架构,该架构具有8倍子采样率,比Transformer或Conformer架构更快。而不是基于聚类的令牌生成,我们诉诸于固定的随机投影,其简单性和有效性。我们还提出了一个广义的嘈杂的语音增强,教模型解开噪声或其他扬声器的主要扬声器。实验表明,所提出的NEST模型改善了现有的自监督模型在各种语音处理任务。代码和检查点将通过NVIDIA NeMo工具包公开提供。摘要:Self-supervised learning has been proved to benefit a wide range of speech processing tasks, such as speech recognition translation, speaker verification and diarization, etc. However, most of these approaches are computationally intensive due to using transformer encoder and lack of sub-sampling. In this paper, we propose a new self-supervised learning model termed as Neural Encoder for Self-supervised Training (NEST). Specifically, we adopt the FastConformer architecture, which has an 8x sub-sampling rate and is faster than Transformer or Conformer architectures. Instead of clustering-based token generation, we resort to fixed random projection for its simplicity and effectiveness. We also propose a generalized noisy speech augmentation that teaches the model to disentangle the main speaker from noise or other speakers. Experiments show that the proposed NEST model improves over existing self-supervised models on a variety of speech processing tasks. Code and checkpoints will be publicly available via NVIDIA NeMo toolkit.

【7】 On Class Separability Pitfalls In Audio-Text Contrastive Zero-Shot Learning
标题: 音频文本对比Zero-Shot学习中的班级可分离性陷阱
作者:Tiago Tavares,Fabio Ayres,Zhepei Wang,Paris Smaragdis
链接:点击下载PDF文件
摘要:音频-文本跨通道对比学习的最新进展显示了其向zero-shot学习方向发展的潜力。一种可能性是将来自预先训练的骨干神经网络的项目嵌入投影到跨模态空间中,在该跨模态空间中,可以在任一域中计算项目相似性。这个过程依赖于骨干网络的强大的单峰预训练,以及投影仪的数据密集型训练任务。这两个过程可能会因无意的数据泄漏而产生偏差,这可能是由于在预训练中使用监督学习或使用来自zero-shot学习评估的标签无意中训练跨模态投影而引起的。在这项研究中,我们表明,测量的zero-shot学习准确性的一个重要部分是由于从音频和文本骨干继承的优势,也就是说,他们不是在跨模态域学习,并没有从一个模态转移到另一个。摘要:Recent advances in audio-text cross-modal contrastive learning have shown its potential towards zero-shot learning. One possibility for this is by projecting item embeddings from pre-trained backbone neural networks into a cross-modal space in which item similarity can be calculated in either domain. This process relies on a strong unimodal pre-training of the backbone networks, and on a data-intensive training task for the projectors. These two processes can be biased by unintentional data leakage, which can arise from using supervised learning in pre-training or from inadvertently training the cross-modal projection using labels from the zero-shot learning evaluation. In this study, we show that a significant part of the measured zero-shot learning accuracy is due to strengths inherited from the audio and text backbones, that is, they are not learned in the cross-modal domain and are not transferred from one modality to another.

【8】 Uncertainty-Aware Mean Opinion Score Prediction
标题: 不确定性意识的平均意见分数预测
作者:Hui Wang,Shiwan Zhao,Jiaming Zhou,Xiguang Zheng,Haoqin Sun,Xuechen Wang,Yong Qin
备注:Accepted by Interspeech 2024, oral
链接:点击下载PDF文件
摘要:平均意见得分(MOS)预测在特定领域取得了重大进展。然而,MOS预测模型在不同样本中的不稳定性能在这些系统的实际应用中提出了持续的挑战。在本文中,我们指出,不确定性建模的缺乏是一个显着的限制,阻碍MOS预测系统应用到现实和开放的世界。我们分析了MOS预测任务中不确定性的来源,并提出建立一个具有不确定性意识的MOS预测系统,该系统分别采用异方差回归和蒙特卡罗退出来建模偶然不确定性和认知不确定性。实验结果表明,该系统能够很好地捕捉不确定性,并能够进行选择性预测和域外检测。这种能力大大增强了MOS系统在各种真实和开放世界环境中的实际效用。摘要:Mean Opinion Score (MOS) prediction has made significant progress in specific domains. However, the unstable performance of MOS prediction models across diverse samples presents ongoing challenges in the practical application of these systems. In this paper, we point out that the absence of uncertainty modeling is a significant limitation hindering MOS prediction systems from applying to the real and open world. We analyze the sources of uncertainty in the MOS prediction task and propose to establish an uncertainty-aware MOS prediction system that models aleatory uncertainty and epistemic uncertainty by heteroscedastic regression and Monte Carlo dropout separately. The experimental results show that the system captures uncertainty well and is capable of performing selective prediction and out-of-domain detection. Such capabilities significantly enhance the practical utility of MOS systems in diverse real and open-world environments.

【9】 Towards measuring fairness in speech recognition: Fair-Speech dataset
标题: 衡量语音识别中的公平性:公平语音数据集
作者:Irina-Elena Veliche,Zhuangqun Huang,Vineeth Ayyat Kochaniyan,Fuchun Peng,Ozlem Kalinli,Michael L. Seltzer
链接:点击下载PDF文件
摘要:目前用于语音识别(ASR)的公共数据集往往不专注于公平性方面,例如不同人口统计群体的性能。本文介绍了一个新的数据集,公平演讲,一个公开发布的语料库,以帮助研究人员评估他们的ASR模型的准确性在一组不同的自我报告的人口统计信息,如年龄,性别,种族,地理变异和参与者是否认为自己的母语是英语。我们的数据集包括美国593人录制的大约26.5K的话语,这些人付费录制并提交自己说语音命令的音频。我们还提供ASR基线,包括在转录和未转录的社交媒体视频和开源模型上训练的模型。摘要:The current public datasets for speech recognition (ASR) tend not to focus specifically on the fairness aspect, such as performance across different demographic groups. This paper introduces a novel dataset, Fair-Speech, a publicly released corpus to help researchers evaluate their ASR models for accuracy across a diverse set of self-reported demographic information, such as age, gender, ethnicity, geographic variation and whether the participants consider themselves native English speakers. Our dataset includes approximately 26.5K utterances in recorded speech by 593 people in the United States, who were paid to record and submit audios of themselves saying voice commands. We also provide ASR baselines, including on models trained on transcribed and untranscribed social media videos and open source models.

【10】 Hierarchical Generative Modeling of Melodic Vocal Contours in Hindustani Classical Music
标题: 印度斯坦古典音乐旋律声乐轮廓的分层生成建模
作者:Nithya Shikarpur,Krishna Maneesha Dendukur,Yusong Wu,Antoine Caillon,Cheng-Zhi Anna Huang
备注:Accepted at International Society for Music Information Retrieval (ISMIR) 2024
链接:点击下载PDF文件
摘要:印度斯坦音乐是一种表演驱动的口头传统,展示了丰富的旋律模式。在本文中,我们专注于生成建模的歌手的声乐旋律提取的录音,因为声音是音乐突出的传统。在印度斯坦音乐中,先前的生成工作将旋律建模为粗糙的离散符号,无法捕捉歌唱中丰富的表达性旋律的复杂性。因此,我们建议使用精细量化的音高轮廓,作为分层音频建模的中间表示。我们提出了GaMaDHaNi,一个模块化的两级层次结构,包括一个生成模型的音高轮廓,和一个音高轮廓音频合成模型。我们比较我们的方法,非分层音频模型和分层模型,使用自监督的中间表示,通过收听测试和定性分析。我们还评估了音频模型的能力,忠实地代表音高轮廓输入使用皮尔逊相关系数。通过使用音高轮廓作为中间表示,我们表明,我们的模型可以通过突出两个潜在的交互用例(1)启动生成和(2)粗音高调节,更好地在人类-AI协作环境中倾听和响应音乐家。摘要:Hindustani music is a performance-driven oral tradition that exhibits the rendition of rich melodic patterns. In this paper, we focus on generative modeling of singers' vocal melodies extracted from audio recordings, as the voice is musically prominent within the tradition. Prior generative work in Hindustani music models melodies as coarse discrete symbols which fails to capture the rich expressive melodic intricacies of singing. Thus, we propose to use a finely quantized pitch contour, as an intermediate representation for hierarchical audio modeling. We propose GaMaDHaNi, a modular two-level hierarchy, consisting of a generative model on pitch contours, and a pitch contour to audio synthesis model. We compare our approach to non-hierarchical audio models and hierarchical models that use a self-supervised intermediate representation, through a listening test and qualitative analysis. We also evaluate audio model's ability to faithfully represent the pitch contour input using Pearson correlation coefficient. By using pitch contours as an intermediate representation, we show that our model may be better equipped to listen and respond to musicians in a human-AI collaborative setting by highlighting two potential interaction use cases (1) primed generation, and (2) coarse pitch conditioning.

【11】 Information and motor constraints shape melodic diversity across cultures
标题: 信息和运动限制塑造了跨文化的旋律多样性
作者:John M McBride,Nahie Kim,Yuri Nishikawa,Mekhmed Saadakeev,Marcus T Pearce,Tsvi Tlusty
链接:点击下载PDF文件
摘要:可能的旋律的数量是相当大的,然而,尽管旋律变化的潜力几乎是无限的,来自不同社会的旋律可以惊人地相似。运动约束假说解释了某些相似性,如标量运动和轮廓形状,但不解释其他主要的共同特征,如重复,歌曲长度和音阶大小。在这里,我们调查的信息限制所产生的人类记忆的局限性,在塑造这些标志性的旋律的作用。我们测量了跨越几大洲的62个民间旋律语料库的信息率的决定因素,发现多个权衡,所有的行为,以限制跨社会的信息率。相比之下,来自欧洲(包括土耳其)的39个艺术音乐语料库显示更长,更复杂的旋律,随着时间的推移,复杂性增加,这表明艺术和民间音乐中不同的文化进化选择压力,可能是由于使用书面与口头传播。我们的无参数模型预测的经验尺度度分布使用标量运动,旋律长度,最重要的是,信息率的信息约束。这提供了强有力的证据表明,在音乐的文化传播过程中的信息约束限制了音阶中音符的数量,并提出了对中间旋律复杂性的偏好是对旋律文化演变的根本约束。摘要:The number of possible melodies is unfathomably large, yet despite this virtually unlimited potential for melodic variation, melodies from different societies can be surprisingly similar. The motor constraint hypothesis accounts for certain similarities, such as scalar motion and contour shape, but not for other major common features, such as repetition, song length, and scale size. Here we investigate the role of information constraints arising from limitations on human memory in shaping these hallmarks of melodies. We measure determinants of information rate in 62 corpora of Folk melodies spanning several continents, finding multiple trade-offs that all act to constrain the information rate across societies. By contrast, 39 corpora of Art music from Europe (including Turkey) show longer, more complex melodies, and increased complexity over time, suggesting different cultural-evolutionary selection pressures in Art and Folk music, possibly due to the use of written versus oral transmission. Our parameter-free model predicts the empirical scale degree distribution using information constraints on scalar motion, melody length, and, most importantly, information rate. This provides strong evidence that information constraints during cultural transmission of music limit the number of notes in a scale, and proposes that preference for intermediate melodic complexity is a fundamental constraint on the cultural evolution of melody.

【12】 Melody predominates over harmony in the evolution of musical scales across 96 countries
标题: 在96个国家的音阶演变中,旋律优先于和声
作者:John M McBride,Elizabeth Phillips,Patrick E Savage,Steven Brown,Tsvi Tlusty
链接:点击下载PDF文件
摘要:自古以来,音阶的标准理论是基于和声,而不是旋律。最近的一些分析支持这两种观点,我们缺乏跨文化数据的比较测试。我们通过对主要理论与来自96个国家的1,314个量表进行严格的计算比较来解决这个长期存在的问题。旋律理论几乎得到了普遍的支持,它预测1-3个音阶的步长。和声解释了一些简单整数比音程的流行,特别是欧亚社会的音乐理论音阶,这可能解释了它们在西方学者中的统治地位。然而,和谐预测从人种学记录,特别是欧亚大陆以外的规模。总的来说,我们表明,历史上强调和谐是误导,旋律是世界音乐音阶的主要决定因素。摘要:The standard theory of musical scales since antiquity has been based on harmony, rather than melody. Some recent analyses support either view, and we lack a comparative test on cross-cultural data. We address this longstanding problem through a rigorous, computational comparison of the main theories against 1,314 scales from 96 countries. There is near-universal support for melodic theories, which predict step-sizes of 1-3 semitones. Harmony accounts for the prevalence of some simple-integer-ratio intervals, particularly for music-theoretic scales from Eurasian societies, which may explain their dominance amongst Western scholars. However, harmony poorly predicts scales measured from ethnographic recordings, particularly outside of Eurasia. Overall, we show that the historical emphasis on harmony is misguided and that melody is the primary determinant of the world's musical scales.


机器翻译,仅供参考