本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
cs.SD语音
【1】Quantification of Tenseness in English and Japanese Tense-Lax Vowels: A Lagrangian Model with Indicator $θ_1$ and Force of Tenseness Ftense(t)
链接:https://arxiv.org/abs/2503.03681
摘要:元音紧张度的概念传统上是通过紧张元音和松弛元音的二元区分来检验的。然而,没有普遍接受的定量定义紧张已建立在任何语言。以前的研究,包括雅各布森,范特和哈雷(1951年)和乔姆斯基和哈雷(1968年),探讨了元音紧张和声道之间的关系。在这些基础上,Ishizaki(2019,2022)提出了一种使用共振峰角$\theta_1 $和$\theta_{F1}$及其一阶和二阶导数$d^Z_1(t)/dt = \lim \tan\theta_1(t$)和$d^2 Z_1(t)/dt^2 = d/dt \lim \tan\theta_1(t)$间接量化元音紧张度的方法。本研究扩展了这种方法,通过调查的潜在作用的力相关的参数在确定元音质量。具体来说,我们引入了一个简化的模型的基础上拉格朗日方程来描述的动态相互作用的舌头和下巴在口腔内的发音闭元音。这个模型提供了一个理论框架,估计在不同语言的元音生产所涉及的力量,提供了新的见解的物理机制,元音发音。研究结果表明,这种基于力量的观点值得进一步探索语音和音系学研究的一个关键因素。
摘要:The concept of vowel tenseness has traditionally been examined through thebinary distinction of tense and lax vowels. However, no universally acceptedquantitative definition of tenseness has been established in any language.Previous studies, including those by Jakobson, Fant, and Halle (1951) andChomsky and Halle (1968), have explored the relationship between voweltenseness and the vocal tract. Building on these foundations, Ishizaki (2019,2022) proposed an indirect quantification of vowel tenseness using formantangles $\theta_1$ and $\theta_{F1}$ and their first and second derivatives,$d^Z_1(t)/dt = \lim \tan \theta_1(t$) and $d^2 Z_1(t)/dt^2 = d/dt \lim \tan\theta_1(t)$. This study extends this approach by investigating the potentialrole of a force-related parameter in determining vowel quality. Specifically,we introduce a simplified model based on the Lagrangian equation to describethe dynamic interaction of the tongue and jaw within the oral cavity during thearticulation of close vowels. This model provides a theoretical framework forestimating the forces involved in vowel production across different languages,offering new insights into the physical mechanisms underlying vowelarticulation. The findings suggest that this force-based perspective warrantsfurther exploration as a key factor in phonetic and phonological studies.
标题:多轨音乐的主要乐器检测
链接:https://arxiv.org/abs/2503.03232
备注:Camera ready version of ICASSP 2025 submission
摘要:引导乐器检测的现有方法主要分析混合音频,限于粗略分类并且缺乏泛化能力。本文提出了一种新的方法,通过精心制作专业注释的数据集,并设计一个新的框架,集成了一个自监督学习模型与一个轨道明智的,帧级的注意力为基础的分类器,在多轨音乐音频中的领先乐器检测。这种注意力机制根据听觉重要性动态提取和聚合特定于音轨的特征,从而能够在不同的乐器类型和组合中进行精确检测。通过轨道分类和排列增强,我们的模型大大优于现有的SVM和CRNN模型,在看不见的仪器和域外测试中表现出鲁棒性。我们相信,我们的探索提供了宝贵的见解,为未来的研究多轨音乐设置的音频内容分析。
摘要:Prior approaches to lead instrument detection primarily analyze mixtureaudio, limited to coarse classifications and lacking generalization ability.This paper presents a novel approach to lead instrument detection in multitrackmusic audio by crafting expertly annotated datasets and designing a novelframework that integrates a self-supervised learning model with a track-wise,frame-level attention-based classifier. This attention mechanism dynamicallyextracts and aggregates track-specific features based on their auditoryimportance, enabling precise detection across varied instrument types andcombinations. Enhanced by track classification and permutation augmentation,our model substantially outperforms existing SVM and CRNN models, showingrobustness on unseen instruments and out-of-domain testing. We believe ourexploration provides valuable insights for future research on audio contentanalysis in multitrack music settings.
标题:用于包容性韵律压力分析的微调耳语器
链接:https://arxiv.org/abs/2503.02907
备注:Appears in Proceedings of the ISCA/ITG Workshop on Diversity in Large Speech and Language Models
摘要:韵律在语音感知中起着至关重要的作用,影响着人类的理解和自动语音识别(ASR)系统。尽管韵律重音很重要,但由于其有效分析的挑战,韵律重音仍然研究不足。本研究探索微调OpenAI的Whisper large-v2 ASR模型,以识别语音中的短语、词汇和对比重音。使用66个英语母语者的数据集,包括男性,女性,神经典型,和神经分歧的个人,我们评估模型的能力,概括压力模式和分类扬声器的神经类型和性别的基础上简短的语音样本。我们的研究结果强调了在所有三种压力类型下ASR表现的接近人类的准确性,以及在分类性别和神经类型方面近乎完美的准确性。通过改进韵律感知的ASR,这项工作有助于为不同人群提供公平和强大的转录技术。
摘要:Prosody plays a crucial role in speech perception, influencing both humanunderstanding and automatic speech recognition (ASR) systems. Despite itsimportance, prosodic stress remains under-studied due to the challenge ofefficiently analyzing it. This study explores fine-tuning OpenAI's Whisperlarge-v2 ASR model to recognize phrasal, lexical, and contrastive stress inspeech. Using a dataset of 66 native English speakers, including male, female,neurotypical, and neurodivergent individuals, we assess the model's ability togeneralize stress patterns and classify speakers by neurotype and gender basedon brief speech samples. Our results highlight near-human accuracy in ASRperformance across all three stress types and near-perfect precision inclassifying gender and neurotype. By improving prosody-aware ASR, this workcontributes to equitable and robust transcription technologies for diversepopulations.
标题:用于综合降噪和声回声消除的广义回声和干扰消除与扩展多通道维纳过滤的比较分析
链接:https://arxiv.org/abs/2503.03593
备注:Accepted for publication in ICASSP 2025
摘要:分析了两种用于声回波消除(AEC)和噪声抑制(NR)的组合算法,即广义回波干扰消除器(GEIC)和扩展多通道维纳滤波器(MWFext)。以前,这些算法已经检查了线性回声路径,并假设访问语音活动检测器(VAD),分别检测所需的语音和回声活动。然而,实现VAD的算法可能引入检测误差。因此,在本文中,前面的分析是通过以下方式扩展的:1)通过广义Bussgang分解对一般非线性回波路径进行建模,以及2)对每个特定算法中的VAD误差效应进行建模,从而也允许对特定VAD假设进行建模。仿真结果表明,MWFext算法具有较高的NR性能,而GEIC算法具有较强的AEC性能。
摘要:Two algorithms for combined acoustic echo cancellation (AEC) and noisereduction (NR) are analysed, namely the generalised echo and interferencecanceller (GEIC) and the extended multichannel Wiener filter (MWFext).Previously, these algorithms have been examined for linear echo paths, andassuming access to voice activity detectors (VADs) that separately detectdesired speech and echo activity. However, algorithms implementing VADs mayintroduce detection errors. Therefore, in this paper, the previous analyses areextended by 1) modelling general nonlinear echo paths by means of thegeneralised Bussgang decomposition, and 2) modelling VAD error effects in eachspecific algorithm, thereby also allowing to model specific VAD assumptions. Itis found and verified with simulations that, generally, the MWFext achieves ahigher NR performance, while the GEIC achieves a more robust AEC performance.
标题:用于综合降噪和声回声消除的广义回声和干扰消除与扩展多通道维纳过滤的比较分析
链接:https://arxiv.org/abs/2503.03593
备注:Accepted for publication in ICASSP 2025
摘要:分析了两种用于声回波消除(AEC)和噪声抑制(NR)的组合算法,即广义回波干扰消除器(GEIC)和扩展多通道维纳滤波器(MWFext)。以前,这些算法已经检查了线性回声路径,并假设访问语音活动检测器(VAD),分别检测所需的语音和回声活动。然而,实现VAD的算法可能引入检测误差。因此,在本文中,前面的分析是通过以下方式扩展的:1)通过广义Bussgang分解对一般非线性回波路径进行建模,以及2)对每个特定算法中的VAD误差效应进行建模,从而也允许对特定VAD假设进行建模。仿真结果表明,MWFext算法具有较高的NR性能,而GEIC算法具有较强的AEC性能。
摘要:Two algorithms for combined acoustic echo cancellation (AEC) and noisereduction (NR) are analysed, namely the generalised echo and interferencecanceller (GEIC) and the extended multichannel Wiener filter (MWFext).Previously, these algorithms have been examined for linear echo paths, andassuming access to voice activity detectors (VADs) that separately detectdesired speech and echo activity. However, algorithms implementing VADs mayintroduce detection errors. Therefore, in this paper, the previous analyses areextended by 1) modelling general nonlinear echo paths by means of thegeneralised Bussgang decomposition, and 2) modelling VAD error effects in eachspecific algorithm, thereby also allowing to model specific VAD assumptions. Itis found and verified with simulations that, generally, the MWFext achieves ahigher NR performance, while the GEIC achieves a more robust AEC performance.
标题:语音质量与神经编解码器量化潜在表示之间的关系
链接:https://arxiv.org/abs/2503.03304
摘要:近年来,神经音频信号编解码器引起了极大的关注。本质上,通过学习捕获编码信号的属性的抽象表示来实现由这种编码器实现的令人印象深刻的低比特率,例如,演讲在这项工作中,我们研究了神经编解码器学习的输入信号的潜在表示与语音信号质量之间的关系。为此,我们引入了潜在表示量化误差比(LQR)度量,该度量量化了给定语音信号与理想化神经编解码器语音信号模型的距离。我们比较建议的指标侵入性措施,以及数据驱动的监督方法,使用两个主观语音质量数据集。该分析表明,所提出的LQR强烈相关(高达0.9皮尔逊相关性)与语音的主观质量。尽管是一个非侵入性的指标,但这产生了与其他预先训练和侵入性措施竞争的性能,甚至更好。这些结果表明,LQR是一个有前途的基础上更复杂的语音质量的措施。
摘要:Neural audio signal codecs have attracted significant attention in recentyears. In essence, the impressive low bitrate achieved by such encoders isenabled by learning an abstract representation that captures the properties ofencoded signals, e.g., speech. In this work, we investigate the relationbetween the latent representation of the input signal learned by a neural codecand the quality of speech signals. To do so, we introduceLatent-representation-to-Quantization error Ratio (LQR) measures, whichquantify the distance from the idealized neural codec's speech signal model fora given speech signal. We compare the proposed metrics to intrusive measures aswell as data-driven supervised methods using two subjective speech qualitydatasets. This analysis shows that the proposed LQR correlates strongly (up to0.9 Pearson's correlation) with the subjective quality of speech. Despite beinga non-intrusive metric, this yields a competitive performance with, or evenbetter than, other pre-trained and intrusive measures. These results show thatLQR is a promising basis for more sophisticated speech quality measures.
标题:合成语音评估的良好实践
链接:https://arxiv.org/abs/2503.03250
摘要:本文档作为语音合成相关论文的审稿指南提供。我们概述了一些关于语音合成的论文的最佳实践和常见陷阱,特别关注评估。我们还建议审稿人检查纸质工具包中的作者指南,并将其视为审稿标准。这是一份活的文件,我们将在收到读者的评论和反馈后进行更新。我们注意到,本文档仅提供指导,评审人员在评估论文时最终应自行决定。
摘要:This document is provided as a guideline for reviewers of papers about speechsynthesis. We outline some best practices and common pitfalls for papers aboutspeech synthesis, with a particular focus on evaluation. We also recommend thatreviewers check the guidelines for authors written in the paper kit andconsider those as reviewing criteria as well. This is intended to be a livingdocument, and it will be updated as we receive comments and feedback fromreaders. We note that this document is meant to provide guidance only, and thatreviewers should ultimately use their own discretion when evaluating papers.
标题:HARP 2.0:扩展托管、同步、远程处理以进行深度学习
链接:https://arxiv.org/abs/2503.02977
备注:ISMIR 2024 Late-Breaking Demo
摘要:HARP 2.0通过托管、异步、远程处理将深度学习模型引入数字音频工作站(Digital Audio Workstation,简称DTS)软件,允许用户通过任何兼容的Gradio端点从插件接口路由音频,以执行任意转换。HARP呈现端点定义的控件和处理的音频插件,这意味着用户可以探索各种尖端的深度学习模型,而无需离开桌面。在2.0版本中,我们引入了对基于MIDI的模型和音频/音频标签模型的支持,为模型开发人员提供了一个精简的pyharp Python API,并实现了许多接口和稳定性改进。通过这项工作,我们希望弥合模型开发人员和创意人员之间的差距,通过将深度学习模型无缝集成到工作流中来改善对深度学习模型的访问。
摘要:HARP 2.0 brings deep learning models to digital audio workstation (DAW)software through hosted, asynchronous, remote processing, allowing users toroute audio from a plug-in interface through any compatible Gradio endpoint toperform arbitrary transformations. HARP renders endpoint-defined controls andprocessed audio in-plugin, meaning users can explore a variety of cutting-edgedeep learning models without ever leaving the DAW. In the 2.0 release weintroduce support for MIDI-based models and audio/MIDI labeling models, providea streamlined pyharp Python API for model developers, and implement numerousinterface and stability improvements. Through this work, we hope to bridge thegap between model developers and creatives, improving access to deep learningmodels by seamlessly integrating them into DAW workflows.
标题:英语和日语时态松弛元音中紧张度的量化:具有指示符$O_1 $和紧张力Ftense(t)的拉格朗日模型
链接:https://arxiv.org/abs/2503.03681
摘要:元音紧张度的概念传统上是通过紧张元音和松弛元音的二元区分来检验的。然而,没有普遍接受的定量定义紧张已建立在任何语言。以前的研究,包括雅各布森,范特和哈雷(1951年)和乔姆斯基和哈雷(1968年),探讨了元音紧张和声道之间的关系。在这些基础上,Ishizaki(2019,2022)提出了一种使用共振峰角$\theta_1 $和$\theta_{F1}$及其一阶和二阶导数$d^Z_1(t)/dt = \lim \tan\theta_1(t$)和$d^2 Z_1(t)/dt^2 = d/dt \lim \tan\theta_1(t)$间接量化元音紧张度的方法。本研究扩展了这种方法,通过调查的潜在作用的力相关的参数在确定元音质量。具体来说,我们引入了一个简化的模型的基础上拉格朗日方程来描述的动态相互作用的舌头和下巴在口腔内的发音闭元音。这个模型提供了一个理论框架,估计在不同语言的元音生产所涉及的力量,提供了新的见解的物理机制,元音发音。研究结果表明,这种基于力量的观点值得进一步探索语音和音系学研究的一个关键因素。
摘要:The concept of vowel tenseness has traditionally been examined through thebinary distinction of tense and lax vowels. However, no universally acceptedquantitative definition of tenseness has been established in any language.Previous studies, including those by Jakobson, Fant, and Halle (1951) andChomsky and Halle (1968), have explored the relationship between voweltenseness and the vocal tract. Building on these foundations, Ishizaki (2019,2022) proposed an indirect quantification of vowel tenseness using formantangles $\theta_1$ and $\theta_{F1}$ and their first and second derivatives,$d^Z_1(t)/dt = \lim \tan \theta_1(t$) and $d^2 Z_1(t)/dt^2 = d/dt \lim \tan\theta_1(t)$. This study extends this approach by investigating the potentialrole of a force-related parameter in determining vowel quality. Specifically,we introduce a simplified model based on the Lagrangian equation to describethe dynamic interaction of the tongue and jaw within the oral cavity during thearticulation of close vowels. This model provides a theoretical framework forestimating the forces involved in vowel production across different languages,offering new insights into the physical mechanisms underlying vowelarticulation. The findings suggest that this force-based perspective warrantsfurther exploration as a key factor in phonetic and phonological studies.
标题:多轨音乐的主要乐器检测
链接:https://arxiv.org/abs/2503.03232
备注:Camera ready version of ICASSP 2025 submission
摘要:引导乐器检测的现有方法主要分析混合音频,限于粗略分类并且缺乏泛化能力。本文提出了一种新的方法,通过精心制作专业注释的数据集,并设计一个新的框架,集成了一个自监督学习模型与一个轨道明智的,帧级的注意力为基础的分类器,在多轨音乐音频中的领先乐器检测。这种注意力机制根据听觉重要性动态提取和聚合特定于音轨的特征,从而能够在不同的乐器类型和组合中进行精确检测。通过轨道分类和排列增强,我们的模型大大优于现有的SVM和CRNN模型,在看不见的仪器和域外测试中表现出鲁棒性。我们相信,我们的探索提供了宝贵的见解,为未来的研究多轨音乐设置的音频内容分析。
摘要:Prior approaches to lead instrument detection primarily analyze mixtureaudio, limited to coarse classifications and lacking generalization ability.This paper presents a novel approach to lead instrument detection in multitrackmusic audio by crafting expertly annotated datasets and designing a novelframework that integrates a self-supervised learning model with a track-wise,frame-level attention-based classifier. This attention mechanism dynamicallyextracts and aggregates track-specific features based on their auditoryimportance, enabling precise detection across varied instrument types andcombinations. Enhanced by track classification and permutation augmentation,our model substantially outperforms existing SVM and CRNN models, showingrobustness on unseen instruments and out-of-domain testing. We believe ourexploration provides valuable insights for future research on audio contentanalysis in multitrack music settings.
标题:用于包容性韵律压力分析的微调耳语器
链接:https://arxiv.org/abs/2503.02907
备注:Appears in Proceedings of the ISCA/ITG Workshop on Diversity in Large Speech and Language Models
摘要:韵律在语音感知中起着至关重要的作用,影响着人类的理解和自动语音识别(ASR)系统。尽管韵律重音很重要,但由于其有效分析的挑战,韵律重音仍然研究不足。本研究探索微调OpenAI的Whisper large-v2 ASR模型,以识别语音中的短语、词汇和对比重音。使用66个英语母语者的数据集,包括男性,女性,神经典型,和神经分歧的个人,我们评估模型的能力,概括压力模式和分类扬声器的神经类型和性别的基础上简短的语音样本。我们的研究结果强调了在所有三种压力类型下ASR表现的接近人类的准确性,以及在分类性别和神经类型方面近乎完美的准确性。通过改进韵律感知的ASR,这项工作有助于为不同人群提供公平和强大的转录技术。
摘要:Prosody plays a crucial role in speech perception, influencing both humanunderstanding and automatic speech recognition (ASR) systems. Despite itsimportance, prosodic stress remains under-studied due to the challenge ofefficiently analyzing it. This study explores fine-tuning OpenAI's Whisperlarge-v2 ASR model to recognize phrasal, lexical, and contrastive stress inspeech. Using a dataset of 66 native English speakers, including male, female,neurotypical, and neurodivergent individuals, we assess the model's ability togeneralize stress patterns and classify speakers by neurotype and gender basedon brief speech samples. Our results highlight near-human accuracy in ASRperformance across all three stress types and near-perfect precision inclassifying gender and neurotype. By improving prosody-aware ASR, this workcontributes to equitable and robust transcription technologies for diversepopulations.
