微信公众号:arXiv_Daily
cs.SD语音
【1】Perceived Femininity in Singing Voice: Analysis and Prediction
标题:歌唱声音中的女性气质:分析与预测
链接:https://arxiv.org/abs/2511.02726
备注:None
摘要:本文着重探讨歌唱声音中经常被忽视的声音女性化方面。虽然现有的研究已经研究了语音中的感知声音女性化,但同样的概念尚未在歌唱声音中进行研究。音乐内容中的性别偏见分析可以从这样的研究中受益。为了解决这一差距,我们设计了一个基于刺激的调查,以衡量感知歌声的女性气质(PSVF),并收集了128名参与者的回应。我们的分析揭示了PSVF在不同人口群体中的变化规律。此外,我们提出了一个自动PSVF预测模型,通过微调的x向量模型,提供了一个新的工具,探索性别刻板印象的声音在音乐内容分析超越二进制性别分类。本研究通过对调查结果的分析,有助于更深入地理解歌唱声音中女性特质感知的复杂性,并为未来的研究提供一个自动化的工具。
摘要:This paper focuses on the often-overlooked aspect of perceived voice femininity in singing voices. While existing research has examined perceived voice femininity in speech, the same concept has not yet been studied in singing voice. The analysis of gender bias in music content could benefit from such study. To address this gap, we design a stimuli-based survey to measure perceived singing voice femininity (PSVF), and collect responses from 128 participants. Our analysis reveals intriguing insights into how PSVF varies across different demographic groups. Furthermore, we propose an automatic PSVF prediction model by fine-tuning an x-vector model, offering a novel tool for exploring gender stereotypes related to voices in music content analysis beyond binary sex classification. This study contributes to a deeper understanding of the complexities surrounding perceived femininity in singing voices by analyzing survey and proposes an automatic tool for future research.
【2】Improving DF-Conformer Using Hydra For High-Fidelity Generative Speech Enhancement on Discrete Codec Token
标题:使用Hydra改进DF-conformer以在离散编解码器令牌上进行高保真生成语音增强
链接:https://arxiv.org/abs/2511.02454
备注:Submitted to ICASSP 2026. Audio samples available at this https URL
摘要:Dilated FAVOR Conformer(DF-Conformer)是为语音增强(SE)设计的Conformer架构的有效变体。它通过正正交随机特征(FAVOR+)来使用快速注意力,以减轻与自我注意力相关的二次复杂度,同时利用扩张卷积来扩大感受野。这种组合在各种SE型号中产生了令人印象深刻的性能。在本文中,我们提出用双向选择性结构化状态空间序列模型取代FAVOR+,以实现两个主要目标:(1)通过消除FAVOR+中固有的近似来增强全局序列建模,以及(2)保持相对于序列长度的线性复杂度。具体来说,我们利用Hydra,Mamba的双向扩展,框架内的结构化矩阵混合器框架。使用生成SE模型对离散编解码器令牌(称为Genhancer)进行的实验表明,所提出的方法超越了DF-Conformer的性能。
摘要:The Dilated FAVOR Conformer (DF-Conformer) is an efficient variant of the Conformer architecture designed for speech enhancement (SE). It employs fast attention through positive orthogonal random features (FAVOR+) to mitigate the quadratic complexity associated with self-attention, while utilizing dilated convolution to expand the receptive field. This combination results in impressive performance across various SE models. In this paper, we propose replacing FAVOR+ with bidirectional selective structured state-space sequence models to achieve two main objectives:(1) enhancing global sequential modeling by eliminating the approximations inherent in FAVOR+, and (2) maintaining linear complexity relative to the sequence length. Specifically, we utilize Hydra, a bidirectional extension of Mamba, framed within the structured matrix mixer framework. Experiments conducted using a generative SE model on discrete codec tokens, known as Genhancer, demonstrate that the proposed method surpasses the performance of the DF-Conformer.
【3】H-Infinity Filter Enhanced CNN-LSTM for Arrhythmia Detection from Heart Sound Recordings
标题:H-Infinity过滤器增强型CNN-LSTM用于从心脏声音记录中检测心律失常
链接:https://arxiv.org/abs/2511.02379
备注:This is a preprint of a paper to appear at the 15th IEEE International Conference on Systems Engineering and Technology (ICSET 2025)
摘要:早期发现心律失常可以预防心脏病患者未来严重的并发症。虽然人工诊断仍然是临床标准,但它严重依赖于视觉解释,并且本质上是主观的。近年来,深度学习已成为自动心律失常检测的强大工具,可提高准确性、一致性和效率。卷积和递归神经网络架构的几种变体已被广泛探索,以捕获生理信号中的空间和时间模式。然而,尽管有这些进步,当前的模型往往很难在现实世界中很好地推广,特别是在处理小型或嘈杂的数据集时,这是生物医学应用中的常见挑战。本文提出了一种新的CNN-H-Infinity-LSTM架构,用于从心音记录中识别心脏信号。该架构引入了受控制理论中H-无穷滤波器启发的可训练参数,增强了鲁棒性和通用性。在PhysioNet CinC Challenge 2016数据集(心脏音频记录的公共基准)上进行的大量实验表明,所提出的模型实现了稳定的收敛,并优于现有的基准,测试准确率为99.42%,F1得分为98.85%。
摘要:Early detection of heart arrhythmia can prevent severe future complications in cardiac patients. While manual diagnosis still remains the clinical standard, it relies heavily on visual interpretation and is inherently subjective. In recent years, deep learning has emerged as a powerful tool to automate arrhythmia detection, offering improved accuracy, consistency, and efficiency. Several variants of convolutional and recurrent neural network architectures have been widely explored to capture spatial and temporal patterns in physiological signals. However, despite these advancements, current models often struggle to generalize well in real-world scenarios, especially when dealing with small or noisy datasets, which are common challenges in biomedical applications. In this paper, a novel CNN-H-Infinity-LSTM architecture is proposed to identify arrhythmic heart signals from heart sound recordings. This architecture introduces trainable parameters inspired by the H-Infinity filter from control theory, enhancing robustness and generalization. Extensive experimentation on the PhysioNet CinC Challenge 2016 dataset, a public benchmark of heart audio recordings, demonstrates that the proposed model achieves stable convergence and outperforms existing benchmarks, with a test accuracy of 99.42% and an F1 score of 98.85%.
【4】An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
标题:音频MLLM中交错指令调整对语义推理性能的评估
链接:https://arxiv.org/abs/2511.02234
摘要:多模态大型语言模型(MLLM)的标准训练涉及将非文本信息(如视觉或音频)与文本提示连接起来。这种方法可能不鼓励模态的深度集成,限制了模型利用核心语言模型的推理能力的能力。这项工作研究了音频MLLM中交错指令调谐的影响,其中音频令牌在提示内交错。使用听,想,理解(LTU)模型作为测试平台,我们进行了一项实验,使用同义词和上位词音频推理数据集(SHARD),我们新创建的基于音频的语义推理的推理基准集中在同义词和上位词识别。我们的研究结果表明,虽然即使是zero-shot交错提示提高了我们的推理任务的性能,少量的微调使用交错训练提示进一步提高了结果,但是,在MLLM的音频标记能力为代价。
摘要:Standard training for Multi-modal Large Language Models (MLLMs) involves concatenating non-textual information, like vision or audio, with a text prompt. This approach may not encourage deep integration of modalities, limiting the model's ability to leverage the core language model's reasoning capabilities. This work examined the impact of interleaved instruction tuning in an audio MLLM, where audio tokens are interleaved within the prompt. Using the Listen, Think, and Understand (LTU) model as a testbed, we conduct an experiment using the Synonym and Hypernym Audio Reasoning Dataset (SHARD), our newly created reasoning benchmark for audio-based semantic reasoning focusing on synonym and hypernym recognition. Our findings show that while even zero-shot interleaved prompting improves performance on our reasoning tasks, a small amount of fine-tuning using interleaved training prompts improves the results further, however, at the expense of the MLLM's audio labeling ability.
【5】From the perspective of perceptual speech quality: The robustness of frequency bands to noise
标题:从感知语音质量的角度:频段对噪音的鲁棒性
链接:https://arxiv.org/abs/2511.02252
备注:Accepted to J. Acoust. Soc. Am. (JASA) 155, 1916-1927 (2024)
摘要:语音质量是语音相关研究的主要焦点之一,其中语音可懂度是另一个重要的测量指标。然而,频带级感知语音可懂度已经被频繁地研究,而语音质量还没有被彻底地分析。本文提出了一种基于MUSHRA(Multiple Stimuli With Hidden Reference and Anchor)的语音增强方法,以语音质量为衡量标准,研究了不同频段语音对噪声的鲁棒性。语音信号被过滤成32个频带,在不同的信噪比下使用真实世界的噪声。基于分配给重构的带噪语音信号的人类评级感知质量分数来计算对各个频带的噪声指数的鲁棒性。结果的趋势表明,在感知语音质量方面,中频区域对噪声的鲁棒性较低。这些发现表明,未来的研究,旨在提高语音质量应更多地关注语音信号的中频区域。
摘要:Speech quality is one of the main foci of speech-related research, where it is frequently studied with speech intelligibility, another essential measurement. Band-level perceptual speech intelligibility, however, has been studied frequently, whereas speech quality has not been thoroughly analyzed. In this paper, a Multiple Stimuli With Hidden Reference and Anchor (MUSHRA) inspired approach was proposed to study the individual robustness of frequency bands to noise with perceptual speech quality as the measure. Speech signals were filtered into thirty-two frequency bands with compromising real-world noise employed at different signal-to-noise ratios. Robustness to noise indices of individual frequency bands was calculated based on the human-rated perceptual quality scores assigned to the reconstructed noisy speech signals. Trends in the results suggest the mid-frequency region appeared less robust to noise in terms of perceptual speech quality. These findings suggest future research aiming at improving speech quality should pay more attention to the mid-frequency region of the speech signals accordingly.
【6】Condition-Invariant fMRI Decoding of Speech Intelligibility with Deep State Space Model
标题:基于深状态空间模型的语音可理解度条件不变fMRI解码
链接:https://arxiv.org/abs/2511.01868
摘要:阐明语音可懂度的神经基础对于计算神经科学和数字语音处理至关重要。最近的神经影像学研究表明,可懂度调制皮质活动超出简单的声学,主要是在颞上回和额下回。然而,以前的研究主要局限于干净的语音,因此不清楚大脑是否在不同的听力环境中使用条件不变的神经代码。为了解决这一差距,我们提出了一种新的架构,建立在深状态空间模型解码清晰度从fMRI信号,专门针对其高维时间结构。我们提出了第一次尝试在声学不同的条件下解码清晰度,显示我们的方法显着优于经典的方法。此外,区域分析突出了听觉,额叶和顶叶区域的贡献,交叉条件转移表明存在条件不变的神经代码,从而促进了对大脑中抽象语言表征的理解。
摘要:Clarifying the neural basis of speech intelligibility is critical for computational neuroscience and digital speech processing. Recent neuroimaging studies have shown that intelligibility modulates cortical activity beyond simple acoustics, primarily in the superior temporal and inferior frontal gyri. However, previous studies have been largely confined to clean speech, leaving it unclear whether the brain employs condition-invariant neural codes across diverse listening environments. To address this gap, we propose a novel architecture built upon a deep state space model for decoding intelligibility from fMRI signals, specifically tailored to their high-dimensional temporal structure. We present the first attempt to decode intelligibility across acoustically distinct conditions, showing our method significantly outperforms classical approaches. Furthermore, region-wise analysis highlights contributions from auditory, frontal, and parietal regions, and cross-condition transfer indicates the presence of condition-invariant neural codes, thereby advancing understanding of abstract linguistic representations in the brain.
【1】An unscented Kalman filter method for real time input-parameter-state estimation
标题:一种用于实时输入参数状态估计的UKF方法
链接:https://arxiv.org/abs/2511.02717
备注:author-accepted manuscript (AAM) published in Mechanical Systems and Signal Processing
摘要:本文研究了一种新型无迹卡尔曼滤波器在线性和非线性系统上的输入参数状态估计能力。在每个时间步长内,在两个阶段中估计未知输入。首先,预测的动态状态和系统参数提供输入的估计。其次,用测量值校正的状态和参数提供最终估计。重要的是,它表明,使用扰动分析,至少有一个零或一个非零的已知输入系统可以潜在地唯一识别。与经典的仅输出参数识别策略相比,这种仅输出方法可以更好地理解系统,因为所有的动态状态,参数和输入都是联合实时估计的。
摘要:The input-parameter-state estimation capabilities of a novel unscented Kalman filter is examined herein on both linear and nonlinear systems. The unknown input is estimated in two stages within each time step. Firstly, the predicted dynamic states and the system parameters provide an estimation of the input. Secondly, the corrected with measurements states and parameters provide a final estimation. Importantly, it is demonstrated using the perturbation analysis that, a system with at least a zero or a non-zero known input can potentially be uniquely identified. This output-only methodology allows for a better understanding of the system compared to classical output-only parameter identification strategies, given that all the dynamic states, the parameters, and the input are estimated jointly and in real-time.
【2】Multiplexing Neural Audio Watermarks
标题:多路传输神经音频水印
链接:https://arxiv.org/abs/2511.02278
备注:Submission of IEEE ICASSP 2026
摘要:音频水印技术是保证语音内容真实性的有效手段。然而,现有的水印方法仍然容易受到更先进的稀释攻击,如有损压缩和神经重建。在本文中,我们提出了多路神经音频水印技术,以利用它们的互补性下不同类型的攻击。具体而言,五种不同的复用设计进行了研究,包括并行,顺序,频分,时分和感知自适应时频复用(PA-TFM)。我们使用11种不同的攻击方法对LibriSpeech数据进行了多路复用技术的评估,其中包括2种新的神经重建攻击,这些攻击具有语音处理的最新进展。因此,所提出的PA-TFM作为一种无训练复用方法,通过清晰的边缘实现了比单个水印基线更好的性能,展示了一种更鲁棒的音频水印使用方法。
摘要:Audio watermarking is a promising tool to ensure authenticity of speech content. However, existing watermarking methods remain vulnerable to more advanced dilution attacks such as lossy compression and neural reconstruction. In this paper, we propose to multiplex neural audio watermarking techniques to leverage their complementarity under different types of attacks. Specifically, five different multiplexing designs are investigated, including parallel, sequential, frequency-division, time-division and perceptual adaptive time-frequency multiplexing (PA-TFM). We evaluate our multiplexing technique on LibriSpeech data with 11 different attack methods, including 2 new neural reconstruction attacks featuring recent advancements in speech processing. As a result, the proposed PA-TFM as a training-free multiplexing method achieves better performance than single watermarking baselines by clear margins, showcasing a more robust way of using watermarks for audio.
【3】Augmenting Open-Vocabulary Dysarthric Speech Assessment with Human Perceptual Supervision
标题:通过人类知觉监督增强开放词汇性发音障碍言语评估
链接:https://arxiv.org/abs/2511.02270
备注:Submission of IEEE ICASSP 2026
摘要:构音障碍是一种言语障碍,其特征是可懂度受损和交流效率降低。自动构音障碍评估提供了一种可扩展的、具有成本效益的方法,用于支持帕金森病、阿尔茨海默病和中风等神经系统疾病的诊断和治疗。本研究探讨利用人类感知注释语音合成评估作为可靠的域外知识构音障碍的语音评估。实验结果表明,这种监督可以在自监督学习预训练模型中产生一致和实质性的性能改进。这些研究结果表明,感知评级与语音合成评估的人类判断一致,是构音障碍语音建模的宝贵资源,能够实现有效的跨领域知识转移。
摘要:Dysarthria is a speech disorder characterized by impaired intelligibility and reduced communicative effectiveness. Automatic dysarthria assessment provides a scalable, cost-effective approach for supporting the diagnosis and treatment of neurological conditions such as Parkinson's disease, Alzheimer's disease, and stroke. This study investigates leveraging human perceptual annotations from speech synthesis assessment as reliable out-of-domain knowledge for dysarthric speech assessment. Experimental results suggest that such supervision can yield consistent and substantial performance improvements in self-supervised learning pre-trained models. These findings suggest that perceptual ratings aligned with human judgments from speech synthesis evaluations represent valuable resources for dysarthric speech modeling, enabling effective cross-domain knowledge transfer.
【4】From the perspective of perceptual speech quality: The robustness of frequency bands to noise
标题:从感知语音质量的角度:频段对噪音的鲁棒性
链接:https://arxiv.org/abs/2511.02252
备注:Accepted to J. Acoust. Soc. Am. (JASA) 155, 1916-1927 (2024)
摘要:语音质量是语音相关研究的主要焦点之一,其中语音可懂度是另一个重要的测量指标。然而,频带级感知语音可懂度已经被频繁地研究,而语音质量还没有被彻底地分析。本文提出了一种基于MUSHRA(Multiple Stimuli With Hidden Reference and Anchor)的语音增强方法,以语音质量为衡量标准,研究了不同频段语音对噪声的鲁棒性。语音信号被过滤成32个频带,在不同的信噪比下使用真实世界的噪声。基于分配给重构的带噪语音信号的人类评级感知质量分数来计算对各个频带的噪声指数的鲁棒性。结果的趋势表明,在感知语音质量方面,中频区域对噪声的鲁棒性较低。这些发现表明,未来的研究,旨在提高语音质量应更多地关注语音信号的中频区域。
摘要:Speech quality is one of the main foci of speech-related research, where it is frequently studied with speech intelligibility, another essential measurement. Band-level perceptual speech intelligibility, however, has been studied frequently, whereas speech quality has not been thoroughly analyzed. In this paper, a Multiple Stimuli With Hidden Reference and Anchor (MUSHRA) inspired approach was proposed to study the individual robustness of frequency bands to noise with perceptual speech quality as the measure. Speech signals were filtered into thirty-two frequency bands with compromising real-world noise employed at different signal-to-noise ratios. Robustness to noise indices of individual frequency bands was calculated based on the human-rated perceptual quality scores assigned to the reconstructed noisy speech signals. Trends in the results suggest the mid-frequency region appeared less robust to noise in terms of perceptual speech quality. These findings suggest future research aiming at improving speech quality should pay more attention to the mid-frequency region of the speech signals accordingly.
【5】Toward Objective and Interpretable Prosody Evaluation in Text-to-Speech: A Linguistically Motivated Approach
标题:文本到语音中客观且可解释的韵律评估:语言动机的方法
链接:https://arxiv.org/abs/2511.02104
摘要:韵律对于言语技术、塑造理解力、自然度和表现力至关重要。然而,目前的文本到语音(TTS)系统仍然难以准确地捕捉人类一样的韵律变化,部分原因是现有的韵律评价方法仍然有限。像平均意见得分(MOS)这样的传统指标是资源密集型的,不一致的,并且很少能深入了解为什么系统听起来不自然。本研究介绍了一种语言学上知情的,半自动的框架评估TTS韵律通过两层架构,反映人类韵律组织。该方法使用定量的语言学标准来评估合成语音对人类语音语料库在多个声学维度。通过整合离散和连续的韵律措施,它提供了客观的和可解释的指标事件的位置和线索的实现,同时考虑到自然变异观察到的扬声器和韵律线索。结果显示,与感知MOS评级的强相关性,同时揭示了传统的感知测试无法单独捕获的模型特定的弱点。这种方法提供了一个原则性的路径诊断,基准测试,并最终提高下一代TTS系统的韵律自然。
摘要:Prosody is essential for speech technology, shaping comprehension, naturalness, and expressiveness. However, current text-to-speech (TTS) systems still struggle to accurately capture human-like prosodic variation, in part because existing evaluation methods for prosody remain limited. Traditional metrics like Mean Opinion Score (MOS) are resource-intensive, inconsistent, and offer little insight into why a system sounds unnatural. This study introduces a linguistically informed, semi-automatic framework for evaluating TTS prosody through a two-tier architecture that mirrors human prosodic organization. The method uses quantitative linguistic criteria to evaluate synthesized speech against human speech corpora across multiple acoustic dimensions. By integrating discrete and continuous prosodic measures, it provides objective and interpretable metrics of both event placement and cue realization, while accounting for the natural variability observed across speakers and prosodic cues. Results show strong correlations with perceptual MOS ratings while revealing model-specific weaknesses that traditional perceptual tests alone cannot capture. This approach provides a principled path toward diagnosing, benchmarking, and ultimately improving the prosodic naturalness of next-generation TTS systems.
【6】Condition-Invariant fMRI Decoding of Speech Intelligibility with Deep State Space Model
标题:基于深状态空间模型的语音可理解度条件不变fMRI解码
链接:https://arxiv.org/abs/2511.01868
摘要:阐明语音可懂度的神经基础对于计算神经科学和数字语音处理至关重要。最近的神经影像学研究表明,可懂度调制皮质活动超出简单的声学,主要是在颞上回和额下回。然而,以前的研究主要局限于干净的语音,因此不清楚大脑是否在不同的听力环境中使用条件不变的神经代码。为了解决这一差距,我们提出了一种新的架构,建立在深状态空间模型解码清晰度从fMRI信号,专门针对其高维时间结构。我们提出了第一次尝试在声学不同的条件下解码清晰度,显示我们的方法显着优于经典的方法。此外,区域分析突出了听觉,额叶和顶叶区域的贡献,交叉条件转移表明存在条件不变的神经代码,从而促进了对大脑中抽象语言表征的理解。
摘要:Clarifying the neural basis of speech intelligibility is critical for computational neuroscience and digital speech processing. Recent neuroimaging studies have shown that intelligibility modulates cortical activity beyond simple acoustics, primarily in the superior temporal and inferior frontal gyri. However, previous studies have been largely confined to clean speech, leaving it unclear whether the brain employs condition-invariant neural codes across diverse listening environments. To address this gap, we propose a novel architecture built upon a deep state space model for decoding intelligibility from fMRI signals, specifically tailored to their high-dimensional temporal structure. We present the first attempt to decode intelligibility across acoustically distinct conditions, showing our method significantly outperforms classical approaches. Furthermore, region-wise analysis highlights contributions from auditory, frontal, and parietal regions, and cross-condition transfer indicates the presence of condition-invariant neural codes, thereby advancing understanding of abstract linguistic representations in the brain.
机器翻译由腾讯交互翻译提供,仅供参考
