微信公众号:arXiv_Daily
cs.SD语音
【1】ASCMamba: Multimodal Time-Frequency Mamba for Acoustic Scene Classification
标题:ASCMamba:用于声学场景分类的多模式时频Mamba
链接:https://arxiv.org/abs/2508.15632
摘要:声场景分类是计算听觉中的一个基本问题,它试图根据不同的声学特征对环境进行分类。在APSIPA ASC 2025 Grand Challenge的ASC任务中,主办方介绍了一项多模式ASC任务。与仅依赖音频输入的传统ASC系统不同,该挑战提供了额外的文本信息作为输入,包括音频记录的位置和记录时间。在本文中,我们提出了我们在APSIPA ASC 2025 Grand Challenge中提出的ASC任务系统。具体来说,我们提出了一个多模态网络,\textbf{ASCMamba},它集成了音频和文本信息,用于细粒度的声学场景理解和有效的多模态ASC。所提出的ASCMamba采用了一个DenseEncoder从频谱图中提取层次化的频谱特征,然后使用基于Mamba的状态空间模型的双路径Mamba块来捕获长范围的时间和频率依赖性。此外,我们提出了一个两步伪标签机制,以产生更可靠的伪标签。结果表明,所提出的系统优于所有参与的团队,并实现了6.2%的改善基线。代码、模型和预先训练的检查点可在https://github.com/S-Orion/ASCMamba.git上获得。
摘要:Acoustic Scene Classification (ASC) is a fundamental problem in computational audition, which seeks to classify environments based on the distinctive acoustic features. In the ASC task of the APSIPA ASC 2025 Grand Challenge, the organizers introduce a multimodal ASC task. Unlike traditional ASC systems that rely solely on audio inputs, this challenge provides additional textual information as inputs, including the location where the audio is recorded and the time of recording. In this paper, we present our proposed system for the ASC task in the APSIPA ASC 2025 Grand Challenge. Specifically, we propose a multimodal network, \textbf{ASCMamba}, which integrates audio and textual information for fine-grained acoustic scene understanding and effective multimodal ASC. The proposed ASCMamba employs a DenseEncoder to extract hierarchical spectral features from spectrograms, followed by a dual-path Mamba blocks that capture long-range temporal and frequency dependencies using Mamba-based state space models. In addition, we present a two-step pseudo-labeling mechanism to generate more reliable pseudo-labels. Results show that the proposed system outperforms all the participating teams and achieves a 6.2% improvement over the baseline. Code, model and pre-trained checkpoints are available at https://github.com/S-Orion/ASCMamba.git.
【2】Any-to-any Speaker Attribute Perturbation for Asynchronous Voice Anonymization
标题:基于Any-to-Any说话人属性扰动的异步语音识别
链接:https://arxiv.org/abs/2508.15565
摘要:说话人属性扰动为异步语音匿名化提供了一种可行的方法,即利用受扰动的语音作为匿名输出。为了提高来自同一说话人的匿名话语之间的身份不可链接性,通常应用针对性攻击训练策略将话语匿名化到共同的指定说话人。然而,这种策略可能会侵犯实际发言人的指定发言人的隐私。为了降低这种风险,本文提出了一种任意对任意的培训策略。它是通过定义一个批量平均损失匿名的话语从不同的扬声器在一个训练的小批量到一个共同的伪扬声器,这是近似的平均扬声器在小批量。在此基础上,提出了一种说话人对抗语音生成模型,该模型结合了无目标攻击和任意对任意策略的监督。生成说话人属性扰动并将其合并到原始语音中以产生其匿名版本。通过在VoxCeleb数据集上进行的实验,证明了该模型在异步语音匿名化中的有效性。进行了额外的实验,以探索说话人对抗语音在语音隐私保护中的潜在局限性。有了它们,我们的目标是为未来的研究提供见解,以防止黑盒扬声器提取器\textcolor{black}{和自适应攻击的保护功效,以及}泛化到域外数据集\textcolor{black}{和稳定性}。音频样本和开源代码发布在https://github.com/VoicePrivacy/any-to-any-speaker-attribute-perturbation上。
摘要:Speaker attribute perturbation offers a feasible approach to asynchronous voice anonymization by employing adversarially perturbed speech as anonymized output. In order to enhance the identity unlinkability among anonymized utterances from the same original speaker, the targeted attack training strategy is usually applied to anonymize the utterances to a common designated speaker. However, this strategy may violate the privacy of the designated speaker who is an actual speaker. To mitigate this risk, this paper proposes an any-to-any training strategy. It is accomplished by defining a batch mean loss to anonymize the utterances from various speakers within a training mini-batch to a common pseudo-speaker, which is approximated as the average speaker in the mini-batch. Based on this, a speaker-adversarial speech generation model is proposed, incorporating the supervision from both the untargeted attack and the any-to-any strategies. The speaker attribute perturbations are generated and incorporated into the original speech to produce its anonymized version. The effectiveness of the proposed model was justified in asynchronous voice anonymization through experiments conducted on the VoxCeleb datasets. Additional experiments were carried out to explore the potential limitations of speaker-adversarial speech in voice privacy protection. With them, we aim to provide insights for future research on its protective efficacy against black-box speaker extractors \textcolor{black}{and adaptive attacks, as well as} generalization to out-of-domain datasets \textcolor{black}{and stability}. Audio samples and open-source code are published in https://github.com/VoicePrivacy/any-to-any-speaker-attribute-perturbation.
【3】DualMark: Identifying Model and Training Data Origins in Generated Audio
标题:DualMark:识别生成的音频中的模型和训练数据来源
链接:https://arxiv.org/abs/2508.15521
备注:13 pages, 5 figures
摘要:现有的音频生成模型的水印方法只能实现模型级属性,允许识别原始生成模型,但无法跟踪底层训练数据集。这一重大限制提出了关键的出处问题,特别是在涉及版权和问责问题的情况下。为了弥合这一根本差距,我们引入了DualMark,这是第一个能够同时编码两个不同属性签名的双来源水印框架,即,模型身份和数据集来源,在训练期间转换为音频生成模型。具体来说,我们提出了一种新的双水印嵌入(DWE)模块无缝嵌入双水印到梅尔频谱表示,伴随着精心设计的水印一致性损失(WCL),这确保了可靠的提取两个水印从生成的音频信号。此外,我们建立了双归因基准(DAB),第一个鲁棒性评估基准专门为联合模型数据属性。大量的实验验证了DualMark实现了出色的属性准确性(模型属性的F1分数为97.01%,数据集属性的AUC为91.51%),同时保持了对侵略性修剪,有损压缩,加性噪声和采样攻击的出色鲁棒性,这些条件严重损害了先前的方法。因此,我们的工作为实现完全负责任的音频生成模型迈出了基础性的一步,显著增强了版权保护和责任追踪能力。
摘要:Existing watermarking methods for audio generative models only enable model-level attribution, allowing the identification of the originating generation model, but are unable to trace the underlying training dataset. This significant limitation raises critical provenance questions, particularly in scenarios involving copyright and accountability concerns. To bridge this fundamental gap, we introduce DualMark, the first dual-provenance watermarking framework capable of simultaneously encoding two distinct attribution signatures, i.e., model identity and dataset origin, into audio generative models during training. Specifically, we propose a novel Dual Watermark Embedding (DWE) module to seamlessly embed dual watermarks into Mel-spectrogram representations, accompanied by a carefully designed Watermark Consistency Loss (WCL), which ensures reliable extraction of both watermarks from generated audio signals. Moreover, we establish the Dual Attribution Benchmark (DAB), the first robustness evaluation benchmark specifically tailored for joint model-data attribution. Extensive experiments validate that DualMark achieves outstanding attribution accuracy (97.01% F1-score for model attribution, and 91.51% AUC for dataset attribution), while maintaining exceptional robustness against aggressive pruning, lossy compression, additive noise, and sampling attacks, conditions that severely compromise prior methods. Our work thus provides a foundational step toward fully accountable audio generative models, significantly enhancing copyright protection and responsibility tracing capabilities.
【4】AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation
标题:AudioSet-R:具有多阶段LLM标签重新注释的精制AudioSet
链接:https://arxiv.org/abs/2508.15429
备注:8 pages, 5 figures, accepted in ACM MM 2025 dataset track
摘要:AudioSet是音频研究社区中广泛使用的基准,并显着推进了各种音频相关的任务。然而,标签的准确性和完整性的持续存在的问题仍然是关键的瓶颈,限制了下游应用程序的性能。为了解决上述挑战,我们提出了一个三阶段的重新注释框架,利用通用的音频语言基础模型,系统地提高AudioSet的标签质量。该框架采用了跨模态提示策略,提示链的概念,其中提示顺序组成执行子任务(音频理解,标签合成,语义对齐)的启发。利用这个框架,我们构建了一个高质量的,结构化的重新标记版本的AudioSet-R。对代表性音频分类模型(包括AST、PANN、SSAST和AudioMAE)进行的大量实验一致地证明了实质性的性能改进,从而验证了所提出的方法在增强标签可靠性方面的通用性和有效性。https://github.com/colaudiolab/AudioSet-R
摘要:AudioSet is a widely used benchmark in the audio research community and has significantly advanced various audio-related tasks. However, persistent issues with label accuracy and completeness remain critical bottlenecks that limit performance in downstream applications.To address the aforementioned challenges, we propose a three-stage reannotation framework that harnesses general-purpose audio-language foundation models to systematically improve the label quality of AudioSet. The framework employs a cross-modal prompting strategy, inspired by the concept of prompt chaining, wherein prompts are sequentially composed to execute subtasks (audio comprehension, label synthesis, and semantic alignment). Leveraging this framework, we construct a high-quality, structured relabeled version of AudioSet-R. Extensive experiments conducted on representative audio classification models--including AST, PANNs, SSAST, and AudioMAE--consistently demonstrate substantial performance improvements, thereby validating the generalizability and effectiveness of the proposed approach in enhancing label reliability.The code is publicly available at: https://github.com/colaudiolab/AudioSet-R.
【5】LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
标题:LLaSO:大型语言和语音模型可重复性研究的基础框架
链接:https://arxiv.org/abs/2508.15418
摘要:大型语音语言模型(LSLMs)的发展由于架构分散和缺乏透明度而放缓,阻碍了研究的系统比较和可重复性。与视觉语言领域不同,LSLM领域的常见做法是在没有相应训练数据和配置的情况下释放模型权重。为了解决这些关键的差距,我们引入了LLaSO,这是第一个完全开放的,端到端的大规模语音语言建模框架。LLaSO为社区提供了三个基本资源:(1)LLaSO-Align,一个12 M实例的语音-文本对齐语料库;(2)LLaSO-Instruct,一个1350万实例的多任务调试数据集;(3)LLaSO-Eval,一个可重复的标准化评估基准。为了验证我们的框架,我们构建并发布了LLaSO-Base,这是一个专门在我们的公共数据上训练的3. 8B参数参考模型。它实现了0.72的标准化得分,建立了一个强大的,可重复的基线,超越了可比模型。我们的分析表明,虽然更广泛的训练覆盖范围可以提高性能,但在看不见的任务上,特别是在纯音频场景中,仍然存在显着的泛化差距。通过发布完整的数据、基准和模型,LLaSO建立了一个基础性的开放标准,以统一研究工作并加速LSLM的社区驱动的进展。我们在https://github.com/EIT-NLP/LLaSO上发布代码、数据集、预训练模型和结果。
摘要:The development of Large Speech-Language Models (LSLMs) has been slowed by fragmented architectures and a lack of transparency, hindering the systematic comparison and reproducibility of research. Unlike in the vision-language domain, the LSLM field suffers from the common practice of releasing model weights without their corresponding training data and configurations. To address these critical gaps, we introduce LLaSO, the first fully open, end-to-end framework for large-scale speech-language modeling. LLaSO provides the community with three essential resources: (1) LLaSO-Align, a 12M-instance speech-text alignment corpus; (2) LLaSO-Instruct, a 13.5M-instance multi-task instruction-tuning dataset; and (3) LLaSO-Eval, a reproducible benchmark for standardized evaluation. To validate our framework, we build and release LLaSO-Base, a 3.8B-parameter reference model trained exclusively on our public data. It achieves a normalized score of 0.72, establishing a strong, reproducible baseline that surpasses comparable models. Our analysis reveals that while broader training coverage enhances performance, significant generalization gaps persist on unseen tasks, particularly in pure audio scenarios. By releasing the complete stack of data, benchmarks, and models, LLaSO establishes a foundational open standard to unify research efforts and accelerate community-driven progress in LSLMs. We release the code, dataset, pretrained models, and results in https://github.com/EIT-NLP/LLaSO.
【6】An Enhanced Audio Feature Tailored for Anomalous Sound Detection Based on Pre-trained Models
标题:基于预训练模型为异常声音检测量身定制的增强音频特征
链接:https://arxiv.org/abs/2508.15334
备注:13 pages, 3 figures, accepted by ICANN2025
摘要:异常声音检测(ASD)旨在识别机器发出的异常声音,已引起学术界和工业界的广泛研究兴趣。然而,异常位置的不确定性和机器声音中的大量冗余信息(如噪声)阻碍了ASD系统性能的提高。本文提出了一种新的音频特征的滤波器组均匀分布的时间间隔,确保平等地关注音频中的所有频率范围,这增强了机器声音中的异常检测。此外,基于预训练的模型,本文提出了一种无参数的特征增强方法来去除机器音频中的冗余信息。据信,这种无参数策略有助于在模型微调期间将通用知识从预先训练的任务有效地转移到ASD任务。声学场景和事件的检测和分类(DCASE)2024挑战数据集的评估结果表明,我们提出的方法在ASD性能方面有显着改善。
摘要:Anomalous Sound Detection (ASD) aims at identifying anomalous sounds from machines and has gained extensive research interests from both academia and industry. However, the uncertainty of anomaly location and much redundant information such as noise in machine sounds hinder the improvement of ASD system performance. This paper proposes a novel audio feature of filter banks with evenly distributed intervals, ensuring equal attention to all frequency ranges in the audio, which enhances the detection of anomalies in machine sounds. Moreover, based on pre-trained models, this paper presents a parameter-free feature enhancement approach to remove redundant information in machine audio. It is believed that this parameter-free strategy facilitates the effective transfer of universal knowledge from pre-trained tasks to the ASD task during model fine-tuning. Evaluation results on the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024 Challenge dataset demonstrate significant improvements in ASD performance with our proposed methods.
【7】UniCoM: A Universal Code-Switching Speech Generator
标题:UniCoM:通用代码切换语音生成器
链接:https://arxiv.org/abs/2508.15244
备注:Accepted to EMNLP 2025 Findings
摘要:语码转换(CS)是一个说话人在两种或两种以上语言之间的转换,在现实世界的会话中很常见,对多语言语音技术提出了重大挑战。然而,能够处理这种现象的系统仍然没有得到充分的探索,主要是由于缺乏合适的数据集。为了解决这个问题,我们提出了通用代码混合器(UniCoM),一种新的管道,用于生成高质量,自然的CS样本,而不改变句子语义。我们的方法利用了一种算法,我们称之为用同义词替换单词(SWORDS),它通过将选定的单词替换为它们的翻译,同时考虑它们的词性来生成CS语音。使用UniCoM,我们构建了代码转换FLEURS(CS-FLEURS),一个多语言CS语料库,专为自动语音识别(ASR)和语音到文本翻译(S2 TT)。实验结果表明,CS-FLEURS实现了较高的可懂度和自然度,对现有数据集的客观和主观指标进行了测试。我们希望我们的方法能够推进CS语音技术,并实现更具包容性的多语言系统。
摘要:Code-switching (CS), the alternation between two or more languages within a single speaker's utterances, is common in real-world conversations and poses significant challenges for multilingual speech technology. However, systems capable of handling this phenomenon remain underexplored, primarily due to the scarcity of suitable datasets. To resolve this issue, we propose Universal Code-Mixer (UniCoM), a novel pipeline for generating high-quality, natural CS samples without altering sentence semantics. Our approach utilizes an algorithm we call Substituting WORDs with Synonyms (SWORDS), which generates CS speech by replacing selected words with their translations while considering their parts of speech. Using UniCoM, we construct Code-Switching FLEURS (CS-FLEURS), a multilingual CS corpus designed for automatic speech recognition (ASR) and speech-to-text translation (S2TT). Experimental results show that CS-FLEURS achieves high intelligibility and naturalness, performing comparably to existing datasets on both objective and subjective metrics. We expect our approach to advance CS speech technology and enable more inclusive multilingual systems.
【8】Comparative Evaluation of Text and Audio Simplification: A Methodological Replication Study
标题:文本和音频简化的比较评估:方法学复制研究
链接:https://arxiv.org/abs/2508.15088
摘要:这项研究是Leroy等人(2022)研究的方法学复制,该研究调查了文本简化对不断发展的多媒体环境中医疗保健信息理解的影响。基于原始研究的见解,我们的复制研究评估了音频内容,认识到其在数字时代传播医疗保健信息的重要性越来越大。具体来说,我们探讨了当用户参与从该文本自动生成的音频内容时,文本简化对感知和实际困难的影响。我们的复制涉及44名参与者,我们评估了他们对使用Leroy等人(2022)原始和简化文本创建的音频形式的医疗保健信息的理解。我们的研究结果强调了文本简化在增强感知可理解性和实际理解方面的有效性,与原始研究的结果一致。此外,我们还研究了教育水平和语言能力的作用,揭示了它们对医疗保健信息获取和理解的潜在影响。这项研究强调了文本简化工具在促进健康素养方面的实用价值。它表明,需要量身定制的沟通策略,以达到不同的受众有效地在医疗保健领域。
摘要:This study serves as a methodological replication of Leroy et al. (2022) research, which investigated the impact of text simplification on healthcare information comprehension in the evolving multimedia landscape. Building upon the original studys insights, our replication study evaluates audio content, recognizing its increasing importance in disseminating healthcare information in the digital age. Specifically, we explored the influence of text simplification on perceived and actual difficulty when users engage with audio content automatically generated from that text. Our replication involved 44 participants for whom we assessed their comprehension of healthcare information presented as audio created using Leroy et al. (2022) original and simplified texts. The findings from our study highlight the effectiveness of text simplification in enhancing perceived understandability and actual comprehension, aligning with the original studys results. Additionally, we examined the role of education level and language proficiency, shedding light on their potential impact on healthcare information access and understanding. This research underscores the practical value of text simplification tools in promoting health literacy. It suggests the need for tailored communication strategies to reach diverse audiences effectively in the healthcare domain.
【9】XAI-Driven Spectral Analysis of Cough Sounds for Respiratory Disease Characterization
标题:XAI驱动的咳嗽声频谱分析用于呼吸系统疾病特征
链接:https://arxiv.org/abs/2508.14949
摘要:本文提出了一种可解释的人工智能(XAI)驱动的方法,以提高对呼吸系统疾病管理咳嗽声分析的理解。我们采用遮挡图来突出显示卷积神经网络(CNN)处理的咳嗽频谱图中的相关频谱区域。随后,由这些遮挡图加权的频谱图的频谱分析揭示了疾病组之间的显著差异,特别是在患有COPD的患者中,其中咳嗽模式在所识别的感兴趣的频谱区域中出现更多变化。这与分析原始光谱图时观察到的缺乏显著差异形成对比。所提出的方法提取和分析了几个频谱特征,展示了XAI技术的潜力,以发现疾病特异性的声学特征,并通过提供更多的可解释的结果来提高咳嗽声分析的诊断能力。
摘要:This paper proposes an eXplainable Artificial Intelligence (XAI)-driven methodology to enhance the understanding of cough sound analysis for respiratory disease management. We employ occlusion maps to highlight relevant spectral regions in cough spectrograms processed by a Convolutional Neural Network (CNN). Subsequently, spectral analysis of spectrograms weighted by these occlusion maps reveals significant differences between disease groups, particularly in patients with COPD, where cough patterns appear more variable in the identified spectral regions of interest. This contrasts with the lack of significant differences observed when analyzing raw spectrograms. The proposed approach extracts and analyzes several spectral features, demonstrating the potential of XAI techniques to uncover disease-specific acoustic signatures and improve the diagnostic capabilities of cough sound analysis by providing more interpretable results.
【10】Human Feedback Driven Dynamic Speech Emotion Recognition
标题:人类反馈驱动的动态语音情感识别
链接:https://arxiv.org/abs/2508.14920
摘要:本文的工作为动态语音情感识别开辟了一个新的领域。与传统方法不同,我们假设每个音轨都与在不同时刻活跃的情绪序列相关联。该研究特别关注情感3D化身的动画。我们提出了一个多阶段的方法,包括一个经典的语音情感识别模型的训练,情感序列的合成生成,并进一步改进模型的基础上,人类的反馈。此外,我们介绍了一种新的方法来建模的基础上的狄利克雷分布的情感混合物。基于从3D面部动画数据集中提取的真实情感对模型进行评估。我们将我们的模型与滑动窗口方法进行比较。我们的实验结果表明,狄利克雷为基础的方法在建模情感的混合物的有效性。简化人工反馈进一步提高了模型质量,同时提供了简化的注释过程。
摘要:This work proposes to explore a new area of dynamic speech emotion recognition. Unlike traditional methods, we assume that each audio track is associated with a sequence of emotions active at different moments in time. The study particularly focuses on the animation of emotional 3D avatars. We propose a multi-stage method that includes the training of a classical speech emotion recognition model, synthetic generation of emotional sequences, and further model improvement based on human feedback. Additionally, we introduce a novel approach to modeling emotional mixtures based on the Dirichlet distribution. The models are evaluated based on ground-truth emotions extracted from a dataset of 3D facial animations. We compare our models against the sliding window approach. Our experimental results show the effectiveness of Dirichlet-based approach in modeling emotional mixtures. Incorporating human feedback further improves the model quality while providing a simplified annotation procedure.
【11】Denoising by neural network for muzzle blast detection
标题:枪口爆炸检测的神经网络去噪
链接:https://arxiv.org/abs/2508.14919
备注:INTER-NOISE 2024, Aug 2024, Nantes (France), France
摘要:Acoem开发枪击检测系统,包括麦克风阵列和软件,可以检测和定位战场上的射手。这种系统的性能显然受到其运行的声学环境的影响:特别是,当安装在移动的军用车辆上时,噪声的存在降低了软件的检测性能。为了限制声学环境的影响,已经开发了神经网络。我们没有使用重型卷积神经网络,而是选择了轻量级神经网络架构,以限制在尽可能多的硬件平台上嵌入算法所需的计算资源。由于两个隐藏层感知器和适当的信号处理技术的结合,脉冲枪口爆炸波形(来自爆炸的波,指示射手的位置)的检测率显着增加。当噪声的均方根值与枪口冲击波峰值幅度同阶时,采用这种去噪处理,检测率提高了一倍以上。
摘要:Acoem develops gunshot detection systems, consisting of a microphone array and software that detects and locates shooters on the battlefield. The performance of such systems is obviously affected by the acoustic environment in which they are operating: in particular, when mounted on a moving military vehicle, the presence of noise reduces the detection performance of the software. To limit the influence of the acoustic environment, a neural network has been developed. Instead of using a heavy convolutional neural network, a lightweight neural network architecture was chosen to limit the computational resources required to embed the algorithm on as many hardware platforms as possible. Thanks to the combination of a two hidden layer perceptron and appropriate signal processing techniques, the detection rate of impulsive muzzle blast waveforms (the wave coming from the detonation and indicating the position of the shooter) is significantly increased. With a rms value of noise of the same order as the muzzle blast peak amplitude, the detect rate is more than doubled with this denoising processing.
【12】Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets
标题:使用GFlowNets通过分布对齐缓解基于LM的TTC模型中的幻觉
链接:https://arxiv.org/abs/2508.15442
摘要:基于语言模型(LM)的文本到语音(TTS)系统通常会生成与输入文本偏离的幻觉语音。现有的缓解策略要么需要过多的训练资源,要么引入显著的推理延迟。在本文中,我们提出了GFlOwNet引导的分布式AlignmentT(GOAT),用于基于LM的TTS,这是一个后训练框架,可以在不依赖大量资源或推理成本的情况下减轻幻觉。具体来说,我们首先进行不确定性分析,揭示了幻觉和模型的不确定性之间的强正相关性。在此基础上,我们将TTS生成重新表示为轨迹流优化问题,并引入增强的Subtrajectory Balance目标以及锐化的内部奖励作为目标分布。我们进一步集成了奖励温度衰减和学习率优化,以实现稳定性和性能平衡。大量的实验表明,GOAT在挑战性测试用例上降低了50%以上的字符错误率,降低了高达58%的不确定性,证明了其强大的泛化能力和有效性。
摘要:Language Model (LM)-based Text-to-Speech (TTS) systems often generate hallucinated speech that deviates from input text. Existing mitigation strategies either demand excessive training resources or introduce significant inference latency. In this paper, we propose GFlOwNet-guided distribution AlignmenT (GOAT) for LM-based TTS, a post-training framework that mitigates hallucinations without relying on massive resources or inference cost. Specifically, we first conduct an uncertainty analysis, revealing a strong positive correlation between hallucination and model uncertainty. Based on this, we reformulate TTS generation as a trajectory flow optimization problem and introduce an enhanced Subtrajectory Balance objective together with a sharpened internal reward as target distribution. We further integrate reward temperature decay and learning rate optimization for stability and performance balance. Extensive experiments show that GOAT reduce over 50% character error rates on challenging test cases and lowering uncertainty by up to 58%, demonstrating its strong generalization ability and effectiveness.
【13】Optimal Interference Signal for Masking an Acoustic Source
标题:用于掩蔽声学源的最佳干扰信号
链接:https://arxiv.org/abs/2508.15023
备注:40 pages, a preprint
摘要:在需要声学隐私或故意信号混淆的环境中,有必要屏蔽在基本操作中生成的声学签名。我们考虑的问题,掩蔽的声源在目标区域中的可能的检测传感器的位置的影响。通过将干扰信号放置在声源附近来实现掩蔽。我们引入了一个理论和计算框架,设计这样的干扰信号的目标是最小化目标区域中的残余振幅。对于具有球对称性的三维强迫波方程,我们导出了几种典型情形的解析拟定常周期解。我们研究了自掩蔽的现象,其中具有一定空间强迫轮廓的声源掩蔽自身,使其强迫足迹之外的检测。然后,我们使用叠加的球对称解决方案,调查在一个给定的目标区域的掩蔽。我们分析和优化的性能,使用一个或两个点的力量部署在声源附近的掩蔽在目标区域。对于声源的空间强迫分布缺乏球对称性的一般情况,我们发展了一种有效的数值方法来求解三维波动方程。这项工作的潜在应用包括海底声学通信安全,海底车辆隐身,以及对声学监视的保护。
摘要:In an environment where acoustic privacy or deliberate signal obfuscation is desired, it is necessary to mask the acoustic signature generated in essential operations. We consider the problem of masking the effect of an acoustic source in a target region where possible detection sensors are located. Masking is achieved by placing interference signals near the acoustic source. We introduce a theoretical and computational framework for designing such interference signals with the goal of minimizing the residual amplitude in the target region. For the three-dimensional (3D) forced wave equation with spherical symmetry, we derive analytical quasi-steady periodic solutions for several canonical cases. We examine the phenomenon of self-masking where an acoustic source with certain spatial forcing profile masks itself from detection outside its forcing footprint. We then use superposition of spherically symmetric solutions to investigate masking in a given target region. We analyze and optimize the performance of using one or two point-forces deployed near the acoustic source for masking in the target region. For the general case where the spatial forcing profile of the acoustic source lacks spherical symmetry, we develop an efficient numerical method for solving the 3D wave equation. Potential applications of this work include undersea acoustic communication security, undersea vehicles stealth, and protection against acoustic surveillance.
【14】A Chinese Heart Failure Status Speech Database with Universal and Personalised Classification
标题:通用个性化分类的中国心力衰竭状态语音数据库
链接:https://arxiv.org/abs/2508.14908
摘要:语音是识别急性和慢性心力衰竭(HF)的一种具有成本效益的非侵入性数据源。然而,汉语音节中是否含有高频相关信息的研究却很少。这项研究提出了第一个中文语音数据库的HF患者,具有配对录音之前和之后的住院。研究结果证实了使用标准的“患者”和个性化的“成对”分类方法进行HF检测的中文语言的有效性,后者可作为未来研究的理想扬声器解耦基线。统计检验和分类结果突出了个体差异是造成不准确的主要因素。此外,自适应频率滤波器(AFF)提出的频率重要性分析。数据和演示发布在https://github.com/panyue1998/Voice_HF上。
摘要:Speech is a cost-effective and non-intrusive data source for identifying acute and chronic heart failure (HF). However, there is a lack of research on whether Chinese syllables contain HF-related information, as observed in other well-studied languages. This study presents the first Chinese speech database of HF patients, featuring paired recordings taken before and after hospitalisation. The findings confirm the effectiveness of the Chinese language in HF detection using both standard 'patient-wise' and personalised 'pair-wise' classification approaches, with the latter serving as an ideal speaker-decoupled baseline for future research. Statistical tests and classification results highlight individual differences as key contributors to inaccuracy. Additionally, an adaptive frequency filter (AFF) is proposed for frequency importance analysis. The data and demonstrations are published at https://github.com/panyue1998/Voice_HF.
【1】EffortNet: A Deep Learning Framework for Objective Assessment of Speech Enhancement Technologies Using EEG-Based Alpha Oscillations
标题:EffortNet:一个深度学习框架,用于使用基于脑电波的Alpha振荡客观评估语音增强技术
链接:https://arxiv.org/abs/2508.15473
摘要:本文介绍了EffortNet,这是一种新型的深度学习框架,用于在语音理解过程中从脑电图(EEG)中解码个人听力努力。听力努力是言语听力研究中的一个重大挑战,特别是对于老龄化人群和听力障碍人群。我们收集了122名参与者在四种条件下的语音理解过程中的64通道EEG数据:干净,嘈杂,MMSE增强和变压器增强语音。统计分析证实,与干净或增强条件相比,阿尔法振荡(8-13 Hz)在嘈杂的语音处理期间表现出显著更高的功率,证实了它们作为听力努力的客观生物标志物的有效性。为了解决EEG信号中的个体间差异,EffortNet集成了三种互补的学习范式:利用未标记数据的自我监督学习,渐进式适应个体特征的增量学习,以及将知识有效转移到新主题的转移学习。我们的实验结果表明,Effort-Net在只有40%来自新主题的训练数据的情况下达到了80.9%的分类准确率,显著优于传统的CNN(62.3%)和STAnet(61.1%)模型。从我们的模型中得出的基于概率的度量显示,Transformer增强的语音引起的神经反应比MMSE增强的语音更类似于干净的语音。这一发现与主观清晰度评级形成对比,但与客观指标一致。建议的框架提供了一个实用的解决方案,个性化评估的听力技术,与设计认知感知语音增强系统的影响。
摘要:This paper presents EffortNet, a novel deep learning framework for decoding individual listening effort from electroencephalography (EEG) during speech comprehension. Listening effort represents a significant challenge in speech-hearing research, particularly for aging populations and those with hearing impairment. We collected 64-channel EEG data from 122 participants during speech comprehension under four conditions: clean, noisy, MMSE-enhanced, and Transformer-enhanced speech. Statistical analyses confirmed that alpha oscillations (8-13 Hz) exhibited significantly higher power during noisy speech processing compared to clean or enhanced conditions, confirming their validity as objective biomarkers of listening effort. To address the substantial inter-individual variability in EEG signals, EffortNet integrates three complementary learning paradigms: self-supervised learning to leverage unlabeled data, incremental learning for progressive adaptation to individual characteristics, and transfer learning for efficient knowledge transfer to new subjects. Our experimental results demonstrate that Effort- Net achieves 80.9% classification accuracy with only 40% training data from new subjects, significantly outperforming conventional CNN (62.3%) and STAnet (61.1%) models. The probability-based metric derived from our model revealed that Transformer-enhanced speech elicited neural responses more similar to clean speech than MMSEenhanced speech. This finding contrasted with subjective intelligibility ratings but aligned with objective metrics. The proposed framework provides a practical solution for personalized assessment of hearing technologies, with implications for designing cognitive-aware speech enhancement systems.
【2】Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets
标题:使用GFlowNets通过分布对齐缓解基于LM的TTC模型中的幻觉
链接:https://arxiv.org/abs/2508.15442
摘要:基于语言模型(LM)的文本到语音(TTS)系统通常会生成与输入文本偏离的幻觉语音。现有的缓解策略要么需要过多的训练资源,要么引入显著的推理延迟。在本文中,我们提出了GFlOwNet引导的分布式AlignmentT(GOAT),用于基于LM的TTS,这是一个后训练框架,可以在不依赖大量资源或推理成本的情况下减轻幻觉。具体来说,我们首先进行不确定性分析,揭示了幻觉和模型的不确定性之间的强正相关性。在此基础上,我们将TTS生成重新表示为轨迹流优化问题,并引入增强的Subtrajectory Balance目标以及锐化的内部奖励作为目标分布。我们进一步整合了奖励温度衰减和学习率优化,以实现稳定性和性能平衡。大量的实验表明,GOAT在挑战性测试用例上降低了50%以上的字符错误率,降低了高达58%的不确定性,证明了其强大的泛化能力和有效性。
摘要:Language Model (LM)-based Text-to-Speech (TTS) systems often generate hallucinated speech that deviates from input text. Existing mitigation strategies either demand excessive training resources or introduce significant inference latency. In this paper, we propose GFlOwNet-guided distribution AlignmenT (GOAT) for LM-based TTS, a post-training framework that mitigates hallucinations without relying on massive resources or inference cost. Specifically, we first conduct an uncertainty analysis, revealing a strong positive correlation between hallucination and model uncertainty. Based on this, we reformulate TTS generation as a trajectory flow optimization problem and introduce an enhanced Subtrajectory Balance objective together with a sharpened internal reward as target distribution. We further integrate reward temperature decay and learning rate optimization for stability and performance balance. Extensive experiments show that GOAT reduce over 50% character error rates on challenging test cases and lowering uncertainty by up to 58%, demonstrating its strong generalization ability and effectiveness.
【3】Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge
标题:针对MLC-SLC 2025挑战赛的Transsion多语言语音识别系统
链接:https://arxiv.org/abs/2508.14916
摘要:本文介绍了一种新的多语言自动语音识别(ASR)系统的架构和性能开发的跨语音团队的MLC-SLM 2025挑战赛的轨道1。所提出的系统包括三个关键组件:1)基于冻结Whisper-large-v3的语音编码器,利用大规模预训练来确保鲁棒的声学特征提取; 2)使用Linear-ReLU-Linear变换机制来有效地对齐语音和文本表示的可训练适配器模块;以及3)冻结的Qwen2.5- 7 B-Instruct大语言模型(LLM)与可训练的LoRA集成,用于优化上下文语言解码。通过系统地将预训练模型与特定任务的微调相结合,该系统在评估集中的11种语言中实现了9.83%的单词/字符错误率(WER/CER),在全球参与者中排名第三。
摘要:This paper presents the architecture and performance of a novel Multilingual Automatic Speech Recognition (ASR) system developed by the Transsion Speech Team for Track 1 of the MLC-SLM 2025 Challenge. The proposed system comprises three key components: 1) a frozen Whisper-large-v3 based speech encoder, leveraging large-scale pretraining to ensure robust acoustic feature extraction; 2) a trainable adaptor module using Linear-ReLU-Linear transformation mechanisms to effectively align speech and text representations; and 3) a frozen Qwen2.5-7B-Instruct large language model (LLM) integrated with trainable LoRA for optimized contextual linguistic decoding. By systematically combining pretrained models with task specific fine-tuning, the system achieved a word/character error rate (WER/CER) of 9.83% across 11 languages in the evaluation set and ranked third place among global participants.
【4】A Chinese Heart Failure Status Speech Database with Universal and Personalised Classification
标题:通用个性化分类的中国心力衰竭状态语音数据库
链接:https://arxiv.org/abs/2508.14908
摘要:语音是识别急性和慢性心力衰竭(HF)的一种具有成本效益的非侵入性数据源。然而,汉语音节中是否含有高频相关信息的研究却很少。这项研究提出了第一个中文语音数据库的HF患者,具有配对录音之前和之后的住院。研究结果证实了使用标准的“患者”和个性化的“成对”分类方法进行HF检测的中文语言的有效性,后者可作为未来研究的理想扬声器解耦基线。统计检验和分类结果突出了个体差异是造成不准确的主要因素。此外,自适应频率滤波器(AFF)提出的频率重要性分析。数据和演示发布在https://github.com/panyue1998/Voice_HF上。
摘要:Speech is a cost-effective and non-intrusive data source for identifying acute and chronic heart failure (HF). However, there is a lack of research on whether Chinese syllables contain HF-related information, as observed in other well-studied languages. This study presents the first Chinese speech database of HF patients, featuring paired recordings taken before and after hospitalisation. The findings confirm the effectiveness of the Chinese language in HF detection using both standard 'patient-wise' and personalised 'pair-wise' classification approaches, with the latter serving as an ideal speaker-decoupled baseline for future research. Statistical tests and classification results highlight individual differences as key contributors to inaccuracy. Additionally, an adaptive frequency filter (AFF) is proposed for frequency importance analysis. The data and demonstrations are published at https://github.com/panyue1998/Voice_HF.
【5】An Enhanced Audio Feature Tailored for Anomalous Sound Detection Based on Pre-trained Models
标题:基于预训练模型为异常声音检测量身定制的增强音频特征
链接:https://arxiv.org/abs/2508.15334
备注:13 pages, 3 figures, accepted by ICANN2025
摘要:异常声音检测(ASD)旨在识别机器发出的异常声音,已引起学术界和工业界的广泛研究兴趣。然而,异常位置的不确定性和机器声音中的大量冗余信息(如噪声)阻碍了ASD系统性能的提高。本文提出了一种新的音频特征的滤波器组均匀分布的时间间隔,确保平等地关注音频中的所有频率范围,这增强了机器声音中的异常检测。此外,基于预训练的模型,本文提出了一种无参数的特征增强方法来去除机器音频中的冗余信息。据信,这种无参数策略有助于在模型微调期间将通用知识从预先训练的任务有效地转移到ASD任务。声学场景和事件的检测和分类(DCASE)2024挑战数据集的评估结果表明,我们提出的方法在ASD性能方面有显着改善。
摘要:Anomalous Sound Detection (ASD) aims at identifying anomalous sounds from machines and has gained extensive research interests from both academia and industry. However, the uncertainty of anomaly location and much redundant information such as noise in machine sounds hinder the improvement of ASD system performance. This paper proposes a novel audio feature of filter banks with evenly distributed intervals, ensuring equal attention to all frequency ranges in the audio, which enhances the detection of anomalies in machine sounds. Moreover, based on pre-trained models, this paper presents a parameter-free feature enhancement approach to remove redundant information in machine audio. It is believed that this parameter-free strategy facilitates the effective transfer of universal knowledge from pre-trained tasks to the ASD task during model fine-tuning. Evaluation results on the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024 Challenge dataset demonstrate significant improvements in ASD performance with our proposed methods.
【6】CUPE: Contextless Universal Phoneme Encoder for Language-Agnostic Speech Processing
标题:CUPE:用于语音不可知语音处理的无上下文通用音素编码器
链接:https://arxiv.org/abs/2508.15316
备注:Accepted in: 8th International Conference on Natural Language and Speech Processing (ICNLSP 2025)
摘要:通用音素识别通常需要分析长语音片段和特定于语言的模式。许多语音处理任务需要不受上下文影响的纯音素表示,这促使我们开发了CUPE -一种轻量级模型,可以在120毫秒内捕获关键音素特征,大约是一个音素的长度。CUPE独立处理短的固定宽度窗口,尽管参数比当前方法少,但通过学习所有语言共同的基本声学模式,实现了具有竞争力的跨语言性能。我们通过对不同语言的监督和自我监督训练进行了广泛的评估,包括对UCLA语音语料库的zero-shot测试,证明了强大的跨语言泛化能力,并揭示了通过在音素长度窗口内建模基本声学模式来实现有效的通用语音处理是可能的。
摘要:Universal phoneme recognition typically requires analyzing long speech segments and language-specific patterns. Many speech processing tasks require pure phoneme representations free from contextual influence, which motivated our development of CUPE - a lightweight model that captures key phoneme features in just 120 milliseconds, about one phoneme's length. CUPE processes short, fixed-width windows independently and, despite fewer parameters than current approaches, achieves competitive cross-lingual performance by learning fundamental acoustic patterns common to all languages. Our extensive evaluation through supervised and self-supervised training on diverse languages, including zero-shot tests on the UCLA Phonetic Corpus, demonstrates strong cross-lingual generalization and reveals that effective universal speech processing is possible through modeling basic acoustic patterns within phoneme-length windows.
【7】UniCoM: A Universal Code-Switching Speech Generator
标题:UniCoM:通用代码切换语音生成器
链接:https://arxiv.org/abs/2508.15244
备注:Accepted to EMNLP 2025 Findings
摘要:语码转换(CS)是一个说话人在两种或两种以上语言之间的转换,在现实世界的会话中很常见,对多语言语音技术提出了重大挑战。然而,能够处理这种现象的系统仍然没有得到充分的探索,主要是由于缺乏合适的数据集。为了解决这个问题,我们提出了通用代码混合器(UniCoM),一种新的管道,用于生成高质量,自然的CS样本,而不改变句子语义。我们的方法利用了一种算法,我们称之为用同义词替换单词(SWORDS),它通过将选定的单词替换为它们的翻译,同时考虑它们的词性来生成CS语音。使用UniCoM,我们构建了代码转换FLEURS(CS-FLEURS),一个多语言CS语料库,专为自动语音识别(ASR)和语音到文本翻译(S2 TT)。实验结果表明,CS-FLEURS实现了高清晰度和自然度,在客观和主观指标上与现有数据集的性能相当。我们希望我们的方法能够推进CS语音技术,并实现更具包容性的多语言系统。
摘要:Code-switching (CS), the alternation between two or more languages within a single speaker's utterances, is common in real-world conversations and poses significant challenges for multilingual speech technology. However, systems capable of handling this phenomenon remain underexplored, primarily due to the scarcity of suitable datasets. To resolve this issue, we propose Universal Code-Mixer (UniCoM), a novel pipeline for generating high-quality, natural CS samples without altering sentence semantics. Our approach utilizes an algorithm we call Substituting WORDs with Synonyms (SWORDS), which generates CS speech by replacing selected words with their translations while considering their parts of speech. Using UniCoM, we construct Code-Switching FLEURS (CS-FLEURS), a multilingual CS corpus designed for automatic speech recognition (ASR) and speech-to-text translation (S2TT). Experimental results show that CS-FLEURS achieves high intelligibility and naturalness, performing comparably to existing datasets on both objective and subjective metrics. We expect our approach to advance CS speech technology and enable more inclusive multilingual systems.
【8】Comparative Evaluation of Text and Audio Simplification: A Methodological Replication Study
标题:文本和音频简化的比较评估:方法学复制研究
链接:https://arxiv.org/abs/2508.15088
摘要:这项研究是Leroy等人(2022)研究的方法学复制,该研究调查了文本简化对不断发展的多媒体环境中医疗保健信息理解的影响。基于原始研究的见解,我们的复制研究评估了音频内容,认识到其在数字时代传播医疗保健信息的重要性越来越大。具体来说,我们探讨了当用户参与从该文本自动生成的音频内容时,文本简化对感知和实际困难的影响。我们的复制涉及44名参与者,我们评估了他们对使用Leroy等人(2022)原始和简化文本创建的音频形式的医疗保健信息的理解。我们的研究结果强调了文本简化在增强感知可理解性和实际理解方面的有效性,与原始研究的结果一致。此外,我们还研究了教育水平和语言能力的作用,揭示了它们对医疗保健信息获取和理解的潜在影响。这项研究强调了文本简化工具在促进健康素养方面的实用价值。它表明,需要量身定制的沟通策略,以达到不同的受众有效地在医疗保健领域。
摘要:This study serves as a methodological replication of Leroy et al. (2022) research, which investigated the impact of text simplification on healthcare information comprehension in the evolving multimedia landscape. Building upon the original studys insights, our replication study evaluates audio content, recognizing its increasing importance in disseminating healthcare information in the digital age. Specifically, we explored the influence of text simplification on perceived and actual difficulty when users engage with audio content automatically generated from that text. Our replication involved 44 participants for whom we assessed their comprehension of healthcare information presented as audio created using Leroy et al. (2022) original and simplified texts. The findings from our study highlight the effectiveness of text simplification in enhancing perceived understandability and actual comprehension, aligning with the original studys results. Additionally, we examined the role of education level and language proficiency, shedding light on their potential impact on healthcare information access and understanding. This research underscores the practical value of text simplification tools in promoting health literacy. It suggests the need for tailored communication strategies to reach diverse audiences effectively in the healthcare domain.
【9】Optimal Interference Signal for Masking an Acoustic Source
标题:用于掩蔽声学源的最佳干扰信号
链接:https://arxiv.org/abs/2508.15023
备注:40 pages, a preprint
摘要:在需要声学隐私或故意信号混淆的环境中,有必要屏蔽在基本操作中生成的声学签名。我们考虑的问题,掩蔽的声源在目标区域中的可能的检测传感器的位置的影响。通过将干扰信号放置在声源附近来实现掩蔽。我们引入了一个理论和计算框架,设计这样的干扰信号的目标是最小化目标区域中的残余振幅。对于具有球对称性的三维强迫波方程,我们导出了几种典型情形的解析拟定常周期解。我们研究了自掩蔽现象,其中具有一定空间强迫轮廓的声源掩盖了自己,使其无法在其强迫足迹之外被检测到。然后,我们使用叠加的球对称解决方案,调查在一个给定的目标区域的掩蔽。我们分析和优化的性能,使用一个或两个点的力量部署在声源附近的掩蔽在目标区域。对于声源的空间强迫分布缺乏球对称性的一般情况,我们发展了一种有效的数值方法来求解三维波动方程。这项工作的潜在应用包括海底声学通信安全,海底车辆隐身,以及对声学监视的保护。
摘要:In an environment where acoustic privacy or deliberate signal obfuscation is desired, it is necessary to mask the acoustic signature generated in essential operations. We consider the problem of masking the effect of an acoustic source in a target region where possible detection sensors are located. Masking is achieved by placing interference signals near the acoustic source. We introduce a theoretical and computational framework for designing such interference signals with the goal of minimizing the residual amplitude in the target region. For the three-dimensional (3D) forced wave equation with spherical symmetry, we derive analytical quasi-steady periodic solutions for several canonical cases. We examine the phenomenon of self-masking where an acoustic source with certain spatial forcing profile masks itself from detection outside its forcing footprint. We then use superposition of spherically symmetric solutions to investigate masking in a given target region. We analyze and optimize the performance of using one or two point-forces deployed near the acoustic source for masking in the target region. For the general case where the spatial forcing profile of the acoustic source lacks spherical symmetry, we develop an efficient numerical method for solving the 3D wave equation. Potential applications of this work include undersea acoustic communication security, undersea vehicles stealth, and protection against acoustic surveillance.
【10】XAI-Driven Spectral Analysis of Cough Sounds for Respiratory Disease Characterization
标题:XAI驱动的咳嗽声频谱分析用于呼吸系统疾病特征
链接:https://arxiv.org/abs/2508.14949
摘要:本文提出了一种可解释的人工智能(XAI)驱动的方法,以提高对呼吸系统疾病管理咳嗽声分析的理解。我们采用遮挡图来突出显示卷积神经网络(CNN)处理的咳嗽频谱图中的相关频谱区域。随后,由这些遮挡图加权的频谱图的频谱分析揭示了疾病组之间的显著差异,特别是在患有COPD的患者中,其中咳嗽模式在所识别的感兴趣的频谱区域中出现更多变化。这与分析原始光谱图时观察到的缺乏显著差异形成对比。所提出的方法提取和分析了几个频谱特征,展示了XAI技术的潜力,以发现疾病特异性的声学特征,并通过提供更多的可解释的结果来提高咳嗽声分析的诊断能力。
摘要:This paper proposes an eXplainable Artificial Intelligence (XAI)-driven methodology to enhance the understanding of cough sound analysis for respiratory disease management. We employ occlusion maps to highlight relevant spectral regions in cough spectrograms processed by a Convolutional Neural Network (CNN). Subsequently, spectral analysis of spectrograms weighted by these occlusion maps reveals significant differences between disease groups, particularly in patients with COPD, where cough patterns appear more variable in the identified spectral regions of interest. This contrasts with the lack of significant differences observed when analyzing raw spectrograms. The proposed approach extracts and analyzes several spectral features, demonstrating the potential of XAI techniques to uncover disease-specific acoustic signatures and improve the diagnostic capabilities of cough sound analysis by providing more interpretable results.
【11】Human Feedback Driven Dynamic Speech Emotion Recognition
标题:人类反馈驱动的动态语音情感识别
链接:https://arxiv.org/abs/2508.14920
摘要:本文的工作为动态语音情感识别开辟了一个新的领域。与传统方法不同,我们假设每个音轨都与在不同时刻活跃的情绪序列相关联。该研究特别关注情感3D化身的动画。我们提出了一个多阶段的方法,包括一个经典的语音情感识别模型的训练,情感序列的合成生成,并进一步改进模型的基础上,人类的反馈。此外,我们介绍了一种新的方法来建模的基础上的狄利克雷分布的情感混合物。基于从3D面部动画数据集中提取的真实情感对模型进行评估。我们将我们的模型与滑动窗口方法进行比较。我们的实验结果表明,狄利克雷为基础的方法在建模情感的混合物的有效性。简化人工反馈进一步提高了模型质量,同时提供了简化的注释过程。
摘要:This work proposes to explore a new area of dynamic speech emotion recognition. Unlike traditional methods, we assume that each audio track is associated with a sequence of emotions active at different moments in time. The study particularly focuses on the animation of emotional 3D avatars. We propose a multi-stage method that includes the training of a classical speech emotion recognition model, synthetic generation of emotional sequences, and further model improvement based on human feedback. Additionally, we introduce a novel approach to modeling emotional mixtures based on the Dirichlet distribution. The models are evaluated based on ground-truth emotions extracted from a dataset of 3D facial animations. We compare our models against the sliding window approach. Our experimental results show the effectiveness of Dirichlet-based approach in modeling emotional mixtures. Incorporating human feedback further improves the model quality while providing a simplified annotation procedure.
【12】Denoising by neural network for muzzle blast detection
标题:枪口爆炸检测的神经网络去噪
链接:https://arxiv.org/abs/2508.14919
备注:INTER-NOISE 2024, Aug 2024, Nantes (France), France
摘要:Acoem开发枪击检测系统,包括麦克风阵列和软件,可以检测和定位战场上的射手。这种系统的性能显然受到其运行的声学环境的影响:特别是,当安装在移动的军用车辆上时,噪声的存在降低了软件的检测性能。为了限制声学环境的影响,已经开发了神经网络。我们没有使用重型卷积神经网络,而是选择了轻量级神经网络架构,以限制在尽可能多的硬件平台上嵌入算法所需的计算资源。由于两个隐藏层感知器和适当的信号处理技术的结合,脉冲枪口爆炸波形(来自爆炸的波,指示射手的位置)的检测率显着增加。当噪声的均方根值与枪口冲击波峰值幅度同阶时,采用这种去噪处理,检测率提高了一倍以上。
摘要:Acoem develops gunshot detection systems, consisting of a microphone array and software that detects and locates shooters on the battlefield. The performance of such systems is obviously affected by the acoustic environment in which they are operating: in particular, when mounted on a moving military vehicle, the presence of noise reduces the detection performance of the software. To limit the influence of the acoustic environment, a neural network has been developed. Instead of using a heavy convolutional neural network, a lightweight neural network architecture was chosen to limit the computational resources required to embed the algorithm on as many hardware platforms as possible. Thanks to the combination of a two hidden layer perceptron and appropriate signal processing techniques, the detection rate of impulsive muzzle blast waveforms (the wave coming from the detonation and indicating the position of the shooter) is significantly increased. With a rms value of noise of the same order as the muzzle blast peak amplitude, the detect rate is more than doubled with this denoising processing.
机器翻译由腾讯交互翻译提供,仅供参考
