今日论文合集:CS.SD语音与音频 | 共 5 篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 语音识别与关键词检测 1 篇

2. 语音合成与声音生成 1 篇

3. 其他/综合语音音频 3 篇

1. 语音识别与关键词检测 | 1 篇

1. SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models

SpeechGuard:针对语音识别模型后门攻击的在线防御

AI 总结:研究针对语音识别模型运行时的后门攻击,提出SpeechGuard在线防御管道,改进STRIP方法为S-STRIP检测过滤中毒样本,利用时频掩蔽等净化样本,经实验验证其能有效减轻后门威胁并维持预测准确率。

链接:https://arxiv.org/abs/2607.15697

作者:Jinwen Xin, Xixiang Lv

英文摘要:Backdoor attacks pose a critical threat to neural network models, allowing attackers to implant a backdoor during the training phase by manipulating a small portion of the training data. In security-sensitive applications such as voice interaction for autonomous driving, the presence of backdoor attacks introduces substantial security risks. This study focuses on implementing backdoor defense measures for speech recognition models in run-time, taking into account the characteristics of audio signals. We propose SpeechGuard, the first online backdoor defense pipeline designed to identify and purify poisoned audio samples. Specifically, we improve STRIP method to perform adaptive perturbation injection to detect and filter poisoned samples, named as S-STRIP. More importantly, we further consider the purification of poisoned samples. We utilize time-frequency (T-F) masking to suppress the expression of trigger signals and autonomously generate masks based on an autoencoder. The two-stage processing prevents the backdoor in the model from being triggered, and even input speech carrying triggers can be accurately predicted. Extensive experimental demonstrate that SpeechGuard can accurately filter out poisoned samples. Through purification, it can significantly mitigate the backdoor threat while maintaining a certain prediction accuracy.

2. 语音合成与声音生成 | 1 篇

2. AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

AuEmoChat:用于对话语音合成的真实情感理解与呈现

AI 总结:针对对话语音合成中情感表达不真实及多模态令牌干扰问题,提出AuEmoChat框架,通过AuEmoCodec学习情感令牌空间、AuEmoToMe合并冗余令牌,集成到模型并结合情感流匹配渲染语音,实验证明其性能优于现有基线。

链接:https://arxiv.org/abs/2607.15755

作者:Zhenqi Jia, Yuan Zhao, Aruukhan, Rui Liu, Haizhou Li

英文摘要:Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions due to limited predefined emotion label spaces (e.g., seven emotion categories), while redundant multimodal tokens in multi-turn dialogue history interfere with context understanding. To address these issues, we propose AuEmoChat, a CSS framework for authentic emotion understanding and rendering. First, we develop AuEmoCodec, which learns a discrete authentic emotion token space from large-scale emotional speech via finite scalar quantization, enabling a more authentic emotion representation than limited basic emotion categories. We further propose AuEmoToMe, an authentic-emotion-guided token merging algorithm that merges redundant tokens in multimodal dialogue history while preserving emotion-relevant context. We integrate it into an autoregressive text-speech model to predict the target authentic emotion token and speech tokens. Finally, we propose Authentic Emotion Flow Matching, which renders speech by jointly conditioning on merged dialogue context, target authentic emotion, and acoustic priors. Extensive experiments on the NCSSD-EmCap dataset demonstrate that AuEmoChat outperforms state-of-the-art CSS baselines and generates more expressive and authentic emotional speech.

3. 其他/综合语音音频 | 3 篇

3. Estimating the Reliability of Dynamic Time Warping Alignments Using Circumstantial Evidence

使用间接证据估计动态时间规整比对的可靠性

AI 总结:研究DTW比对路径不确定性问题,提出基于间接证据的可靠性度量指标,通过FlexDTW重新估计比对并衡量路径一致性来计算指标,在音频比对任务中评估,能无监督地准确估计DTW比对路径可靠性。

链接:https://arxiv.org/abs/2607.15443

作者:Aanya Pratapneni, Alice Yuan, TJ Tsai

英文摘要:Recent works have explored ways to handle uncertainty in dynamic time warping (DTW) alignment paths through the use of differentiable variants of DTW like Soft-DTW. In this paper, we approach the issue of uncertainty in DTW alignment paths in a different way. Given a DTW alignment path, we propose a metric that indicates how reliable a local segment of the alignment path is. The intuition for our metric is based on the idea of circumstantial evidence. If DTW has found a very prominent path, then if we re-run the alignment with relaxed boundary conditions, it will still pick the same path. If, on the other hand, DTW has found a "weak" path, then re-running the alignment with relaxed boundary conditions will likely yield a different path. Accordingly, our reliability metric is computed by picking a local section of the DTW alignment path, re-estimating the alignment with FlexDTW (which allows flexibility in the boundary conditions), and then measuring how well the DTW and FlexDTW paths agree. We assess the proposed reliability metric on DTW alignment paths containing both matching and non-matching regions across a range of scenarios on an audio-audio alignment task. We find that the reliability metric correctly identifies reliable regions of the alignment path with an aggregate AUROC of 0.97. This approach provides an unsupervised method for estimating the reliability of a DTW alignment path.

4. Segmental DTW: A Parallelizable Alternative to Dynamic Time Warping

分段动态时间规整:动态时间规整的一种可并行化替代方法

AI 总结:研究探索DTW可并行化替代方法,介绍分段DTW算法变体,其将全局成本矩阵分解为子矩阵计算,在音频对齐任务中评估性能,结果接近常规DTW,且几乎所有计算可并行化,一个变体表现更优。

链接:https://arxiv.org/abs/2607.15475

作者:TJ Tsai

英文摘要:In this work we explore parallelizable alternatives to DTW for globally aligning two feature sequences. One of the main practical limitations of DTW is its quadratic computation and memory cost. Previous works have sought to reduce the computational cost in various ways, such as imposing bands in the cost matrix or using a multiresolution approach. In this work, we utilize the fact that computation is an abundant resource and focus instead on exploring alternatives that approximate the inherently sequential DTW algorithm with one that is parallelizable. We describe two variations of an algorithm called Segmental DTW, in which the global cost matrix is broken into smaller sub-matrices, subsequence DTW is performed on each sub-matrix, and the results are used to solve a segment-level dynamic programming problem that specifies a globally optimal alignment path. We evaluate the proposed alignment algorithms on an audio-audio alignment task using the Chopin Mazurka dataset, and we show that they closely match the performance of regular DTW. We further demonstrate that almost all of the computations in Segmental DTW are parallelizable, and that one of the variants is unilaterally better than the other for both empirical and theoretical reasons.

5. StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems

StemFX:通过对源分离音轨进行自回归FX链预测来学习混音风格表示

AI 总结:研究如何学习混音风格表示,提出StemFX框架,通过自回归预测源分离音轨上的FX链,Transformer解码器与多频段CNN编码器协同工作,经大规模训练后在混音风格检索和转换任务中表现出色,远超基线模型和迭代优化。

链接:https://arxiv.org/abs/2607.15634

作者:Yuan-Chiao Cheng, Jui-Te Wu, Brian Chen, Yen-Tung Yeh, Yu-Hua Chen, Yi-Hsuan Yang

英文摘要:Audio mixing style encompasses the artistic and technical decisions a mix engineer makes, including level balancing, spatialization, and the choice, ordering, and parameterization of audio effects (FX) on each stem. FX chains are a key determinant of this style, yet existing approaches to modeling them remain limited. Some operate on stereo mixtures without explicit per-stem FX chain modeling, others fix the number or type of effects per track, and many require differentiable effect implementations or scarce multitrack datasets. We present StemFX, a framework that learns mixing style representations by autoregressively predicting variable-length FX chains on source-separated stems. A Transformer decoder predicts tokenized FX chains autoregressively, while a band-split multi-band CNN encoder with FiLM conditioning captures per-stem spectral structure. To enable large-scale paired training, we extract pseudo-stems from about 105K songs via source separation and augment them using MultiAFx, a toolkit unifying 85 audio effects from 7 Python libraries. Evaluated on mixing style retrieval, StemFX outperforms all baseline models across all tested chain lengths. On paired mixing style transfer, StemFX achieves the best spectral fidelity and the highest listener preference, over 4000 times faster than iterative optimization.