微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音识别与关键词检测 2 篇
2. 语音合成与声音生成 2 篇
3. 音频事件检测与场景理解 2 篇
4. 音乐信息检索与音乐生成 1 篇
5. 语音翻译与语音语言模型 1 篇
6. 安全、隐私与深度伪造音频 2 篇
7. 其他/综合语音音频 3 篇
1. 语音识别与关键词检测 | 2 篇
1. Multi-Level Privacy-Preserving Dementia Detection from Speech via Targeted Adversarial Obfuscation and Representation Learning
通过有针对性的对抗性混淆和表征学习实现基于语音的多级隐私保护痴呆症检测
AI 总结:研究针对痴呆症检测语音记录的隐私问题,提出多级框架,信号层用CSA、特征层用带MI引导噪声注入的GRL,在痴呆症银行匹兹堡语料库上评估,兼顾隐私保护与痴呆症分类性能。
链接:https://arxiv.org/abs/2607.17098
机构:University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校)
作者:Henriette Flore Kenne, Raphael Anaadumba, Mohammad Arif Ul Alam
英文摘要:Speech recordings used for dementia detection inherently expose speaker identity, raising critical privacy concerns. Existing methods typically address only singular threats and fail to resolve the privacy--utility trade-off. We propose a multi-level framework designed to neutralize two distinct eavesdropping vectors. At the signal level, a Cumulative Signal Attack (CSA) concentrates perturbations in keyword-aligned regions to maximize transcription error (Word Error Rate WER = 1.00) while preserving vital prosodic biomarkers. At the feature level, a Gradient Reversal Layer (GRL) with Mutual Information (MI)-guided noise injection suppresses speaker-discriminative dimensions while retaining dementia-relevant diagnostic structure. Evaluated on the DementiaBank Pitt Corpus, our framework achieves near-chance speaker identification (Equal Error Rate EER = 0.59, F1 = 0.003) while maintaining strong dementia classification performance (F1 = 0.78, AUC = 0.86).
2. Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture
共振:基于三级级联ASR-LLM-TTS架构的构音障碍异步实时语音转换系统
AI 总结:针对构音障碍患者在专业演讲场景的沟通问题,提出基于三级级联ASR-LLM-TTS架构的Re-Sonance系统,集成多种技术,经评估其能提高轻度至中度构音障碍患者语音清晰度与自然度,验证了基于LLM方法增强语音驱动AAC系统的潜力。
链接:https://arxiv.org/abs/2607.17615
机构:Southeast University(东南大学); School of Biological Science & Medical Engineering, Southeast University(东南大学生物科学与医学工程学院)
作者:Yuxuan Wu, Yifan Xu, Junkun Wang, Jiayong Jiang, Xin Zhao, Zhaojie Luo
英文摘要:Individuals with dysarthria face significant challenges in professional speaking scenarios such as conferences, presentations, and meetings, where real-time communication is crucial. While existing Augmentative and Alternative Communication (AAC) systems provide basic support, they often fail to meet the demands of professional speaking environments due to high latency and unnatural speech patterns. This paper presents Re-Sonance, a novel LLM-enhanced speech-driven AAC system designed for real-time professional speaking scenarios. By integrating Whisper ASR, Qwen LLM, and CosyVoice TTS, Re-Sonance achieves improved speech intelligibility and naturalness while maintaining real-time performance. Both subjective and objective evaluations using a Mandarin dysarthric speech dataset demonstrate that our speech reconstruction approach significantly improved intelligibility while preserving semantic coherence for speakers with mild to moderate dysarthria. Although performance remains limited for severe dysarthria cases, our findings validate the potential of LLM-based methods for enhancing speech-driven AAC systems, paving the way for more effective and accessible communication technologies.
2. 语音合成与声音生成 | 2 篇
3. HARP: Harmonic-Aware Residual Partitioning for Neural Audio Codecs
HARP:用于神经音频编解码器的谐波感知残差划分
AI 总结:研究针对神经音频编解码器中码本频谱纠缠等问题,提出HARP训练策略,将RVQ阶段分组,在解码器能访问低频时各小组细化目标频带,重建泛音保留连贯性,该策略无需架构改变,性能优于标准RVQ和并行分解。
链接:https://arxiv.org/abs/2607.16657
机构:Georgia Institute of Technology(佐治亚理工学院); The Chinese University of Hong Kong(香港中文大学); Tencent Music Entertainment(腾讯音乐娱乐集团)
作者:Qiaoyu Yang, Lixing He, Binyue Deng, Weifeng Zhao
英文摘要:Neural audio codecs with residual vector quantization (RVQ) normally treat all frequencies uniformly, so their codebooks become spectrally entangled. Truncating stages then removes an unpredictable mix of frequencies. Parallel band decomposition addresses this by splitting audio into independent bands, but fragments the latent space and loses cross-frequency coherence. We introduce HARP (Harmonic-Aware Residual Partitioning), a training strategy that partitions RVQ stages into frequency-ordered groups where each group refines its target band while the decoder retains access to all lower frequencies. Overtones are reconstructed in the context of their fundamentals, preserving coherence that parallel methods lose. HARP requires no architectural changes; it only modifies the training loss, leaving inference identical to standard RVQ. On speech, music, and general audio, HARP outperforms both standard RVQ and parallel decomposition. MUSHRA listening tests also show perceptual improvements.
4. Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer
利用TTS:借助利用层实现上下文感知的富有表现力的语音合成
AI 总结:研究针对语音助手富有表现力的语音合成中灵活风格控制的需求,提出Harness TTS轻量级控制层,通过封闭集提示 - 工具路由实现风格控制,实验证明其在路由和合成任务中表现出色,为语音助手表达控制提供实用方案。
链接:https://arxiv.org/abs/2607.17900
机构:Xiaomi Inc.(小米公司); Nanjing University(南京大学)
作者:Shengfan Shen, Di Wu, Xingchen Song, Dinghao Zhou, Pengyu Cheng, Sixiang Lyu, Jian Luan, Shuai Wang
英文摘要:Expressive speech synthesis for voice assistants requires flexible style control that adapts to explicit requests and broader interaction context. We propose Harness TTS, a lightweight control layer that wraps around a TTS engine to externalize and govern its expressive behavior. It reformulates style control as closed-set prompt-tool routing: offline, a compact registry of stylistic prompt tools is constructed with structured metadata; online, an LLM planner selects the appropriate tool based on a priority-aware observation schema, and the TTS executor synthesizes speech using the corresponding prompt audio. We evaluate Harness TTS on both routing and synthesis tasks. In routing, Qwen3-4B achieves Top-1 accuracies of 74.3%, 43.0%, and 64.6% on explicit, implicit, and conflict subsets. For synthesis, experiments on CosyVoice3 and VoxCPM2 show that Harness TTS outperforms instruction-only control, achieving higher instruction-following win rates (margins of 23.1-35.6 points on CosyVoice3 and 13.8-20.0 points on VoxCPM2) and improving UTMOSv2 scores by 0.11-0.38. Moreover, the 4B planner delivers its first tool recommendation in under 50 ms in standard mode, introducing negligible latency for real-time interaction. These results demonstrate that equipping TTS engines with a dedicated Harness layer offers a practical, auditable, and context-aware solution for voice assistant expression control.
3. 音频事件检测与场景理解 | 2 篇
5. Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection
用于语音和环境声音深度伪造检测的组件级集成融合
AI 总结:针对ICME 2026 ESDD2挑战赛的音频五类分类任务,提出基于四个预训练模型的组件级集成系统,经微调、训练增强变体及融合选定检查点等操作,取得较好成绩,在31个团队中排第5,超越官方基线。
链接:https://arxiv.org/abs/2607.16369
作者:André Runewicz, Karla Schäfer, Martin Steinebach
英文摘要:This paper describes our submission to the ICME 2026 ESDD2 challenge on environment-aware speech and sound deepfake detection. The task requires five-class classification of audio clips in which speech, environmental sound, both components, or neither component may be spoofed. We propose a component-level ensemble system based on four publicly available pre-trained anti-spoofing models: XLSR-Mamba, DF-Arena, SLS, and TCM-ADD. Each model is fine-tuned on the official CompSpoofV2 development data using three binary heads for original, speech, and environmental sound detection. We further train RawBoost-augmented variants and combine selected checkpoints using margin-space score fusion. A component-wise fusion strategy with lightweight head- and class-bias calibration yields our best configuration, reaching 0.7715 macro-F1 on the evaluation set and 0.7828 macro-F1 on the test set, ranking 5th out of 31 teams in the final ranking phase and substantially outperforming the official baseline.
6. Explainable Lightweight Compact Deep Models for Speech Emotion Recognition
用于语音情感识别的可解释轻量级紧凑深度模型
AI 总结:针对语音情感识别中现有方法存在的问题,提出基于紧凑卷积神经网络架构的可解释轻量级框架,利用对数梅尔频谱图和注意力统计池化,结合基于梯度的类激活映射,在SAVEE数据集上实验,平衡了识别精度、计算效率和模型透明度。
链接:https://arxiv.org/abs/2607.16803
作者:Nelly Elsayed
英文摘要:Speech Emotion Recognition (SER) is an important component in a wide range of human-centered applications, including healthcare, customer service, and human-omputer interaction. In medical and decision-support settings, there is increasing interest in models that not only achieve accurate emotion recognition but also support transparent predictions and efficient deployment. However, many existing SER approaches rely on complex deep learning architectures that limit interpretability and increase computational cost. This paper presents an explainable and lightweight speech emotion recognition framework based on a compact convolutional neural network architecture. The proposed approach utilizes log-Mel spectrogram representations to capture spectro-temporal speech characteristics and employs attentive statistics pooling to emphasize emotionally salient temporal segments. To improve model transparency, gradient-based class activation mapping (Grad-CAM) is incorporated to visualize the time-frequency regions that influence the model's predictions. Experimental evaluation on the SAVEE emotional speech dataset demonstrates that the proposed framework achieves competitive recognition performance while maintaining a compact architecture with significantly fewer parameters than many existing SER models. The results indicate that efficient convolutional architectures combined with interpretable analysis can provide a practical balance between recognition accuracy, computational efficiency, and model transparency.
4. 音乐信息检索与音乐生成 | 1 篇
7. FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration
FlowSonic:通过高阶轨迹积分实现稳定的零样本音乐编辑
AI 总结:研究针对真实音乐录音零样本编辑难题,提出FlowSonic框架,基于预训练扩散变压器,通过重用交叉注意力表示保留结构,引入高阶ODE求解器,实验表明其在多方面优于现有方法,能提高潜在轨迹稳定性实现可靠编辑。
链接:https://arxiv.org/abs/2607.17526
5. 语音翻译与语音语言模型 | 1 篇
8. Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
语音令牌会泄露声纹吗?针对端到端语音语言模型的说话人反转攻击
AI 总结:研究端到端语音语言模型中语音令牌是否泄露声纹,提出用Audio BERT构建令牌嵌入及SpInv两阶段反转方法,在VoxCeleb数据集实验显示,仅三秒前端输出,SpInv在指定空间能达高于0.70的余弦相似度。
链接:https://arxiv.org/abs/2607.16870
6. 安全、隐私与深度伪造音频 | 2 篇
9. SSTMark: Robust Training-Free Semantic-Level Speech Watermarking
SSTMark:鲁棒的免训练语义级语音水印
AI 总结:针对合成语音的相关问题,提出免训练的SSTMark语音水印框架,通过文本水印在语义级别操作,将水印信息编码到语音语义内容中并检测。实验表明其平均鲁棒性最强,相比基线在不同编辑下检测率有显著提升。
链接:https://arxiv.org/abs/2607.17592
10. Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
用于鲁棒语音深度伪造检测的时频一致性学习
AI 总结:研究语音深度伪造检测在实际部署中因声学前端处理管道失真导致鲁棒性不足的问题,提出时频一致性学习框架,通过注意力驱动软对齐机制和频域结构一致性约束,提升检测模型在实际场景中的鲁棒性。
链接:https://arxiv.org/abs/2607.17761
7. 其他/综合语音音频 | 3 篇
11. Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves
一个分数够吗?用时间分数曲线评估歌曲的演唱质量
AI 总结:研究针对全长歌曲演唱质量评估难题,提出SongSQA两阶段框架。第一阶段用伪标签训练段分数预测器,第二阶段聚合器整合特征与分数生成嵌入并捕捉联系,有效提升评估效果,在KTAU上相对提高13.95%。
链接:https://arxiv.org/abs/2607.16599
12. Dense-Sparse Dynamic Time Warping for Customizing Piano Concerto Accompaniments
用于定制钢琴协奏曲伴奏的密集-稀疏动态时间规整
AI 总结:研究钢琴家定制MMO协奏曲伴奏的方法,提出密集-稀疏DTW,通过使用三种音频数据,解决频谱不匹配问题,建立数据收集与评估框架,实验表明该方法在定制伴奏录音上比复杂方法性能更好或相当。
链接:https://arxiv.org/abs/2607.18189
13. Audio Cross Verification Using Dual Alignment Likelihood Ratio Test
使用双对齐似然比检验的音频交叉验证
AI 总结:研究针对新闻短视频音频的交叉验证方法,通过定义假设、计算对齐并进行似然比检验来实现,该方法计算快、更稳健且具可解释性,为音频篡改检测提供新视角与补充工具。
链接:https://arxiv.org/abs/2607.18190
