微信公众号:arXiv_Daily
cs.SD语音
标题: 来自原始波的端到端音频深度伪造检测:一种基于RawNet的方法,具有跨数据集评估
链接:https://arxiv.org/abs/2504.20923
【2】 Effect of Avatar Head Movement on Communication Behaviour, Experience of Presence and Conversation Success in Triadic Conversations
标题: 阿凡达头部运动对三重对话中沟通行为、在场体验和对话成功的影响链接:https://arxiv.org/abs/2504.20844
【3】 Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT Reasoning
标题: 通过半隐式跨语言CoT推理增强非核心语言教学-言语中的LLM链接:https://arxiv.org/abs/2504.20835
备注:10 pages, 6 figures, Submitted to ACM MM 2025
【4】 ECOSoundSet: a finely annotated dataset for the automated acoustic identification of Orthoptera and Cicadidae in North, Central and temperate Western Europe
标题: ECOSoundSet:一个经过精心注释的数据集,用于西欧北部、中部和温带直翅目和蝉科的自动声学识别链接:https://arxiv.org/abs/2504.20776
备注:3 Figures + 2 Supplementary Figures, 2 Tables + 3 Supplementary Tables
【5】 DiffusionRIR: Room Impulse Response Interpolation using Diffusion Models
标题: 扩散RIR:使用扩散模型的房间脉冲响应插值链接:https://arxiv.org/abs/2504.20625
【6】 TriniMark: A Robust Generative Speech Watermarking Method for Trinity-Level Attribution
标题: TriniMark:一种用于三位一体化属性的鲁棒生成语音水印方法链接:https://arxiv.org/abs/2504.20532
【7】 APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech
标题: APG-MOS:合成语音的听觉感知引导MOS预测器链接:https://arxiv.org/abs/2504.20447
【8】 Pediatric Asthma Detection with Googles HeAR Model: An AI-Driven Respiratory Sound Classifier
标题: 使用Googles HeAR模型检测儿科哮喘:人工智能驱动的呼吸声分类器链接:https://arxiv.org/abs/2504.20124
【9】 ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
标题: ISDrama:通过多模式投影生成沉浸式空间戏剧链接:https://arxiv.org/abs/2504.20630
标题: ISDrama:通过多模式投影生成沉浸式空间戏剧
链接:https://arxiv.org/abs/2504.20630
【2】 Towards Flow-Matching-based TTS without Classifier-Free Guidance
标题: 在没有分类器指导的情况下迈向基于流匹配的TTC链接:https://arxiv.org/abs/2504.20334
【3】 End-to-end Audio Deepfake Detection from RAW Waveforms: a RawNet-Based Approach with Cross-Dataset Evaluation
标题: 来自原始波的端到端音频深度伪造检测:一种基于RawNet的方法,具有跨数据集评估链接:https://arxiv.org/abs/2504.20923
【4】 Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT Reasoning
标题: 通过半隐式跨语言CoT推理增强非核心语言教学-言语中的LLM链接:https://arxiv.org/abs/2504.20835
备注:10 pages, 6 figures, Submitted to ACM MM 2025
【5】 ECOSoundSet: a finely annotated dataset for the automated acoustic identification of Orthoptera and Cicadidae in North, Central and temperate Western Europe
标题: ECOSoundSet:一个经过精心注释的数据集,用于西欧北部、中部和温带直翅目和蝉科的自动声学识别链接:https://arxiv.org/abs/2504.20776
备注:3 Figures + 2 Supplementary Figures, 2 Tables + 3 Supplementary Tables
【6】 Non-native Children's Automatic Speech Assessment Challenge (NOCASA)
标题: 非本地儿童自动言语评估挑战赛(NOCASA)链接:https://arxiv.org/abs/2504.20678
备注:First draft of the baseline paper for the NOCASA competition (this https URL), 5 pages
【7】 DiffusionRIR: Room Impulse Response Interpolation using Diffusion Models
标题: 扩散RIR:使用扩散模型的房间脉冲响应插值链接:https://arxiv.org/abs/2504.20625
【8】 TriniMark: A Robust Generative Speech Watermarking Method for Trinity-Level Attribution
标题: TriniMark:一种用于三位一体化属性的鲁棒生成语音水印方法链接:https://arxiv.org/abs/2504.20532
【9】 APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech
标题: APG-MOS:合成语音的听觉感知引导MOS预测器链接:https://arxiv.org/abs/2504.20447
机器翻译由腾讯交互翻译提供,仅供参考
