今日论文合集:cs.SD语音9篇,eess.AS音频处理9篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 End-to-end Audio Deepfake Detection from RAW Waveforms: a RawNet-Based  Approach with Cross-Dataset Evaluation
标题: 来自原始波的端到端音频深度伪造检测:一种基于RawNet的方法,具有跨数据集评估
链接:https://arxiv.org/abs/2504.20923
作者: Andrea Di Pierno (1 and 2),  Luca Guarnera (2),  Dario Allegra (2),  Sebastiano Battiato (2) ((1) IMT School of Advanced Studies, Lucca, Italy, (2) Department of Mathematics and Computer Science, University of Catania, Italy) 


【2】 Effect of Avatar Head Movement on Communication Behaviour, Experience of  Presence and Conversation Success in Triadic Conversations

标题: 阿凡达头部运动对三重对话中沟通行为、在场体验和对话成功的影响
链接:https://arxiv.org/abs/2504.20844
作者: Angelika Kothe,  Volker Hohmann,  Giso Grimm 


【3】 Enhancing Non-Core Language Instruction-Following in Speech LLMs via  Semi-Implicit Cross-Lingual CoT Reasoning

标题: 通过半隐式跨语言CoT推理增强非核心语言教学-言语中的LLM
链接:https://arxiv.org/abs/2504.20835
作者: Hongfei Xue,  Yufeng Tang,  Hexin Liu,  Jun Zhang,  Xuelong Geng,  Lei Xie 
备注:10 pages, 6 figures, Submitted to ACM MM 2025


【4】 ECOSoundSet: a finely annotated dataset for the automated acoustic  identification of Orthoptera and Cicadidae in North, Central and temperate  Western Europe

标题: ECOSoundSet:一个经过精心注释的数据集,用于西欧北部、中部和温带直翅目和蝉科的自动声学识别
链接:https://arxiv.org/abs/2504.20776
作者: David Funosas,  Elodie Massol,  Yves Bas,  Svenja Schmidt,  Dominik Arend,  Alexander Gebhard,  Luc Barbaro,  Sebastian König,  Rafael Carbonell Font,  David Sannier,  Fernand Deroussen,  Jérôme Sueur,  Christian Roesti,  Tomi Trilar,  Wolfgang Forstmeier,  Lucas Roger,  Eloïsa Matheu,  Piotr Guzik,  Julien Barataud,  Laurent Pelozuelo,  Stéphane Puissant,  Sandra Mueller,  Björn Schuller,  Jose M. Montoya,  Andreas Triantafyllopoulos,  Maxime Cauchoix 
备注:3 Figures + 2 Supplementary Figures, 2 Tables + 3 Supplementary Tables


【5】 DiffusionRIR: Room Impulse Response Interpolation using Diffusion Models

标题: 扩散RIR:使用扩散模型的房间脉冲响应插值
链接:https://arxiv.org/abs/2504.20625
作者: Sagi Della Torre,  Mirco Pezzoli,  Fabio Antonacci,  Sharon Gannot 


【6】 TriniMark: A Robust Generative Speech Watermarking Method for  Trinity-Level Attribution

标题: TriniMark:一种用于三位一体化属性的鲁棒生成语音水印方法
链接:https://arxiv.org/abs/2504.20532
作者: Yue Li,  Weizhi Liu,  Dongdong Lin 


【7】 APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech

标题: APG-MOS:合成语音的听觉感知引导MOS预测器
链接:https://arxiv.org/abs/2504.20447
作者: Zhicheng Lian,  Lizhi Wang,  Hua Huang 


【8】 Pediatric Asthma Detection with Googles HeAR Model: An AI-Driven  Respiratory Sound Classifier

标题: 使用Googles HeAR模型检测儿科哮喘:人工智能驱动的呼吸声分类器
链接:https://arxiv.org/abs/2504.20124
作者: Abul Ehtesham,  Saket Kumar,  Aditi Singh,  Tala Talaei Khoei 


【9】 ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting

标题: ISDrama:通过多模式投影生成沉浸式空间戏剧
链接:https://arxiv.org/abs/2504.20630
作者: Yu Zhang,  Wenxiang Guo,  Changhao Pan,  Zhiyuan Zhu,  Tao Jin,  Zhou Zhao 


eess.AS音频处理
【1】 ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
标题: ISDrama:通过多模式投影生成沉浸式空间戏剧
链接:https://arxiv.org/abs/2504.20630
作者: Yu Zhang,  Wenxiang Guo,  Changhao Pan,  Zhiyuan Zhu,  Tao Jin,  Zhou Zhao 


【2】 Towards Flow-Matching-based TTS without Classifier-Free Guidance

标题: 在没有分类器指导的情况下迈向基于流匹配的TTC
链接:https://arxiv.org/abs/2504.20334
作者: Yuzhe Liang,  Wenzhe Liu,  Chunyu Qiang,  Zhikang Niu,  Yushen Chen,  Ziyang Ma,  Wenxi Chen,  Nan Li,  Chen Zhang,  Xie Chen 


【3】 End-to-end Audio Deepfake Detection from RAW Waveforms: a RawNet-Based  Approach with Cross-Dataset Evaluation

标题: 来自原始波的端到端音频深度伪造检测:一种基于RawNet的方法,具有跨数据集评估
链接:https://arxiv.org/abs/2504.20923
作者: Andrea Di Pierno (1 and 2),  Luca Guarnera (2),  Dario Allegra (2),  Sebastiano Battiato (2) ((1) IMT School of Advanced Studies, Lucca, Italy, (2) Department of Mathematics and Computer Science, University of Catania, Italy) 


【4】 Enhancing Non-Core Language Instruction-Following in Speech LLMs via  Semi-Implicit Cross-Lingual CoT Reasoning

标题: 通过半隐式跨语言CoT推理增强非核心语言教学-言语中的LLM
链接:https://arxiv.org/abs/2504.20835
作者: Hongfei Xue,  Yufeng Tang,  Hexin Liu,  Jun Zhang,  Xuelong Geng,  Lei Xie 
备注:10 pages, 6 figures, Submitted to ACM MM 2025


【5】 ECOSoundSet: a finely annotated dataset for the automated acoustic  identification of Orthoptera and Cicadidae in North, Central and temperate  Western Europe

标题: ECOSoundSet:一个经过精心注释的数据集,用于西欧北部、中部和温带直翅目和蝉科的自动声学识别
链接:https://arxiv.org/abs/2504.20776
作者: David Funosas,  Elodie Massol,  Yves Bas,  Svenja Schmidt,  Dominik Arend,  Alexander Gebhard,  Luc Barbaro,  Sebastian König,  Rafael Carbonell Font,  David Sannier,  Fernand Deroussen,  Jérôme Sueur,  Christian Roesti,  Tomi Trilar,  Wolfgang Forstmeier,  Lucas Roger,  Eloïsa Matheu,  Piotr Guzik,  Julien Barataud,  Laurent Pelozuelo,  Stéphane Puissant,  Sandra Mueller,  Björn Schuller,  Jose M. Montoya,  Andreas Triantafyllopoulos,  Maxime Cauchoix 
备注:3 Figures + 2 Supplementary Figures, 2 Tables + 3 Supplementary Tables


【6】 Non-native Children's Automatic Speech Assessment Challenge (NOCASA)

标题: 非本地儿童自动言语评估挑战赛(NOCASA)
链接:https://arxiv.org/abs/2504.20678
作者: Yaroslav Getman,  Tamás Grósz,  Mikko Kurimo,  Giampiero Salvi 
备注:First draft of the baseline paper for the NOCASA competition (this https URL), 5 pages


【7】 DiffusionRIR: Room Impulse Response Interpolation using Diffusion Models

标题: 扩散RIR:使用扩散模型的房间脉冲响应插值
链接:https://arxiv.org/abs/2504.20625
作者: Sagi Della Torre,  Mirco Pezzoli,  Fabio Antonacci,  Sharon Gannot 


【8】 TriniMark: A Robust Generative Speech Watermarking Method for  Trinity-Level Attribution

标题: TriniMark:一种用于三位一体化属性的鲁棒生成语音水印方法
链接:https://arxiv.org/abs/2504.20532
作者: Yue Li,  Weizhi Liu,  Dongdong Lin 


【9】 APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech

标题: APG-MOS:合成语音的听觉感知引导MOS预测器
链接:https://arxiv.org/abs/2504.20447
作者: Zhicheng Lian,  Lizhi Wang,  Hua Huang 


机器翻译由腾讯交互翻译提供,仅供参考