本文经arXiv每日学术速递授权转载
【1】 Learning Spatially-Aware Language and Audio Embedding
标题: 学习空间感知语言和音频嵌入
作者:Bhavika Devnani,Skyler Seto,Zakaria Aldeneh,Alessandro Toso,Elena Menyaylenko,Barry-John Theobald,Jonathan Sheaffer,Miguel Sarabia
备注:25 pages, 7 figures
链接:点击下载PDF文件
【2】 LC-Protonets: Multi-label Few-shot learning for world music audio tagging
标题: LC-Protonets:用于世界音乐音频标记的多标签Few-Shot学习
作者:Charilaos Papaioannou,Emmanouil Benetos,Alexandros Potamianos
链接:点击下载PDF文件
【3】 The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
标题: 家庭之声:用于声音事件检测的语音删除住宅音频数据集
作者:Gabriel Bibbó,Thomas Deacon,Arshdeep Singh,Mark D. Plumbley
链接:点击下载PDF文件
【4】 WER We Stand: Benchmarking Urdu ASR Models
标题: WER我们的立场:乌尔都语ASB模型基准
作者:Samee Arif,Aamina Jamal Khan,Mustafa Abbas,Agha Ali Raza,Awais Athar
链接:点击下载PDF文件
【5】 Spontaneous Informal Speech Dataset for Punctuation Restoration
标题: 用于标点符号恢复的自发非正式语音数据集
作者:Xing Yi Liu,Homayoon Beigi
Journal-ref:Recognition Technologies, Inc. Technical Report, 2024
链接:点击下载PDF文件
【6】 Learning Source Disentanglement in Neural Audio Codec
标题: 神经音频编解码器中的学习源解纠缠
作者:Xiaoyu Bie,Xubo Liu,Gaël Richard
备注:project page: this https URL
链接:点击下载PDF文件
【7】 High-Resolution Speech Restoration with Latent Diffusion Model
标题: 基于潜在扩散模型的高分辨率语音恢复
作者:Tushar Dhyani,Florian Lux,Michele Mancusi,Giorgio Fabbro,Fritz Hohl,Ngoc Thang Vu
链接:点击下载PDF文件
【8】 Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
标题: 具有掩蔽音频令牌建模和语义知识提炼的单级TTC
作者:Gerard I. Gállego,Roy Fejgin,Chunghsin Yeh,Xiaoyu Liu,Gautam Bhattacharya
备注:Demo page: see this https URL
链接:点击下载PDF文件
【9】 Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models
标题: 增强音频语言模型的低资源语言和教学遵循能力
作者:Potsawee Manakul,Guangzhi Sun,Warit Sirichotedumrong,Kasima Tharnpipitchai,Kunat Pipatanakul
备注:5 pages. Preprint under review
链接:点击下载PDF文件
【10】 Adaptive Large Language Models By Layerwise Attention Shortcuts
标题: 通过分层注意力快捷方式自适应大型语言模型
作者:Prateek Verma,Mert Pilanci
备注:6 pages, 3 figures
链接:点击下载PDF文件
【11】 Speech Recognition for Analysis of Police Radio Communication
标题: 用于警用无线电通信分析的语音识别
作者:Tejes Srivastava,Ju-Chieh Chou,Priyank Shroff,Karen Livescu,Christopher Graziul
备注:Accepted by SLT 2024
链接:点击下载PDF文件
【12】 3DFacePolicy: Speech-Driven 3D Facial Animation with Diffusion Policy
标题: 3DFacePolicy:具有扩散策略的语音驱动3D面部动画
作者:Xuanmeng Sha,Liyun Zhang,Tomohiro Mashita,Yuki Uranishi
链接:点击下载PDF文件
【13】 PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
标题: PMX:用于符号音乐处理的大规模公共领域音乐ML数据集
作者:Phillip Long,Zachary Novack,Taylor Berg-Kirkpatrick,Julian McAuley
链接:点击下载PDF文件
【14】 Mitigating Sex Bias in Audio Data-driven COPD and COVID-19 Breathing Pattern Detection Models
标题: 缓解音频数据驱动的COPD和COVID-19呼吸模式检测模型中的性别偏见
作者:Rachel Pfeifer,Sudip Vhaduri,James Eric Dietz
备注:Accepted at 2024 IEEE-EMBS International Conference on Body Sensor Networks (IEEE BSN 2024)
链接:点击下载PDF文件
【15】 Learning Co-Speech Gesture Representations in Dialogue through Contrastive Learning: An Intrinsic Evaluation
标题: 通过对比学习学习对话中的同语手势表示:内在评价
作者:Esam Ghaleb,Bulat Khaertdinov,Wim Pouw,Marlou Rasenberg,Judith Holler,Aslı Özyürek,Raquel Fernández
Journal-ref:INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION (ICMI 2024)
链接:点击下载PDF文件
【16】 Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text
标题: Ideal-LLM:集成双编码器和扩展自适应LLM,用于多语言语音到文本
作者:Hongfei Xue,Wei Ren,Xuelong Geng,Kun Wei,Longhao Li,Qijie Shao,Linju Yang,Kai Diao,Lei Xie
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【17】 Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
标题: Zero-Shot文本到语音增强,用于低资源语音库上的自动语音识别
作者:Francesco Nespoli,Daniel Barreda,Patrick A. Naylor
备注:Accepted to the Asilomar 2023 Conference
链接:点击下载PDF文件
【18】 An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech Characterization
标题: 欺骗语音特征的可解释概率属性嵌入方法
作者:Manasi Chhibber,Jagabandhu Mishra,Hyejin Shim,Tomi H. Kinnunen
备注:Submitted to ICASSP-2025
链接:点击下载PDF文件
【19】 SynthSOD: Developing an Heterogeneous Dataset for Orchestra Music Source Separation
标题: SynthSOC:开发用于管弦乐队音乐源分离的异类数据集
作者:Jaime Garcia-Martinez,David Diaz-Guerra,Archontis Politis,Tuomas Virtanen,Julio J. Carabias-Orti,Pedro Vera-Candeas
备注:Submitted to the OJSP - ICASSP 2025
链接:点击下载PDF文件
【20】 Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
标题: 通过具有引导数据选择的语音到语音翻译来改善资源不足的语言中的语音情感识别
作者:Hsi-Che Lin,Yi-Cheng Lin,Huang-Cheng Chou,Hung-yi Lee
备注:5 pages, 2 figures, Submitted to ICASSP 2025
链接:点击下载PDF文件
【21】 Enhancing Multilingual Speech Generation and Recognition Abilities in LLMs with Constructed Code-switched Data
标题: 利用构造代码交换数据增强LLM中的多语言语音生成和识别能力
作者:Jing Xu,Daxin Tan,Jiaqi Wang,Xiao Chen
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【22】 EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
标题: EzAudio:通过高效的扩散Transformer增强文本到音频的生成
作者:Jiarui Hai,Yong Xu,Hao Zhang,Chenxing Li,Helin Wang,Mounya Elhilali,Dong Yu
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【23】 Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
标题: Speaker-IPL:使用基于i-Vector的伪标签进行说话者特征的无监督学习
作者:Zakaria Aldeneh,Takuya Higuchi,Jee-weon Jung,Li-Wei Chen,Stephen Shum,Ahmed Hussen Abdelaziz,Shinji Watanabe,Tatiana Likhomanenko,Barry-John Theobald
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【24】 Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
标题: 语音基础模型掩蔽预训练中的预测目标探索
作者:Li-Wei Chen,Takuya Higuchi,He Bai,Ahmed Hussen Abdelaziz,Alexander Rudnicky,Shinji Watanabe,Tatiana Likhomanenko,Barry-John Theobald,Zakaria Aldeneh
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【25】 Towards Automatic Assessment of Self-Supervised Speech Models using Rank
标题: 使用排名自动评估自我监督语音模型
作者:Zakaria Aldeneh,Vimal Thilak,Takuya Higuchi,Barry-John Theobald,Tatiana Likhomanenko
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【26】 Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
标题: 刺激模式很重要:不同模式的感知评估对语音情感识别系统性能的影响
作者:Huang-Cheng Chou,Haibin Wu,Chi-Chun Lee
备注:5 pages, 2 figures, 4 tables, submission for ICASSP 2025
链接:点击下载PDF文件
【27】 Investigating Training Objectives for Generative Speech Enhancement
标题: 研究生成性语音增强的训练目标
作者:Julius Richter,Danilo de Oliveira,Timo Gerkmann
链接:点击下载PDF文件
【28】 Self-supervised Speech Models for Word-Level Stuttered Speech Detection
标题: 用于词级口吃语音检测的自我监督语音模型
作者:Yi-Jen Shih,Zoi Gkalitsiou,Alexandros G. Dimakis,David Harwath
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
【29】 Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
标题: 使用视觉变形器的人机交互中的个性化语音情感识别
作者:Ruchik Mishra,Andrew Frye,Madan Mohan Rayguru,Dan O. Popa
备注:Will be submitted to IEEE for possible publication
链接:点击下载PDF文件
【30】 FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
标题: FakeMusicCaps:用于检测和归因通过文本到音乐模型生成的合成音乐的数据集
作者:Luca Comanducci,Paolo Bestagini,Stefano Tubaro
链接:点击下载PDF文件
【31】 A Real-Time Platform for Portable and Scalable Active Noise Mitigation for Construction Machinery
标题: 便携式、可扩展的工程机械主动降噪实时平台
作者:Woon-Seng Gan,Santi Peksi,Chung Kwan Lai,Yen Theng Lee,Dongyuan Shi,Bhan Lam
Journal-ref:2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
链接:点击下载PDF文件
标题: Ideal-LLM:集成双编码器和扩展自适应LLM,用于多语言语音到文本
作者:Hongfei Xue,Wei Ren,Xuelong Geng,Kun Wei,Longhao Li,Qijie Shao,Linju Yang,Kai Diao,Lei Xie
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【2】 Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
标题: Zero-Shot文本到语音增强,用于低资源语音库上的自动语音识别
作者:Francesco Nespoli,Daniel Barreda,Patrick A. Naylor
备注:Accepted to the Asilomar 2023 Conference
链接:点击下载PDF文件
【3】 An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech Characterization
标题: 欺骗语音特征的可解释概率属性嵌入方法
作者:Manasi Chhibber,Jagabandhu Mishra,Hyejin Shim,Tomi H. Kinnunen
备注:Submitted to ICASSP-2025
链接:点击下载PDF文件
【4】 SynthSOD: Developing an Heterogeneous Dataset for Orchestra Music Source Separation
标题: SynthSOC:开发用于管弦乐队音乐源分离的异类数据集
作者:Jaime Garcia-Martinez,David Diaz-Guerra,Archontis Politis,Tuomas Virtanen,Julio J. Carabias-Orti,Pedro Vera-Candeas
备注:Submitted to the OJSP - ICASSP 2025
链接:点击下载PDF文件
【5】 Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
标题: 通过具有引导数据选择的语音到语音翻译来改善资源不足的语言中的语音情感识别
作者:Hsi-Che Lin,Yi-Cheng Lin,Huang-Cheng Chou,Hung-yi Lee
备注:5 pages, 2 figures, Submitted to ICASSP 2025
链接:点击下载PDF文件
【6】 Enhancing Multilingual Speech Generation and Recognition Abilities in LLMs with Constructed Code-switched Data
标题: 利用构造代码交换数据增强LLM中的多语言语音生成和识别能力
作者:Jing Xu,Daxin Tan,Jiaqi Wang,Xiao Chen
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【7】 EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
标题: EzAudio:通过高效的扩散Transformer增强文本到音频的生成
作者:Jiarui Hai,Yong Xu,Hao Zhang,Chenxing Li,Helin Wang,Mounya Elhilali,Dong Yu
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【8】 Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
标题: Speaker-IPL:使用基于i-Vector的伪标签进行说话者特征的无监督学习
作者:Zakaria Aldeneh,Takuya Higuchi,Jee-weon Jung,Li-Wei Chen,Stephen Shum,Ahmed Hussen Abdelaziz,Shinji Watanabe,Tatiana Likhomanenko,Barry-John Theobald
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【9】 Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
标题: 语音基础模型掩蔽预训练中的预测目标探索
作者:Li-Wei Chen,Takuya Higuchi,He Bai,Ahmed Hussen Abdelaziz,Alexander Rudnicky,Shinji Watanabe,Tatiana Likhomanenko,Barry-John Theobald,Zakaria Aldeneh
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【10】 Towards Automatic Assessment of Self-Supervised Speech Models using Rank
标题: 使用排名自动评估自我监督语音模型
作者:Zakaria Aldeneh,Vimal Thilak,Takuya Higuchi,Barry-John Theobald,Tatiana Likhomanenko
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【11】 Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
标题: 刺激模式很重要:不同模式的感知评估对语音情感识别系统性能的影响
作者:Huang-Cheng Chou,Haibin Wu,Chi-Chun Lee
备注:5 pages, 2 figures, 4 tables, submission for ICASSP 2025
链接:点击下载PDF文件
【12】 Investigating Training Objectives for Generative Speech Enhancement
标题: 研究生成性语音增强的训练目标
作者:Julius Richter,Danilo de Oliveira,Timo Gerkmann
链接:点击下载PDF文件
【13】 Self-supervised Speech Models for Word-Level Stuttered Speech Detection
标题: 用于词级口吃语音检测的自我监督语音模型
作者:Yi-Jen Shih,Zoi Gkalitsiou,Alexandros G. Dimakis,David Harwath
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
【14】 Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
标题: 使用视觉变形器的人机交互中的个性化语音情感识别
作者:Ruchik Mishra,Andrew Frye,Madan Mohan Rayguru,Dan O. Popa
备注:Will be submitted to IEEE for possible publication
链接:点击下载PDF文件
【15】 FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
标题: FakeMusicCaps:用于检测和归因通过文本到音乐模型生成的合成音乐的数据集
作者:Luca Comanducci,Paolo Bestagini,Stefano Tubaro
链接:点击下载PDF文件
【16】 A Real-Time Platform for Portable and Scalable Active Noise Mitigation for Construction Machinery
标题: 便携式、可扩展的工程机械主动降噪实时平台
作者:Woon-Seng Gan,Santi Peksi,Chung Kwan Lai,Yen Theng Lee,Dongyuan Shi,Bhan Lam
Journal-ref:2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
链接:点击下载PDF文件
【17】 Learning Spatially-Aware Language and Audio Embedding
标题: 学习空间感知语言和音频嵌入
作者:Bhavika Devnani,Skyler Seto,Zakaria Aldeneh,Alessandro Toso,Elena Menyaylenko,Barry-John Theobald,Jonathan Sheaffer,Miguel Sarabia
备注:25 pages, 7 figures
链接:点击下载PDF文件
【18】 LC-Protonets: Multi-label Few-shot learning for world music audio tagging
标题: LC-Protonets:用于世界音乐音频标记的多标签Few-Shot学习
作者:Charilaos Papaioannou,Emmanouil Benetos,Alexandros Potamianos
链接:点击下载PDF文件
【19】 The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
标题: 家庭之声:用于声音事件检测的语音删除住宅音频数据集
作者:Gabriel Bibbó,Thomas Deacon,Arshdeep Singh,Mark D. Plumbley
链接:点击下载PDF文件
【20】 WER We Stand: Benchmarking Urdu ASR Models
标题: WER我们的立场:乌尔都语ASB模型基准
作者:Samee Arif,Aamina Jamal Khan,Mustafa Abbas,Agha Ali Raza,Awais Athar
链接:点击下载PDF文件
【21】 Spontaneous Informal Speech Dataset for Punctuation Restoration
标题: 用于标点符号恢复的自发非正式语音数据集
作者:Xing Yi Liu,Homayoon Beigi
Journal-ref:Recognition Technologies, Inc. Technical Report, 2024
链接:点击下载PDF文件
【22】 Learning Source Disentanglement in Neural Audio Codec
标题: 神经音频编解码器中的学习源解纠缠
作者:Xiaoyu Bie,Xubo Liu,Gaël Richard
备注:project page: this https URL
链接:点击下载PDF文件
【23】 High-Resolution Speech Restoration with Latent Diffusion Model
标题: 基于潜在扩散模型的高分辨率语音恢复
作者:Tushar Dhyani,Florian Lux,Michele Mancusi,Giorgio Fabbro,Fritz Hohl,Ngoc Thang Vu
链接:点击下载PDF文件
【24】 Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
标题: 具有掩蔽音频令牌建模和语义知识提炼的单级TTC
作者:Gerard I. Gállego,Roy Fejgin,Chunghsin Yeh,Xiaoyu Liu,Gautam Bhattacharya
备注:Demo page: see this https URL
链接:点击下载PDF文件
【25】 Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models
标题: 增强音频语言模型的低资源语言和教学遵循能力
作者:Potsawee Manakul,Guangzhi Sun,Warit Sirichotedumrong,Kasima Tharnpipitchai,Kunat Pipatanakul
备注:5 pages. Preprint under review
链接:点击下载PDF文件
【26】 Adaptive Large Language Models By Layerwise Attention Shortcuts
标题: 通过分层注意力快捷方式自适应大型语言模型
作者:Prateek Verma,Mert Pilanci
备注:6 pages, 3 figures
链接:点击下载PDF文件
【27】 Speech Recognition for Analysis of Police Radio Communication
标题: 用于警用无线电通信分析的语音识别
作者:Tejes Srivastava,Ju-Chieh Chou,Priyank Shroff,Karen Livescu,Christopher Graziul
备注:Accepted by SLT 2024
链接:点击下载PDF文件
【28】 3DFacePolicy: Speech-Driven 3D Facial Animation with Diffusion Policy
标题: 3DFacePolicy:具有扩散策略的语音驱动3D面部动画
作者:Xuanmeng Sha,Liyun Zhang,Tomohiro Mashita,Yuki Uranishi
链接:点击下载PDF文件
【29】 PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
标题: PMX:用于符号音乐处理的大规模公共领域音乐ML数据集
作者:Phillip Long,Zachary Novack,Taylor Berg-Kirkpatrick,Julian McAuley
链接:点击下载PDF文件
【30】 Mitigating Sex Bias in Audio Data-driven COPD and COVID-19 Breathing Pattern Detection Models
标题: 缓解音频数据驱动的COPD和COVID-19呼吸模式检测模型中的性别偏见
作者:Rachel Pfeifer,Sudip Vhaduri,James Eric Dietz
备注:Accepted at 2024 IEEE-EMBS International Conference on Body Sensor Networks (IEEE BSN 2024)
链接:点击下载PDF文件
【31】 Learning Co-Speech Gesture Representations in Dialogue through Contrastive Learning: An Intrinsic Evaluation
标题: 通过对比学习学习对话中的同语手势表示:内在评价
作者:Esam Ghaleb,Bulat Khaertdinov,Wim Pouw,Marlou Rasenberg,Judith Holler,Aslı Özyürek,Raquel Fernández
Journal-ref:INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION (ICMI 2024)
链接:点击下载PDF文件
标题: 学习空间感知语言和音频嵌入
作者:Bhavika Devnani,Skyler Seto,Zakaria Aldeneh,Alessandro Toso,Elena Menyaylenko,Barry-John Theobald,Jonathan Sheaffer,Miguel Sarabia
备注:25 pages, 7 figures
链接:点击下载PDF文件
摘要:人类可以通过不精确的自然语言描述来描绘声音场景。例如,很容易想象一个声学环境,给出一个短语,如“狮子的咆哮声来自我身后!".为了让机器具有相同程度的理解能力,机器必须知道狮子是什么(语义属性),“后面”的概念是什么(空间属性),以及这些语言信息如何与声音的语义和空间属性相匹配(当它从后面传来时,咆哮声听起来像什么)。最先进的音频基础模型学习在音频场景和自然文本描述之间进行映射,这些模型是在非空间音频和文本对上训练的,因此缺乏空间意识。相比之下,声音事件定位和检测模型限于从固定数量的类别中识别声音,并且它们将源定位到绝对位置(例如,0.2m)而不是使用自然语言描述的位置(例如,“在我旁边”)。为了解决这些差距,我们提出了ELSA空间感知音频和文本嵌入模型训练使用多模态对比学习。ELSA支持非空间音频,空间音频和开放词汇文本标题,描述声音的空间和语义成分。训练ELSA:(a)我们在空间上增强了三个开源音频数据集的音频和字幕,总计4,738小时的音频,以及(b)我们设计了一个编码器来捕获非空间音频的语义,以及使用对比学习的空间音频的语义和空间属性。ELSA在语义检索和3D源定位方面具有竞争力。特别地,ELSA实现了高于基线的+2.8%的平均音频到文本和文本到音频R@1 ,并且在3D源定位中优于基线的-11.6{ deg}平均绝对误差。摘要:Humans can picture a sound scene given an imprecise natural language description. For example, it is easy to imagine an acoustic environment given a phrase like "the lion roar came from right behind me!". For a machine to have the same degree of comprehension, the machine must know what a lion is (semantic attribute), what the concept of "behind" is (spatial attribute) and how these pieces of linguistic information align with the semantic and spatial attributes of the sound (what a roar sounds like when its coming from behind). State-of-the-art audio foundation models which learn to map between audio scenes and natural textual descriptions, are trained on non-spatial audio and text pairs, and hence lack spatial awareness. In contrast, sound event localization and detection models are limited to recognizing sounds from a fixed number of classes, and they localize the source to absolute position (e.g., 0.2m) rather than a position described using natural language (e.g., "next to me"). To address these gaps, we present ELSA a spatially aware-audio and text embedding model trained using multimodal contrastive learning. ELSA supports non-spatial audio, spatial audio, and open vocabulary text captions describing both the spatial and semantic components of sound. To train ELSA: (a) we spatially augment the audio and captions of three open-source audio datasets totaling 4,738 hours of audio, and (b) we design an encoder to capture the semantics of non-spatial audio, and the semantics and spatial attributes of spatial audio using contrastive learning. ELSA is competitive with state-of-the-art for both semantic retrieval and 3D source localization. In particular, ELSA achieves +2.8% mean audio-to-text and text-to-audio R@1 above the baseline, and outperforms by -11.6{ deg} mean-absolute-error in 3D source localization over the baseline.
【2】 LC-Protonets: Multi-label Few-shot learning for world music audio tagging
标题: LC-Protonets:用于世界音乐音频标记的多标签Few-Shot学习
作者:Charilaos Papaioannou,Emmanouil Benetos,Alexandros Potamianos
链接:点击下载PDF文件
摘要:我们引入标签组合原型网络(LC-Protonets)来解决多标签Few-Shot分类问题,其中模型必须仅基于少数可用示例推广到新的类别。扩展原型网络,LC-Protonets每个标签组合生成一个原型,来自有限训练项中存在的标签的幂集,而不是每个标签一个原型。我们的方法应用于跨不同音乐数据集的自动音频标记,涵盖各种文化,包括现代和传统音乐,并针对文献中的现有方法进行评估。结果表明,当使用LC-Protonets进行多标签分类时,几乎所有领域和训练设置都有显着的性能改善。除了从头开始训练Few-Shot学习模型之外,我们还探索使用通过监督学习获得的预训练模型将项嵌入特征空间中。微调提高了所有方法的泛化能力,但LC-质子网络实现高水平的性能,即使没有微调,相比之下,比较方法。最后,我们分析了所提出的方法的可扩展性,从我们的实验提供详细的定量指标。实施和实验设置公开提供,为未来的研究提供了一个基准。摘要:We introduce Label-Combination Prototypical Networks (LC-Protonets) to address the problem of multi-label few-shot classification, where a model must generalize to new classes based on only a few available examples. Extending Prototypical Networks, LC-Protonets generate one prototype per label combination, derived from the power set of labels present in the limited training items, rather than one prototype per label. Our method is applied to automatic audio tagging across diverse music datasets, covering various cultures and including both modern and traditional music, and is evaluated against existing approaches in the literature. The results demonstrate a significant performance improvement in almost all domains and training setups when using LC-Protonets for multi-label classification. In addition to training a few-shot learning model from scratch, we explore the use of a pre-trained model, obtained via supervised learning, to embed items in the feature space. Fine-tuning improves the generalization ability of all methods, yet LC-Protonets achieve high-level performance even without fine-tuning, in contrast to the comparative approaches. We finally analyze the scalability of the proposed method, providing detailed quantitative metrics from our experiments. The implementation and experimental setup are made publicly available, offering a benchmark for future research.
【3】 The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
标题: 家庭之声:用于声音事件检测的语音删除住宅音频数据集
作者:Gabriel Bibbó,Thomas Deacon,Arshdeep Singh,Mark D. Plumbley
链接:点击下载PDF文件
摘要:本文提出了一个住宅音频数据集,以支持智能家居应用的声音事件检测研究,旨在促进老年人的健康。该数据集是通过在8名年龄在55-80岁之间的参与者的家中部署音频记录系统来构建的,为期7天。声学特性通过详细的平面图和建筑材料信息进行记录,以复制AI模型部署的记录环境。开发了一种新的自动语音去除流水线,使用预训练的音频神经网络来检测和去除包含口语的片段,同时保留包含其他声音事件的片段。由此产生的数据集由符合隐私的音频记录组成,可准确捕捉住宅空间内的音景和日常生活活动。本文详细介绍了数据集的创建方法,语音去除流水线利用级联模型架构,并分析了声乐标签分布,以验证语音去除过程。该数据集支持专门为家庭应用定制的声音事件检测模型的开发和基准测试。摘要:This paper presents a residential audio dataset to support sound event detection research for smart home applications aimed at promoting wellbeing for older adults. The dataset is constructed by deploying audio recording systems in the homes of 8 participants aged 55-80 years for a 7-day period. Acoustic characteristics are documented through detailed floor plans and construction material information to enable replication of the recording environments for AI model deployment. A novel automated speech removal pipeline is developed, using pre-trained audio neural networks to detect and remove segments containing spoken voice, while preserving segments containing other sound events. The resulting dataset consists of privacy-compliant audio recordings that accurately capture the soundscapes and activities of daily living within residential spaces. The paper details the dataset creation methodology, the speech removal pipeline utilizing cascaded model architectures, and an analysis of the vocal label distribution to validate the speech removal process. This dataset enables the development and benchmarking of sound event detection models tailored specifically for in-home applications.
【4】 WER We Stand: Benchmarking Urdu ASR Models
标题: WER我们的立场:乌尔都语ASB模型基准
作者:Samee Arif,Aamina Jamal Khan,Mustafa Abbas,Agha Ali Raza,Awais Athar
链接:点击下载PDF文件
摘要:本文提出了一个全面的评价乌尔都语自动语音识别(ASR)模型。我们使用单词错误率(WER)分析了三种ASR模型系列的性能:Whisper,MMS和无障碍M4 T,并详细检查了最常见的错误单词和错误类型,包括插入,删除和替换。我们的分析是使用两种类型的数据集,阅读语音和会话语音。值得注意的是,我们提出了第一个会话语音数据集,旨在基准乌尔都语ASR模型。我们发现,seamless-large在阅读语音数据集上的性能优于其他ASR模型,而whisper-large在会话语音数据集上的性能最好。此外,该评估强调了仅使用定量指标评估乌尔都语等低资源语言的ASR模型的复杂性,并强调了对强大的乌尔都语文本规范化系统的需求。我们的研究结果为为乌尔都语等低资源语言开发强大的ASR系统提供了宝贵的见解。摘要:This paper presents a comprehensive evaluation of Urdu Automatic Speech Recognition (ASR) models. We analyze the performance of three ASR model families: Whisper, MMS, and Seamless-M4T using Word Error Rate (WER), along with a detailed examination of the most frequent wrong words and error types including insertions, deletions, and substitutions. Our analysis is conducted using two types of datasets, read speech and conversational speech. Notably, we present the first conversational speech dataset designed for benchmarking Urdu ASR models. We find that seamless-large outperforms other ASR models on the read speech dataset, while whisper-large performs best on the conversational speech dataset. Furthermore, this evaluation highlights the complexities of assessing ASR models for low-resource languages like Urdu using quantitative metrics alone and emphasizes the need for a robust Urdu text normalization system. Our findings contribute valuable insights for developing robust ASR systems for low-resource languages like Urdu.
【5】 Spontaneous Informal Speech Dataset for Punctuation Restoration
标题: 用于标点符号恢复的自发非正式语音数据集
作者:Xing Yi Liu,Homayoon Beigi
Journal-ref:Recognition Technologies, Inc. Technical Report, 2024
链接:点击下载PDF文件
摘要:目前,标点符号恢复模型几乎完全在结构良好的脚本语料库上进行评估。另一方面,现实世界的ASR系统和后处理管道通常适用于具有显著不规则性、口吃和偏离完美语法的自发语音。为了解决这一差异,我们引入SponSpeech,一个来自非正式语音来源的标点符号恢复数据集,其中包括标点符号和大小写信息。除了公开发布数据集外,我们还提供了一个过滤管道,可用于生成更多数据。我们的过滤管道检查语音音频和转录文本的质量。我们还仔细构建了一个“具有挑战性”的测试集,旨在评估模型利用音频信息来预测语法模糊标点符号的能力。SponSpeech可以在https: github.com GitHubAccountAnonymous PR上找到,以及用于数据集构建和模型运行的所有代码。摘要:Presently, punctuation restoration models are evaluated almost solely on well-structured, scripted corpora. On the other hand, real-world ASR systems and post-processing pipelines typically apply towards spontaneous speech with significant irregularities, stutters, and deviations from perfect grammar. To address this discrepancy, we introduce SponSpeech, a punctuation restoration dataset derived from informal speech sources, which includes punctuation and casing information. In addition to publicly releasing the dataset, we contribute a filtering pipeline that can be used to generate more data. Our filtering pipeline examines the quality of both speech audio and transcription text. We also carefully construct a challenging" test set, aimed at evaluating models' ability to leverage audio information to predict otherwise grammatically ambiguous punctuation. SponSpeech is available at https: github.com GitHubAccountAnonymous PR, along with all code for dataset building and model runs.
【6】 Learning Source Disentanglement in Neural Audio Codec
标题: 神经音频编解码器中的学习源解纠缠
作者:Xiaoyu Bie,Xubo Liu,Gaël Richard
备注:project page: this https URL
链接:点击下载PDF文件
摘要:神经音频编解码器通过有效地将连续音频信号转换为离散令牌来显著提高音频压缩。这些编解码器保留了高质量的声音,并通过在这些令牌上训练的生成模型实现了复杂的声音生成。然而,现有的神经编解码器模型通常是在大型的、无差别的音频数据集上训练的,忽略了语音、音乐和环境音效等声音域之间的本质差异。这种疏忽使数据建模复杂化,并对声音生成的可控性提出了额外的挑战。为了解决这些问题,我们引入了源解纠缠神经音频编解码器(SD-Codec),这是一种结合音频编码和源分离的新方法。通过联合学习音频再合成和分离,SD-Codec将来自不同域的音频信号明确分配给不同的码本,即离散表示集。实验结果表明,SD-Codec不仅保持了有竞争力的再合成质量,而且在分离结果的支持下,成功地将潜在空间中的不同源解纠缠,从而增强了音频编解码器的可解释性,并提供了对音频生成过程的潜在更精细控制。摘要:Neural audio codecs have significantly advanced audio compression by efficiently converting continuous audio signals into discrete tokens. These codecs preserve high-quality sound and enable sophisticated sound generation through generative models trained on these tokens. However, existing neural codec models are typically trained on large, undifferentiated audio datasets, neglecting the essential discrepancies between sound domains like speech, music, and environmental sound effects. This oversight complicates data modeling and poses additional challenges to the controllability of sound generation. To tackle these issues, we introduce the Source-Disentangled Neural Audio Codec (SD-Codec), a novel approach that combines audio coding and source separation. By jointly learning audio resynthesis and separation, SD-Codec explicitly assigns audio signals from different domains to distinct codebooks, sets of discrete representations. Experimental results indicate that SD-Codec not only maintains competitive resynthesis quality but also, supported by the separation results, demonstrates successful disentanglement of different sources in the latent space, thereby enhancing interpretability in audio codec and providing potential finer control over the audio generation process.
【7】 High-Resolution Speech Restoration with Latent Diffusion Model
标题: 基于潜在扩散模型的高分辨率语音恢复
作者:Tushar Dhyani,Florian Lux,Michele Mancusi,Giorgio Fabbro,Fritz Hohl,Ngoc Thang Vu
链接:点击下载PDF文件
摘要:传统的语音增强方法往往过于简化的恢复任务,专注于单一类型的失真。处理多重失真的生成模型经常与音素重建和高频谐波作斗争,导致呼吸和喘息伪影,降低重建语音的可懂度。这些模型在计算上也要求很高,许多解决方案仅限于在宽带频率范围内产生输出,这限制了它们对专业应用的适用性。为了解决这些挑战,我们提出了Hi-ResLDM,这是一种基于潜在扩散的新型生成模型,旨在消除多种失真并将语音录音恢复到以48 kHz采样的录音室质量。我们将Hi-ResLDM与利用GAN和条件流匹配(CFM)组件的最先进方法进行基准测试,在重新生成高频带细节方面表现出卓越的性能。Hi-ResLDM不仅在非侵入性指标方面表现出色,而且在人工评估中也一直是首选,并且在侵入性评估中表现出色,使其成为高分辨率语音恢复的理想选择。摘要:Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions frequently struggle with phone reconstruction and high-frequency harmonics, leading to breathing and gasping artifacts that reduce the intelligibility of reconstructed speech. These models are also computationally demanding, and many solutions are restricted to producing outputs in the wide-band frequency range, which limits their suitability for professional applications. To address these challenges, we propose Hi-ResLDM, a novel generative model based on latent diffusion designed to remove multiple distortions and restore speech recordings to studio quality, sampled at 48kHz. We benchmark Hi-ResLDM against state-of-the-art methods that leverage GAN and Conditional Flow Matching (CFM) components, demonstrating superior performance in regenerating high-frequency-band details. Hi-ResLDM not only excels in non-instrusive metrics but is also consistently preferred in human evaluation and performs competitively on intrusive evaluations, making it ideal for high-resolution speech restoration.
【8】 Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
标题: 具有掩蔽音频令牌建模和语义知识提炼的单级TTC
作者:Gerard I. Gállego,Roy Fejgin,Chunghsin Yeh,Xiaoyu Liu,Gautam Bhattacharya
备注:Demo page: see this https URL
链接:点击下载PDF文件
摘要:音频令牌建模已经成为语音合成的一个强大的框架,采用语义令牌的两阶段方法仍然很流行。在本文中,我们的目标是简化这一过程,通过引入语义知识蒸馏方法,使高质量的语音生成在一个单一的阶段。与单阶段基线相比,我们提出的模型提高了语音质量,可懂度和说话人相似性。虽然两阶段系统仍然领先的可懂度,我们的模型显着缩小差距,同时提供可比的语音质量。这些发现展示了单级模型的潜力,以实现高效,高质量的TTS与更紧凑和精简的架构。摘要:Audio token modeling has become a powerful framework for speech synthesis, with two-stage approaches employing semantic tokens remaining prevalent. In this paper, we aim to simplify this process by introducing a semantic knowledge distillation method that enables high-quality speech generation in a single stage. Our proposed model improves speech quality, intelligibility, and speaker similarity compared to a single-stage baseline. Although two-stage systems still lead in intelligibility, our model significantly narrows the gap while delivering comparable speech quality. These findings showcase the potential of single-stage models to achieve efficient, high-quality TTS with a more compact and streamlined architecture.
【9】 Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models
标题: 增强音频语言模型的低资源语言和教学遵循能力
作者:Potsawee Manakul,Guangzhi Sun,Warit Sirichotedumrong,Kasima Tharnpipitchai,Kunat Pipatanakul
备注:5 pages. Preprint under review
链接:点击下载PDF文件
摘要:音频语言模型可以理解音频输入,并基于指令执行一系列与音频相关的任务,例如语音识别和音频字幕,其中指令通常是文本提示。音频语言模型大多从预训练的音频编码器和大型语言模型(LLM)初始化。虽然这些预先训练的组件是为了支持多种语言而开发的,但音频语言模型主要是在英语数据上训练的,这可能会将其可用性限制在英语指令或英语语音输入上。首先,本文以泰语为例,研究了现有音频语言模型在服务不足的语言中的性能。本文表明,尽管建立在多语言的骨干,音频语言模型不表现出跨语言的紧急能力,低资源的语言。其次,本文研究了开发音频语言模型的数据混合,这些模型针对目标语言以及英语进行了优化。另外。本文将音频理解和语音解释跟随能力集成到单个统一模型中。我们的实验提供了洞察数据混合,以提高在低资源的语言和英语的解释以下的能力。我们的模型Typhoon-Audio比现有的开源音频语言模型性能好得多,并且在英语和泰语中与最先进的Gemini-1.5-Pro相当。摘要:Audio language models can understand audio inputs and perform a range of audio-related tasks based on instructions, such as speech recognition and audio captioning, where the instructions are usually textual prompts. Audio language models are mostly initialized from pre-trained audio encoders and large language models (LLMs). Although these pre-trained components were developed to support multiple languages, audio-language models are trained predominantly on English data, which may limit their usability to only English instructions or English speech inputs. First, this paper examines the performance of existing audio language models in an underserved language using Thai as an example. This paper demonstrates that, despite being built on multilingual backbones, audio language models do not exhibit cross-lingual emergent abilities to low-resource languages. Second, this paper studies data mixture for developing audio language models that are optimized for a target language as well as English. In addition. this paper integrates audio comprehension and speech instruction-following capabilities into a single unified model. Our experiments provide insights into data mixture for enhancing instruction-following capabilities in both a low-resource language and English. Our model, Typhoon-Audio, outperforms existing open-source audio language models by a considerable margin, and it is comparable to state-of-the-art Gemini-1.5-Pro in both English and Thai languages.
【10】 Adaptive Large Language Models By Layerwise Attention Shortcuts
标题: 通过分层注意力快捷方式自适应大型语言模型
作者:Prateek Verma,Mert Pilanci
备注:6 pages, 3 figures
链接:点击下载PDF文件
摘要:Transformer架构是现代AI革命的支柱。然而,它们是基于简单地将相同的块堆叠成几十层,并从一个块到另一个块顺序地处理信息。在本文中,我们建议挑战这一点,并为类似LLM的设置引入自适应计算,这使得最终层能够通过注意力机制来关注所有中间层,从而引入计算 textbf{attention shortcuts}。因此,这些快捷方式可以使架构深度和上下文自适应。我们展示了四个不同的数据集,即声学令牌,自然语言和符号音乐,我们实现了类似GPT架构的卓越性能。我们通过注意力地图提供证据,证明模型可以跨层学习复杂的依赖关系,这些依赖关系在上下文和深度方面都是自适应的,这取决于输入的令牌。摘要:Transformer architectures are the backbone of the modern AI revolution. However, they are based on simply stacking the same blocks in dozens of layers and processing information sequentially from one block to another. In this paper, we propose to challenge this and introduce adaptive computations for LLM-like setups, which allow the final layer to attend to all of the intermediate layers as it deems fit through the attention mechanism, thereby introducing computational textbf{attention shortcuts}. These shortcuts can thus make the architecture depth and context adaptive. We showcase four different datasets, namely acoustic tokens, natural language, and symbolic music, and we achieve superior performance for GPT-like architecture. We give evidence via attention maps that the models learn complex dependencies across layers that are adaptive in context and depth depending on the input tokens.
【11】 Speech Recognition for Analysis of Police Radio Communication
标题: 用于警用无线电通信分析的语音识别
作者:Tejes Srivastava,Ju-Chieh Chou,Priyank Shroff,Karen Livescu,Christopher Graziul
备注:Accepted by SLT 2024
链接:点击下载PDF文件
摘要:世界各地的警察部门使用双向无线电进行协调。这些广播警察通信(BPC)是关于日常警察活动和紧急反应的独特信息来源。然而,BPC不被转录,它们的自然主义音频特性使自动转录具有挑战性。我们收集了大约62,000个手动转录的无线电传输(约46小时的音频)的语料库,以评估使用现代识别模型进行自动语音识别(ASR)的可行性。我们评估了现成的语音识别器,模型微调的BPC数据和定制的端到端模型的性能。我们发现,人类和机器转录在这一领域是具有挑战性的。大型现成的ASR模型表现不佳,但经过微调的模型可以达到人类表现的近似范围。我们的工作为未来的工作提出了方向,包括分析警察无线电互动中的短话语和潜在的误解。我们将我们的语料库和数据注释管道提供给其他研究人员,以便进一步研究警察通信的识别和分析。摘要:Police departments around the world use two-way radio for coordination. These broadcast police communications (BPC) are a unique source of information about everyday police activity and emergency response. Yet BPC are not transcribed, and their naturalistic audio properties make automatic transcription challenging. We collect a corpus of roughly 62,000 manually transcribed radio transmissions (~46 hours of audio) to evaluate the feasibility of automatic speech recognition (ASR) using modern recognition models. We evaluate the performance of off-the-shelf speech recognizers, models fine-tuned on BPC data, and customized end-to-end models. We find that both human and machine transcription is challenging in this domain. Large off-the-shelf ASR models perform poorly, but fine-tuned models can reach the approximate range of human performance. Our work suggests directions for future work, including analysis of short utterances and potential miscommunication in police radio interactions. We make our corpus and data annotation pipeline available to other researchers, to enable further research on recognition and analysis of police communication.
【12】 3DFacePolicy: Speech-Driven 3D Facial Animation with Diffusion Policy
标题: 3DFacePolicy:具有扩散策略的语音驱动3D面部动画
作者:Xuanmeng Sha,Liyun Zhang,Tomohiro Mashita,Yuki Uranishi
链接:点击下载PDF文件
摘要:音频驱动的3D人脸动画在研究和应用开发方面都取得了令人身临其境的进展。最新的方法主要集中在基于变换器的方法和基于扩散的方法,然而,生成的动画与真实的人脸在生动性和情感表达方面仍然存在差距。为了解决这个问题,我们提出了3DFacePolicy,一个用于3D面部动画预测的扩散策略模型。该方法利用扩散策略预测三维人脸模板上的三维顶点轨迹,而不是逐帧生成人脸,从而生成可变的、逼真的人脸运动。它以音频和顶点状态作为观测值,预测顶点轨迹,模仿真实的人类面部表情,保持了人类情感的连续和自然流动。实验结果表明,该方法在可变动态人脸运动合成中是有效的。摘要:Audio-driven 3D facial animation has made immersive progress both in research and application developments. The newest approaches focus on Transformer-based methods and diffusion-based methods, however, there is still gap in the vividness and emotional expression between the generated animation and real human face. To tackle this limitation, we propose 3DFacePolicy, a diffusion policy model for 3D facial animation prediction. This method generates variable and realistic human facial movements by predicting the 3D vertex trajectory on the 3D facial template with diffusion policy instead of facial generation for every frame. It takes audio and vertex states as observations to predict the vertex trajectory and imitate real human facial expressions, which keeps the continuous and natural flow of human emotions. The experiments show that our approach is effective in variable and dynamic facial motion synthesizing.
【13】 PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
标题: PMX:用于符号音乐处理的大规模公共领域音乐ML数据集
作者:Phillip Long,Zachary Novack,Taylor Berg-Kirkpatrick,Julian McAuley
链接:点击下载PDF文件
摘要:最近生成式AI音乐系统的爆炸式增长引发了人们对数据版权、音乐家音乐许可以及开源AI与大型知名公司之间冲突的担忧。这些问题凸显了对公开可用、无版权的音乐数据的需求,而这些数据存在很大短缺,特别是符号音乐数据。为了缓解这个问题,我们提出了PDMX:一个从分数共享论坛MuseScore收集的超过25万公共领域MusicXML分数的大规模开源数据集,使其成为我们所知的最大的无版权符号音乐数据集。PDmx还包括大量的标签和用户交互元数据,使我们能够有效地分析数据集并过滤高质量的用户生成的分数。由于我们的数据收集过程中提供的额外的元数据,我们进行多轨音乐生成实验,评估不同的代表性子集的PDMX如何导致下游模型中的不同行为,以及如何用户评级统计数据可以用作数据质量的有效措施。例子可在https: pnlong.github.io PDMX.demo 上找到。摘要:The recent explosion of generative AI-Music systems has raised numerous concerns over data copyright, licensing music from musicians, and the conflict between open-source AI and large prestige companies. Such issues highlight the need for publicly available, copyright-free musical data, in which there is a large shortage, particularly for symbolic music data. To alleviate this issue, we present PDMX: a large-scale open-source dataset of over 250K public domain MusicXML scores collected from the score-sharing forum MuseScore, making it the largest available copyright-free symbolic music dataset to our knowledge. PDMX additionally includes a wealth of both tag and user interaction metadata, allowing us to efficiently analyze the dataset and filter for high quality user-generated scores. Given the additional metadata afforded by our data collection process, we conduct multitrack music generation experiments evaluating how different representative subsets of PDMX lead to different behaviors in downstream models, and how user-rating statistics can be used as an effective measure of data quality. Examples can be found at https: pnlong.github.io PDMX.demo .
【14】 Mitigating Sex Bias in Audio Data-driven COPD and COVID-19 Breathing Pattern Detection Models
标题: 缓解音频数据驱动的COPD和COVID-19呼吸模式检测模型中的性别偏见
作者:Rachel Pfeifer,Sudip Vhaduri,James Eric Dietz
备注:Accepted at 2024 IEEE-EMBS International Conference on Body Sensor Networks (IEEE BSN 2024)
链接:点击下载PDF文件
摘要:在医疗保健行业,研究人员一直在开发机器学习模型,根据呼吸模式自动诊断呼吸系统疾病患者。然而,这些模型没有考虑人口统计学偏见,特别是性别偏见,这在使用倾斜的患者数据集训练模型时经常发生。因此,在如此重要的行业中,减少这种偏见至关重要,以便模型能够做出公平的诊断。在这项工作中,我们研究了用于检测两种主要呼吸系统疾病的呼吸模式的模型中的偏差,即,慢性阻塞性肺疾病(COPD)和COVID-19。我们使用由29名COPD和680名COVID-19阳性患者组成的两个开源数据集获得的呼吸模式音频记录训练的决策树模型,分析了性别偏见对模型的影响。使用阈值优化器和两个约束(人口统计学奇偶性和均衡赔率)来减轻偏差,我们见证了81.43%(人口统计学奇偶性差异)和71.81%(均衡赔率差异)的改善。这些发现具有统计学意义。摘要:In the healthcare industry, researchers have been developing machine learning models to automate diagnosing patients with respiratory illnesses based on their breathing patterns. However, these models do not consider the demographic biases, particularly sex bias, that often occur when models are trained with a skewed patient dataset. Hence, it is essential in such an important industry to reduce this bias so that models can make fair diagnoses. In this work, we examine the bias in models used to detect breathing patterns of two major respiratory diseases, i.e., chronic obstructive pulmonary disease (COPD) and COVID-19. Using decision tree models trained with audio recordings of breathing patterns obtained from two open-source datasets consisting of 29 COPD and 680 COVID-19-positive patients, we analyze the effect of sex bias on the models. With a threshold optimizer and two constraints (demographic parity and equalized odds) to mitigate the bias, we witness 81.43% (demographic parity difference) and 71.81% (equalized odds difference) improvements. These findings are statistically significant.
【15】 Learning Co-Speech Gesture Representations in Dialogue through Contrastive Learning: An Intrinsic Evaluation
标题: 通过对比学习学习对话中的同语手势表示:内在评价
作者:Esam Ghaleb,Bulat Khaertdinov,Wim Pouw,Marlou Rasenberg,Judith Holler,Aslı Özyürek,Raquel Fernández
Journal-ref:INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION (ICMI 2024)
链接:点击下载PDF文件
摘要:在面对面的对话中,共语手势的形式-意义关系取决于语境因素,如手势所指的是什么和说话者的个人特征。这些因素使得协同语音手势表示学习具有挑战性。考虑到手势的可变性和与语音的关系,我们如何学习有意义的手势表示?本文通过采用自监督对比学习技术从骨骼和语音信息中学习手势表示来应对这一挑战。我们提出了一种方法,包括单模态和多模态预训练地面手势表示共现语音。为了训练,我们利用了一个面对面的对话数据集丰富的代表性的标志性手势。我们通过与人类注释的成对手势相似性进行比较,对所学习的表示进行彻底的内在评估。此外,我们进行了诊断探测分析,以评估从学习的表示恢复可解释的手势功能的可能性。我们的研究结果显示出显着的正相关性与人类注释的手势相似性,并揭示了学习表示之间的相似性是一致的动机良好的模式相关的动态对话互动。此外,我们的研究结果表明,一些功能的形式的手势可以恢复从潜在的表征。总的来说,这项研究表明,多模态对比学习是一种很有前途的方法,学习手势表示,这打开了大门,使用这种表示在大规模的手势分析研究。摘要:In face-to-face dialogues, the form-meaning relationship of co-speech gestures varies depending on contextual factors such as what the gestures refer to and the individual characteristics of speakers. These factors make co-speech gesture representation learning challenging. How can we learn meaningful gestures representations considering gestures' variability and relationship with speech? This paper tackles this challenge by employing self-supervised contrastive learning techniques to learn gesture representations from skeletal and speech information. We propose an approach that includes both unimodal and multimodal pre-training to ground gesture representations in co-occurring speech. For training, we utilize a face-to-face dialogue dataset rich with representational iconic gestures. We conduct thorough intrinsic evaluations of the learned representations through comparison with human-annotated pairwise gesture similarity. Moreover, we perform a diagnostic probing analysis to assess the possibility of recovering interpretable gesture features from the learned representations. Our results show a significant positive correlation with human-annotated gesture similarity and reveal that the similarity between the learned representations is consistent with well-motivated patterns related to the dynamics of dialogue interaction. Moreover, our findings demonstrate that several features concerning the form of gestures can be recovered from the latent representations. Overall, this study shows that multimodal contrastive learning is a promising approach for learning gesture representations, which opens the door to using such representations in larger-scale gesture analysis studies.
【16】 Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text
标题: Ideal-LLM:集成双编码器和扩展自适应LLM,用于多语言语音到文本
作者:Hongfei Xue,Wei Ren,Xuelong Geng,Kun Wei,Longhao Li,Qijie Shao,Linju Yang,Kai Diao,Lei Xie
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:通过连接器将音频编码器与LLM集成,使这些模型能够处理和理解音频模态,显着增强语音到文本的任务,包括自动语音识别(ASR)和自动语音翻译(AST)。然而,这些方法往往忽略了在多语言环境中语言适应的关键方面,而是依赖于多语言数据,而没有充分解决语言差异。为了解决这一差距,我们提出了Ideal-LLM模型,该模型采用双多语言编码器来丰富语言特征信息,并利用语言适应连接器来专门针对每种语言的适应。通过利用Whisper和MMS编码器的互补优势,我们的方法确保了更丰富的多语言表示。此外,语言适配连接器通过为每种语言定制的语言权重选择器来增强模态转换。实验结果表明,Ideal-LLM显着提高ASR性能,实现了32.6%的平均字错误率相对减少与LLM集成的标准语音编码器相比,并产生平均BLEU得分为36.78 AST任务。摘要:Integrating audio encoders with LLMs through connectors has enabled these models to process and comprehend audio modalities, significantly enhancing speech-to-text tasks, including automatic speech recognition (ASR) and automatic speech translation (AST). However, these methods often overlook the critical aspect of language adaptation in multilingual settings, relying instead on multilingual data without adequately addressing language differences. To address this gap, we propose the Ideal-LLM model, which employs dual multilingual encoders to enrich language feature information and utilizes a language-adapted connector to target the adaptation of each language specifically. By leveraging the complementary strengths of Whisper and MMS encoders, our approach ensures richer multilingual representations. Additionally, the language-adapted connector enhances modal transformation via a language weight selector tailored for each language. Experimental results demonstrate that Ideal-LLM significantly improves ASR performance, achieving a 32.6% relative reduction in average word error rates compared to the standard speech encoder integrated with LLMs and yields an average BLEU score of 36.78 for AST task.
【17】 Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
标题: Zero-Shot文本到语音增强,用于低资源语音库上的自动语音识别
作者:Francesco Nespoli,Daniel Barreda,Patrick A. Naylor
备注:Accepted to the Asilomar 2023 Conference
链接:点击下载PDF文件
摘要:近年来,自动语音识别(ASR)模型大大提高了在干净,低噪声,声学条件和混响环境中的转录性能。然而,所有这些系统都依赖于在特定声学条件下数百小时的标记训练数据的可用性。当这样的训练数据集不可用时,系统的性能会受到严重影响。例如,当特定的声学环境或特定的说话者群体在训练数据集中表现不足时,就会发生这种情况。具体而言,在本文中,我们调查的影响口音语音数据上的现成的ASR系统。此外,我们提出了一种基于zero-shot的文本到语音的策略来增强重音语音语料库。我们表明,这种增强方法是能够减轻损失的ASR系统的重音数据高达5%的字错误率降低(WERR)的性能。总之,我们证明,通过将适度的真实分数与合成生成的数据相结合,ASR系统表现出优越的性能相比,一个模型专门训练的真实口音的语音高达14%WERR。摘要:In recent years, automatic speech recognition (ASR) models greatly improved transcription performance both in clean, low noise, acoustic conditions and in reverberant environments. However, all these systems rely on the availability of hundreds of hours of labelled training data in specific acoustic conditions. When such a training dataset is not available, the performance of the system is heavily impacted. For example, this happens when a specific acoustic environment or a particular population of speakers is under-represented in the training dataset. Specifically, in this paper we investigate the effect of accented speech data on an off-the-shelf ASR system. Furthermore, we suggest a strategy based on zero-shot text-to-speech to augment the accented speech corpora. We show that this augmentation method is able to mitigate the loss in performance of the ASR system on accented data up to 5% word error rate reduction (WERR). In conclusion, we demonstrate that by incorporating a modest fraction of real with synthetically generated data, the ASR system exhibits superior performance compared to a model trained exclusively on authentic accented speech with up to 14% WERR.
【18】 An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech Characterization
标题: 欺骗语音特征的可解释概率属性嵌入方法
作者:Manasi Chhibber,Jagabandhu Mishra,Hyejin Shim,Tomi H. Kinnunen
备注:Submitted to ICASSP-2025
链接:点击下载PDF文件
摘要:我们提出了一种新的方法,通过可解释的概率属性嵌入欺骗语音特征。与从其维度不容易解释的欺骗对策(CM)中提取的高维原始嵌入相比,概率属性被设计为测量构成特定欺骗攻击的子组件的存在或不存在。然后将这些属性应用于两个下游任务:欺骗检测和攻击归因。为了增强后端的可解释性,我们采用了决策树分类器。我们在ASVspoof 2019数据集上的实验表明,从三个模型(AASIST,Rawboost-AASIST,SSL-AASIST)中提取的欺骗CM嵌入的属性嵌入的性能与两个任务的原始欺骗CM嵌入相当。与使用原始CM嵌入的99.7%和94.7%相比,所提出的方法在欺骗检测和攻击归因方面的最佳性能分别为99.7%和99.2%。为了分析每个属性的相对贡献,我们估计它们的Shapley值。声学特征预测,波形生成(声码器),扬声器建模相关的属性被发现是重要的欺骗检测,而持续时间建模,声码器,输入类型发挥作用,欺骗攻击属性。摘要:We propose a novel approach for spoofed speech characterization through explainable probabilistic attribute embeddings. In contrast to high-dimensional raw embeddings extracted from a spoofing countermeasure (CM) whose dimensions are not easy to interpret, the probabilistic attributes are designed to gauge the presence or absence of sub-components that make up a specific spoofing attack. These attributes are then applied to two downstream tasks: spoofing detection and attack attribution. To enforce interpretability also to the back-end, we adopt a decision tree classifier. Our experiments on the ASVspoof2019 dataset with spoof CM embeddings extracted from three models (AASIST, Rawboost-AASIST, SSL-AASIST) suggest that the performance of the attribute embeddings are on par with the original raw spoof CM embeddings for both tasks. The best performance achieved with the proposed approach for spoofing detection and attack attribution, in terms of accuracy, is 99.7% and 99.2%, respectively, compared to 99.7% and 94.7% using the raw CM embeddings. To analyze the relative contribution of each attribute, we estimate their Shapley values. Attributes related to acoustic feature prediction, waveform generation (vocoder), and speaker modeling are found important for spoofing detection; while duration modeling, vocoder, and input type play a role in spoofing attack attribution.
【19】 SynthSOD: Developing an Heterogeneous Dataset for Orchestra Music Source Separation
标题: SynthSOC:开发用于管弦乐队音乐源分离的异类数据集
作者:Jaime Garcia-Martinez,David Diaz-Guerra,Archontis Politis,Tuomas Virtanen,Julio J. Carabias-Orti,Pedro Vera-Candeas
备注:Submitted to the OJSP - ICASSP 2025
链接:点击下载PDF文件
摘要:最近在音乐源分离方面取得了显著进展,特别是在从混合音轨中分离人声、鼓和低音元素方面。这些发展在很大程度上归功于专门针对这些特定组成部分的大规模多轨数据集的创建和使用。然而,从管弦乐队录音中提取类似声音源的挑战尚未得到广泛探讨,主要是由于缺乏全面和干净的(即无出血)多轨数据集。在本文中,我们介绍了一种名为SynthSOD的新型多轨数据集,该数据集使用一组模拟技术开发,以创建一个逼真的(即使用高质量的soundfonts),音乐动机和异构的训练集,包括不同的动态,自然节奏变化,风格和条件。此外,我们展示了在我们的合成数据集w.r.t上训练的广泛使用的基线音乐分离模型在着名的EnsembleSet上的应用,并评估了其在合成和真实世界条件下的性能。摘要:Recent advancements in music source separation have significantly progressed, particularly in isolating vocals, drums, and bass elements from mixed tracks. These developments owe much to the creation and use of large-scale, multitrack datasets dedicated to these specific components. However, the challenge of extracting similarly sounding sources from orchestra recordings has not been extensively explored, largely due to a scarcity of comprehensive and clean (i.e bleed-free) multitrack datasets. In this paper, we introduce a novel multitrack dataset called SynthSOD, developed using a set of simulation techniques to create a realistic (i.e. using high-quality soundfonts), musically motivated, and heterogeneous training set comprising different dynamics, natural tempo changes, styles, and conditions. Moreover, we demonstrate the application of a widely used baseline music separation model trained on our synthesized dataset w.r.t to the well-known EnsembleSet, and evaluate its performance under both synthetic and real-world conditions.
【20】 Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
标题: 通过具有引导数据选择的语音到语音翻译来改善资源不足的语言中的语音情感识别
作者:Hsi-Che Lin,Yi-Cheng Lin,Huang-Cheng Chou,Hung-yi Lee
备注:5 pages, 2 figures, Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:语音情感识别(SER)是开发能够进行自然人机交互的通用AI代理的关键组成部分。然而,建立强大的多语言SER系统仍然是具有挑战性的,由于缺乏标记的数据,而不是英语和中文。在本文中,我们提出了一种方法来提高SER性能低SER资源语言,利用数据从高资源语言。具体来说,我们采用表达语音到语音翻译(S2ST)结合一种新的自举数据选择管道,以生成标记的数据在目标语言。大量的实验表明,我们的方法是有效的,可推广到不同的上游模型和语言。我们的研究结果表明,这种方法可以促进更可扩展和强大的多语言SER系统的发展。摘要:Speech Emotion Recognition (SER) is a crucial component in developing general-purpose AI agents capable of natural human-computer interaction. However, building robust multilingual SER systems remains challenging due to the scarcity of labeled data in languages other than English and Chinese. In this paper, we propose an approach to enhance SER performance in low SER resource languages by leveraging data from high-resource languages. Specifically, we employ expressive Speech-to-Speech translation (S2ST) combined with a novel bootstrapping data selection pipeline to generate labeled data in the target language. Extensive experiments demonstrate that our method is both effective and generalizable across different upstream models and languages. Our results suggest that this approach can facilitate the development of more scalable and robust multilingual SER systems.
【21】 Enhancing Multilingual Speech Generation and Recognition Abilities in LLMs with Constructed Code-switched Data
标题: 利用构造代码交换数据增强LLM中的多语言语音生成和识别能力
作者:Jing Xu,Daxin Tan,Jiaqi Wang,Xiao Chen
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:虽然大型语言模型(LLM)已经在语音领域的生成和识别任务进行了探索,但它们的应用主要局限于单语场景,在多语言和代码切换(CS)环境中的探索有限。此外,语音生成和识别任务通常是分开处理的,例如VALL-E和Qwen-Audio。在本文中,我们提出了一个多语言多任务(MLMT)模型,集成多语言语音生成和识别任务内的单一LLM。此外,我们开发了一种有效的数据构建方法,拆分和连接来自不同语言的单词,使LLM具有CS合成能力,而不依赖于CS数据。实验结果表明,我们的模型优于其他基线具有可比的数据规模。此外,我们的数据构建方法不仅装备LLM与CS语音合成能力与可比的扬声器的一致性和相似性的任何给定的扬声器,但也提高了性能的LLM在多语言语音生成和识别任务。摘要:While large language models (LLMs) have been explored in the speech domain for both generation and recognition tasks, their applications are predominantly confined to the monolingual scenario, with limited exploration in multilingual and code-switched (CS) contexts. Additionally, speech generation and recognition tasks are often handled separately, such as VALL-E and Qwen-Audio. In this paper, we propose a MutltiLingual MultiTask (MLMT) model, integrating multilingual speech generation and recognition tasks within the single LLM. Furthermore, we develop an effective data construction approach that splits and concatenates words from different languages to equip LLMs with CS synthesis ability without relying on CS data. The experimental results demonstrate that our model outperforms other baselines with a comparable data scale. Furthermore, our data construction approach not only equips LLMs with CS speech synthesis capability with comparable speaker consistency and similarity to any given speaker, but also improves the performance of LLMs in multilingual speech generation and recognition tasks.
【22】 EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
标题: EzAudio:通过高效的扩散Transformer增强文本到音频的生成
作者:Jiarui Hai,Yong Xu,Hao Zhang,Chenxing Li,Helin Wang,Mounya Elhilali,Dong Yu
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:潜在扩散模型在文本到音频(T2A)生成任务中已经显示出有希望的结果,然而先前的模型在生成质量、计算成本、扩散采样和数据准备方面遇到了困难。在本文中,我们引入了EzAudio,一个基于transformer的T2A扩散模型,来应对这些挑战。我们的方法包括几个关键的创新:(1)我们在一维波形变分自动编码器(VAE)的潜在空间上构建T2A模型,避免了处理2D频谱图表示的复杂性,并使用了额外的神经声码器。(2)我们设计了一个优化的扩散Transformer架构,专门针对音频潜在表示和扩散建模,提高了收敛速度,训练稳定性和内存使用,使训练过程更简单,更高效。(3)为了解决数据稀缺问题,我们采用了一种数据高效的训练策略,利用未标记的数据来学习声学依赖性,利用音频语言模型注释的音频字幕数据进行文本到音频对齐学习,并利用人工标记的数据进行微调。(4)我们引入了一种无分类器指导(CFG)重新缩放方法,该方法通过实现强大的即时对齐来简化EzAudio,同时在使用较大的CFG分数时保持出色的音频质量,从而消除了寻找最佳CFG分数以平衡这种权衡的需要。EzAudio在客观指标和主观评估方面都超越了现有的开源模型,提供逼真的聆听体验,同时保持精简的模型结构,低培训成本和易于遵循的培训管道。代码、数据和预训练模型发布于:https: haidog-yaqub.github.io EzAudio-Page 。摘要:Latent diffusion models have shown promising results in text-to-audio (T2A) generation tasks, yet previous models have encountered difficulties in generation quality, computational cost, diffusion sampling, and data preparation. In this paper, we introduce EzAudio, a transformer-based T2A diffusion model, to handle these challenges. Our approach includes several key innovations: (1) We build the T2A model on the latent space of a 1D waveform Variational Autoencoder (VAE), avoiding the complexities of handling 2D spectrogram representations and using an additional neural vocoder. (2) We design an optimized diffusion transformer architecture specifically tailored for audio latent representations and diffusion modeling, which enhances convergence speed, training stability, and memory usage, making the training process easier and more efficient. (3) To tackle data scarcity, we adopt a data-efficient training strategy that leverages unlabeled data for learning acoustic dependencies, audio caption data annotated by audio-language models for text-to-audio alignment learning, and human-labeled data for fine-tuning. (4) We introduce a classifier-free guidance (CFG) rescaling method that simplifies EzAudio by achieving strong prompt alignment while preserving great audio quality when using larger CFG scores, eliminating the need to struggle with finding the optimal CFG score to balance this trade-off. EzAudio surpasses existing open-source models in both objective metrics and subjective evaluations, delivering realistic listening experiences while maintaining a streamlined model structure, low training costs, and an easy-to-follow training pipeline. Code, data, and pre-trained models are released at: https: haidog-yaqub.github.io EzAudio-Page .
【23】 Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
标题: Speaker-IPL:使用基于i-Vector的伪标签进行说话者特征的无监督学习
作者:Zakaria Aldeneh,Takuya Higuchi,Jee-weon Jung,Li-Wei Chen,Stephen Shum,Ahmed Hussen Abdelaziz,Shinji Watanabe,Tatiana Likhomanenko,Barry-John Theobald
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:迭代自训练,或迭代伪标签(IPL)-使用当前迭代的改进模型为下一次迭代提供伪标签-已被证明是增强说话人表示质量的强大方法。IPL在无监督说话人识别中的最新应用始于从非常精细的自监督方法(例如,DINO)。然而,训练这种强自监督模型并不简单(它们需要超参数调整,并且可能无法推广到域外数据),而且可能根本不需要。为此,我们证明了简单的,经过充分研究的,并建立了i-向量生成模型是足够的引导IPL过程的无监督学习的发言人表示。我们还系统地研究了IPL过程中的其他组件的影响,其中包括初始模型,编码器,增强,集群的数量,和聚类算法。值得注意的是,我们发现即使使用简单且明显较弱的初始模型(如i-vector),IPL仍然可以实现与最先进方法相媲美的说话人验证性能。摘要:Iterative self-training, or iterative pseudo-labeling (IPL)--using an improved model from the current iteration to provide pseudo-labels for the next iteration--has proven to be a powerful approach to enhance the quality of speaker representations. Recent applications of IPL in unsupervised speaker recognition start with representations extracted from very elaborate self-supervised methods (e.g., DINO). However, training such strong self-supervised models is not straightforward (they require hyper-parameters tuning and may not generalize to out-of-domain data) and, moreover, may not be needed at all. To this end, we show the simple, well-studied, and established i-vector generative model is enough to bootstrap the IPL process for unsupervised learning of speaker representations. We also systematically study the impact of other components on the IPL process, which includes the initial model, the encoder, augmentations, the number of clusters, and the clustering algorithm. Remarkably, we find that even with a simple and significantly weaker initial model like i-vector, IPL can still achieve speaker verification performance that rivals state-of-the-art methods.
【24】 Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
标题: 语音基础模型掩蔽预训练中的预测目标探索
作者:Li-Wei Chen,Takuya Higuchi,He Bai,Ahmed Hussen Abdelaziz,Alexander Rudnicky,Shinji Watanabe,Tatiana Likhomanenko,Barry-John Theobald,Zakaria Aldeneh
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:语音基础模型,如HuBERT及其变体,在大量未标记语音上进行预训练,用于各种下游任务。这些模型使用掩蔽的预测目标,其中模型学习从未掩蔽的上下文预测关于掩蔽的输入片段的信息。该框架中预测目标的选择可能会影响下游任务的性能。例如,对韵律进行编码的目标对于说话者相关的任务是有益的,而对语音进行编码的目标更适合于内容相关的任务。此外,预测目标可以在它们编码的细节水平上有所不同;编码细粒度声学细节的目标有利于去噪任务,而编码更高级别抽象的目标更适合与内容相关的任务。尽管预测目标的重要性,影响他们的设计选择还没有得到彻底的研究。这项工作探讨了设计选择及其对下游任务绩效的影响。我们的研究结果表明,HuBERT常用的设计选择可能是次优的。我们提出了新的方法来创建更多信息的预测目标,并通过各种下游任务的改进来证明其有效性。摘要:Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech for various downstream tasks. These models use a masked prediction objective, where the model learns to predict information about masked input segments from the unmasked context. The choice of prediction targets in this framework can influence performance on downstream tasks. For example, targets that encode prosody are beneficial for speaker-related tasks, while targets that encode phonetics are more suited for content-related tasks. Additionally, prediction targets can vary in the level of detail they encode; targets that encode fine-grained acoustic details are beneficial for denoising tasks, while targets that encode higher-level abstractions are more suited for content-related tasks. Despite the importance of prediction targets, the design choices that affect them have not been thoroughly studied. This work explores the design choices and their impact on downstream task performance. Our results indicate that the commonly used design choices for HuBERT can be suboptimal. We propose novel approaches to create more informative prediction targets and demonstrate their effectiveness through improvements across various downstream tasks.
【25】 Towards Automatic Assessment of Self-Supervised Speech Models using Rank
标题: 使用排名自动评估自我监督语音模型
作者:Zakaria Aldeneh,Vimal Thilak,Takuya Higuchi,Barry-John Theobald,Tatiana Likhomanenko
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:本研究探讨使用嵌入秩作为通过自监督学习(SSL)训练的通用语音编码器的无监督评估指标。传统上,评估这些编码器的性能是资源密集型的,并且需要来自下游任务的标记数据。受视觉域的启发,嵌入秩已经显示出在不调整标记的下游数据的情况下评估图像编码器的前景,考虑到信号的时间性质,这项工作研究了其在语音域中的适用性。研究结果表明,排名与下游性能编码器层内的各种下游任务和域内和域外的情况。然而,排名并不能可靠地预测特定下游任务的最佳性能层,因为排名较低的层可能优于排名较高的层。尽管存在这种限制,但结果表明,嵌入秩可以成为监控SSL语音模型训练进度的有价值的工具,为传统评估方法提供了一种资源需求较少的替代方法。摘要:This study explores using embedding rank as an unsupervised evaluation metric for general-purpose speech encoders trained via self-supervised learning (SSL). Traditionally, assessing the performance of these encoders is resource-intensive and requires labeled data from the downstream tasks. Inspired by the vision domain, where embedding rank has shown promise for evaluating image encoders without tuning on labeled downstream data, this work examines its applicability in the speech domain, considering the temporal nature of the signals. The findings indicate rank correlates with downstream performance within encoder layers across various downstream tasks and for in- and out-of-domain scenarios. However, rank does not reliably predict the best-performing layer for specific downstream tasks, as lower-ranked layers can outperform higher-ranked ones. Despite this limitation, the results suggest that embedding rank can be a valuable tool for monitoring training progress in SSL speech models, offering a less resource-demanding alternative to traditional evaluation methods.
【26】 Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
标题: 刺激模式很重要:不同模式的感知评估对语音情感识别系统性能的影响
作者:Huang-Cheng Chou,Haibin Wu,Chi-Chun Lee
备注:5 pages, 2 figures, 4 tables, submission for ICASSP 2025
链接:点击下载PDF文件
摘要:语音情感识别(SER)系统依赖于语音输入和人类注释的情感标签。然而,各种情感数据库以不同的方式收集感知评价。例如,IEMOCAP数据集使用带有声音的视频剪辑来为注释者提供他们的情感感知。然而,最重要的英语情感数据集,MSP-PODCAST,只提供语音评分员选择的情感评级。然而,使用语音作为输入是训练SER系统的标准方法。因此,开放的问题是情感标签引起的场景是最有效的训练SER系统。我们全面比较了SER系统的有效性与标签引起的不同模态的刺激和评估SER系统在各种测试条件。此外,我们引入了一个包罗万象的标签,结合了各种方式引起的所有标签。我们表明,使用标签引起的语音刺激训练产生更好的性能测试集上,而标签引起的语音刺激。摘要:Speech Emotion Recognition (SER) systems rely on speech input and emotional labels annotated by humans. However, various emotion databases collect perceptional evaluations in different ways. For instance, the IEMOCAP dataset uses video clips with sounds for annotators to provide their emotional perceptions. However, the most significant English emotion dataset, the MSP-PODCAST, only provides speech for raters to choose the emotional ratings. Nevertheless, using speech as input is the standard approach to training SER systems. Therefore, the open question is the emotional labels elicited by which scenarios are the most effective for training SER systems. We comprehensively compare the effectiveness of SER systems trained with labels elicited by different modality stimuli and evaluate the SER systems on various testing conditions. Also, we introduce an all-inclusive label that combines all labels elicited by various modalities. We show that using labels elicited by voice-only stimuli for training yields better performance on the test set, whereas labels elicited by voice-only stimuli.
【27】 Investigating Training Objectives for Generative Speech Enhancement
标题: 研究生成性语音增强的训练目标
作者:Julius Richter,Danilo de Oliveira,Timo Gerkmann
链接:点击下载PDF文件
摘要:生成式语音增强技术在提高噪声环境下的语音质量方面取得了可喜的进展。存在多个基于扩散的框架,每个框架采用不同的训练目标和学习技术。本文旨在通过重点研究基于分数的生成模型和薛定谔桥来解释这些框架之间的差异。我们进行了一系列综合实验,以比较他们的表现,并突出不同的训练行为。此外,我们提出了一种针对Schr “odinger桥框架定制的新型感知损失函数,证明了增强语音信号的增强性能和改善的感知质量。所有实验代码和预训练模型都是公开的,以促进进一步的研究和开发。摘要:Generative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learning techniques. This paper aims at explaining the differences between these frameworks by focusing our investigation on score-based generative models and Schr "odinger bridge. We conduct a series of comprehensive experiments to compare their performance and highlight differing training behaviors. Furthermore, we propose a novel perceptual loss function tailored for the Schr "odinger bridge framework, demonstrating enhanced performance and improved perceptual quality of the enhanced speech signals. All experimental code and pre-trained models are publicly available to facilitate further research and development in this.
【28】 Self-supervised Speech Models for Word-Level Stuttered Speech Detection
标题: 用于词级口吃语音检测的自我监督语音模型
作者:Yi-Jen Shih,Zoi Gkalitsiou,Alexandros G. Dimakis,David Harwath
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
摘要:口吃的临床诊断需要有执照的言语语言病理学家的评估。然而,这个过程是耗时的,需要临床医生与培训和经验,口吃和流畅性障碍。不幸的是,只有一小部分的语言病理学家报告说,他们对与口吃者一起工作感到舒适,这不足以适应全世界8000万口吃者。开发用于检测口吃的机器学习模型将实现对口吃的通用和自动化筛查,使语言病理学家能够识别和随访最有可能被诊断患有口吃性语言障碍的患者。在这一领域以前的研究主要集中在话语水平的检测,这是不够的临床设置中,单词水平的口吃注释是常态。在本研究中,我们策划了一个具有单词级注释的口吃语音数据集,并引入了一个利用自监督语音模型的单词级口吃语音检测模型。我们的评估表明,我们的模型超越了以前的方法在字级口吃语音检测。此外,我们对我们的方法进行了广泛的消融分析,深入了解了自监督语音模型用于口吃语音检测的最重要方面。摘要:Clinical diagnosis of stuttering requires an assessment by a licensed speech-language pathologist. However, this process is time-consuming and requires clinicians with training and experience in stuttering and fluency disorders. Unfortunately, only a small percentage of speech-language pathologists report being comfortable working with individuals who stutter, which is inadequate to accommodate for the 80 million individuals who stutter worldwide. Developing machine learning models for detecting stuttered speech would enable universal and automated screening for stuttering, enabling speech pathologists to identify and follow up with patients who are most likely to be diagnosed with a stuttering speech disorder. Previous research in this area has predominantly focused on utterance-level detection, which is not sufficient for clinical settings where word-level annotation of stuttering is the norm. In this study, we curated a stuttered speech dataset with word-level annotations and introduced a word-level stuttering speech detection model leveraging self-supervised speech models. Our evaluation demonstrates that our model surpasses previous approaches in word-level stuttering speech detection. Additionally, we conducted an extensive ablation analysis of our method, providing insight into the most important aspects of adapting self-supervised speech models for stuttered speech detection.
【29】 Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
标题: 使用视觉变形器的人机交互中的个性化语音情感识别
作者:Ruchik Mishra,Andrew Frye,Madan Mohan Rayguru,Dan O. Popa
备注:Will be submitted to IEEE for possible publication
链接:点击下载PDF文件
摘要:情感是语言交流中的一个基本要素,因此了解人与机器人交互过程中个体的情感变得至关重要。本文研究了Vision Transformer模型,即ViT(Vision Transformers)和BEiT(BERT Pre-Training of Image Transformers)管道在HRI语音情感识别中的应用。重点是通过在基准数据集上微调这些模型并利用集成方法来概括单个语音特征的SER模型。为此,我们收集了来自不同人类主体的音频数据,这些人类主体与NAO机器人进行伪自然主义对话。然后,我们对基于ViT和BEiT的模型进行了微调,并在参与者未见过的语音样本上测试了这些模型。在结果中,我们表明,在基准数据集上微调Vision Transformers,然后使用这些已经微调的模型或集成ViT BEiT模型,当涉及到从他们的语音中识别四种主要情绪时,每个人的分类准确率最高:中性,快乐,悲伤和愤怒,与微调vanilla-ViTs或BEiTs相比。摘要:Emotions are an essential element in verbal communication, so understanding individuals' affect during a human-robot interaction (HRI) becomes imperative. This paper investigates the application of vision transformer models, namely ViT (Vision Transformers) and BEiT (BERT Pre-Training of Image Transformers) pipelines, for Speech Emotion Recognition (SER) in HRI. The focus is to generalize the SER models for individual speech characteristics by fine-tuning these models on benchmark datasets and exploiting ensemble methods. For this purpose, we collected audio data from different human subjects having pseudo-naturalistic conversations with the NAO robot. We then fine-tuned our ViT and BEiT-based models and tested these models on unseen speech samples from the participants. In the results, we show that fine-tuning vision transformers on benchmark datasets and and then using either these already fine-tuned models or ensembling ViT BEiT models gets us the highest classification accuracies per individual when it comes to identifying four primary emotions from their speech: neutral, happy, sad, and angry, as compared to fine-tuning vanilla-ViTs or BEiTs.
【30】 FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
标题: FakeMusicCaps:用于检测和归因通过文本到音乐模型生成的合成音乐的数据集
作者:Luca Comanducci,Paolo Bestagini,Stefano Tubaro
链接:点击下载PDF文件
摘要:文本到音乐(Text-To-Music,TTM)模型是近年来音乐自动生成研究领域的一次革命。具体而言,通过达到优于所有先前最先进型号的性能,并通过降低使用它们所需的技术熟练程度。由于这些原因,它们已经很容易地开始被用于商业用途和音乐制作实践。这种广泛传播的TTM提出了一些关于侵犯版权和合法归属的问题,提出了音频取证社区需要认真考虑的问题。在本文中,我们解决了TTM生成的数据的检测和属性的问题。我们提出了一个数据集,FakeMusicCaps,其中包含几个版本的音乐字幕对数据集MusicCaps通过几个国家的最先进的TTM技术重新生成。我们通过对TTM生成的音频的检测和归属进行初始实验来评估所提出的数据集。摘要:Text-To-Music (TTM) models have recently revolutionized the automatic music generation research field. Specifically, by reaching superior performances to all previous state-of-the-art models and by lowering the technical proficiency needed to use them. Due to these reasons, they have readily started to be adopted for commercial uses and music production practices. This widespread diffusion of TTMs poses several concerns regarding copyright violation and rightful attribution, posing the need of serious consideration of them by the audio forensics community. In this paper, we tackle the problem of detection and attribution of TTM-generated data. We propose a dataset, FakeMusicCaps that contains several versions of the music-caption pairs dataset MusicCaps re-generated via several state-of-the-art TTM techniques. We evaluate the proposed dataset by performing initial experiments regarding the detection and attribution of TTM-generated audio.
【31】 A Real-Time Platform for Portable and Scalable Active Noise Mitigation for Construction Machinery
标题: 便携式、可扩展的工程机械主动降噪实时平台
作者:Woon-Seng Gan,Santi Peksi,Chung Kwan Lai,Yen Theng Lee,Dongyuan Shi,Bhan Lam
Journal-ref:2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
链接:点击下载PDF文件
摘要:本文介绍了一种新型的便携式和可扩展的主动降噪(PSANM)系统,旨在降低工程机械的低频噪声。PSANM系统由具有自主功能的便携式设备组成,在特定功率范围内优化了稳定的性能。具有可变惩罚因子的自适应控制算法防止自适应滤波器过度驱动抗噪声致动器,从而避免非线性操作和不稳定性。这一特性确保PSANM系统可以自主控制噪声源,从而实现连续运行,无需人为干预。此外,该系统还包括一个用于远程管理的网络服务器,并配备了耐候传感器和执行器,增强了其在户外条件下的可用性。实验室和现场实验证明了PSANM系统在全球范围内减少与建筑相关的低频噪音的有效性。为了进一步扩大降噪区域,可以在噪声源前面战略性地放置额外的PSANM单元,从而增强系统的可扩展性。PSANM系统还提供了一个宝贵的原型平台,用于在部署之前开发自适应算法。不同于许多研究,仅仅依赖于理想条件下的仿真结果,本文提供了一个整体的评价,直接在噪声源应用有源噪声控制技术的有效性,展示现实和可感知的降噪。这项工作通过为建筑业提供创新的噪音管理解决方案,支持可持续的城市发展,为更安静,更宜居的城市环境做出贡献。摘要:This paper introduces a novel portable and scalable Active Noise Mitigation (PSANM) system designed to reduce low-frequency noise from construction machinery. The PSANM system consists of portable units with autonomous capabilities, optimized for stable performance within a specific power range. An adaptive control algorithm with a variable penalty factor prevents the adaptive filter from over-driving the anti-noise actuators, avoiding non-linear operation and instability. This feature ensures the PSANM system can autonomously control noise at its source, allowing for continuous operation without human intervention. Additionally, the system includes a web server for remote management and is equipped with weather-resistant sensors and actuators, enhancing its usability in outdoor conditions. Laboratory and in-situ experiments demonstrate the PSANM system's effectiveness in reducing construction-related low-frequency noise on a global scale. To further expand the noise reduction zone, additional PSANM units can be strategically positioned in front of noise sources, enhancing the system's scalability.The PSANM system also provides a valuable prototyping platform for developing adaptive algorithms prior to deployment. Unlike many studies that rely solely on simulation results under ideal conditions, this paper offers a holistic evaluation of the effectiveness of applying active noise control techniques directly at the noise source, demonstrating realistic and perceptible noise reduction. This work supports sustainable urban development by offering innovative noise management solutions for the construction industry, contributing to a quieter and more livable urban environment.
eess.AS音频处理
【1】 Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text标题: Ideal-LLM:集成双编码器和扩展自适应LLM,用于多语言语音到文本
作者:Hongfei Xue,Wei Ren,Xuelong Geng,Kun Wei,Longhao Li,Qijie Shao,Linju Yang,Kai Diao,Lei Xie
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:通过连接器将音频编码器与LLM集成,使这些模型能够处理和理解音频模态,显着增强语音到文本的任务,包括自动语音识别(ASR)和自动语音翻译(AST)。然而,这些方法往往忽视了多语言环境中语言适应的关键方面,而是依赖于多语言数据,而没有充分解决语言差异。为了解决这一差距,我们提出了理想LLM模型,它采用双多语言编码器来丰富语言特征信息,并利用语言适应连接器来专门针对每种语言的适应。通过利用Whisper和MMS编码器的互补优势,我们的方法确保了更丰富的多语言表示。此外,语言适配连接器通过为每种语言定制的语言权重选择器来增强模态转换。实验结果表明,Ideal-LLM显着提高ASR性能,实现了32.6%的平均字错误率相对减少与LLM集成的标准语音编码器相比,并产生平均BLEU得分为36.78 AST任务。摘要:Integrating audio encoders with LLMs through connectors has enabled these models to process and comprehend audio modalities, significantly enhancing speech-to-text tasks, including automatic speech recognition (ASR) and automatic speech translation (AST). However, these methods often overlook the critical aspect of language adaptation in multilingual settings, relying instead on multilingual data without adequately addressing language differences. To address this gap, we propose the Ideal-LLM model, which employs dual multilingual encoders to enrich language feature information and utilizes a language-adapted connector to target the adaptation of each language specifically. By leveraging the complementary strengths of Whisper and MMS encoders, our approach ensures richer multilingual representations. Additionally, the language-adapted connector enhances modal transformation via a language weight selector tailored for each language. Experimental results demonstrate that Ideal-LLM significantly improves ASR performance, achieving a 32.6% relative reduction in average word error rates compared to the standard speech encoder integrated with LLMs and yields an average BLEU score of 36.78 for AST task.
【2】 Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
标题: Zero-Shot文本到语音增强,用于低资源语音库上的自动语音识别
作者:Francesco Nespoli,Daniel Barreda,Patrick A. Naylor
备注:Accepted to the Asilomar 2023 Conference
链接:点击下载PDF文件
摘要:近年来,自动语音识别(ASR)模型大大提高了在干净,低噪声,声学条件和混响环境中的转录性能。然而,所有这些系统都依赖于在特定声学条件下数百小时的标记训练数据的可用性。当这样的训练数据集不可用时,系统的性能会受到严重影响。例如,当特定的声学环境或特定的说话者群体在训练数据集中表现不足时,就会发生这种情况。具体而言,在本文中,我们调查的影响口音语音数据上的现成的ASR系统。此外,我们提出了一种基于zero-shot的文本到语音的策略来增强重音语音语料库。我们表明,这种增强方法是能够减轻损失的ASR系统的重音数据高达5%的字错误率降低(WERR)的性能。总之,我们证明,通过将适度的真实分数与合成生成的数据相结合,ASR系统表现出优越的性能相比,一个模型专门训练的真实口音的语音高达14%WERR。摘要:In recent years, automatic speech recognition (ASR) models greatly improved transcription performance both in clean, low noise, acoustic conditions and in reverberant environments. However, all these systems rely on the availability of hundreds of hours of labelled training data in specific acoustic conditions. When such a training dataset is not available, the performance of the system is heavily impacted. For example, this happens when a specific acoustic environment or a particular population of speakers is under-represented in the training dataset. Specifically, in this paper we investigate the effect of accented speech data on an off-the-shelf ASR system. Furthermore, we suggest a strategy based on zero-shot text-to-speech to augment the accented speech corpora. We show that this augmentation method is able to mitigate the loss in performance of the ASR system on accented data up to 5% word error rate reduction (WERR). In conclusion, we demonstrate that by incorporating a modest fraction of real with synthetically generated data, the ASR system exhibits superior performance compared to a model trained exclusively on authentic accented speech with up to 14% WERR.
【3】 An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech Characterization
标题: 欺骗语音特征的可解释概率属性嵌入方法
作者:Manasi Chhibber,Jagabandhu Mishra,Hyejin Shim,Tomi H. Kinnunen
备注:Submitted to ICASSP-2025
链接:点击下载PDF文件
摘要:我们提出了一种新的方法,通过可解释的概率属性嵌入欺骗语音特征。与从其维度不容易解释的欺骗对策(CM)中提取的高维原始嵌入相比,概率属性被设计为测量构成特定欺骗攻击的子组件的存在或不存在。然后将这些属性应用于两个下游任务:欺骗检测和攻击归因。为了增强后端的可解释性,我们采用了决策树分类器。我们在ASVspoof 2019数据集上的实验表明,从三个模型(AASIST,Rawboost-AASIST,SSL-AASIST)中提取的欺骗CM嵌入的属性嵌入的性能与两个任务的原始欺骗CM嵌入相当。与使用原始CM嵌入的99.7%和94.7%相比,所提出的方法在欺骗检测和攻击归因方面的最佳性能分别为99.7%和99.2%。为了分析每个属性的相对贡献,我们估计它们的Shapley值。声学特征预测,波形生成(声码器),扬声器建模相关的属性被发现是重要的欺骗检测,而持续时间建模,声码器,输入类型发挥作用,欺骗攻击属性。摘要:We propose a novel approach for spoofed speech characterization through explainable probabilistic attribute embeddings. In contrast to high-dimensional raw embeddings extracted from a spoofing countermeasure (CM) whose dimensions are not easy to interpret, the probabilistic attributes are designed to gauge the presence or absence of sub-components that make up a specific spoofing attack. These attributes are then applied to two downstream tasks: spoofing detection and attack attribution. To enforce interpretability also to the back-end, we adopt a decision tree classifier. Our experiments on the ASVspoof2019 dataset with spoof CM embeddings extracted from three models (AASIST, Rawboost-AASIST, SSL-AASIST) suggest that the performance of the attribute embeddings are on par with the original raw spoof CM embeddings for both tasks. The best performance achieved with the proposed approach for spoofing detection and attack attribution, in terms of accuracy, is 99.7% and 99.2%, respectively, compared to 99.7% and 94.7% using the raw CM embeddings. To analyze the relative contribution of each attribute, we estimate their Shapley values. Attributes related to acoustic feature prediction, waveform generation (vocoder), and speaker modeling are found important for spoofing detection; while duration modeling, vocoder, and input type play a role in spoofing attack attribution.
【4】 SynthSOD: Developing an Heterogeneous Dataset for Orchestra Music Source Separation
标题: SynthSOC:开发用于管弦乐队音乐源分离的异类数据集
作者:Jaime Garcia-Martinez,David Diaz-Guerra,Archontis Politis,Tuomas Virtanen,Julio J. Carabias-Orti,Pedro Vera-Candeas
备注:Submitted to the OJSP - ICASSP 2025
链接:点击下载PDF文件
摘要:最近在音乐源分离方面取得了显著进展,特别是在从混合音轨中分离人声、鼓和低音元素方面。这些发展在很大程度上归功于专门针对这些特定组成部分的大规模多轨数据集的创建和使用。然而,从管弦乐队录音中提取类似声音源的挑战尚未得到广泛探讨,主要是由于缺乏全面和干净的(即无出血)多轨数据集。在本文中,我们介绍了一种名为SynthSOD的新型多轨数据集,该数据集使用一组模拟技术开发,以创建一个逼真的(即使用高质量的soundfonts)、音乐动机和异构的训练集,包括不同的动态、自然节奏变化、风格和条件。此外,我们展示了在我们的合成数据集w.r.t上训练的广泛使用的基线音乐分离模型在着名的EnsembleSet上的应用,并评估了其在合成和真实世界条件下的性能。摘要:Recent advancements in music source separation have significantly progressed, particularly in isolating vocals, drums, and bass elements from mixed tracks. These developments owe much to the creation and use of large-scale, multitrack datasets dedicated to these specific components. However, the challenge of extracting similarly sounding sources from orchestra recordings has not been extensively explored, largely due to a scarcity of comprehensive and clean (i.e bleed-free) multitrack datasets. In this paper, we introduce a novel multitrack dataset called SynthSOD, developed using a set of simulation techniques to create a realistic (i.e. using high-quality soundfonts), musically motivated, and heterogeneous training set comprising different dynamics, natural tempo changes, styles, and conditions. Moreover, we demonstrate the application of a widely used baseline music separation model trained on our synthesized dataset w.r.t to the well-known EnsembleSet, and evaluate its performance under both synthetic and real-world conditions.
【5】 Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
标题: 通过具有引导数据选择的语音到语音翻译来改善资源不足的语言中的语音情感识别
作者:Hsi-Che Lin,Yi-Cheng Lin,Huang-Cheng Chou,Hung-yi Lee
备注:5 pages, 2 figures, Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:语音情感识别(SER)是开发能够进行自然人机交互的通用AI代理的关键组成部分。然而,建立强大的多语言SER系统仍然是具有挑战性的,由于缺乏标记的数据,而不是英语和中文。在本文中,我们提出了一种方法来提高SER性能低SER资源语言,利用数据从高资源语言。具体来说,我们采用表达语音到语音翻译(S2ST)结合一种新的自举数据选择管道,以生成标记的数据在目标语言。大量的实验表明,我们的方法是有效的,可推广到不同的上游模型和语言。我们的研究结果表明,这种方法可以促进更可扩展和强大的多语言SER系统的发展。摘要:Speech Emotion Recognition (SER) is a crucial component in developing general-purpose AI agents capable of natural human-computer interaction. However, building robust multilingual SER systems remains challenging due to the scarcity of labeled data in languages other than English and Chinese. In this paper, we propose an approach to enhance SER performance in low SER resource languages by leveraging data from high-resource languages. Specifically, we employ expressive Speech-to-Speech translation (S2ST) combined with a novel bootstrapping data selection pipeline to generate labeled data in the target language. Extensive experiments demonstrate that our method is both effective and generalizable across different upstream models and languages. Our results suggest that this approach can facilitate the development of more scalable and robust multilingual SER systems.
【6】 Enhancing Multilingual Speech Generation and Recognition Abilities in LLMs with Constructed Code-switched Data
标题: 利用构造代码交换数据增强LLM中的多语言语音生成和识别能力
作者:Jing Xu,Daxin Tan,Jiaqi Wang,Xiao Chen
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:虽然大型语言模型(LLM)已经在语音领域的生成和识别任务进行了探索,但它们的应用主要局限于单语场景,在多语言和代码切换(CS)环境中的探索有限。此外,语音生成和识别任务通常是分开处理的,例如VALL-E和Qwen-Audio。在本文中,我们提出了一个多语言多任务(MLMT)模型,集成多语言语音生成和识别任务内的单一LLM。此外,我们开发了一种有效的数据构建方法,拆分和连接来自不同语言的单词,使LLM具有CS合成能力,而不依赖于CS数据。实验结果表明,我们的模型优于其他基线具有可比的数据规模。此外,我们的数据构建方法不仅装备LLM与CS语音合成能力与可比的扬声器的一致性和相似性的任何给定的扬声器,但也提高了性能的LLM在多语言语音生成和识别任务。摘要:While large language models (LLMs) have been explored in the speech domain for both generation and recognition tasks, their applications are predominantly confined to the monolingual scenario, with limited exploration in multilingual and code-switched (CS) contexts. Additionally, speech generation and recognition tasks are often handled separately, such as VALL-E and Qwen-Audio. In this paper, we propose a MutltiLingual MultiTask (MLMT) model, integrating multilingual speech generation and recognition tasks within the single LLM. Furthermore, we develop an effective data construction approach that splits and concatenates words from different languages to equip LLMs with CS synthesis ability without relying on CS data. The experimental results demonstrate that our model outperforms other baselines with a comparable data scale. Furthermore, our data construction approach not only equips LLMs with CS speech synthesis capability with comparable speaker consistency and similarity to any given speaker, but also improves the performance of LLMs in multilingual speech generation and recognition tasks.
【7】 EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
标题: EzAudio:通过高效的扩散Transformer增强文本到音频的生成
作者:Jiarui Hai,Yong Xu,Hao Zhang,Chenxing Li,Helin Wang,Mounya Elhilali,Dong Yu
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:潜在扩散模型在文本到音频(T2A)生成任务中已经显示出有希望的结果,然而先前的模型在生成质量、计算成本、扩散采样和数据准备方面遇到了困难。在本文中,我们引入了EzAudio,一个基于transformer的T2A扩散模型,来应对这些挑战。我们的方法包括几个关键的创新:(1)我们在一维波形变分自动编码器(VAE)的潜在空间上构建T2A模型,避免了处理2D频谱图表示的复杂性,并使用了额外的神经声码器。(2)我们设计了一个优化的扩散Transformer架构,专门针对音频潜在表示和扩散建模,提高了收敛速度,训练稳定性和内存使用,使训练过程更简单,更高效。(3)为了解决数据稀缺问题,我们采用了一种数据高效的训练策略,利用未标记的数据来学习声学依赖性,利用音频语言模型注释的音频字幕数据进行文本到音频对齐学习,并利用人工标记的数据进行微调。(4)我们引入了一种无分类器指导(CFG)重新缩放方法,该方法通过实现强大的即时对齐来简化EzAudio,同时在使用较大的CFG分数时保持出色的音频质量,从而消除了寻找最佳CFG分数以平衡这种权衡的需要。EzAudio在客观指标和主观评估方面都超越了现有的开源模型,提供逼真的聆听体验,同时保持精简的模型结构,低培训成本和易于遵循的培训管道。代码、数据和预训练模型发布于:https: haidog-yaqub.github.io EzAudio-Page 。摘要:Latent diffusion models have shown promising results in text-to-audio (T2A) generation tasks, yet previous models have encountered difficulties in generation quality, computational cost, diffusion sampling, and data preparation. In this paper, we introduce EzAudio, a transformer-based T2A diffusion model, to handle these challenges. Our approach includes several key innovations: (1) We build the T2A model on the latent space of a 1D waveform Variational Autoencoder (VAE), avoiding the complexities of handling 2D spectrogram representations and using an additional neural vocoder. (2) We design an optimized diffusion transformer architecture specifically tailored for audio latent representations and diffusion modeling, which enhances convergence speed, training stability, and memory usage, making the training process easier and more efficient. (3) To tackle data scarcity, we adopt a data-efficient training strategy that leverages unlabeled data for learning acoustic dependencies, audio caption data annotated by audio-language models for text-to-audio alignment learning, and human-labeled data for fine-tuning. (4) We introduce a classifier-free guidance (CFG) rescaling method that simplifies EzAudio by achieving strong prompt alignment while preserving great audio quality when using larger CFG scores, eliminating the need to struggle with finding the optimal CFG score to balance this trade-off. EzAudio surpasses existing open-source models in both objective metrics and subjective evaluations, delivering realistic listening experiences while maintaining a streamlined model structure, low training costs, and an easy-to-follow training pipeline. Code, data, and pre-trained models are released at: https: haidog-yaqub.github.io EzAudio-Page .
【8】 Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
标题: Speaker-IPL:使用基于i-Vector的伪标签进行说话者特征的无监督学习
作者:Zakaria Aldeneh,Takuya Higuchi,Jee-weon Jung,Li-Wei Chen,Stephen Shum,Ahmed Hussen Abdelaziz,Shinji Watanabe,Tatiana Likhomanenko,Barry-John Theobald
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:迭代自训练,或迭代伪标签(IPL)-使用当前迭代的改进模型为下一次迭代提供伪标签-已被证明是增强说话人表示质量的强大方法。IPL在无监督说话人识别中的最新应用始于从非常精细的自监督方法(例如,DINO)。然而,训练这种强自监督模型并不简单(它们需要超参数调整,并且可能无法推广到域外数据),而且可能根本不需要。为此,我们证明了简单的,经过充分研究的,并建立了i-向量生成模型是足够的引导IPL过程的无监督学习的发言人表示。我们还系统地研究了IPL过程中的其他组件的影响,其中包括初始模型,编码器,增强,集群的数量,和聚类算法。值得注意的是,我们发现即使使用简单且明显较弱的初始模型(如i-vector),IPL仍然可以实现与最先进方法相媲美的说话人验证性能。摘要:Iterative self-training, or iterative pseudo-labeling (IPL)--using an improved model from the current iteration to provide pseudo-labels for the next iteration--has proven to be a powerful approach to enhance the quality of speaker representations. Recent applications of IPL in unsupervised speaker recognition start with representations extracted from very elaborate self-supervised methods (e.g., DINO). However, training such strong self-supervised models is not straightforward (they require hyper-parameters tuning and may not generalize to out-of-domain data) and, moreover, may not be needed at all. To this end, we show the simple, well-studied, and established i-vector generative model is enough to bootstrap the IPL process for unsupervised learning of speaker representations. We also systematically study the impact of other components on the IPL process, which includes the initial model, the encoder, augmentations, the number of clusters, and the clustering algorithm. Remarkably, we find that even with a simple and significantly weaker initial model like i-vector, IPL can still achieve speaker verification performance that rivals state-of-the-art methods.
【9】 Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
标题: 语音基础模型掩蔽预训练中的预测目标探索
作者:Li-Wei Chen,Takuya Higuchi,He Bai,Ahmed Hussen Abdelaziz,Alexander Rudnicky,Shinji Watanabe,Tatiana Likhomanenko,Barry-John Theobald,Zakaria Aldeneh
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:语音基础模型,如HuBERT及其变体,在大量未标记语音上进行预训练,用于各种下游任务。这些模型使用掩蔽的预测目标,其中模型学习从未掩蔽的上下文预测关于掩蔽的输入片段的信息。该框架中预测目标的选择可能会影响下游任务的性能。例如,对韵律进行编码的目标对于说话者相关的任务是有益的,而对语音进行编码的目标更适合于内容相关的任务。此外,预测目标可以在它们编码的细节水平上有所不同;编码细粒度声学细节的目标有利于去噪任务,而编码更高级别抽象的目标更适合与内容相关的任务。尽管预测目标的重要性,影响他们的设计选择还没有得到彻底的研究。这项工作探讨了设计选择及其对下游任务性能的影响。我们的研究结果表明,HuBERT常用的设计选择可能是次优的。我们提出了新的方法来创建更多信息的预测目标,并通过各种下游任务的改进来证明其有效性。摘要:Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech for various downstream tasks. These models use a masked prediction objective, where the model learns to predict information about masked input segments from the unmasked context. The choice of prediction targets in this framework can influence performance on downstream tasks. For example, targets that encode prosody are beneficial for speaker-related tasks, while targets that encode phonetics are more suited for content-related tasks. Additionally, prediction targets can vary in the level of detail they encode; targets that encode fine-grained acoustic details are beneficial for denoising tasks, while targets that encode higher-level abstractions are more suited for content-related tasks. Despite the importance of prediction targets, the design choices that affect them have not been thoroughly studied. This work explores the design choices and their impact on downstream task performance. Our results indicate that the commonly used design choices for HuBERT can be suboptimal. We propose novel approaches to create more informative prediction targets and demonstrate their effectiveness through improvements across various downstream tasks.
【10】 Towards Automatic Assessment of Self-Supervised Speech Models using Rank
标题: 使用排名自动评估自我监督语音模型
作者:Zakaria Aldeneh,Vimal Thilak,Takuya Higuchi,Barry-John Theobald,Tatiana Likhomanenko
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:本研究探讨使用嵌入秩作为通过自监督学习(SSL)训练的通用语音编码器的无监督评估指标。传统上,评估这些编码器的性能是资源密集型的,并且需要来自下游任务的标记数据。受视觉域的启发,嵌入秩已经显示出在不调整标记的下游数据的情况下评估图像编码器的前景,考虑到信号的时间性质,这项工作研究了其在语音域中的适用性。研究结果表明,排名与下游性能编码器层内的各种下游任务和域内和域外的情况。然而,排名并不能可靠地预测特定下游任务的最佳性能层,因为排名较低的层可能优于排名较高的层。尽管存在这种限制,但结果表明,嵌入秩可以成为监控SSL语音模型训练进度的有价值的工具,为传统评估方法提供了一种资源需求较少的替代方法。摘要:This study explores using embedding rank as an unsupervised evaluation metric for general-purpose speech encoders trained via self-supervised learning (SSL). Traditionally, assessing the performance of these encoders is resource-intensive and requires labeled data from the downstream tasks. Inspired by the vision domain, where embedding rank has shown promise for evaluating image encoders without tuning on labeled downstream data, this work examines its applicability in the speech domain, considering the temporal nature of the signals. The findings indicate rank correlates with downstream performance within encoder layers across various downstream tasks and for in- and out-of-domain scenarios. However, rank does not reliably predict the best-performing layer for specific downstream tasks, as lower-ranked layers can outperform higher-ranked ones. Despite this limitation, the results suggest that embedding rank can be a valuable tool for monitoring training progress in SSL speech models, offering a less resource-demanding alternative to traditional evaluation methods.
【11】 Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
标题: 刺激模式很重要:不同模式的感知评估对语音情感识别系统性能的影响
作者:Huang-Cheng Chou,Haibin Wu,Chi-Chun Lee
备注:5 pages, 2 figures, 4 tables, submission for ICASSP 2025
链接:点击下载PDF文件
摘要:语音情感识别(SER)系统依赖于语音输入和人类注释的情感标签。然而,各种情感数据库以不同的方式收集感知评价。例如,IEMOCAP数据集使用带有声音的视频剪辑来为注释者提供他们的情感感知。然而,最重要的英语情感数据集MSP-PODCAST仅为评分者提供选择情感评分的语音。然而,使用语音作为输入是训练SER系统的标准方法。因此,悬而未决的问题是哪些场景对训练SER系统最有效,由此引发的情感标签。我们全面比较了SER系统的有效性与标签引起的不同模态的刺激和评估SER系统在各种测试条件。此外,我们引入了一个包罗万象的标签,结合了各种方式引起的所有标签。我们表明,使用标签引起的语音刺激训练产生更好的性能测试集上,而标签引起的语音刺激。摘要:Speech Emotion Recognition (SER) systems rely on speech input and emotional labels annotated by humans. However, various emotion databases collect perceptional evaluations in different ways. For instance, the IEMOCAP dataset uses video clips with sounds for annotators to provide their emotional perceptions. However, the most significant English emotion dataset, the MSP-PODCAST, only provides speech for raters to choose the emotional ratings. Nevertheless, using speech as input is the standard approach to training SER systems. Therefore, the open question is the emotional labels elicited by which scenarios are the most effective for training SER systems. We comprehensively compare the effectiveness of SER systems trained with labels elicited by different modality stimuli and evaluate the SER systems on various testing conditions. Also, we introduce an all-inclusive label that combines all labels elicited by various modalities. We show that using labels elicited by voice-only stimuli for training yields better performance on the test set, whereas labels elicited by voice-only stimuli.
【12】 Investigating Training Objectives for Generative Speech Enhancement
标题: 研究生成性语音增强的训练目标
作者:Julius Richter,Danilo de Oliveira,Timo Gerkmann
链接:点击下载PDF文件
摘要:生成式语音增强技术在提高噪声环境下的语音质量方面取得了可喜的进展。存在多个基于扩散的框架,每个框架采用不同的训练目标和学习技术。本文旨在通过对基于分数的生成模型和薛定谔桥的研究来解释这些框架之间的差异。我们进行了一系列综合实验,以比较他们的表现,并突出不同的训练行为。此外,我们提出了一种针对Schr “odinger桥框架定制的新型感知损失函数,证明了增强语音信号的增强性能和改善的感知质量。所有实验代码和预训练模型都是公开的,以促进进一步的研究和开发。摘要:Generative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learning techniques. This paper aims at explaining the differences between these frameworks by focusing our investigation on score-based generative models and Schr "odinger bridge. We conduct a series of comprehensive experiments to compare their performance and highlight differing training behaviors. Furthermore, we propose a novel perceptual loss function tailored for the Schr "odinger bridge framework, demonstrating enhanced performance and improved perceptual quality of the enhanced speech signals. All experimental code and pre-trained models are publicly available to facilitate further research and development in this.
【13】 Self-supervised Speech Models for Word-Level Stuttered Speech Detection
标题: 用于词级口吃语音检测的自我监督语音模型
作者:Yi-Jen Shih,Zoi Gkalitsiou,Alexandros G. Dimakis,David Harwath
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
摘要:口吃的临床诊断需要有执照的言语语言病理学家的评估。然而,这个过程是耗时的,需要临床医生与培训和经验,口吃和流畅性障碍。不幸的是,只有一小部分的语言病理学家报告说,他们对与口吃者一起工作感到舒适,这不足以适应全世界8000万口吃者。开发用于检测口吃的机器学习模型将实现对口吃的通用和自动化筛查,使语言病理学家能够识别和随访最有可能被诊断患有口吃性语言障碍的患者。在这一领域以前的研究主要集中在话语水平的检测,这是不够的临床设置中,单词水平的口吃注释是常态。在本研究中,我们策划了一个具有单词级注释的口吃语音数据集,并引入了一个利用自监督语音模型的单词级口吃语音检测模型。我们的评估表明,我们的模型超越了以前的方法在字级口吃语音检测。此外,我们对我们的方法进行了广泛的消融分析,深入了解了自监督语音模型用于口吃语音检测的最重要方面。摘要:Clinical diagnosis of stuttering requires an assessment by a licensed speech-language pathologist. However, this process is time-consuming and requires clinicians with training and experience in stuttering and fluency disorders. Unfortunately, only a small percentage of speech-language pathologists report being comfortable working with individuals who stutter, which is inadequate to accommodate for the 80 million individuals who stutter worldwide. Developing machine learning models for detecting stuttered speech would enable universal and automated screening for stuttering, enabling speech pathologists to identify and follow up with patients who are most likely to be diagnosed with a stuttering speech disorder. Previous research in this area has predominantly focused on utterance-level detection, which is not sufficient for clinical settings where word-level annotation of stuttering is the norm. In this study, we curated a stuttered speech dataset with word-level annotations and introduced a word-level stuttering speech detection model leveraging self-supervised speech models. Our evaluation demonstrates that our model surpasses previous approaches in word-level stuttering speech detection. Additionally, we conducted an extensive ablation analysis of our method, providing insight into the most important aspects of adapting self-supervised speech models for stuttered speech detection.
【14】 Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
标题: 使用视觉变形器的人机交互中的个性化语音情感识别
作者:Ruchik Mishra,Andrew Frye,Madan Mohan Rayguru,Dan O. Popa
备注:Will be submitted to IEEE for possible publication
链接:点击下载PDF文件
摘要:情感是语言交流中的一个基本要素,因此了解人与机器人交互过程中个体的情感变得至关重要。本文研究了Vision Transformer模型,即ViT(Vision Transformers)和BEiT(BERT Pre-Training of Image Transformers)管道在HRI语音情感识别中的应用。重点是通过在基准数据集上微调这些模型并利用集成方法来概括单个语音特征的SER模型。为此,我们收集了来自不同人类主体的音频数据,这些人类主体与NAO机器人进行伪自然主义对话。然后,我们对基于ViT和BEiT的模型进行了微调,并在参与者未见过的语音样本上测试了这些模型。在结果中,我们表明,在基准数据集上微调Vision Transformers,然后使用这些已经微调的模型或集成ViT BEiT模型,当涉及到从他们的语音中识别四种主要情绪时,每个人的分类准确率最高:中性,快乐,悲伤和愤怒,与微调vanilla-ViTs或BEiTs相比。摘要:Emotions are an essential element in verbal communication, so understanding individuals' affect during a human-robot interaction (HRI) becomes imperative. This paper investigates the application of vision transformer models, namely ViT (Vision Transformers) and BEiT (BERT Pre-Training of Image Transformers) pipelines, for Speech Emotion Recognition (SER) in HRI. The focus is to generalize the SER models for individual speech characteristics by fine-tuning these models on benchmark datasets and exploiting ensemble methods. For this purpose, we collected audio data from different human subjects having pseudo-naturalistic conversations with the NAO robot. We then fine-tuned our ViT and BEiT-based models and tested these models on unseen speech samples from the participants. In the results, we show that fine-tuning vision transformers on benchmark datasets and and then using either these already fine-tuned models or ensembling ViT BEiT models gets us the highest classification accuracies per individual when it comes to identifying four primary emotions from their speech: neutral, happy, sad, and angry, as compared to fine-tuning vanilla-ViTs or BEiTs.
【15】 FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
标题: FakeMusicCaps:用于检测和归因通过文本到音乐模型生成的合成音乐的数据集
作者:Luca Comanducci,Paolo Bestagini,Stefano Tubaro
链接:点击下载PDF文件
摘要:文本到音乐(Text-To-Music,TTM)模型是近年来音乐自动生成研究领域的一次革命。具体而言,通过达到优于所有先前最先进型号的性能,并通过降低使用它们所需的技术熟练程度。由于这些原因,它们已经很容易地开始被用于商业用途和音乐制作实践。这种广泛传播的TTM提出了一些关于侵犯版权和合法归属的问题,提出了音频取证社区需要认真考虑的问题。在本文中,我们解决了TTM生成的数据的检测和属性的问题。我们提出了一个数据集,FakeMusicCaps,其中包含几个版本的音乐字幕对数据集MusicCaps通过几个国家的最先进的TTM技术重新生成。我们通过对TTM生成的音频的检测和归属进行初始实验来评估所提出的数据集。摘要:Text-To-Music (TTM) models have recently revolutionized the automatic music generation research field. Specifically, by reaching superior performances to all previous state-of-the-art models and by lowering the technical proficiency needed to use them. Due to these reasons, they have readily started to be adopted for commercial uses and music production practices. This widespread diffusion of TTMs poses several concerns regarding copyright violation and rightful attribution, posing the need of serious consideration of them by the audio forensics community. In this paper, we tackle the problem of detection and attribution of TTM-generated data. We propose a dataset, FakeMusicCaps that contains several versions of the music-caption pairs dataset MusicCaps re-generated via several state-of-the-art TTM techniques. We evaluate the proposed dataset by performing initial experiments regarding the detection and attribution of TTM-generated audio.
【16】 A Real-Time Platform for Portable and Scalable Active Noise Mitigation for Construction Machinery
标题: 便携式、可扩展的工程机械主动降噪实时平台
作者:Woon-Seng Gan,Santi Peksi,Chung Kwan Lai,Yen Theng Lee,Dongyuan Shi,Bhan Lam
Journal-ref:2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
链接:点击下载PDF文件
摘要:本文介绍了一种新型的便携式和可扩展的主动降噪(PSANM)系统,旨在降低工程机械的低频噪声。PSANM系统由具有自主功能的便携式设备组成,在特定功率范围内优化了稳定的性能。具有可变惩罚因子的自适应控制算法防止自适应滤波器过度驱动抗噪声致动器,从而避免非线性操作和不稳定性。这一特性确保PSANM系统可以自主控制噪声源,从而实现连续运行,无需人为干预。此外,该系统还包括一个用于远程管理的网络服务器,并配备了耐候传感器和执行器,增强了其在户外条件下的可用性。实验室和现场实验证明了PSANM系统在全球范围内减少与建筑相关的低频噪音的有效性。为了进一步扩大降噪区域,可以在噪声源前面战略性地放置额外的PSANM单元,从而增强系统的可扩展性。PSANM系统还提供了一个宝贵的原型平台,用于在部署之前开发自适应算法。不同于许多研究,仅仅依赖于理想条件下的仿真结果,本文提供了一个整体的评价,直接在噪声源应用有源噪声控制技术的有效性,展示现实和可感知的降噪。这项工作通过为建筑业提供创新的噪音管理解决方案,支持可持续的城市发展,为更安静,更宜居的城市环境做出贡献。摘要:This paper introduces a novel portable and scalable Active Noise Mitigation (PSANM) system designed to reduce low-frequency noise from construction machinery. The PSANM system consists of portable units with autonomous capabilities, optimized for stable performance within a specific power range. An adaptive control algorithm with a variable penalty factor prevents the adaptive filter from over-driving the anti-noise actuators, avoiding non-linear operation and instability. This feature ensures the PSANM system can autonomously control noise at its source, allowing for continuous operation without human intervention. Additionally, the system includes a web server for remote management and is equipped with weather-resistant sensors and actuators, enhancing its usability in outdoor conditions. Laboratory and in-situ experiments demonstrate the PSANM system's effectiveness in reducing construction-related low-frequency noise on a global scale. To further expand the noise reduction zone, additional PSANM units can be strategically positioned in front of noise sources, enhancing the system's scalability.The PSANM system also provides a valuable prototyping platform for developing adaptive algorithms prior to deployment. Unlike many studies that rely solely on simulation results under ideal conditions, this paper offers a holistic evaluation of the effectiveness of applying active noise control techniques directly at the noise source, demonstrating realistic and perceptible noise reduction. This work supports sustainable urban development by offering innovative noise management solutions for the construction industry, contributing to a quieter and more livable urban environment.
【17】 Learning Spatially-Aware Language and Audio Embedding
标题: 学习空间感知语言和音频嵌入
作者:Bhavika Devnani,Skyler Seto,Zakaria Aldeneh,Alessandro Toso,Elena Menyaylenko,Barry-John Theobald,Jonathan Sheaffer,Miguel Sarabia
备注:25 pages, 7 figures
链接:点击下载PDF文件
摘要:人类可以通过不精确的自然语言描述来描绘声音场景。例如,很容易想象一个声学环境,给出一个短语,如“狮子的咆哮声来自我身后!".为了让机器具有相同程度的理解能力,机器必须知道狮子是什么(语义属性),“后面”的概念是什么(空间属性),以及这些语言信息如何与声音的语义和空间属性相匹配(当它从后面传来时,咆哮声听起来像什么)。最先进的音频基础模型学习在音频场景和自然文本描述之间进行映射,这些模型是在非空间音频和文本对上训练的,因此缺乏空间意识。相比之下,声音事件定位和检测模型限于从固定数量的类别中识别声音,并且它们将源定位到绝对位置(例如,0.2m)而不是使用自然语言描述的位置(例如,“在我旁边”)。为了解决这些差距,我们提出了ELSA空间感知音频和文本嵌入模型训练使用多模态对比学习。ELSA支持非空间音频,空间音频和开放词汇文本标题,描述声音的空间和语义成分。训练ELSA:(a)我们在空间上增强了三个开源音频数据集的音频和字幕,总计4,738小时的音频,以及(b)我们设计了一个编码器来捕获非空间音频的语义,以及使用对比学习的空间音频的语义和空间属性。ELSA在语义检索和3D源定位方面具有竞争力。特别地,ELSA实现了高于基线的+2.8%的平均音频到文本和文本到音频R@1,并且在3D源定位中优于基线的-11.6{ deg}平均绝对误差。摘要:Humans can picture a sound scene given an imprecise natural language description. For example, it is easy to imagine an acoustic environment given a phrase like "the lion roar came from right behind me!". For a machine to have the same degree of comprehension, the machine must know what a lion is (semantic attribute), what the concept of "behind" is (spatial attribute) and how these pieces of linguistic information align with the semantic and spatial attributes of the sound (what a roar sounds like when its coming from behind). State-of-the-art audio foundation models which learn to map between audio scenes and natural textual descriptions, are trained on non-spatial audio and text pairs, and hence lack spatial awareness. In contrast, sound event localization and detection models are limited to recognizing sounds from a fixed number of classes, and they localize the source to absolute position (e.g., 0.2m) rather than a position described using natural language (e.g., "next to me"). To address these gaps, we present ELSA a spatially aware-audio and text embedding model trained using multimodal contrastive learning. ELSA supports non-spatial audio, spatial audio, and open vocabulary text captions describing both the spatial and semantic components of sound. To train ELSA: (a) we spatially augment the audio and captions of three open-source audio datasets totaling 4,738 hours of audio, and (b) we design an encoder to capture the semantics of non-spatial audio, and the semantics and spatial attributes of spatial audio using contrastive learning. ELSA is competitive with state-of-the-art for both semantic retrieval and 3D source localization. In particular, ELSA achieves +2.8% mean audio-to-text and text-to-audio R@1 above the baseline, and outperforms by -11.6{ deg} mean-absolute-error in 3D source localization over the baseline.
【18】 LC-Protonets: Multi-label Few-shot learning for world music audio tagging
标题: LC-Protonets:用于世界音乐音频标记的多标签Few-Shot学习
作者:Charilaos Papaioannou,Emmanouil Benetos,Alexandros Potamianos
链接:点击下载PDF文件
摘要:我们引入标签组合原型网络(LC-Protonets)来解决多标签Few-Shot分类问题,其中模型必须仅基于少数可用示例推广到新的类别。扩展原型网络,LC-Protonets每个标签组合生成一个原型,来自有限训练项中存在的标签的幂集,而不是每个标签一个原型。我们的方法被应用于自动音频标记在不同的音乐数据集,涵盖各种文化,包括现代和传统音乐,并对现有的方法在文献中进行评估。结果表明,当使用LC-Protonets进行多标签分类时,几乎所有领域和训练设置都有显着的性能改善。除了从头开始训练Few-Shot学习模型之外,我们还探索了使用通过监督学习获得的预训练模型来将项嵌入特征空间中。微调提高了所有方法的泛化能力,但LC-质子网络实现高水平的性能,即使没有微调,相比之下,比较方法。最后,我们分析了所提出的方法的可扩展性,从我们的实验提供详细的定量指标。实施和实验设置公开提供,为未来的研究提供了一个基准。摘要:We introduce Label-Combination Prototypical Networks (LC-Protonets) to address the problem of multi-label few-shot classification, where a model must generalize to new classes based on only a few available examples. Extending Prototypical Networks, LC-Protonets generate one prototype per label combination, derived from the power set of labels present in the limited training items, rather than one prototype per label. Our method is applied to automatic audio tagging across diverse music datasets, covering various cultures and including both modern and traditional music, and is evaluated against existing approaches in the literature. The results demonstrate a significant performance improvement in almost all domains and training setups when using LC-Protonets for multi-label classification. In addition to training a few-shot learning model from scratch, we explore the use of a pre-trained model, obtained via supervised learning, to embed items in the feature space. Fine-tuning improves the generalization ability of all methods, yet LC-Protonets achieve high-level performance even without fine-tuning, in contrast to the comparative approaches. We finally analyze the scalability of the proposed method, providing detailed quantitative metrics from our experiments. The implementation and experimental setup are made publicly available, offering a benchmark for future research.
【19】 The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
标题: 家庭之声:用于声音事件检测的语音删除住宅音频数据集
作者:Gabriel Bibbó,Thomas Deacon,Arshdeep Singh,Mark D. Plumbley
链接:点击下载PDF文件
摘要:本文提出了一个住宅音频数据集,以支持智能家居应用的声音事件检测研究,旨在促进老年人的健康。该数据集是通过在8名年龄在55-80岁之间的参与者的家中部署音频记录系统来构建的,为期7天。声学特性通过详细的平面图和建筑材料信息进行记录,以复制AI模型部署的记录环境。开发了一种新的自动语音去除流水线,使用预训练的音频神经网络来检测和去除包含口语的片段,同时保留包含其他声音事件的片段。由此产生的数据集由符合隐私的音频记录组成,可准确捕捉住宅空间内的音景和日常生活活动。本文详细介绍了数据集的创建方法,语音去除流水线利用级联模型架构,并分析了声乐标签分布,以验证语音去除过程。该数据集支持专门为家庭应用定制的声音事件检测模型的开发和基准测试。摘要:This paper presents a residential audio dataset to support sound event detection research for smart home applications aimed at promoting wellbeing for older adults. The dataset is constructed by deploying audio recording systems in the homes of 8 participants aged 55-80 years for a 7-day period. Acoustic characteristics are documented through detailed floor plans and construction material information to enable replication of the recording environments for AI model deployment. A novel automated speech removal pipeline is developed, using pre-trained audio neural networks to detect and remove segments containing spoken voice, while preserving segments containing other sound events. The resulting dataset consists of privacy-compliant audio recordings that accurately capture the soundscapes and activities of daily living within residential spaces. The paper details the dataset creation methodology, the speech removal pipeline utilizing cascaded model architectures, and an analysis of the vocal label distribution to validate the speech removal process. This dataset enables the development and benchmarking of sound event detection models tailored specifically for in-home applications.
【20】 WER We Stand: Benchmarking Urdu ASR Models
标题: WER我们的立场:乌尔都语ASB模型基准
作者:Samee Arif,Aamina Jamal Khan,Mustafa Abbas,Agha Ali Raza,Awais Athar
链接:点击下载PDF文件
摘要:本文提出了一个全面的评价乌尔都语自动语音识别(ASR)模型。我们使用单词错误率(WER)分析了三种ASR模型系列的性能:Whisper,MMS和无障碍M4 T,并详细检查了最常见的错误单词和错误类型,包括插入,删除和替换。我们的分析是使用两种类型的数据集,阅读语音和会话语音。值得注意的是,我们提出了第一个会话语音数据集,旨在基准乌尔都语ASR模型。我们发现,seamless-large在阅读语音数据集上的性能优于其他ASR模型,而whisper-large在会话语音数据集上的性能最好。此外,这项评估强调了单独使用定量指标评估乌尔都语等低资源语言的ASR模型的复杂性,并强调了对强大的乌尔都语文本规范化系统的需求。我们的研究结果为为乌尔都语等低资源语言开发强大的ASR系统提供了宝贵的见解。摘要:This paper presents a comprehensive evaluation of Urdu Automatic Speech Recognition (ASR) models. We analyze the performance of three ASR model families: Whisper, MMS, and Seamless-M4T using Word Error Rate (WER), along with a detailed examination of the most frequent wrong words and error types including insertions, deletions, and substitutions. Our analysis is conducted using two types of datasets, read speech and conversational speech. Notably, we present the first conversational speech dataset designed for benchmarking Urdu ASR models. We find that seamless-large outperforms other ASR models on the read speech dataset, while whisper-large performs best on the conversational speech dataset. Furthermore, this evaluation highlights the complexities of assessing ASR models for low-resource languages like Urdu using quantitative metrics alone and emphasizes the need for a robust Urdu text normalization system. Our findings contribute valuable insights for developing robust ASR systems for low-resource languages like Urdu.
【21】 Spontaneous Informal Speech Dataset for Punctuation Restoration
标题: 用于标点符号恢复的自发非正式语音数据集
作者:Xing Yi Liu,Homayoon Beigi
Journal-ref:Recognition Technologies, Inc. Technical Report, 2024
链接:点击下载PDF文件
摘要:目前,标点符号恢复模型几乎完全在结构良好的脚本语料库上进行评估。另一方面,现实世界的ASR系统和后处理管道通常适用于具有显著不规则性、口吃和偏离完美语法的自发语音。为了解决这一差异,我们引入SponSpeech,一个来自非正式语音来源的标点符号恢复数据集,其中包括标点符号和大小写信息。除了公开发布数据集外,我们还提供了一个过滤管道,可用于生成更多数据。我们的过滤管道检查语音音频和转录文本的质量。我们还仔细构建了一个“具有挑战性”的测试集,旨在评估模型利用音频信息预测语法模糊标点符号的能力。SponSpeech可以在https: github.com GitHubAccountAnonymous PR上找到,以及用于数据集构建和模型运行的所有代码。摘要:Presently, punctuation restoration models are evaluated almost solely on well-structured, scripted corpora. On the other hand, real-world ASR systems and post-processing pipelines typically apply towards spontaneous speech with significant irregularities, stutters, and deviations from perfect grammar. To address this discrepancy, we introduce SponSpeech, a punctuation restoration dataset derived from informal speech sources, which includes punctuation and casing information. In addition to publicly releasing the dataset, we contribute a filtering pipeline that can be used to generate more data. Our filtering pipeline examines the quality of both speech audio and transcription text. We also carefully construct a challenging" test set, aimed at evaluating models' ability to leverage audio information to predict otherwise grammatically ambiguous punctuation. SponSpeech is available at https: github.com GitHubAccountAnonymous PR, along with all code for dataset building and model runs.
【22】 Learning Source Disentanglement in Neural Audio Codec
标题: 神经音频编解码器中的学习源解纠缠
作者:Xiaoyu Bie,Xubo Liu,Gaël Richard
备注:project page: this https URL
链接:点击下载PDF文件
摘要:神经音频编解码器通过有效地将连续音频信号转换为离散令牌来显著提高音频压缩。这些编解码器保留了高质量的声音,并通过在这些令牌上训练的生成模型实现了复杂的声音生成。然而,现有的神经编解码器模型通常是在大型的、无差别的音频数据集上训练的,忽略了语音、音乐和环境音效等声音域之间的本质差异。这种疏忽使数据建模复杂化,并对声音生成的可控性提出了额外的挑战。为了解决这些问题,我们引入了源解纠缠神经音频编解码器(SD-Codec),这是一种结合音频编码和源分离的新方法。通过联合学习音频再合成和分离,SD-Codec将来自不同域的音频信号明确分配给不同的码本,即离散表示集。实验结果表明,SD-Codec不仅保持了有竞争力的再合成质量,而且在分离结果的支持下,成功地将潜在空间中的不同源解纠缠,从而增强了音频编解码器的可解释性,并提供了对音频生成过程的潜在更精细控制。摘要:Neural audio codecs have significantly advanced audio compression by efficiently converting continuous audio signals into discrete tokens. These codecs preserve high-quality sound and enable sophisticated sound generation through generative models trained on these tokens. However, existing neural codec models are typically trained on large, undifferentiated audio datasets, neglecting the essential discrepancies between sound domains like speech, music, and environmental sound effects. This oversight complicates data modeling and poses additional challenges to the controllability of sound generation. To tackle these issues, we introduce the Source-Disentangled Neural Audio Codec (SD-Codec), a novel approach that combines audio coding and source separation. By jointly learning audio resynthesis and separation, SD-Codec explicitly assigns audio signals from different domains to distinct codebooks, sets of discrete representations. Experimental results indicate that SD-Codec not only maintains competitive resynthesis quality but also, supported by the separation results, demonstrates successful disentanglement of different sources in the latent space, thereby enhancing interpretability in audio codec and providing potential finer control over the audio generation process.
【23】 High-Resolution Speech Restoration with Latent Diffusion Model
标题: 基于潜在扩散模型的高分辨率语音恢复
作者:Tushar Dhyani,Florian Lux,Michele Mancusi,Giorgio Fabbro,Fritz Hohl,Ngoc Thang Vu
链接:点击下载PDF文件
摘要:传统的语音增强方法往往过于简化的恢复任务,专注于单一类型的失真。处理多重失真的生成模型经常与音素重建和高频谐波作斗争,导致呼吸和喘息伪影,降低重建语音的可懂度。这些模型在计算上也要求很高,许多解决方案仅限于在宽带频率范围内产生输出,这限制了它们对专业应用的适用性。为了解决这些挑战,我们提出了Hi-ResLDM,这是一种基于潜在扩散的新型生成模型,旨在消除多种失真并将语音录音恢复到以48 kHz采样的录音室质量。我们将Hi-ResLDM与利用GAN和条件流匹配(CFM)组件的最先进方法进行基准测试,在重新生成高频带细节方面表现出卓越的性能。Hi-ResLDM不仅在非侵入性指标方面表现出色,而且在人工评估中也一直是首选,并且在侵入性评估中表现出色,使其成为高分辨率语音恢复的理想选择。摘要:Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions frequently struggle with phone reconstruction and high-frequency harmonics, leading to breathing and gasping artifacts that reduce the intelligibility of reconstructed speech. These models are also computationally demanding, and many solutions are restricted to producing outputs in the wide-band frequency range, which limits their suitability for professional applications. To address these challenges, we propose Hi-ResLDM, a novel generative model based on latent diffusion designed to remove multiple distortions and restore speech recordings to studio quality, sampled at 48kHz. We benchmark Hi-ResLDM against state-of-the-art methods that leverage GAN and Conditional Flow Matching (CFM) components, demonstrating superior performance in regenerating high-frequency-band details. Hi-ResLDM not only excels in non-instrusive metrics but is also consistently preferred in human evaluation and performs competitively on intrusive evaluations, making it ideal for high-resolution speech restoration.
【24】 Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
标题: 具有掩蔽音频令牌建模和语义知识提炼的单级TTC
作者:Gerard I. Gállego,Roy Fejgin,Chunghsin Yeh,Xiaoyu Liu,Gautam Bhattacharya
备注:Demo page: see this https URL
链接:点击下载PDF文件
摘要:音频令牌建模已经成为语音合成的一个强大的框架,采用语义令牌的两阶段方法仍然流行。在本文中,我们的目标是简化这一过程,通过引入语义知识蒸馏方法,使高质量的语音生成在一个单一的阶段。与单阶段基线相比,我们提出的模型提高了语音质量,可懂度和说话人相似性。虽然两阶段系统仍然领先的可懂度,我们的模型显着缩小差距,同时提供可比的语音质量。这些发现展示了单级模型的潜力,以实现高效,高质量的TTS与更紧凑和精简的架构。摘要:Audio token modeling has become a powerful framework for speech synthesis, with two-stage approaches employing semantic tokens remaining prevalent. In this paper, we aim to simplify this process by introducing a semantic knowledge distillation method that enables high-quality speech generation in a single stage. Our proposed model improves speech quality, intelligibility, and speaker similarity compared to a single-stage baseline. Although two-stage systems still lead in intelligibility, our model significantly narrows the gap while delivering comparable speech quality. These findings showcase the potential of single-stage models to achieve efficient, high-quality TTS with a more compact and streamlined architecture.
【25】 Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models
标题: 增强音频语言模型的低资源语言和教学遵循能力
作者:Potsawee Manakul,Guangzhi Sun,Warit Sirichotedumrong,Kasima Tharnpipitchai,Kunat Pipatanakul
备注:5 pages. Preprint under review
链接:点击下载PDF文件
摘要:音频语言模型可以理解音频输入,并基于指令执行一系列与音频相关的任务,例如语音识别和音频字幕,其中指令通常是文本提示。音频语言模型大多从预训练的音频编码器和大型语言模型(LLM)初始化。虽然这些预先训练的组件是为了支持多种语言而开发的,但音频语言模型主要是在英语数据上训练的,这可能会将其可用性限制在英语指令或英语语音输入上。首先,本文以泰语为例,考察了现有音频语言模型在服务不足的语言中的性能。本文表明,尽管建立在多语言的骨干,音频语言模型不表现出跨语言的紧急能力,低资源的语言。其次,本文研究了开发音频语言模型的数据混合,这些模型针对目标语言以及英语进行了优化。另外。本文将音频理解和语音解释跟随能力集成到单个统一模型中。我们的实验提供了洞察数据混合,以提高在低资源的语言和英语的解释以下的能力。我们的模型Typhoon-Audio比现有的开源音频语言模型性能好得多,并且在英语和泰语中与最先进的Gemini-1.5-Pro相当。摘要:Audio language models can understand audio inputs and perform a range of audio-related tasks based on instructions, such as speech recognition and audio captioning, where the instructions are usually textual prompts. Audio language models are mostly initialized from pre-trained audio encoders and large language models (LLMs). Although these pre-trained components were developed to support multiple languages, audio-language models are trained predominantly on English data, which may limit their usability to only English instructions or English speech inputs. First, this paper examines the performance of existing audio language models in an underserved language using Thai as an example. This paper demonstrates that, despite being built on multilingual backbones, audio language models do not exhibit cross-lingual emergent abilities to low-resource languages. Second, this paper studies data mixture for developing audio language models that are optimized for a target language as well as English. In addition. this paper integrates audio comprehension and speech instruction-following capabilities into a single unified model. Our experiments provide insights into data mixture for enhancing instruction-following capabilities in both a low-resource language and English. Our model, Typhoon-Audio, outperforms existing open-source audio language models by a considerable margin, and it is comparable to state-of-the-art Gemini-1.5-Pro in both English and Thai languages.
【26】 Adaptive Large Language Models By Layerwise Attention Shortcuts
标题: 通过分层注意力快捷方式自适应大型语言模型
作者:Prateek Verma,Mert Pilanci
备注:6 pages, 3 figures
链接:点击下载PDF文件
摘要:Transformer架构是现代AI革命的支柱。然而,它们是基于简单地将相同的块堆叠成几十层,并从一个块到另一个块顺序地处理信息。在本文中,我们建议挑战这一点,并为类似LLM的设置引入自适应计算,这使得最终层能够通过注意力机制来关注所有中间层,从而引入计算 textbf{attention shortcuts}。因此,这些快捷方式可以使架构深度和上下文自适应。我们展示了四个不同的数据集,即声学令牌,自然语言和符号音乐,我们实现了类似GPT架构的卓越性能。我们通过注意力地图提供证据表明,模型可以跨层学习复杂的依赖关系,这些依赖关系在上下文和深度方面具有自适应性,具体取决于输入标记。摘要:Transformer architectures are the backbone of the modern AI revolution. However, they are based on simply stacking the same blocks in dozens of layers and processing information sequentially from one block to another. In this paper, we propose to challenge this and introduce adaptive computations for LLM-like setups, which allow the final layer to attend to all of the intermediate layers as it deems fit through the attention mechanism, thereby introducing computational textbf{attention shortcuts}. These shortcuts can thus make the architecture depth and context adaptive. We showcase four different datasets, namely acoustic tokens, natural language, and symbolic music, and we achieve superior performance for GPT-like architecture. We give evidence via attention maps that the models learn complex dependencies across layers that are adaptive in context and depth depending on the input tokens.
【27】 Speech Recognition for Analysis of Police Radio Communication
标题: 用于警用无线电通信分析的语音识别
作者:Tejes Srivastava,Ju-Chieh Chou,Priyank Shroff,Karen Livescu,Christopher Graziul
备注:Accepted by SLT 2024
链接:点击下载PDF文件
摘要:世界各地的警察部门使用双向无线电进行协调。这些广播警察通信(BPC)是关于日常警察活动和紧急反应的独特信息来源。然而,BPC不被转录,它们的自然主义音频特性使自动转录具有挑战性。我们收集了大约62,000个手动转录的无线电传输(约46小时的音频)的语料库,以评估使用现代识别模型进行自动语音识别(ASR)的可行性。我们评估了现成的语音识别器,模型微调的BPC数据和定制的端到端模型的性能。我们发现,人类和机器转录在这一领域是具有挑战性的。大型现成的ASR模型表现不佳,但经过微调的模型可以达到人类表现的近似范围。我们的工作为未来的工作提出了方向,包括分析警察无线电互动中的短话语和潜在的误解。我们将我们的语料库和数据注释管道提供给其他研究人员,以便进一步研究警察通信的识别和分析。摘要:Police departments around the world use two-way radio for coordination. These broadcast police communications (BPC) are a unique source of information about everyday police activity and emergency response. Yet BPC are not transcribed, and their naturalistic audio properties make automatic transcription challenging. We collect a corpus of roughly 62,000 manually transcribed radio transmissions (~46 hours of audio) to evaluate the feasibility of automatic speech recognition (ASR) using modern recognition models. We evaluate the performance of off-the-shelf speech recognizers, models fine-tuned on BPC data, and customized end-to-end models. We find that both human and machine transcription is challenging in this domain. Large off-the-shelf ASR models perform poorly, but fine-tuned models can reach the approximate range of human performance. Our work suggests directions for future work, including analysis of short utterances and potential miscommunication in police radio interactions. We make our corpus and data annotation pipeline available to other researchers, to enable further research on recognition and analysis of police communication.
【28】 3DFacePolicy: Speech-Driven 3D Facial Animation with Diffusion Policy
标题: 3DFacePolicy:具有扩散策略的语音驱动3D面部动画
作者:Xuanmeng Sha,Liyun Zhang,Tomohiro Mashita,Yuki Uranishi
链接:点击下载PDF文件
摘要:音频驱动的3D人脸动画在研究和应用开发方面都取得了令人身临其境的进展。最新的方法主要集中在基于变换器的方法和基于扩散的方法,然而,生成的动画与真实的人脸在生动性和情感表达方面仍然存在差距。为了解决这个问题,我们提出了3DFacePolicy,一个用于3D面部动画预测的扩散策略模型。该方法利用扩散策略预测三维人脸模板上的三维顶点轨迹,而不是逐帧生成人脸,从而生成可变的、逼真的人脸运动。它以音频和顶点状态作为观测值,预测顶点轨迹,模仿真实的人类面部表情,保持了人类情感的连续和自然流动。实验结果表明,该方法在可变动态人脸运动合成中是有效的。摘要:Audio-driven 3D facial animation has made immersive progress both in research and application developments. The newest approaches focus on Transformer-based methods and diffusion-based methods, however, there is still gap in the vividness and emotional expression between the generated animation and real human face. To tackle this limitation, we propose 3DFacePolicy, a diffusion policy model for 3D facial animation prediction. This method generates variable and realistic human facial movements by predicting the 3D vertex trajectory on the 3D facial template with diffusion policy instead of facial generation for every frame. It takes audio and vertex states as observations to predict the vertex trajectory and imitate real human facial expressions, which keeps the continuous and natural flow of human emotions. The experiments show that our approach is effective in variable and dynamic facial motion synthesizing.
【29】 PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
标题: PMX:用于符号音乐处理的大规模公共领域音乐ML数据集
作者:Phillip Long,Zachary Novack,Taylor Berg-Kirkpatrick,Julian McAuley
链接:点击下载PDF文件
摘要:最近生成式AI音乐系统的爆炸式增长引发了人们对数据版权、音乐家音乐许可以及开源AI与大型知名公司之间冲突的担忧。这些问题突出表明,需要公开的、无版权的音乐数据,而这方面的数据,特别是象征性的音乐数据,非常短缺。为了缓解这个问题,我们提出了PDMX:一个从分数共享论坛MuseScore收集的超过25万公共领域MusicXML分数的大规模开源数据集,使其成为我们所知的最大的无版权符号音乐数据集。PDmx还包括大量的标签和用户交互元数据,使我们能够有效地分析数据集并过滤高质量的用户生成的分数。由于我们的数据收集过程中提供的额外的元数据,我们进行多轨音乐生成实验,评估不同的代表性子集的PDMX如何导致下游模型中的不同行为,以及如何用户评级统计数据可以用作数据质量的有效措施。示例可在https: pnlong.github.io PDMX.demo 上找到。摘要:The recent explosion of generative AI-Music systems has raised numerous concerns over data copyright, licensing music from musicians, and the conflict between open-source AI and large prestige companies. Such issues highlight the need for publicly available, copyright-free musical data, in which there is a large shortage, particularly for symbolic music data. To alleviate this issue, we present PDMX: a large-scale open-source dataset of over 250K public domain MusicXML scores collected from the score-sharing forum MuseScore, making it the largest available copyright-free symbolic music dataset to our knowledge. PDMX additionally includes a wealth of both tag and user interaction metadata, allowing us to efficiently analyze the dataset and filter for high quality user-generated scores. Given the additional metadata afforded by our data collection process, we conduct multitrack music generation experiments evaluating how different representative subsets of PDMX lead to different behaviors in downstream models, and how user-rating statistics can be used as an effective measure of data quality. Examples can be found at https: pnlong.github.io PDMX.demo .
【30】 Mitigating Sex Bias in Audio Data-driven COPD and COVID-19 Breathing Pattern Detection Models
标题: 缓解音频数据驱动的COPD和COVID-19呼吸模式检测模型中的性别偏见
作者:Rachel Pfeifer,Sudip Vhaduri,James Eric Dietz
备注:Accepted at 2024 IEEE-EMBS International Conference on Body Sensor Networks (IEEE BSN 2024)
链接:点击下载PDF文件
摘要:在医疗保健行业,研究人员一直在开发机器学习模型,根据呼吸模式自动诊断呼吸系统疾病患者。然而,这些模型没有考虑人口统计学偏见,特别是性别偏见,这在使用倾斜的患者数据集训练模型时经常发生。因此,在这样一个重要的行业中,减少这种偏见是至关重要的,这样模型就可以做出公平的诊断。在这项工作中,我们研究了用于检测两种主要呼吸系统疾病的呼吸模式的模型中的偏差,即,慢性阻塞性肺疾病(COPD)和COVID-19。我们使用由29名COPD和680名COVID-19阳性患者组成的两个开源数据集获得的呼吸模式音频记录训练的决策树模型,分析了性别偏见对模型的影响。使用阈值优化器和两个约束(人口统计学奇偶性和均衡赔率)来减轻偏差,我们见证了81.43%(人口统计学奇偶性差异)和71.81%(均衡赔率差异)的改善。这些发现具有统计学意义。摘要:In the healthcare industry, researchers have been developing machine learning models to automate diagnosing patients with respiratory illnesses based on their breathing patterns. However, these models do not consider the demographic biases, particularly sex bias, that often occur when models are trained with a skewed patient dataset. Hence, it is essential in such an important industry to reduce this bias so that models can make fair diagnoses. In this work, we examine the bias in models used to detect breathing patterns of two major respiratory diseases, i.e., chronic obstructive pulmonary disease (COPD) and COVID-19. Using decision tree models trained with audio recordings of breathing patterns obtained from two open-source datasets consisting of 29 COPD and 680 COVID-19-positive patients, we analyze the effect of sex bias on the models. With a threshold optimizer and two constraints (demographic parity and equalized odds) to mitigate the bias, we witness 81.43% (demographic parity difference) and 71.81% (equalized odds difference) improvements. These findings are statistically significant.
【31】 Learning Co-Speech Gesture Representations in Dialogue through Contrastive Learning: An Intrinsic Evaluation
标题: 通过对比学习学习对话中的同语手势表示:内在评价
作者:Esam Ghaleb,Bulat Khaertdinov,Wim Pouw,Marlou Rasenberg,Judith Holler,Aslı Özyürek,Raquel Fernández
Journal-ref:INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION (ICMI 2024)
链接:点击下载PDF文件
摘要:在面对面的对话中,共语手势的形式-意义关系取决于语境因素,如手势所指的是什么和说话者的个人特征。这些因素使得协同语音手势表示学习具有挑战性。考虑到手势的可变性和与语音的关系,我们如何学习有意义的手势表示?本文通过采用自监督对比学习技术从骨骼和语音信息中学习手势表示来应对这一挑战。我们提出了一种方法,包括单模态和多模态预训练地面手势表示共现语音。为了训练,我们利用了一个面对面的对话数据集丰富的代表性的标志性手势。我们通过与人类注释的成对手势相似性进行比较,对所学习的表示进行彻底的内在评估。此外,我们进行了诊断探测分析,以评估从学习的表示恢复可解释的手势功能的可能性。我们的研究结果显示出显着的正相关性与人类注释的手势相似性,并揭示了学习表示之间的相似性是一致的动机良好的模式相关的动态对话互动。此外,我们的研究结果表明,一些功能的形式的手势可以恢复从潜在的表征。总的来说,这项研究表明,多模态对比学习是一种很有前途的方法,学习手势表示,这打开了大门,使用这种表示在大规模的手势分析研究。摘要:In face-to-face dialogues, the form-meaning relationship of co-speech gestures varies depending on contextual factors such as what the gestures refer to and the individual characteristics of speakers. These factors make co-speech gesture representation learning challenging. How can we learn meaningful gestures representations considering gestures' variability and relationship with speech? This paper tackles this challenge by employing self-supervised contrastive learning techniques to learn gesture representations from skeletal and speech information. We propose an approach that includes both unimodal and multimodal pre-training to ground gesture representations in co-occurring speech. For training, we utilize a face-to-face dialogue dataset rich with representational iconic gestures. We conduct thorough intrinsic evaluations of the learned representations through comparison with human-annotated pairwise gesture similarity. Moreover, we perform a diagnostic probing analysis to assess the possibility of recovering interpretable gesture features from the learned representations. Our results show a significant positive correlation with human-annotated gesture similarity and reveal that the similarity between the learned representations is consistent with well-motivated patterns related to the dynamics of dialogue interaction. Moreover, our findings demonstrate that several features concerning the form of gestures can be recovered from the latent representations. Overall, this study shows that multimodal contrastive learning is a promising approach for learning gesture representations, which opens the door to using such representations in larger-scale gesture analysis studies.
机器翻译,仅供参考
