本文经arXiv每日学术速递授权转载
【1】 Benchmarking Sub-Genre Classification For Mainstage Dance Music
标题: 主流舞台舞曲亚流派分类基准
作者:Hongzhi Shu,Xinglin Li,Hongyu Jiang,Minghao Fu,Xinyu Li
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【2】 LLaMA-Omni: Seamless Speech Interaction with Large Language Models
标题: LLaMA-Omni:与大型语言模型的无缝语音交互
作者:Qingkai Fang,Shoutao Guo,Yan Zhou,Zhengrui Ma,Shaolei Zhang,Yang Feng
备注:Preprint. Project: this https URL
链接:点击下载PDF文件
【3】 MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
标题: MoWE-音频:具有混合弱编码器的多任务音频LLM
作者:Wenyu Zhang,Shuo Sun,Bin Wang,Xunlong Zou,Zhuohan Liu,Yingxu He,Geyu Lin,Nancy F. Chen,Ai Ti Aw
链接:点击下载PDF文件
【4】 Sine, Transient, Noise Neural Modeling of Piano Notes
标题: 钢琴音符的弦、瞬时、噪音神经建模
作者:Riccardo Simionato,Stefano Fasciani
链接:点击下载PDF文件
【5】 An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition
标题: 一种有效的长尾语音识别上下文平衡自适应方法
作者:Yi-Cheng Wang,Li-Ting Pai,Bi-Cheng Yan,Hsin-Wei Wang,Chi-Han Lin,Berlin Chen
备注:Accepted by SLT 2024
链接:点击下载PDF文件
【6】 Attention-Based Beamformer For Multi-Channel Speech Enhancement
标题: 用于多通道语音增强的基于注意力的束形成器
作者:Jinglin Bai,Hao Li,Xueliang Zhang,Fei Chen
链接:点击下载PDF文件
【7】 Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
标题: 通过对比学习和扩散模型通过自然语言指导增强情感文本到语音的可控制性
作者:Xin Jing,Kun Zhou,Andreas Triantafyllopoulos,Björn W. Schuller
链接:点击下载PDF文件
【8】 Human-mimetic binaural ear design and sound source direction estimation for task realization of musculoskeletal humanoids
标题: 用于肌肉骨骼人形机器人任务实现的仿人双耳设计和声音方向估计
作者:Yusuke Omura,Kento Kawaharazuka,Yuya Nagamatsu,Yuya Koga,Manabu Nishiura,Yasunori Toshimitsu,Yuki Asano,Kei Okada,Koji Kawasaki,Masayuki Inaba
备注:Accepted at ROBOMECH Journal
链接:点击下载PDF文件
【9】 Soft Acoustic Curvature Sensor: Design and Development
标题: 软声学弯曲传感器:设计与开发
作者:Mohammad Sheikh Sofla,Hanita Golshanian,Vishnu Rajendran S,Amir Ghalamzan E
备注:To appear in Robotics and Automation Letter
链接:点击下载PDF文件
【10】 SpeechTaxi: On Multilingual Semantic Speech Classification
标题: SpeechTaxi:多语言语音语义分类
作者:Lennart Keller,Goran Glavaš
链接:点击下载PDF文件
【11】 VoiceWukong: Benchmarking Deepfake Voice Detection
标题: Voice悟空:Deepfake语音检测基准
作者:Ziwei Yan,Yanjie Zhao,Haoyu Wang
链接:点击下载PDF文件
【12】 An End-to-End Approach for Chord-Conditioned Song Generation
标题: 和弦条件歌曲生成的端到端方法
作者:Shuochen Gao,Shun Lei,Fan Zhuo,Hangyu Liu,Feng Liu,Boshi Tang,Qiaochu Huang,Shiyin Kang,Zhiyong Wu
链接:点击下载PDF文件
【13】 Spectral oversubtraction? An approach for speech enhancement after robot ego speech filtering in semi-real-time
标题: 光谱过度减法?半实时机器人自我语音过滤后的语音增强方法
作者:Yue Li,Koen V. Hindriks,Florian A. Kunneman
备注:6 pages, 2 figures, submitted to 2025 IEEE ICASSP
链接:点击下载PDF文件
【14】 A Two-Stage Band-Split Mamba-2 Network for Music Separation
标题: 用于音乐分离的两级带宽Mamba-2网络
作者:Jinglin Bai,Yuan Fang,Jiajie Wang,Xueliang Zhang
链接:点击下载PDF文件
【15】 RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
标题: RobustCSV:基于HuBERT的旋律提取器和对抗学习,用于稳健的歌唱声音转换
作者:Wei Chen,Xintao Zhao,Jun Chen,Binzhu Sha,Zhiwei Lin,Zhiyong Wu
备注:Accepted by ISCSLP 2024
链接:点击下载PDF文件
【16】 Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models
标题: 增强大型音频语言模型音频问题回答中的时间理解
作者:Arvind Krishna Sridhar,Yinyi Guo,Erik Visser
备注:5 pages, 3 figures
链接:点击下载PDF文件
【17】 Advancing Topic Segmentation of Broadcasted Speech with Multilingual Semantic Embeddings
标题: 利用多语言语义嵌入推进广播语音的主题分割
作者:Sakshi Deo Shukla,Pavel Denisov,Tugtekin Turan
链接:点击下载PDF文件
【18】 MTDA-HSED: Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection
标题: MTDA-HMED:用于异类声音事件检测的互助调谐和双分支聚集
作者:Zehao Wang,Haobo Yue,Zhicheng Zhang,Da Mu,Jin Tang,Jianqin Yin
备注:Submit to Icassp2025
链接:点击下载PDF文件
【19】 DENSE: Dynamic Embedding Causal Target Speech Extraction
标题: DENSE:动态嵌入因果目标语音提取
作者:Yiwen Wang,Zeyu Yuan,Xihong Wu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【20】 Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis
标题: 绘制音频:利用多指令进行视频到音频合成
作者:Qi Yang,Binjie Mao,Zili Wang,Xing Nie,Pengfei Gao,Ying Guo,Cheng Zhen,Pengfei Yan,Shiming Xiang
备注:14 pages, 11 figures
链接:点击下载PDF文件
【21】 Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
标题: 无监督音乐音频音色传输的潜在扩散桥梁
作者:Michele Mancusi,Yurii Halychansky,Kin Wai Cheuk,Chieh-Hsin Lai,Stefan Uhlich,Junghyun Koo,Marco A. Martínez-Ramírez,Wei-Hsiang Liao,Giorgio Fabbro,Yuhki Mitsufuji
链接:点击下载PDF文件
【22】 Investigating Causal Cues: Strengthening Spoofed Audio Detection with Human-Discernible Linguistic Features
标题: 调查因果线索:利用人类可辨别的语言特征加强欺骗音频检测
作者:Zahra Khanjani,Tolulope Ale,Jianwu Wang,Lavon Davis,Christine Mallinson,Vandana P. Janeja
链接:点击下载PDF文件
【23】 SongCreator: Lyrics-based Universal Song Generation
标题: SongCreator:基于歌词的通用歌曲一代
作者:Shun Lei,Yixuan Zhou,Boshi Tang,Max W. Y. Lam,Feng Liu,Hangyu Liu,Jingcheng Wu,Shiyin Kang,Zhiyong Wu,Helen Meng
备注:work in progress
链接:点击下载PDF文件
【24】 Musical Chords: A Novel Java Algorithm and App Utility to Enumerate Chord-Progressions Adhering to Music Theory Guidelines
标题: 音乐和弦:一种新颖的Java算法和应用程序实用程序,用于列举符合音乐理论准则的和弦进行
作者:Aditya Lakshminarasimhan
备注:10 Pages, 5 Figures
链接:点击下载PDF文件
【25】 Continuous Learning of Transformer-based Audio Deepfake Detection
标题: 基于变形器的音频深度伪造检测的持续学习
作者:Tuan Duy Nguyen Le,Kah Kuan Teh,Huy Dat Tran
备注:Submitted to INTERSPEECH 2024
链接:点击下载PDF文件
【26】 Sortformer: Seamless Integration of Speaker Diarization and ASR by Bridging Timestamps and Tokens
标题: 排序器:通过桥梁时间戳和令牌无缝集成说话者拨号和ASB
作者:Taejin Park,Ivan Medennikov,Kunal Dhawan,Weiqing Wang,He Huang,Nithin Rao Koluguri,Krishna C. Puvvada,Jagadeesh Balam,Boris Ginsburg
链接:点击下载PDF文件
【27】 Exploring Differences between Human Perception and Model Inference in Audio Event Recognition
标题: 探讨音频事件识别中人类感知和模型推理之间的差异
作者:Yizhou Tan,Yanru Wu,Yuanbo Hou,Xin Xu,Hui Bu,Shengchen Li,Dick Botteldooren,Mark D. Plumbley
备注:Dataset homepage: this https URL
链接:点击下载PDF文件
【28】 Janssen 2.0: Audio Inpainting in the Time-frequency Domain
标题: Janssen 2.0:时频域中的音频修复
作者:Ondřej Mokrý,Peter Balušík,Pavel Rajmic
链接:点击下载PDF文件
【29】 InstructSing: High-Fidelity Singing Voice Generation via Instructing Yourself
标题: DirectSing:通过指导自己产生高保真歌唱声音
作者:Chang Zeng,Chunhui Wang,Xiaoxiao Miao,Jian Zhao,Zhonglin Jiang,Yong Chen
备注:To appear in 2024 IEEE Spoken Language Technology Workshop, Dec 02-05, 2024, Macao, China
链接:点击下载PDF文件
【30】 Spoofing-Aware Speaker Verification Robust Against Domain and Channel Mismatches
标题: 针对域和通道不匹配的强有力的欺骗感知说话者验证
作者:Chang Zeng,Xiaoxiao Miao,Xin Wang,Erica Cooper,Junichi Yamagishi
备注:To appear in 2024 IEEE Spoken Language Technology Workshop, Dec 02-05, 2024, Macao, China
链接:点击下载PDF文件
【31】 Multi-Source Music Generation with Latent Diffusion
标题: 具有潜在扩散的多来源音乐生成
作者:Zhongweiyang Xu,Debottam Dutta,Yu-Lin Wei,Romit Roy Choudhury
备注:ICASSP 2025 in Submission
链接:点击下载PDF文件
【32】 DeWinder: Single-Channel Wind Noise Reduction using Ultrasound Sensing
标题: DeWinder:使用超声波传感降低单通道风噪音
作者:Kuang Yuan,Shuo Han,Swarun Kumar,Bhiksha Raj
链接:点击下载PDF文件
【33】 VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion
标题: VC-ENHANCE:具有集成噪音抑制和语音转换的语音恢复
作者:Kyungguen Byun,Jason Filos,Erik Visser,Sunkuk Moon
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【34】 Retrieval Augmented Correction of Named Entity Speech Recognition Errors
标题: 命名实体语音识别错误的检索增强纠正
作者:Ernest Pusateri,Anmol Walia,Anirudh Kashi,Bortik Bandyopadhyay,Nadia Hyder,Sayantan Mahinder,Raviteja Anantha,Daben Liu,Sashank Gondala
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
标题: 排序器:通过桥梁时间戳和令牌无缝集成说话者拨号和ASB
作者:Taejin Park,Ivan Medennikov,Kunal Dhawan,Weiqing Wang,He Huang,Nithin Rao Koluguri,Krishna C. Puvvada,Jagadeesh Balam,Boris Ginsburg
链接:点击下载PDF文件
【2】 Exploring Differences between Human Perception and Model Inference in Audio Event Recognition
标题: 探讨音频事件识别中人类感知和模型推理之间的差异
作者:Yizhou Tan,Yanru Wu,Yuanbo Hou,Xin Xu,Hui Bu,Shengchen Li,Dick Botteldooren,Mark D. Plumbley
备注:Dataset homepage: this https URL
链接:点击下载PDF文件
【3】 Janssen 2.0: Audio Inpainting in the Time-frequency Domain
标题: Janssen 2.0:时频域中的音频修复
作者:Ondřej Mokrý,Peter Balušík,Pavel Rajmic
链接:点击下载PDF文件
【4】 InstructSing: High-Fidelity Singing Voice Generation via Instructing Yourself
标题: DirectSing:通过指导自己产生高保真歌唱声音
作者:Chang Zeng,Chunhui Wang,Xiaoxiao Miao,Jian Zhao,Zhonglin Jiang,Yong Chen
备注:To appear in 2024 IEEE Spoken Language Technology Workshop, Dec 02-05, 2024, Macao, China
链接:点击下载PDF文件
【5】 Spoofing-Aware Speaker Verification Robust Against Domain and Channel Mismatches
标题: 针对域和通道不匹配的强有力的欺骗感知说话者验证
作者:Chang Zeng,Xiaoxiao Miao,Xin Wang,Erica Cooper,Junichi Yamagishi
备注:To appear in 2024 IEEE Spoken Language Technology Workshop, Dec 02-05, 2024, Macao, China
链接:点击下载PDF文件
【6】 Multi-Source Music Generation with Latent Diffusion
标题: 具有潜在扩散的多来源音乐生成
作者:Zhongweiyang Xu,Debottam Dutta,Yu-Lin Wei,Romit Roy Choudhury
备注:ICASSP 2025 in Submission
链接:点击下载PDF文件
【7】 DeWinder: Single-Channel Wind Noise Reduction using Ultrasound Sensing
标题: DeWinder:使用超声波传感降低单通道风噪音
作者:Kuang Yuan,Shuo Han,Swarun Kumar,Bhiksha Raj
链接:点击下载PDF文件
【8】 VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion
标题: VC-ENHANCE:具有集成噪音抑制和语音转换的语音恢复
作者:Kyungguen Byun,Jason Filos,Erik Visser,Sunkuk Moon
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【9】 Retrieval Augmented Correction of Named Entity Speech Recognition Errors
标题: 命名实体语音识别错误的检索增强纠正
作者:Ernest Pusateri,Anmol Walia,Anirudh Kashi,Bortik Bandyopadhyay,Nadia Hyder,Sayantan Mahinder,Raviteja Anantha,Daben Liu,Sashank Gondala
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【10】 Property Neurons in Self-Supervised Speech Transformers
标题: 自我监督语音Transformer中的属性神经元
作者:Tzu-Quan Lin,Guan-Ting Lin,Hung-yi Lee,Hao Tang
备注:Accepted by SLT 2024
链接:点击下载PDF文件
【11】 LLaMA-Omni: Seamless Speech Interaction with Large Language Models
标题: LLaMA-Omni:与大型语言模型的无缝语音交互
作者:Qingkai Fang,Shoutao Guo,Yan Zhou,Zhengrui Ma,Shaolei Zhang,Yang Feng
备注:Preprint. Project: this https URL
链接:点击下载PDF文件
【12】 MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
标题: MoWE-音频:具有混合弱编码器的多任务音频LLM
作者:Wenyu Zhang,Shuo Sun,Bin Wang,Xunlong Zou,Zhuohan Liu,Yingxu He,Geyu Lin,Nancy F. Chen,Ai Ti Aw
链接:点击下载PDF文件
【13】 Sine, Transient, Noise Neural Modeling of Piano Notes
标题: 钢琴音符的弦、瞬时、噪音神经建模
作者:Riccardo Simionato,Stefano Fasciani
链接:点击下载PDF文件
【14】 An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition
标题: 一种有效的长尾语音识别上下文平衡自适应方法
作者:Yi-Cheng Wang,Li-Ting Pai,Bi-Cheng Yan,Hsin-Wei Wang,Chi-Han Lin,Berlin Chen
备注:Accepted by SLT 2024
链接:点击下载PDF文件
标题: 主流舞台舞曲亚流派分类基准
作者:Hongzhi Shu,Xinglin Li,Hongyu Jiang,Minghao Fu,Xinyu Li
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:音乐分类是音乐信息检索中最突出的任务之一,有着广泛的应用。为了解决缺乏全面的数据集和高性能的方法在主舞台舞曲的分类,这项工作引入了一个新的基准,包括一个新的数据集和基线。我们的数据集扩展了子流派的数量,以涵盖全球顶级DJ在音乐节上的最新主舞台现场演出。一个连续的软标签的方法是占轨道,跨越多个子流派,保持固有的复杂性。对于基线,我们开发了优于当前最先进的多模型语言模型的深度学习模型,这些模型很难识别家庭音乐子流派,强调需要在细粒度数据集上训练的专用模型。我们的基准适用于服务于音乐推荐、DJ集策展、互动多媒体等应用场景,并提供视频演示。我们的代码位于 url{https: anonymous.4open.science r Mainstage-EDM-Benchmark }。摘要:Music classification, with a wide range of applications, is one of the most prominent tasks in music information retrieval. To address the absence of comprehensive datasets and high-performing methods in the classification of mainstage dance music, this work introduces a novel benchmark comprising a new dataset and a baseline. Our dataset extends the number of sub-genres to cover most recent mainstage live sets by top DJs worldwide in music festivals. A continuous soft labeling approach is employed to account for tracks that span multiple sub-genres, preserving the inherent sophistication. For the baseline, we developed deep learning models that outperform current state-of-the-art multimodel language models, which struggle to identify house music sub-genres, emphasizing the need for specialized models trained on fine-grained datasets. Our benchmark is applicable to serve for application scenarios such as music recommendation, DJ set curation, and interactive multimedia, where we also provide video demos. Our code is on url{https: anonymous.4open.science r Mainstage-EDM-Benchmark }.
【2】 LLaMA-Omni: Seamless Speech Interaction with Large Language Models
标题: LLaMA-Omni:与大型语言模型的无缝语音交互
作者:Qingkai Fang,Shoutao Guo,Yan Zhou,Zhengrui Ma,Shaolei Zhang,Yang Feng
备注:Preprint. Project: this https URL
链接:点击下载PDF文件
摘要:像GPT-4 o这样的模型可以通过语音与大型语言模型(LLM)进行实时交互,与传统的基于文本的交互相比,显著增强了用户体验。然而,如何构建基于开源LLM的语音交互模型,目前还缺乏探索。为了解决这个问题,我们提出了LLaMA-Omni,一种新的模型架构,设计用于与LLM进行低延迟和高质量的语音交互。LLaMA-Omni集成了一个预训练的语音编码器,一个语音适配器,一个LLM和一个流式语音解码器。它消除了对语音转录的需要,并且可以以极低的延迟直接从语音指令同时生成文本和语音响应。我们基于最新的Llama-3.1-8B-Instruct模型构建我们的模型。为了使模型与语音交互场景相匹配,我们构建了一个名为InstructS 2S-200 K的数据集,其中包括200 K语音指令和相应的语音响应。实验结果表明,与以往的语音语言模型相比,LLaMA-Omni提供了更好的响应在内容和风格,响应延迟低至226毫秒。此外,在4个GPU上训练LLaMA-Omni需要不到3天的时间,为未来高效开发语音语言模型铺平了道路。摘要:Models like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how to build speech interaction models based on open-source LLMs. To address this, we propose LLaMA-Omni, a novel model architecture designed for low-latency and high-quality speech interaction with LLMs. LLaMA-Omni integrates a pretrained speech encoder, a speech adaptor, an LLM, and a streaming speech decoder. It eliminates the need for speech transcription, and can simultaneously generate text and speech responses directly from speech instructions with extremely low latency. We build our model based on the latest Llama-3.1-8B-Instruct model. To align the model with speech interaction scenarios, we construct a dataset named InstructS2S-200K, which includes 200K speech instructions and corresponding speech responses. Experimental results show that compared to previous speech-language models, LLaMA-Omni provides better responses in both content and style, with a response latency as low as 226ms. Additionally, training LLaMA-Omni takes less than 3 days on just 4 GPUs, paving the way for the efficient development of speech-language models in the future.
【3】 MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
标题: MoWE-音频:具有混合弱编码器的多任务音频LLM
作者:Wenyu Zhang,Shuo Sun,Bin Wang,Xunlong Zou,Zhuohan Liu,Yingxu He,Geyu Lin,Nancy F. Chen,Ai Ti Aw
链接:点击下载PDF文件
摘要:大型语言模型(LLM)的快速发展显着增强了自然语言处理能力,促进了AudioLLM的开发,它可以处理和理解语音和音频输入以及文本。现有的AudioLLM通常将预训练的音频编码器与预训练的LLM相结合,随后在特定的音频任务上进行微调。然而,预先训练的音频编码器具有有限的能力来捕获新任务和数据集的特征。为了解决这个问题,我们建议将混合的“弱”编码器(MoWE)到AudioLLM框架。MoWE用相对较轻权重的编码器池补充基本编码器,基于音频输入选择性地激活以增强特征提取而不显著增加模型大小。我们的实证结果表明,MoWE有效地提高了多任务性能,扩大了AudioLLM的适用性,以更多样化的音频任务。摘要:The rapid advancements in large language models (LLMs) have significantly enhanced natural language processing capabilities, facilitating the development of AudioLLMs that process and understand speech and audio inputs alongside text. Existing AudioLLMs typically combine a pre-trained audio encoder with a pre-trained LLM, which are subsequently finetuned on specific audio tasks. However, the pre-trained audio encoder has constrained capacity to capture features for new tasks and datasets. To address this, we propose to incorporate mixtures of weak' encoders (MoWE) into the AudioLLM framework. MoWE supplements a base encoder with a pool of relatively light weight encoders, selectively activated based on the audio input to enhance feature extraction without significantly increasing model size. Our empirical results demonstrate that MoWE effectively improves multi-task performance, broadening the applicability of AudioLLMs to more diverse audio tasks.
【4】 Sine, Transient, Noise Neural Modeling of Piano Notes
标题: 钢琴音符的弦、瞬时、噪音神经建模
作者:Riccardo Simionato,Stefano Fasciani
链接:点击下载PDF文件
摘要:介绍了一种新颖的钢琴音色仿真方法。我们建议利用正弦,瞬态,和噪声分解设计一个微分频谱建模合成器复制钢琴音符。三个子模块从钢琴录音中学习这些分量,并生成相应的谐波、瞬态和噪声信号。将仿真分为三个独立的可训练模型,降低了建模任务的复杂性。准谐波内容是使用由物理导出的公式指导的可微分正弦模型产生的,其参数是从音频记录中自动估计的。噪声子模块使用可学习的时变滤波器,瞬态使用深度卷积网络生成。从单个音符开始,我们使用基于卷积的网络模拟毛键中不同键之间的耦合。结果表明,该模型匹配的部分分布的目标,而预测的能量在较高的部分的频谱提出了更多的挑战。瞬态和噪声分量的频谱中的能量分布总体上是准确的。虽然该模型在计算和内存效率上更高,但感知测试显示,在准确建模音符的攻击阶段方面存在局限性。尽管如此,它通常在模拟单音符和三弦琴时达到感知准确性。摘要:This paper introduces a novel method for emulating piano sounds. We propose to exploit the sine, transient, and noise decomposition to design a differentiable spectral modeling synthesizer replicating piano notes. Three sub-modules learn these components from piano recordings and generate the corresponding harmonic, transient, and noise signals. Splitting the emulation into three independently trainable models reduces the modeling tasks' complexity. The quasi-harmonic content is produced using a differentiable sinusoidal model guided by physics-derived formulas, whose parameters are automatically estimated from audio recordings. The noise sub-module uses a learnable time-varying filter, and the transients are generated using a deep convolutional network. From singular notes, we emulate the coupling between different keys in trichords with a convolutional-based network. Results show the model matches the partial distribution of the target while predicting the energy in the higher part of the spectrum presents more challenges. The energy distribution in the spectra of the transient and noise components is accurate overall. While the model is more computationally and memory efficient, perceptual tests reveal limitations in accurately modeling the attack phase of notes. Despite this, it generally achieves perceptual accuracy in emulating single notes and trichords.
【5】 An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition
标题: 一种有效的长尾语音识别上下文平衡自适应方法
作者:Yi-Cheng Wang,Li-Ting Pai,Bi-Cheng Yan,Hsin-Wei Wang,Chi-Han Lin,Berlin Chen
备注:Accepted by SLT 2024
链接:点击下载PDF文件
摘要:端到端(E2 E)自动语音识别(ASR)模型已经成为各种商业应用的标准实践。然而,在现实世界中,单词分布的长尾性质通常导致E2 E ASR模型在常见单词上表现良好,但在识别不常见单词方面表现不佳。最近,上下文适配器(CA)的概念被提出来注入外部知识表示的上下文单词列表到E2 E ASR模型。虽然CA可以提高对稀有词的识别性能,但仍然存在两个关键的数据不平衡问题。首先,当在训练期间使用低频词作为上下文词时,由于这些词很少出现在话语中,CA在关注令牌摘要:End-to-end (E2E) automatic speech recognition (ASR) models have become standard practice for various commercial applications. However, in real-world scenarios, the long-tailed nature of word distribution often leads E2E ASR models to perform well on common words but fall short in recognizing uncommon ones. Recently, the notion of a contextual adapter (CA) was proposed to infuse external knowledge represented by a context word list into E2E ASR models. Although CA can improve recognition performance on rare words, two crucial data imbalance problems remain. First, when using low-frequency words as context words during training, since these words rarely occur in the utterance, CA becomes prone to overfit on attending to the
标题: 用于多通道语音增强的基于注意力的束形成器
作者:Jinglin Bai,Hao Li,Xueliang Zhang,Fei Chen
链接:点击下载PDF文件
摘要:最小方差无失真响应(MVDR)是一种经典的自适应波束形成器,理论上可以保证信号在目标方向上的无失真传输。它的降噪性能实际上取决于噪声空间协方差矩阵(SCM)估计的准确性。尽管最近的深度学习在多通道语音增强方面表现出了卓越的性能,但无失真响应的特性仍然使MVDR在实际应用中非常受欢迎。在本文中,我们提出了一种基于注意力的机制来计算语音和噪声SCM,然后应用MVDR来获得增强的语音。此外,使用原地卷积算子和频率无关LSTM的深度学习架构已被证明在促进SCM估计方面是有效的。该模型以端到端的方式进行优化。实验结果表明,该方法是非常有效的跟踪运动或静止的说话人在非因果和因果条件下,优于其他基线。值得一提的是,我们的模型只有35万个参数,很容易部署在边缘设备上。摘要:Minimum Variance Distortionless Response (MVDR) is a classical adaptive beamformer that theoretically ensures the distortionless transmission of signals in the target direction. Its performance in noise reduction actually depends on the accuracy of the noise spatial covariance matrix (SCM) estimate. Although recent deep learning has shown remarkable performance in multi-channel speech enhancement, the property of distortionless response still makes MVDR highly popular in real applications. In this paper, we propose an attention-based mechanism to calculate the speech and noise SCM and then apply MVDR to obtain the enhanced speech. Moreover, a deep learning architecture using the inplace convolution operator and frequency-independent LSTM has proven effective in facilitating SCM estimation. The model is optimized in an end-to-end manner. Experimental results indicate that the proposed method is extremely effective in tracking moving or stationary speakers under non-causal and causal conditions, outperforming other baselines. It is worth mentioning that our model has only 0.35 million parameters, making it easy to be deployed on edge devices.
【7】 Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
标题: 通过对比学习和扩散模型通过自然语言指导增强情感文本到语音的可控制性
作者:Xin Jing,Kun Zhou,Andreas Triantafyllopoulos,Björn W. Schuller
链接:点击下载PDF文件
摘要:虽然目前的情感文本到语音(TTS)系统可以生成高度可理解的情感语音,实现输出语音的情感渲染的精细控制仍然是一个重大的挑战。在本文中,我们介绍了ParaEVITS,一种新的情感TTS框架,利用自然语言的组合性,以提高对情感渲染的控制。通过结合一个文本音频编码器的启发ParaCLAP,一个对比语言音频预训练(CLAP)模型计算语言学,扩散模型进行训练,以生成基于文本情感风格描述的情感嵌入。我们的框架首先使用音频编码器对参考音频进行训练,然后微调扩散模型以处理来自ParaCLAP文本编码器的文本输入。在推理过程中,语音属性,如音高,抖动和响度操纵只使用文本条件。我们的实验表明,ParaEVITS有效地控制情感渲染,而不影响语音质量。演讲演示是公开的。摘要:While current emotional text-to-speech (TTS) systems can generate highly intelligible emotional speech, achieving fine control over emotion rendering of the output speech still remains a significant challenge. In this paper, we introduce ParaEVITS, a novel emotional TTS framework that leverages the compositionality of natural language to enhance control over emotional rendering. By incorporating a text-audio encoder inspired by ParaCLAP, a contrastive language-audio pretraining (CLAP) model for computational paralinguistics, the diffusion model is trained to generate emotional embeddings based on textual emotional style descriptions. Our framework first trains on reference audio using the audio encoder, then fine-tunes a diffusion model to process textual inputs from ParaCLAP's text encoder. During inference, speech attributes such as pitch, jitter, and loudness are manipulated using only textual conditioning. Our experiments demonstrate that ParaEVITS effectively control emotion rendering without compromising speech quality. Speech demos are publicly available.
【8】 Human-mimetic binaural ear design and sound source direction estimation for task realization of musculoskeletal humanoids
标题: 用于肌肉骨骼人形机器人任务实现的仿人双耳设计和声音方向估计
作者:Yusuke Omura,Kento Kawaharazuka,Yuya Nagamatsu,Yuya Koga,Manabu Nishiura,Yasunori Toshimitsu,Yuki Asano,Kei Okada,Koji Kawasaki,Masayuki Inaba
备注:Accepted at ROBOMECH Journal
链接:点击下载PDF文件
摘要:肌肉骨骼类人机器人的类人环境识别对于在真实复杂环境中实现任务和用作测试对象的假人是重要的。人类整合各种感官信息来感知周围环境,听觉对于识别视线外或触摸不到的物体特别有用。在本研究中,我们的目标是实现类人听觉环境识别和任务实现的肌肉骨骼类人通过配备一个类似人类的听觉处理系统。人类通过估计声源的方向并基于传入声音的时域和频域的变化以及听觉信息在中枢神经系统中的整合来检测环境声音,从而实现基于声音的环境识别。我们提出了一个仿人听觉信息处理系统,它由三个部分组成:仿人双耳耳单元,它模仿人耳的结构和特性,声源方向估计系统,和环境声音检测系统,它模仿在中枢神经系统中的处理。我们将其应用于Musashi,一个人类模仿肌肉骨骼类人动物,并让它执行任务,需要在真实的嘈杂环境中的视图之外的声音信息,以确认所提出的方法的有用性。摘要:Human-like environment recognition by musculoskeletal humanoids is important for task realization in real complex environments and for use as dummies for test subjects. Humans integrate various sensory information to perceive their surroundings, and hearing is particularly useful for recognizing objects out of view or out of touch. In this research, we aim to realize human-like auditory environmental recognition and task realization for musculoskeletal humanoids by equipping them with a human-like auditory processing system. Humans realize sound-based environmental recognition by estimating directions of the sound sources and detecting environmental sounds based on changes in the time and frequency domain of incoming sounds and the integration of auditory information in the central nervous system. We propose a human mimetic auditory information processing system, which consists of three components: the human mimetic binaural ear unit, which mimics human ear structure and characteristics, the sound source direction estimation system, and the environmental sound detection system, which mimics processing in the central nervous system. We apply it to Musashi, a human mimetic musculoskeletal humanoid, and have it perform tasks that require sound information outside of view in real noisy environments to confirm the usefulness of the proposed methods.
【9】 Soft Acoustic Curvature Sensor: Design and Development
标题: 软声学弯曲传感器:设计与开发
作者:Mohammad Sheikh Sofla,Hanita Golshanian,Vishnu Rajendran S,Amir Ghalamzan E
备注:To appear in Robotics and Automation Letter
链接:点击下载PDF文件
摘要:介绍了一种新型的软声曲率传感器。SAC集成了音频组件,并在灵活的结构中具有声学通道。由通道一端的扬声器产生的参考声波传播并由另一通道端的麦克风接收。我们以前的研究表明,声波能量耗散随声道变形而变化,这使我们设计了一种能够因弯曲而大变形的新型声道。然后,我们使用机器学习(ML)模型来建立通道变形和声音调制之间的复杂映射。各种声音频率和ML模型进行了评估,以提高曲率检测精度。该传感器采用软材料和3D打印制成,经过实验验证,在0至60 m-1曲率范围内,曲率测量误差保持在3.5 m-1以内。这些结果证明了所提出的方法估计曲率的有效性。由于其灵活的结构,SAC传感器在软机器人中具有应用潜力,包括连续体机械手,软夹持器和可穿戴设备的形状测量。摘要:This paper introduces a novel Soft Acoustic Curvature (SAC) sensor. SAC incorporates integrated audio components and features an acoustic channel within a flexible structure. A reference acoustic wave, generated by a speaker at one end of the channel, propagates and is received by a microphone at the other channel's end. Our previous study revealed that acoustic wave energy dissipation varies with acoustic channel deformation, leading us to design a novel channel capable of large deformation due to bending. We then use Machine Learning (ML) models to establish a complex mapping between channel deformations and sound modulation. Various sound frequencies and ML models were evaluated to enhance curvature detection accuracy. The sensor, constructed using soft material and 3D printing, was validated experimentally, with curvature measurement errors remaining within 3.5 m-1 for a range of 0 to 60 m-1 curvatures. These results demonstrate the effectiveness of the proposed method for estimating curvatures. With its flexible structure, the SAC sensor holds potential for applications in soft robotics, including shape measurement for continuum manipulators, soft grippers, and wearable devices.
【10】 SpeechTaxi: On Multilingual Semantic Speech Classification
标题: SpeechTaxi:多语言语音语义分类
作者:Lennart Keller,Goran Glavaš
链接:点击下载PDF文件
摘要:多语言语音编码和转录的最新进展提出了最有效的语义语音分类方法的问题。具体而言,(1)通过微调最先进的多语言语音编码器(MSE)获得的端到端(E2 E)分类器是否可以匹配或超过(2)级联(CA)的性能,其中语音首先被转录为文本,并将分类委托给基于文本的分类器。为了回答这个问题,我们首先构建了SpeechTaxi,这是一个80小时的多语言数据集,用于圣经经文的语义语音分类,涵盖28种不同的语言。然后,我们利用SpeechTaxi进行了广泛的实验,比较E2 E和CA在单语语义语音分类以及跨语言迁移。我们发现,基于MSE的E2 E在单语设置中优于CA,即,在语言数据上训练时。然而,MSE似乎具有较差的跨语言迁移能力,E2 E在(1)zero-shot迁移到训练中未见过的语言和(2)多语言训练(即,多语种联合培训。最后,我们设计了一种新的CA方法的基础上转录罗马化文本作为一种语言不可知的中间表示,并表明它代表了一个强大的解决方案,没有本地ASR支持的语言。我们的SpeechTaxi数据集可在以下网址公开获取:https: huggingface.co datasets LennartKeller SpeechTaxi 。摘要:Recent advancements in multilingual speech encoding as well as transcription raise the question of the most effective approach to semantic speech classification. Concretely, can (1) end-to-end (E2E) classifiers obtained by fine-tuning state-of-the-art multilingual speech encoders (MSEs) match or surpass the performance of (2) cascading (CA), where speech is first transcribed into text and classification is delegated to a text-based classifier. To answer this, we first construct SpeechTaxi, an 80-hour multilingual dataset for semantic speech classification of Bible verses, covering 28 diverse languages. We then leverage SpeechTaxi to conduct a wide range of experiments comparing E2E and CA in monolingual semantic speech classification as well as in cross-lingual transfer. We find that E2E based on MSEs outperforms CA in monolingual setups, i.e., when trained on in-language data. However, MSEs seem to have poor cross-lingual transfer abilities, with E2E substantially lagging CA both in (1) zero-shot transfer to languages unseen in training and (2) multilingual training, i.e., joint training on multiple languages. Finally, we devise a novel CA approach based on transcription to Romanized text as a language-agnostic intermediate representation and show that it represents a robust solution for languages without native ASR support. Our SpeechTaxi dataset is publicly available at: https: huggingface.co datasets LennartKeller SpeechTaxi .
【11】 VoiceWukong: Benchmarking Deepfake Voice Detection
标题: Voice悟空:Deepfake语音检测基准
作者:Ziwei Yan,Yanjie Zhao,Haoyu Wang
链接:点击下载PDF文件
摘要:随着文本到语音(TTS)和语音转换(VC)等技术的快速发展,检测深度伪造语音变得越来越重要。然而,学术界和工业界都缺乏一个全面和直观的基准来评估探测器。现有的数据集在语言多样性方面受到限制,并且缺乏在现实世界的生产环境中遇到的许多操作。 为了填补这一空白,我们提出了VoiceWukong,这是一个旨在评估deepfake语音检测器性能的基准测试。为了构建数据集,我们首先收集了由19种先进且广泛认可的商业工具和15种开源工具生成的deepfake语音。然后,我们创建了38个数据变量,涵盖了六种类型的操作,构建了用于deepfake语音检测的评估数据集。因此,VoiceWukong包括265,200个英语和148,200个中文deepfake语音样本。使用VoiceWukong,我们评估了12个最先进的探测器。AASIST 2的最佳等误率(EER)为13.50%,而其他所有算法均超过20%。我们的研究结果表明,这些探测器在现实世界的应用中面临着巨大的挑战,性能急剧下降。此外,我们还进行了一项有300多名参与者的用户研究。将结果与12个检测器和多模型大语言模型(MLLM)的性能进行比较,即,Qwen 2-Audio,其中不同的检测器和人类在不同的欺骗水平下表现出不同的识别能力,而LALM则完全没有检测能力。此外,我们还提供了deepfake语音检测的排行榜,可在{https: www.example.com}上公开获取。voicewukong.github.io摘要:With the rapid advancement of technologies like text-to-speech (TTS) and voice conversion (VC), detecting deepfake voices has become increasingly crucial. However, both academia and industry lack a comprehensive and intuitive benchmark for evaluating detectors. Existing datasets are limited in language diversity and lack many manipulations encountered in real-world production environments. To fill this gap, we propose VoiceWukong, a benchmark designed to evaluate the performance of deepfake voice detectors. To build the dataset, we first collected deepfake voices generated by 19 advanced and widely recognized commercial tools and 15 open-source tools. We then created 38 data variants covering six types of manipulations, constructing the evaluation dataset for deepfake voice detection. VoiceWukong thus includes 265,200 English and 148,200 Chinese deepfake voice samples. Using VoiceWukong, we evaluated 12 state-of-the-art detectors. AASIST2 achieved the best equal error rate (EER) of 13.50%, while all others exceeded 20%. Our findings reveal that these detectors face significant challenges in real-world applications, with dramatically declining performance. In addition, we conducted a user study with more than 300 participants. The results are compared with the performance of the 12 detectors and a multimodel large language model (MLLM), i.e., Qwen2-Audio, where different detectors and humans exhibit varying identification capabilities for deepfake voices at different deception levels, while the LALM demonstrates no detection ability at all. Furthermore, we provide a leaderboard for deepfake voice detection, publicly available at {https: voicewukong.github.io}.
【12】 An End-to-End Approach for Chord-Conditioned Song Generation
标题: 和弦条件歌曲生成的端到端方法
作者:Shuochen Gao,Shun Lei,Fan Zhuo,Hangyu Liu,Feng Liu,Boshi Tang,Qiaochu Huang,Shiyin Kang,Zhiyong Wu
链接:点击下载PDF文件
摘要:歌曲生成任务的目标是从给定的歌词合成由人声和伴奏组成的音乐。虽然现有的方法Jukebox已经探索了这一任务,但其对世代的限制控制往往导致音乐表现的不足。为了缓解这个问题,我们引入了一个重要的概念,从音乐创作,即和弦,歌曲生成网络。和弦构成了伴奏的基础,并为声乐旋律提供了相关的和声。针对自动和弦提取器的不准确性,设计了一种基于动态权值序列的交叉注意机制,将提取的和弦信息整合到歌曲生成中,减少了帧级错误,并在此基础上提出了一种新的和弦条件歌曲生成器模型CSG,实验结果表明,该方法在音乐性能和生成歌曲的控制精度方面优于其他方法.摘要:The Song Generation task aims to synthesize music composed of vocals and accompaniment from given lyrics. While the existing method, Jukebox, has explored this task, its constrained control over the generations often leads to deficiency in music performance. To mitigate the issue, we introduce an important concept from music composition, namely chords, to song generation networks. Chords form the foundation of accompaniment and provide vocal melody with associated harmony. Given the inaccuracy of automatic chord extractors, we devise a robust cross-attention mechanism augmented with dynamic weight sequence to integrate extracted chord information into song generations and reduce frame-level flaws, and propose a novel model termed Chord-Conditioned Song Generator (CSG) based on it. Experimental evidence demonstrates our proposed method outperforms other approaches in terms of musical performance and control precision of generated songs.
【13】 Spectral oversubtraction? An approach for speech enhancement after robot ego speech filtering in semi-real-time
标题: 光谱过度减法?半实时机器人自我语音过滤后的语音增强方法
作者:Yue Li,Koen V. Hindriks,Florian A. Kunneman
备注:6 pages, 2 figures, submitted to 2025 IEEE ICASSP
链接:点击下载PDF文件
摘要:谱减法,广泛使用的简单,已被用来解决机器人自我语音过滤(RESF)的问题,从机器人的单通道麦克风录音时,它正在说话的人中断检测语音内容。然而,这种方法遭受在基频范围(FFR)中的过减法,从而导致语音内容识别降级。为了解决这个问题,我们提出了一个基于双掩码一致性的度量生成对抗网络(CMGAN),以增强检测到的语音,提高识别结果。我们的模型用高频信息和长期特征补偿减影过多的血流储备分数值,然后对新的频谱图进行去噪。此外,我们还介绍了一种增量处理方法,该方法允许在长固定长度输入上训练的网络上使用流式输入进行半实时音频处理。两个数据集的评估,包括一个看不见的噪音,证明显着提高识别准确性和有效性的建议的双掩模方法和增量处理,提高鲁棒性的建议RESF管道在现实世界的HRI场景。摘要:Spectral subtraction, widely used for its simplicity, has been employed to address the Robot Ego Speech Filtering (RESF) problem for detecting speech contents of human interruption from robot's single-channel microphone recordings when it is speaking. However, this approach suffers from oversubtraction in the fundamental frequency range (FFR), leading to degraded speech content recognition. To address this, we propose a Two-Mask Conformer-based Metric Generative Adversarial Network (CMGAN) to enhance the detected speech and improve recognition results. Our model compensates for oversubtracted FFR values with high-frequency information and long-term features and then de-noises the new spectrogram. In addition, we introduce an incremental processing method that allows semi-real-time audio processing with streaming input on a network trained on long fixed-length input. Evaluations of two datasets, including one with unseen noise, demonstrate significant improvements in recognition accuracy and the effectiveness of the proposed two-mask approach and incremental processing, enhancing the robustness of the proposed RESF pipeline in real-world HRI scenarios.
【14】 A Two-Stage Band-Split Mamba-2 Network for Music Separation
标题: 用于音乐分离的两级带宽Mamba-2网络
作者:Jinglin Bai,Yuan Fang,Jiajie Wang,Xueliang Zhang
链接:点击下载PDF文件
摘要:音乐源分离(MSS)旨在将混合音乐分离成不同的音轨,如人声,低音,鼓等。由于音乐信号的复杂性,MSS被认为是一项具有挑战性的音频分离任务。虽然RNN和Transformer架构并不完美,但它们通常用于为MSS建模音乐序列。最近,Mamba-2已经在各种顺序建模任务中表现出了高效率,但其优越性尚未在MSS中得到研究。本文应用Mamba-2算法,采用两阶段策略,在掩模方法的基础上引入残差映射,有效地补偿了掩模中缺失的细节,进一步提高了分离性能。实验证明了双向Mamba-2的优越性和两级网络在MSS中的有效性。源代码可在https: github.com baijinglin TS-BSmamba2上公开访问。摘要:Music source separation (MSS) aims to separate mixed music into its distinct tracks, such as vocals, bass, drums, and more. MSS is considered to be a challenging audio separation task due to the complexity of music signals. Although the RNN and Transformer architecture are not perfect, they are commonly used to model the music sequence for MSS. Recently, Mamba-2 has already demonstrated high efficiency in various sequential modeling tasks, but its superiority has not been investigated in MSS. This paper applies Mamba-2 with a two-stage strategy, which introduces residual mapping based on the mask method, effectively compensating for the details absent in the mask and further improving separation performance. Experiments confirm the superiority of bidirectional Mamba-2 and the effectiveness of the two-stage network in MSS. The source code is publicly accessible at https: github.com baijinglin TS-BSmamba2.
【15】 RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
标题: RobustCSV:基于HuBERT的旋律提取器和对抗学习,用于稳健的歌唱声音转换
作者:Wei Chen,Xintao Zhao,Jun Chen,Binzhu Sha,Zhiwei Lin,Zhiyong Wu
备注:Accepted by ISCSLP 2024
链接:点击下载PDF文件
摘要:由于在推断过程中使用非鲁棒的方法来提取音高和能量,因此歌唱声音转换(SVC)受到噪声敏感性的阻碍。由于干净的信号是SVC中源音频的关键,因此音乐源分离预处理为处理嘈杂音频提供了可行的解决方案,如背景音乐歌唱(BGM)。然而,目前的分离方法难以完全去除噪声或过度抑制信号分量,影响了处理后音频的自然度和相似性。为了解决这个问题,我们的研究引入了RobustSVC,这是一种新颖的任意对一SVC框架,可以将嘈杂的人声转换为目标歌手演唱的干净人声。我们用基于HuBERT的旋律提取器替换非鲁棒特征,并使用具有三个鉴别器的对抗训练机制来减少自监督表示中的信息泄漏。实验结果表明,RobustSVC算法具有较好的抗噪性,在有噪和无噪的情况下都能获得比基线算法更高的相似度和自然度。摘要:Singing voice conversion (SVC) is hindered by noise sensitivity due to the use of non-robust methods for extracting pitch and energy during the inference. As clean signals are key for the source audio in SVC, music source separation preprocessing offers a viable solution for handling noisy audio, like singing with background music (BGM). However, current separating methods struggle to fully remove noise or excessively suppress signal components, affecting the naturalness and similarity of the processed audio. To tackle this, our study introduces RobustSVC, a novel any-to-one SVC framework that converts noisy vocals into clean vocals sung by the target singer. We replace the non-robust feature with a HuBERT-based melody extractor and use adversarial training mechanisms with three discriminators to reduce information leakage in self-supervised representations. Experimental results show that RobustSVC is noise-robust and achieves higher similarity and naturalness than baseline methods in both noisy and clean vocal conditions.
【16】 Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models
标题: 增强大型音频语言模型音频问题回答中的时间理解
作者:Arvind Krishna Sridhar,Yinyi Guo,Erik Visser
备注:5 pages, 3 figures
链接:点击下载PDF文件
摘要:音频问题分类任务包括音频事件分类、音频字幕和开放式推理。最近,由于大型音频语言模型的出现,音频问题分类引起了人们的关注。当前的文献集中于通过投影模块将音频编码器与仅文本的大型语言模型集成来构建LALM。虽然大型音频语言模型在一般音频理解方面表现出色,但它们在时间推理方面受到限制,这可能会阻碍其商业应用和设备部署。本文讨论了这些挑战和限制音频时间推理。首先,我们介绍了一种数据增强技术,用于使用LLM生成可靠的音频时间问题和答案。其次,我们提出了一个持续的微调课程学习策略,专注于时间推理,而不影响微调任务的性能。最后,我们开发了一个可靠和透明的自动化指标,在LLM的帮助下,智能地测量大型音频语言模型响应和地面真实数据之间的相关性。我们证明了我们提出的技术使用SOTA LALM在公共音频基准数据集上的有效性。摘要:The Audio Question Answering task includes audio event classification, audio captioning, and open ended reasoning. Recently, Audio Question Answering has garnered attention due to the advent of Large Audio Language Models. Current literature focuses on constructing LALMs by integrating audio encoders with text only Large Language Models through a projection module. While Large Audio Language Models excel in general audio understanding, they are limited in temporal reasoning which may hinder their commercial applications and on device deployment. This paper addresses these challenges and limitations in audio temporal reasoning. First, we introduce a data augmentation technique for generating reliable audio temporal questions and answers using an LLM. Second, we propose a continued finetuning curriculum learning strategy to specialize in temporal reasoning without compromising performance on finetuned tasks. Finally, we develop a reliable and transparent automated metric, assisted by an LLM, to measure the correlation between Large Audio Language Model responses and ground truth data intelligently. We demonstrate the effectiveness of our proposed techniques using SOTA LALMs on public audio benchmark datasets.
【17】 Advancing Topic Segmentation of Broadcasted Speech with Multilingual Semantic Embeddings
标题: 利用多语言语义嵌入推进广播语音的主题分割
作者:Sakshi Deo Shukla,Pavel Denisov,Tugtekin Turan
链接:点击下载PDF文件
摘要:基于语音的主题分割的最新进展突出了预训练语音编码器直接从语音中捕获语义表示的潜力。传统上,主题分割依赖于流水线方法,其中自动语音识别系统的成绩单被生成,随后是基于文本的分割算法。在本文中,我们介绍了一个端到端的计划,绕过这个传统的两步过程中,直接采用语义语音编码器进行分割。专注于广播新闻领域,这带来了独特的挑战,由于扬声器的多样性和单一录音中的主题,我们解决了以端到端的方式有效地访问话题变化点的挑战。此外,我们提出了一个新的基准口语新闻主题分割,利用数据集具有约1000小时的公开可用的录音在六种欧洲语言,包括在印地语的评估集,以测试模型的跨域性能在跨语言,zero-shot的情况。这种设置反映了现实世界的多样性以及对适应各种语言环境的模型的需求。我们的结果表明,虽然传统的管道方法实现了最先进的$P_k$得分为0.2431的英语,我们的端到端模型提供了一个有竞争力的$P_k$得分为0.2564。当进行多语言训练时,这些分数分别进一步提高到0.1988和0.2370。为了支持进一步的研究,我们发布了我们的模型以及数据准备脚本,促进了对多语言口语新闻主题分割的开放研究。摘要:Recent advancements in speech-based topic segmentation have highlighted the potential of pretrained speech encoders to capture semantic representations directly from speech. Traditionally, topic segmentation has relied on a pipeline approach in which transcripts of the automatic speech recognition systems are generated, followed by text-based segmentation algorithms. In this paper, we introduce an end-to-end scheme that bypasses this conventional two-step process by directly employing semantic speech encoders for segmentation. Focused on the broadcasted news domain, which poses unique challenges due to the diversity of speakers and topics within single recordings, we address the challenge of accessing topic change points efficiently in an end-to-end manner. Furthermore, we propose a new benchmark for spoken news topic segmentation by utilizing a dataset featuring approximately 1000 hours of publicly available recordings across six European languages and including an evaluation set in Hindi to test the model's cross-domain performance in a cross-lingual, zero-shot scenario. This setup reflects real-world diversity and the need for models adapting to various linguistic settings. Our results demonstrate that while the traditional pipeline approach achieves a state-of-the-art $P_k$ score of 0.2431 for English, our end-to-end model delivers a competitive $P_k$ score of 0.2564. When trained multilingually, these scores further improve to 0.1988 and 0.2370, respectively. To support further research, we release our model along with data preparation scripts, facilitating open research on multilingual spoken news topic segmentation.
【18】 MTDA-HSED: Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection
标题: MTDA-HMED:用于异类声音事件检测的互助调谐和双分支聚集
作者:Zehao Wang,Haobo Yue,Zhicheng Zhang,Da Mu,Jin Tang,Jianqin Yin
备注:Submit to Icassp2025
链接:点击下载PDF文件
摘要:声音事件检测(SED)在理解和感知声学场景中起着至关重要的作用。以前的方法已经证明了令人印象深刻的能力。然而,它们在从异构数据集学习复杂场景的特征方面存在不足。在本文中,我们介绍了一种新的双分支架构称为互助调谐和双分支聚合异构声音事件检测(MTDA-HSED)。MTDA-HSED架构采用互助音频适配器(M3 A)来有效地解决多场景问题,并使用双分支中间融合(DBMF)模块来解决多粒度问题。具体来说,M3 A作为适配器集成到BEAT块中,通过在多场景数据集上对其进行微调来提高BEAT的性能。DBMF模块连接BEAT和CNN分支,这有助于来自BEAT和CNN分支的信息的深度融合。实验结果表明,在DESED和MAESTRO Real数据集上,该方法比mpAUC的基线提高了 textbf{$5 %$}.代码为 href{https: github.com Visitor-W MTDA}{here}。摘要:Sound Event Detection (SED) plays a vital role in comprehending and perceiving acoustic scenes. Previous methods have demonstrated impressive capabilities. However, they are deficient in learning features of complex scenes from heterogeneous dataset. In this paper, we introduce a novel dual-branch architecture named Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection (MTDA-HSED). The MTDA-HSED architecture employs the Mutual-Assistance Audio Adapter (M3A) to effectively tackle the multi-scenario problem and uses the Dual-Branch Mid-Fusion (DBMF) module to tackle the multi-granularity problem. Specifically, M3A is integrated into the BEATs block as an adapter to improve the BEATs' performance by fine-tuning it on the multi-scenario dataset. The DBMF module connects BEATs and CNN branches, which facilitates the deep fusion of information from the BEATs and the CNN branches. Experimental results show that the proposed methods exceed the baseline of mpAUC by textbf{$5 %$} on the DESED and MAESTRO Real datasets. Code is href{https: github.com Visitor-W MTDA}{here}.
【19】 DENSE: Dynamic Embedding Causal Target Speech Extraction
标题: DENSE:动态嵌入因果目标语音提取
作者:Yiwen Wang,Zeyu Yuan,Xihong Wu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:目标语音提取(TSE)的重点是从混合信号中提取特定目标说话人的语音。现有的TSE模型通常利用静态嵌入作为用于提取目标说话者的语音的条件。然而,静态嵌入往往无法捕捉提取的语音信号的上下文信息,这可能会限制模型的性能。我们提出了一种新的动态嵌入因果目标语音提取模型,以解决这一限制。我们的方法采用了自回归机制,根据提取的语音生成上下文相关的嵌入,从而实现实时的帧级提取。实验结果表明,该模型提高了短时目标可懂度(STOI)和信号失真比(SDR),为复杂场景下的目标语音提取提供了一种有前途的解决方案。摘要:Target speech extraction (TSE) focuses on extracting the speech of a specific target speaker from a mixture of signals. Existing TSE models typically utilize static embeddings as conditions for extracting the target speaker's voice. However, the static embeddings often fail to capture the contextual information of the extracted speech signal, which may limit the model's performance. We propose a novel dynamic embedding causal target speech extraction model to address this limitation. Our approach incorporates an autoregressive mechanism to generate context-dependent embeddings based on the extracted speech, enabling real-time, frame-level extraction. Experimental results demonstrate that the proposed model enhances short-time objective intelligibility (STOI) and signal-to-distortion ratio (SDR), offering a promising solution for target speech extraction in challenging scenarios.
【20】 Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis
标题: 绘制音频:利用多指令进行视频到音频合成
作者:Qi Yang,Binjie Mao,Zili Wang,Xing Nie,Pengfei Gao,Ying Guo,Cheng Zhen,Pengfei Yan,Shiming Xiang
备注:14 pages, 11 figures
链接:点击下载PDF文件
摘要:Foley是电影制作中常用的一个术语,指的是在无声电影或视频中添加日常音效,以增强听觉体验。视频到音频(V2 A),作为一种特殊类型的自动福利任务,提出了与视听同步相关的固有挑战。这些挑战包括保持输入视频和生成的音频之间的内容一致性,以及视频内的时间和响度属性的对齐。为了解决这些问题,我们构建了一个可控的视频到音频的合成模型,称为绘制音频,它支持多个输入指令,通过绘制掩码和响度信号。为了确保合成的音频和目标视频之间的内容一致性,我们引入了掩蔽注意模块(MAM),它采用掩蔽视频指令,使模型能够专注于感兴趣的区域。此外,我们还实现了时间响度模块(TLM),它使用辅助响度信号来确保声音的合成在响度和时间维度上与视频保持一致。此外,我们扩展了一个大规模的V2 A数据集,命名为VGGSound-Caption,通过注释字幕提示。在两个大规模V2 A数据集上对具有挑战性的基准进行了大量实验,验证了Draw an Audio达到了最先进的水平。项目页面:https: yannqi.github.io Draw-an-Audio 。摘要:Foley is a term commonly used in filmmaking, referring to the addition of daily sound effects to silent films or videos to enhance the auditory experience. Video-to-Audio (V2A), as a particular type of automatic foley task, presents inherent challenges related to audio-visual synchronization. These challenges encompass maintaining the content consistency between the input video and the generated audio, as well as the alignment of temporal and loudness properties within the video. To address these issues, we construct a controllable video-to-audio synthesis model, termed Draw an Audio, which supports multiple input instructions through drawn masks and loudness signals. To ensure content consistency between the synthesized audio and target video, we introduce the Mask-Attention Module (MAM), which employs masked video instruction to enable the model to focus on regions of interest. Additionally, we implement the Time-Loudness Module (TLM), which uses an auxiliary loudness signal to ensure the synthesis of sound that aligns with the video in both loudness and temporal dimensions. Furthermore, we have extended a large-scale V2A dataset, named VGGSound-Caption, by annotating caption prompts. Extensive experiments on challenging benchmarks across two large-scale V2A datasets verify Draw an Audio achieves the state-of-the-art. Project page: https: yannqi.github.io Draw-an-Audio .
【21】 Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
标题: 无监督音乐音频音色传输的潜在扩散桥梁
作者:Michele Mancusi,Yurii Halychansky,Kin Wai Cheuk,Chieh-Hsin Lai,Stefan Uhlich,Junghyun Koo,Marco A. Martínez-Ramírez,Wei-Hsiang Liao,Giorgio Fabbro,Yuhki Mitsufuji
链接:点击下载PDF文件
摘要:音乐音色转移是一项具有挑战性的任务,涉及修改音频信号的音色特性,同时保持其旋律结构。在本文中,我们提出了一种基于双扩散桥的新方法,该方法使用CocoChorales数据集进行训练,该数据集由未配对的单声道单乐器音频数据组成。每个扩散模型都是在具有高斯先验的特定仪器上训练的。在推断期间,将一个模型指定为源模型以将输入音频映射到其对应的高斯先验,并且将另一个模型指定为目标模型以从该高斯先验重构目标音频,从而促进音色传递。我们将我们的方法与现有的无监督音色传输模型(如VAEGAN和高斯流桥(GFB))进行比较。实验结果表明,我们的方法实现了更好的Fr 'echet音频距离(FAD)和旋律保存,反映了较低的音高距离(DPD)相比,VAEGAN和GFB。此外,我们发现,从高斯先验,$ sigma$的噪声水平,可以进行调整,以控制旋律保存的程度和转移的音色量。摘要:Music timbre transfer is a challenging task that involves modifying the timbral characteristics of an audio signal while preserving its melodic structure. In this paper, we propose a novel method based on dual diffusion bridges, trained using the CocoChorales Dataset, which consists of unpaired monophonic single-instrument audio data. Each diffusion model is trained on a specific instrument with a Gaussian prior. During inference, a model is designated as the source model to map the input audio to its corresponding Gaussian prior, and another model is designated as the target model to reconstruct the target audio from this Gaussian prior, thereby facilitating timbre transfer. We compare our approach against existing unsupervised timbre transfer models such as VAEGAN and Gaussian Flow Bridges (GFB). Experimental results demonstrate that our method achieves both better Fr 'echet Audio Distance (FAD) and melody preservation, as reflected by lower pitch distances (DPD) compared to VAEGAN and GFB. Additionally, we discover that the noise level from the Gaussian prior, $ sigma$, can be adjusted to control the degree of melody preservation and amount of timbre transferred.
【22】 Investigating Causal Cues: Strengthening Spoofed Audio Detection with Human-Discernible Linguistic Features
标题: 调查因果线索:利用人类可辨别的语言特征加强欺骗音频检测
作者:Zahra Khanjani,Tolulope Ale,Jianwu Wang,Lavon Davis,Christine Mallinson,Vandana P. Janeja
链接:点击下载PDF文件
摘要:几种类型的欺骗音频,如模仿,重放攻击和深度伪造,对信息完整性造成了社会挑战。最近,研究人员与社会语言学专家合作,用人类耳朵可以识别的专家定义语言特征(EDLF)来标记欺骗性音频样本:音调,停顿,辅音停止的单词首字母和单词结尾释放,可听到的吸气或呼气以及整体音频质量。已经确定,当几种深度伪造检测算法使用这些EDLF增强音频数据的传统和常见特征时,它们会有所改进。在本文中,使用由多种类型的欺骗音频增强社会语言学注释组成的混合数据集,我们研究了因果发现和可辨别的语言特征和音频片段中的标签之间的推断,将因果模型的结果与专家地面实况验证标记过程进行比较。我们的研究结果表明,因果模型表明了结合语言特征以帮助识别欺骗音频的实用性,以及将人类知识纳入模型和技术以加强AI模型的整体需求和机会。因果发现和推理可以用作训练人类识别欺骗音频以及自动化EDLF标记的基础,以提高常见的基于AI的欺骗音频检测器的性能。摘要:Several types of spoofed audio, such as mimicry, replay attacks, and deepfakes, have created societal challenges to information integrity. Recently, researchers have worked with sociolinguistics experts to label spoofed audio samples with Expert Defined Linguistic Features (EDLFs) that can be discerned by the human ear: pitch, pause, word-initial and word-final release bursts of consonant stops, audible intake or outtake of breath, and overall audio quality. It is established that there is an improvement in several deepfake detection algorithms when they augmented the traditional and common features of audio data with these EDLFs. In this paper, using a hybrid dataset comprised of multiple types of spoofed audio augmented with sociolinguistic annotations, we investigate causal discovery and inferences between the discernible linguistic features and the label in the audio clips, comparing the findings of the causal models with the expert ground truth validation labeling process. Our findings suggest that the causal models indicate the utility of incorporating linguistic features to help discern spoofed audio, as well as the overall need and opportunity to incorporate human knowledge into models and techniques for strengthening AI models. The causal discovery and inference can be used as a foundation of training humans to discern spoofed audio as well as automating EDLFs labeling for the purpose of performance improvement of the common AI-based spoofed audio detectors.
【23】 SongCreator: Lyrics-based Universal Song Generation
标题: SongCreator:基于歌词的通用歌曲一代
作者:Shun Lei,Yixuan Zhou,Boshi Tang,Max W. Y. Lam,Feng Liu,Hangyu Liu,Jingcheng Wu,Shiyin Kang,Zhiyong Wu,Helen Meng
备注:work in progress
链接:点击下载PDF文件
摘要:音乐是人类文化的重要组成部分,是人类智慧和创造力的集中体现,而歌曲是其中不可或缺的一部分。虽然以前的作品已经探讨了歌曲生成的各个方面,如歌唱声音,声乐创作和乐器安排等,在给定歌词的情况下生成具有人声和伴奏的歌曲仍然是一个重大的挑战,这阻碍了音乐生成模型在现实世界中的应用。在这种情况下,我们提出了SongCreator,一个歌曲生成系统,旨在应对这一挑战。该模型具有两个新颖的设计:精心设计的双序列语言模型(DSLM),以捕获用于歌曲生成的人声和伴奏信息,以及DSLM的额外注意力掩码策略,该策略允许我们的模型理解,生成和编辑歌曲,使其适合于各种与歌曲相关的生成任务。大量的实验证明了SongCreator的有效性,在所有八个任务上都实现了最先进或有竞争力的性能。值得注意的是,它在歌词到歌曲和歌词到人声方面大大超过了以前的作品。此外,它能够通过不同的提示独立地控制生成的歌曲中的人声和伴奏的声学条件,显示出其潜在的适用性。我们的样品可在https: songcreator.github.io 上获得。摘要:Music is an integral part of human culture, embodying human intelligence and creativity, of which songs compose an essential part. While various aspects of song generation have been explored by previous works, such as singing voice, vocal composition and instrumental arrangement, etc., generating songs with both vocals and accompaniment given lyrics remains a significant challenge, hindering the application of music generation models in the real world. In this light, we propose SongCreator, a song-generation system designed to tackle this challenge. The model features two novel designs: a meticulously designed dual-sequence language model (DSLM) to capture the information of vocals and accompaniment for song generation, and an additional attention mask strategy for DSLM, which allows our model to understand, generate and edit songs, making it suitable for various song-related generation tasks. Extensive experiments demonstrate the effectiveness of SongCreator by achieving state-of-the-art or competitive performances on all eight tasks. Notably, it surpasses previous works by a large margin in lyrics-to-song and lyrics-to-vocals. Additionally, it is able to independently control the acoustic conditions of the vocals and accompaniment in the generated song through different prompts, exhibiting its potential applicability. Our samples are available at https: songcreator.github.io .
【24】 Musical Chords: A Novel Java Algorithm and App Utility to Enumerate Chord-Progressions Adhering to Music Theory Guidelines
标题: 音乐和弦:一种新颖的Java算法和应用程序实用程序,用于列举符合音乐理论准则的和弦进行
作者:Aditya Lakshminarasimhan
备注:10 Pages, 5 Figures
链接:点击下载PDF文件
摘要:一首歌的骨干是它的和弦进行,一系列的和弦,提高和谐和增加到整体组成。对于从初学者到创意艺术家的个人来说,理解和实施音乐理论语法可以扼杀音乐创作过程并导致歌曲作者的障碍。市场上现有的Chord Progression方法仅限于产生预先选择的进行,并且通常无法符合音乐理论指南或为其他音乐家提供API以进行构建。由于四和弦和八和弦进行尚未被枚举,因此在和弦进行上进行训练的机器学习用例有限,并且移动应用程序无法为用户提供独特或未开发的进行。为了解决这些限制,已经开发了一种新颖的Java算法和自动音乐理论和弦进行和变化生成器App。这个应用程序提供了一个钢琴用户界面,应用音乐理论来生成所有可能的四和弦和八和弦进行,并产生三个由用户选择的生成进行的交替变化。该算法阐明了总共3,297个4弦进行式和总共405,216个8弦进行式。在4和弦进行池中,有1,533个主要4和弦进行和1,764个次要4和弦进行。在8和弦进行池中,有182,094个主要进行和223,122个次要进行。这种创新的方法为音乐家提供了一个全面和可定制的音乐创作工具,使他们能够开发自己的签名声音。摘要:A song's backbone is its chord progressions, a series of chords that improve the harmony and add to the overall composition. For individuals ranging from beginners to creative artists, comprehending and implementing music theory grammar for their own compositions can stifle the music creation process and cause song-writer's block. The existing Chord Progression approaches in the marketplace are limited on producing only pre-selected progressions and often fail to conform to music theory guidelines or provide APIs for other musicians to build on. Because four-chord and eight-chord progressions are yet to be enumerated, Machine learning use-cases that train on chord progressions are limited, and mobile applications don't provide users with unique or unexplored progressions. To address these limitations, a novel Java Algorithm and automated music theory chord progression and variations generator App has been developed. This App offers a piano user interface, that applies music theory to generate all possible four-chord and eight-chord progressions and produces three alternate variations of the generated progressions selected by the user. The Algorithm elucidates 3,297 Total 4-Chord Progressions and 405,216 Total 8-Chord Progressions. Within the 4-Chord Progression pool, there are 1,533 Major 4-chord Progressions and 1,764 Minor 4-Chord Progressions. Within the 8-chord Progression pool, there are 182,094 Major Progressions and 223,122 Minor Progressions. This innovative approach provides musicians with a comprehensive and customizable tool for their music creation, allowing them to develop their signature sounds.
【25】 Continuous Learning of Transformer-based Audio Deepfake Detection
标题: 基于变形器的音频深度伪造检测的持续学习
作者:Tuan Duy Nguyen Le,Kah Kuan Teh,Huy Dat Tran
备注:Submitted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:本文提出了一种新的音频deepfake检测框架,其主要目标有两个:i)在可用的假数据上获得尽可能高的准确性,以及ii)以Few-Shot学习方式有效地对新的假数据进行连续学习。具体来说,我们使用各种深度音频生成方法进行大型音频deepfake收集。数据通过额外的增强方法进一步增强,以增加压缩、远场记录、噪声和其他失真中的变化。然后,我们采用音频谱图Transformer来进行音频深度伪造检测模型。因此,所提出的方法在各种基准数据集上实现了有前途的性能。此外,我们提出了一个持续学习插件模块,以最有效地更新训练模型,使用最少的新假类型的标记数据点。所提出的方法优于传统的直接微调方法,标记的数据点少得多。摘要:This paper proposes a novel framework for audio deepfake detection with two main objectives: i) attaining the highest possible accuracy on available fake data, and ii) effectively performing continuous learning on new fake data in a few-shot learning manner. Specifically, we conduct a large audio deepfake collection using various deep audio generation methods. The data is further enhanced with additional augmentation methods to increase variations amidst compressions, far-field recordings, noise, and other distortions. We then adopt the Audio Spectrogram Transformer for the audio deepfake detection model. Accordingly, the proposed method achieves promising performance on various benchmark datasets. Furthermore, we present a continuous learning plugin module to update the trained model most effectively with the fewest possible labeled data points of the new fake type. The proposed method outperforms the conventional direct fine-tuning approach with much fewer labeled data points.
【26】 Sortformer: Seamless Integration of Speaker Diarization and ASR by Bridging Timestamps and Tokens
标题: 排序器:通过桥梁时间戳和令牌无缝集成说话者拨号和ASB
作者:Taejin Park,Ivan Medennikov,Kunal Dhawan,Weiqing Wang,He Huang,Nithin Rao Koluguri,Krishna C. Puvvada,Jagadeesh Balam,Boris Ginsburg
链接:点击下载PDF文件
摘要:我们提出了Sortformer,一种用于说话人日记化的新型神经模型,与现有的端到端日记化模型相比,它采用非常规目标进行训练。说话人日记中的置换问题一直被认为是一个重要的挑战。大多数现有的端到端日志化系统采用置换不变损失(PIL),其优化产生最低误差的置换。相比之下,我们引入了排序损失,它使日记模型能够在有或没有PIL的情况下自主解析排列。我们证明,结合排序损失和PIL实现的性能与最先进的端到端的日记模型完全用PIL训练。至关重要的是,我们提出了一种简化的多扬声器ASR架构,该架构利用Sortformer作为扬声器监督模型,使用正弦核函数将扬声器标签估计嵌入到ASR编码器状态中。这种方法通过排序目标解决了说话人置换问题,有效地桥接了说话人标签时间戳和说话人令牌。在我们的实验中,我们表明,建议的多扬声器ASR架构,增强扬声器监督,通过适配器技术提高性能。代码和经过训练的模型将通过NVIDIA NeMo框架公开提供摘要:We propose Sortformer, a novel neural model for speaker diarization, trained with unconventional objectives compared to existing end-to-end diarization models. The permutation problem in speaker diarization has long been regarded as a critical challenge. Most prior end-to-end diarization systems employ permutation invariant loss (PIL), which optimizes for the permutation that yields the lowest error. In contrast, we introduce Sort Loss, which enables a diarization model to autonomously resolve permutation, with or without PIL. We demonstrate that combining Sort Loss and PIL achieves performance competitive with state-of-the-art end-to-end diarization models trained exclusively with PIL. Crucially, we present a streamlined multispeaker ASR architecture that leverages Sortformer as a speaker supervision model, embedding speaker label estimation within the ASR encoder state using a sinusoidal kernel function. This approach resolves the speaker permutation problem through sorted objectives, effectively bridging speaker-label timestamps and speaker tokens. In our experiments, we show that the proposed multispeaker ASR architecture, enhanced with speaker supervision, improves performance via adapter techniques. Code and trained models will be made publicly available via the NVIDIA NeMo framework
【27】 Exploring Differences between Human Perception and Model Inference in Audio Event Recognition
标题: 探讨音频事件识别中人类感知和模型推理之间的差异
作者:Yizhou Tan,Yanru Wu,Yuanbo Hou,Xin Xu,Hui Bu,Shengchen Li,Dick Botteldooren,Mark D. Plumbley
备注:Dataset homepage: this https URL
链接:点击下载PDF文件
摘要:音频事件识别(AER)传统上专注于检测和识别音频事件。大多数现有的AER模型倾向于检测所有潜在的事件,而不考虑它们在不同背景下的不同意义。这使得现有模型检测出的AER结果往往与人的听觉感知存在较大的差异。虽然这是一个关键和重要的问题,但它还没有被声音场景和事件的检测和分类(DCASE)社区广泛研究,因为解决它是耗时和劳动密集型的。为了解决这个问题,本文在AER中引入了语义重要性的概念,重点探讨人类感知和模型推理之间的差异。本文构建了一个多注释前景音频事件识别(MAFAR)数据集,其中包括由10个专业注释者标记的音频记录。通过标记频率和方差,MAFAR数据集有助于量化语义重要性和分析人类感知。通过将人类注释与集成预训练模型的预测进行比较,本文揭示了人类感知和模型推理在音频事件的语义识别和存在性检测方面的显著差距。实验结果表明,在事件语义识别中,人类感知往往会忽略细微或琐碎的事件,而模型推理则容易受到带有噪声的事件的影响。同时,在事件存在检测中,模型通常比人类更敏感。摘要:Audio Event Recognition (AER) traditionally focuses on detecting and identifying audio events. Most existing AER models tend to detect all potential events without considering their varying significance across different contexts. This makes the AER results detected by existing models often have a large discrepancy with human auditory perception. Although this is a critical and significant issue, it has not been extensively studied by the Detection and Classification of Sound Scenes and Events (DCASE) community because solving it is time-consuming and labour-intensive. To address this issue, this paper introduces the concept of semantic importance in AER, focusing on exploring the differences between human perception and model inference. This paper constructs a Multi-Annotated Foreground Audio Event Recognition (MAFAR) dataset, which comprises audio recordings labelled by 10 professional annotators. Through labelling frequency and variance, the MAFAR dataset facilitates the quantification of semantic importance and analysis of human perception. By comparing human annotations with the predictions of ensemble pre-trained models, this paper uncovers a significant gap between human perception and model inference in both semantic identification and existence detection of audio events. Experimental results reveal that human perception tends to ignore subtle or trivial events in the event semantic identification, while model inference is easily affected by events with noises. Meanwhile, in event existence detection, models are usually more sensitive than humans.
【28】 Janssen 2.0: Audio Inpainting in the Time-frequency Domain
标题: Janssen 2.0:时频域中的音频修复
作者:Ondřej Mokrý,Peter Balušík,Pavel Rajmic
链接:点击下载PDF文件
摘要:本文主要研究音频信号频谱图中缺失部分的修复问题。首先,最近成功的方法的基础上,未经训练的神经网络进行了修订,并提出了几个修改,提高恢复的音频的信噪比。其次,Janssen算法,基于自回归的时域音频修复的最新技术,适用于时频设置。这种新的方法,创造Janssen-TF,相比神经网络的方法,使用客观指标和主观的听力测试,证明Janssen-TF是优越的,在所有考虑的措施。摘要:The paper focuses on inpainting missing parts of an audio signal spectrogram. First, a recent successful approach based on an untrained neural network is revised and its several modifications are proposed, improving the signal-to-noise ratio of the restored audio. Second, the Janssen algorithm, the autoregression-based state-of-the-art for time-domain audio inpainting, is adapted for the time-frequency setting. This novel method, coined Janssen-TF, is compared to the neural network approach using both objective metrics and a subjective listening test, proving Janssen-TF to be superior in all the considered measures.
【29】 InstructSing: High-Fidelity Singing Voice Generation via Instructing Yourself
标题: DirectSing:通过指导自己产生高保真歌唱声音
作者:Chang Zeng,Chunhui Wang,Xiaoxiao Miao,Jian Zhao,Zhonglin Jiang,Yong Chen
备注:To appear in 2024 IEEE Spoken Language Technology Workshop, Dec 02-05, 2024, Macao, China
链接:点击下载PDF文件
摘要:加快训练过程,同时确保高质量的生成语音和可接受的推理速度是具有挑战性的。在本文中,我们提出了一种新的神经声码器称为InstructSing,它可以收敛速度比其他神经声码器,同时保持良好的性能,通过集成可微数字信号处理和对抗训练。它包括一个发生器和两个鉴别器。具体而言,发生器结合了谐波加噪声(HN)模块,以产生8 kHz的音频作为指导信号。随后,HN模块通过基于UNet的模块与扩展WaveNet连接,该模块将HN模块的输出转换为包含基本周期和非周期信息的潜在变量序列。除了潜在序列,扩展的WaveNet还将mel频谱图作为输入,以生成48 kHz高保真歌声。在鉴别器方面,我们将HiFiGAN中最初提出的多周期鉴别器与多分辨率多频带STFT鉴别器相结合。值得注意的是,InstructSing实现了与其他神经声码器相当的语音质量,但在4个NVIDIA V100 GPU机器上的训练步骤只有十分之一 footnote{{演示页面: href{https: wavelandspeech.github.io inst ructsing }。我们计划在论文被接受后开源我们的代码和预训练模型。摘要:It is challenging to accelerate the training process while ensuring both high-quality generated voices and acceptable inference speed. In this paper, we propose a novel neural vocoder called InstructSing, which can converge much faster compared with other neural vocoders while maintaining good performance by integrating differentiable digital signal processing and adversarial training. It includes one generator and two discriminators. Specifically, the generator incorporates a harmonic-plus-noise (HN) module to produce 8kHz audio as an instructive signal. Subsequently, the HN module is connected with an extended WaveNet by an UNet-based module, which transforms the output of the HN module to a latent variable sequence containing essential periodic and aperiodic information. In addition to the latent sequence, the extended WaveNet also takes the mel-spectrogram as input to generate 48kHz high-fidelity singing voices. In terms of discriminators, we combine a multi-period discriminator, as originally proposed in HiFiGAN, with a multi-resolution multi-band STFT discriminator. Notably, InstructSing achieves comparable voice quality to other neural vocoders but with only one-tenth of the training steps on a 4 NVIDIA V100 GPU machine footnote{{Demo page: href{https: wavelandspeech.github.io instructsing }{ texttt{https: wavelandspeech.github.io inst ructsing }}}}. We plan to open-source our code and pretrained model once the paper get accepted.
【30】 Spoofing-Aware Speaker Verification Robust Against Domain and Channel Mismatches
标题: 针对域和通道不匹配的强有力的欺骗感知说话者验证
作者:Chang Zeng,Xiaoxiao Miao,Xin Wang,Erica Cooper,Junichi Yamagishi
备注:To appear in 2024 IEEE Spoken Language Technology Workshop, Dec 02-05, 2024, Macao, China
链接:点击下载PDF文件
摘要:在现实世界的应用中,建立一个说话人验证系统,同时对常见的威胁,包括欺骗攻击,通道不匹配,域不匹配是具有挑战性的。传统的自动说话人确认(ASV)系统通常单独解决这些问题,导致在同时面临挑战时性能不佳。在本文中,我们提出了一个集成的框架,将成对学习和欺骗攻击模拟到元学习范式,以提高对这些多方面的威胁的鲁棒性。该方法采用非对称双路径模型和多任务学习策略来同时处理ASV、反欺骗和欺骗感知ASV任务。引入了一个新的测试数据集CNComplex来评估这些组合威胁下的系统性能。实验结果表明,我们的集成模型显着提高了在各种情况下,传统的ASV系统的性能,展示了其在现实世界中部署的潜力。此外,所提出的框架的能力,在不同的条件下概括突出了其鲁棒性和可靠性,使其成为一个有前途的解决方案,为实际ASV应用。摘要:In real-world applications, it is challenging to build a speaker verification system that is simultaneously robust against common threats, including spoofing attacks, channel mismatch, and domain mismatch. Traditional automatic speaker verification (ASV) systems often tackle these issues separately, leading to suboptimal performance when faced with simultaneous challenges. In this paper, we propose an integrated framework that incorporates pair-wise learning and spoofing attack simulation into the meta-learning paradigm to enhance robustness against these multifaceted threats. This novel approach employs an asymmetric dual-path model and a multi-task learning strategy to handle ASV, anti-spoofing, and spoofing-aware ASV tasks concurrently. A new testing dataset, CNComplex, is introduced to evaluate system performance under these combined threats. Experimental results demonstrate that our integrated model significantly improves performance over traditional ASV systems across various scenarios, showcasing its potential for real-world deployment. Additionally, the proposed framework's ability to generalize across different conditions highlights its robustness and reliability, making it a promising solution for practical ASV applications.
【31】 Multi-Source Music Generation with Latent Diffusion
标题: 具有潜在扩散的多来源音乐生成
作者:Zhongweiyang Xu,Debottam Dutta,Yu-Lin Wei,Romit Roy Choudhury
备注:ICASSP 2025 in Submission
链接:点击下载PDF文件
摘要:大多数音乐生成模型直接生成单个音乐混合。为了允许更灵活和可控的生成,已经提出了多源扩散模型(MSDM)来将音乐建模为多个乐器源的混合物(例如,钢琴、鼓、贝司和吉他)。它的目标是使用一个单一的扩散模型来生成一致的音乐源,这些音乐源进一步混合以形成音乐。尽管它的能力,MSDM是无法产生丰富的旋律歌曲,往往产生空洞的声音。此外,其波形扩散引入显著的高斯噪声伪影,这损害了音频质量。为此,我们引入了一个多源潜在扩散模型(MSLDM),该模型采用变分自编码器(VAE)将每个工具源编码为不同的潜在表示。通过在所有音乐源上训练VAE,我们有效地捕获了我们的扩散模型联合建模的源潜伏中每个源的独特特征。这种方法通过利用VAE的潜在压缩和噪声鲁棒性显著增强了音乐的整体和部分生成。压缩的潜在源还促进更有效的生成。主观听力测试和Frechet音频距离(FAD)分数证实,我们的模型优于MSDM,展示了其在音乐生成系统中的实用性和增强的适用性。我们还强调,建模的来源是更有效的比直接的音乐混合建模。有关代码和型号,请访问https: github.com XZWY MSLDM。演示可在https: xzwy.github.io MSLDMDemo上获得。摘要:Most music generation models directly generate a single music mixture. To allow for more flexible and controllable generation, the Multi-Source Diffusion Model (MSDM) has been proposed to model music as a mixture of multiple instrumental sources (e.g., piano, drums, bass, and guitar). Its goal is to use one single diffusion model to generate consistent music sources, which are further mixed to form the music. Despite its capabilities, MSDM is unable to generate songs with rich melodies and often generates empty sounds. Also, its waveform diffusion introduces significant Gaussian noise artifacts, which compromises audio quality. In response, we introduce a multi-source latent diffusion model (MSLDM) that employs Variational Autoencoders (VAEs) to encode each instrumental source into a distinct latent representation. By training a VAE on all music sources, we efficiently capture each source's unique characteristics in a source latent that our diffusion model models jointly. This approach significantly enhances the total and partial generation of music by leveraging the VAE's latent compression and noise-robustness. The compressed source latent also facilitates more efficient generation. Subjective listening tests and Frechet Audio Distance (FAD) scores confirm that our model outperforms MSDM, showcasing its practical and enhanced applicability in music generation systems. We also emphasize that modeling sources is more effective than direct music mixture modeling. Codes and models are available at https: github.com XZWY MSLDM. Demos are available at https: xzwy.github.io MSLDMDemo.
【32】 DeWinder: Single-Channel Wind Noise Reduction using Ultrasound Sensing
标题: DeWinder:使用超声波传感降低单通道风噪音
作者:Kuang Yuan,Shuo Han,Swarun Kumar,Bhiksha Raj
链接:点击下载PDF文件
摘要:在室外环境中的音频记录的质量通常由于风的存在而降低。由于单通道语音的非平稳特性,减轻风噪声对语音感知质量的影响仍然是一个重大挑战。噪声抑制中的先前工作将风噪声视为一般背景噪声,而没有对其特性进行显式建模。在本文中,我们利用超声波作为一种辅助方式来明确地感测气流和风噪声的特征。我们提出了一个多模态深度学习框架,用于融合超声多普勒特征和语音信号以降低风噪声。我们的研究结果表明,DeWinder可以显着提高国家的最先进的语音增强模型的降噪能力。摘要:The quality of audio recordings in outdoor environments is often degraded by the presence of wind. Mitigating the impact of wind noise on the perceptual quality of single-channel speech remains a significant challenge due to its non-stationary characteristics. Prior work in noise suppression treats wind noise as a general background noise without explicit modeling of its characteristics. In this paper, we leverage ultrasound as an auxiliary modality to explicitly sense the airflow and characterize the wind noise. We propose a multi-modal deep-learning framework to fuse the ultrasonic Doppler features and speech signals for wind noise reduction. Our results show that DeWinder can significantly improve the noise reduction capabilities of state-of-the-art speech enhancement models.
【33】 VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion
标题: VC-ENHANCE:具有集成噪音抑制和语音转换的语音恢复
作者:Kyungguen Byun,Jason Filos,Erik Visser,Sunkuk Moon
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:噪声抑制(NS)算法在许多情况下可以有效地改善语音质量。然而,积极的噪声抑制可能会损害目标语音,降低语音清晰度和质量,尽管消除了噪声。本研究提出了一种显式语音恢复方法,使用语音转换(VC)技术进行噪声抑制后的恢复。我们观察到,高质量的语音可以通过基于扩散的语音转换阶段恢复,条件是从去噪语音中提取的目标说话人嵌入和语音内容信息。这种语音恢复可以实现诸如带宽扩展、去混响和图像修复的增强效果。我们的实验结果表明,这两个阶段的NS+VC框架优于单级增强模型的输出语音质量,客观指标衡量,而得分略低的语音清晰度。为了进一步提高组合系统的可懂度,我们提出了一种内容编码器自适应方法,用于在噪声条件下进行鲁棒的内容提取。摘要:Noise suppression (NS) algorithms are effective in improving speech quality in many cases. However, aggressive noise suppression can damage the target speech, reducing both speech intelligibility and quality despite removing the noise. This study proposes an explicit speech restoration method using a voice conversion (VC) technique for restoration after noise suppression. We observed that high-quality speech can be restored through a diffusion-based voice conversion stage, conditioned on the target speaker embedding and speech content information extracted from the de-noised speech. This speech restoration can achieve enhancement effects such as bandwidth extension, de-reverberation, and in-painting. Our experimental results demonstrate that this two-stage NS+VC framework outperforms single-stage enhancement models in terms of output speech quality, as measured by objective metrics, while scoring slightly lower in speech intelligibility. To further improve the intelligibility of the combined system, we propose a content encoder adaptation method for robust content extraction in noisy conditions.
【34】 Retrieval Augmented Correction of Named Entity Speech Recognition Errors
标题: 命名实体语音识别错误的检索增强纠正
作者:Ernest Pusateri,Anmol Walia,Anirudh Kashi,Bortik Bandyopadhyay,Nadia Hyder,Sayantan Mahinder,Raviteja Anantha,Daben Liu,Sashank Gondala
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:近年来,端到端自动语音识别(ASR)系统已经证明了自己非常准确和高性能,但这些系统仍然有一个显着的错误率为实体名称,很少出现在他们的训练数据。在端到端ASR系统兴起的同时,大型语言模型(LLM)已被证明是各种自然语言处理(NLP)任务的通用工具。在有相关知识数据库的NLP任务中,检索增强生成(RAG)与LLM一起使用时取得了令人印象深刻的结果。在这项工作中,我们提出了一个类似RAG的技术,用于纠正语音识别实体名称错误。我们的方法使用矢量数据库来索引一组相关的实体。在运行时,从可能错误的文本ASR假设生成数据库查询,并且使用这些查询检索的实体与ASR假设一起被馈送到已经被适配为校正ASR错误的LLM。总的来说,我们最好的系统实现了33%-39%的相对字错误率降低合成测试集集中在语音助理查询罕见的音乐实体没有倒退的STOP测试集,一个公开的语音助理测试集,涵盖了许多领域。摘要:In recent years, end-to-end automatic speech recognition (ASR) systems have proven themselves remarkably accurate and performant, but these systems still have a significant error rate for entity names which appear infrequently in their training data. In parallel to the rise of end-to-end ASR systems, large language models (LLMs) have proven to be a versatile tool for various natural language processing (NLP) tasks. In NLP tasks where a database of relevant knowledge is available, retrieval augmented generation (RAG) has achieved impressive results when used with LLMs. In this work, we propose a RAG-like technique for correcting speech recognition entity name errors. Our approach uses a vector database to index a set of relevant entities. At runtime, database queries are generated from possibly errorful textual ASR hypotheses, and the entities retrieved using these queries are fed, along with the ASR hypotheses, to an LLM which has been adapted to correct ASR errors. Overall, our best system achieves 33%-39% relative word error rate reductions on synthetic test sets focused on voice assistant queries of rare music entities without regressing on the STOP test set, a publicly available voice assistant test set covering many domains.
eess.AS音频处理
【1】 Sortformer: Seamless Integration of Speaker Diarization and ASR by Bridging Timestamps and Tokens标题: 排序器:通过桥梁时间戳和令牌无缝集成说话者拨号和ASB
作者:Taejin Park,Ivan Medennikov,Kunal Dhawan,Weiqing Wang,He Huang,Nithin Rao Koluguri,Krishna C. Puvvada,Jagadeesh Balam,Boris Ginsburg
链接:点击下载PDF文件
摘要:我们提出了Sortformer,一种用于说话人日记化的新型神经模型,与现有的端到端日记化模型相比,它采用非常规目标进行训练。说话人日记中的置换问题一直被认为是一个重要的挑战。大多数现有的端到端日志化系统采用置换不变损失(PIL),其优化产生最低误差的置换。相比之下,我们引入了排序损失,它使日记模型能够在有或没有PIL的情况下自主解析排列。我们证明,结合排序损失和PIL实现的性能与最先进的端到端的日记模型完全用PIL训练。至关重要的是,我们提出了一个精简的多扬声器ASR架构,利用Sortformer作为扬声器监督模型,使用正弦核函数在ASR编码器状态中嵌入扬声器标签估计。这种方法通过排序目标解决了说话人置换问题,有效地桥接了说话人标签时间戳和说话人令牌。在我们的实验中,我们表明,建议的多扬声器ASR架构,增强扬声器监督,通过适配器技术提高性能。代码和经过训练的模型将通过NVIDIA NeMo框架公开提供摘要:We propose Sortformer, a novel neural model for speaker diarization, trained with unconventional objectives compared to existing end-to-end diarization models. The permutation problem in speaker diarization has long been regarded as a critical challenge. Most prior end-to-end diarization systems employ permutation invariant loss (PIL), which optimizes for the permutation that yields the lowest error. In contrast, we introduce Sort Loss, which enables a diarization model to autonomously resolve permutation, with or without PIL. We demonstrate that combining Sort Loss and PIL achieves performance competitive with state-of-the-art end-to-end diarization models trained exclusively with PIL. Crucially, we present a streamlined multispeaker ASR architecture that leverages Sortformer as a speaker supervision model, embedding speaker label estimation within the ASR encoder state using a sinusoidal kernel function. This approach resolves the speaker permutation problem through sorted objectives, effectively bridging speaker-label timestamps and speaker tokens. In our experiments, we show that the proposed multispeaker ASR architecture, enhanced with speaker supervision, improves performance via adapter techniques. Code and trained models will be made publicly available via the NVIDIA NeMo framework
【2】 Exploring Differences between Human Perception and Model Inference in Audio Event Recognition
标题: 探讨音频事件识别中人类感知和模型推理之间的差异
作者:Yizhou Tan,Yanru Wu,Yuanbo Hou,Xin Xu,Hui Bu,Shengchen Li,Dick Botteldooren,Mark D. Plumbley
备注:Dataset homepage: this https URL
链接:点击下载PDF文件
摘要:音频事件识别(AER)传统上专注于检测和识别音频事件。大多数现有的AER模型倾向于检测所有潜在的事件,而不考虑它们在不同背景下的不同意义。这使得现有模型检测出的AER结果往往与人的听觉感知存在较大的差异。虽然这是一个关键和重要的问题,但它还没有被声音场景和事件的检测和分类(DCASE)社区广泛研究,因为解决它是耗时和劳动密集型的。为了解决这个问题,本文在AER中引入了语义重要性的概念,重点探讨人类感知和模型推理之间的差异。本文构建了一个多注释前景音频事件识别(MAFAR)数据集,其中包括由10个专业注释者标记的音频记录。通过标记频率和方差,MAFAR数据集有助于量化语义重要性和分析人类感知。通过将人类注释与集成预训练模型的预测进行比较,本文揭示了人类感知和模型推理在音频事件的语义识别和存在性检测方面的显著差距。实验结果表明,在事件语义识别中,人类感知往往会忽略细微或琐碎的事件,而模型推理则容易受到带有噪声的事件的影响。同时,在事件存在检测中,模型通常比人类更敏感。摘要:Audio Event Recognition (AER) traditionally focuses on detecting and identifying audio events. Most existing AER models tend to detect all potential events without considering their varying significance across different contexts. This makes the AER results detected by existing models often have a large discrepancy with human auditory perception. Although this is a critical and significant issue, it has not been extensively studied by the Detection and Classification of Sound Scenes and Events (DCASE) community because solving it is time-consuming and labour-intensive. To address this issue, this paper introduces the concept of semantic importance in AER, focusing on exploring the differences between human perception and model inference. This paper constructs a Multi-Annotated Foreground Audio Event Recognition (MAFAR) dataset, which comprises audio recordings labelled by 10 professional annotators. Through labelling frequency and variance, the MAFAR dataset facilitates the quantification of semantic importance and analysis of human perception. By comparing human annotations with the predictions of ensemble pre-trained models, this paper uncovers a significant gap between human perception and model inference in both semantic identification and existence detection of audio events. Experimental results reveal that human perception tends to ignore subtle or trivial events in the event semantic identification, while model inference is easily affected by events with noises. Meanwhile, in event existence detection, models are usually more sensitive than humans.
【3】 Janssen 2.0: Audio Inpainting in the Time-frequency Domain
标题: Janssen 2.0:时频域中的音频修复
作者:Ondřej Mokrý,Peter Balušík,Pavel Rajmic
链接:点击下载PDF文件
摘要:本文主要研究音频信号频谱图中缺失部分的修复问题。首先,最近成功的方法的基础上,未经训练的神经网络进行了修订,并提出了几个修改,提高恢复的音频的信噪比。其次,Janssen算法,基于自回归的时域音频修复的最新技术,适用于时频设置。这种新的方法,创造Janssen-TF,相比神经网络的方法,使用客观指标和主观的听力测试,证明Janssen-TF是优越的,在所有考虑的措施。摘要:The paper focuses on inpainting missing parts of an audio signal spectrogram. First, a recent successful approach based on an untrained neural network is revised and its several modifications are proposed, improving the signal-to-noise ratio of the restored audio. Second, the Janssen algorithm, the autoregression-based state-of-the-art for time-domain audio inpainting, is adapted for the time-frequency setting. This novel method, coined Janssen-TF, is compared to the neural network approach using both objective metrics and a subjective listening test, proving Janssen-TF to be superior in all the considered measures.
【4】 InstructSing: High-Fidelity Singing Voice Generation via Instructing Yourself
标题: DirectSing:通过指导自己产生高保真歌唱声音
作者:Chang Zeng,Chunhui Wang,Xiaoxiao Miao,Jian Zhao,Zhonglin Jiang,Yong Chen
备注:To appear in 2024 IEEE Spoken Language Technology Workshop, Dec 02-05, 2024, Macao, China
链接:点击下载PDF文件
摘要:在确保高质量生成的声音和可接受的推理速度的同时加快训练过程是一项挑战。在本文中,我们提出了一种新的神经声码器称为InstructSing,它可以收敛速度比其他神经声码器,同时保持良好的性能,通过集成可微数字信号处理和对抗训练。它包括一个发生器和两个鉴别器。具体而言,发生器结合了谐波加噪声(HN)模块,以产生8 kHz的音频作为指导信号。随后,HN模块通过基于UNet的模块与扩展WaveNet连接,该模块将HN模块的输出转换为包含基本周期和非周期信息的潜在变量序列。除了潜在序列,扩展的WaveNet还将mel频谱图作为输入,以生成48 kHz高保真歌声。在鉴别器方面,我们将HiFiGAN中最初提出的多周期鉴别器与多分辨率多频带STFT鉴别器相结合。值得注意的是,InstructSing实现了与其他神经声码器相当的语音质量,但在4个NVIDIA V100 GPU机器上的训练步骤只有十分之一 footnote{{演示页面: href{https: wavelandspeech.github.io inst ructsing }。我们计划在论文被接受后开源我们的代码和预训练模型。摘要:It is challenging to accelerate the training process while ensuring both high-quality generated voices and acceptable inference speed. In this paper, we propose a novel neural vocoder called InstructSing, which can converge much faster compared with other neural vocoders while maintaining good performance by integrating differentiable digital signal processing and adversarial training. It includes one generator and two discriminators. Specifically, the generator incorporates a harmonic-plus-noise (HN) module to produce 8kHz audio as an instructive signal. Subsequently, the HN module is connected with an extended WaveNet by an UNet-based module, which transforms the output of the HN module to a latent variable sequence containing essential periodic and aperiodic information. In addition to the latent sequence, the extended WaveNet also takes the mel-spectrogram as input to generate 48kHz high-fidelity singing voices. In terms of discriminators, we combine a multi-period discriminator, as originally proposed in HiFiGAN, with a multi-resolution multi-band STFT discriminator. Notably, InstructSing achieves comparable voice quality to other neural vocoders but with only one-tenth of the training steps on a 4 NVIDIA V100 GPU machine footnote{{Demo page: href{https: wavelandspeech.github.io instructsing }{ texttt{https: wavelandspeech.github.io inst ructsing }}}}. We plan to open-source our code and pretrained model once the paper get accepted.
【5】 Spoofing-Aware Speaker Verification Robust Against Domain and Channel Mismatches
标题: 针对域和通道不匹配的强有力的欺骗感知说话者验证
作者:Chang Zeng,Xiaoxiao Miao,Xin Wang,Erica Cooper,Junichi Yamagishi
备注:To appear in 2024 IEEE Spoken Language Technology Workshop, Dec 02-05, 2024, Macao, China
链接:点击下载PDF文件
摘要:在现实世界的应用中,建立一个说话人验证系统,同时对常见的威胁,包括欺骗攻击,通道不匹配,域不匹配是具有挑战性的。传统的自动说话人确认(ASV)系统通常单独解决这些问题,导致在同时面临挑战时性能不佳。在本文中,我们提出了一个集成的框架,将成对学习和欺骗攻击模拟到元学习范式,以提高对这些多方面的威胁的鲁棒性。该方法采用非对称双路径模型和多任务学习策略来同时处理ASV、反欺骗和欺骗感知ASV任务。引入了一个新的测试数据集CNComplex来评估这些组合威胁下的系统性能。实验结果表明,我们的集成模型显着提高了在各种情况下,传统的ASV系统的性能,展示了其在现实世界中部署的潜力。此外,所提出的框架的能力,在不同的条件下概括突出了其鲁棒性和可靠性,使其成为一个有前途的解决方案,为实际ASV应用。摘要:In real-world applications, it is challenging to build a speaker verification system that is simultaneously robust against common threats, including spoofing attacks, channel mismatch, and domain mismatch. Traditional automatic speaker verification (ASV) systems often tackle these issues separately, leading to suboptimal performance when faced with simultaneous challenges. In this paper, we propose an integrated framework that incorporates pair-wise learning and spoofing attack simulation into the meta-learning paradigm to enhance robustness against these multifaceted threats. This novel approach employs an asymmetric dual-path model and a multi-task learning strategy to handle ASV, anti-spoofing, and spoofing-aware ASV tasks concurrently. A new testing dataset, CNComplex, is introduced to evaluate system performance under these combined threats. Experimental results demonstrate that our integrated model significantly improves performance over traditional ASV systems across various scenarios, showcasing its potential for real-world deployment. Additionally, the proposed framework's ability to generalize across different conditions highlights its robustness and reliability, making it a promising solution for practical ASV applications.
【6】 Multi-Source Music Generation with Latent Diffusion
标题: 具有潜在扩散的多来源音乐生成
作者:Zhongweiyang Xu,Debottam Dutta,Yu-Lin Wei,Romit Roy Choudhury
备注:ICASSP 2025 in Submission
链接:点击下载PDF文件
摘要:大多数音乐生成模型直接生成单个音乐混合。为了允许更灵活和可控的生成,已经提出了多源扩散模型(MSDM)来将音乐建模为多个乐器源的混合物(例如,钢琴、鼓、贝司和吉他)。它的目标是使用一个单一的扩散模型来生成一致的音乐源,这些音乐源进一步混合以形成音乐。尽管它的能力,MSDM是无法产生丰富的旋律歌曲,往往产生空洞的声音。此外,其波形扩散引入显著的高斯噪声伪影,这损害了音频质量。为此,我们引入了一个多源潜在扩散模型(MSLDM),该模型采用变分自编码器(VAE)将每个工具源编码为不同的潜在表示。通过在所有音乐源上训练VAE,我们有效地捕获了我们的扩散模型联合建模的源潜伏中每个源的独特特征。这种方法通过利用VAE的潜在压缩和噪声鲁棒性显著增强了音乐的整体和部分生成。压缩的潜在源还促进更有效的生成。主观听力测试和Frechet音频距离(FAD)分数证实,我们的模型优于MSDM,展示了其在音乐生成系统中的实用性和增强的适用性。我们还强调,对源进行建模比直接音乐混合建模更有效。有关代码和型号,请访问https: github.com XZWY MSLDM。演示可在https: xzwy.github.io MSLDMDemo上获得。摘要:Most music generation models directly generate a single music mixture. To allow for more flexible and controllable generation, the Multi-Source Diffusion Model (MSDM) has been proposed to model music as a mixture of multiple instrumental sources (e.g., piano, drums, bass, and guitar). Its goal is to use one single diffusion model to generate consistent music sources, which are further mixed to form the music. Despite its capabilities, MSDM is unable to generate songs with rich melodies and often generates empty sounds. Also, its waveform diffusion introduces significant Gaussian noise artifacts, which compromises audio quality. In response, we introduce a multi-source latent diffusion model (MSLDM) that employs Variational Autoencoders (VAEs) to encode each instrumental source into a distinct latent representation. By training a VAE on all music sources, we efficiently capture each source's unique characteristics in a source latent that our diffusion model models jointly. This approach significantly enhances the total and partial generation of music by leveraging the VAE's latent compression and noise-robustness. The compressed source latent also facilitates more efficient generation. Subjective listening tests and Frechet Audio Distance (FAD) scores confirm that our model outperforms MSDM, showcasing its practical and enhanced applicability in music generation systems. We also emphasize that modeling sources is more effective than direct music mixture modeling. Codes and models are available at https: github.com XZWY MSLDM. Demos are available at https: xzwy.github.io MSLDMDemo.
【7】 DeWinder: Single-Channel Wind Noise Reduction using Ultrasound Sensing
标题: DeWinder:使用超声波传感降低单通道风噪音
作者:Kuang Yuan,Shuo Han,Swarun Kumar,Bhiksha Raj
链接:点击下载PDF文件
摘要:在室外环境中的音频记录的质量通常由于风的存在而降低。由于单通道语音的非平稳特性,减轻风噪声对语音感知质量的影响仍然是一个重大挑战。噪声抑制中的先前工作将风噪声视为一般背景噪声,而没有对其特性进行显式建模。在本文中,我们利用超声波作为一种辅助方式来明确地感测气流和风噪声的特征。我们提出了一个多模态深度学习框架,用于融合超声多普勒特征和语音信号以降低风噪声。我们的研究结果表明,DeWinder可以显着提高国家的最先进的语音增强模型的降噪能力。摘要:The quality of audio recordings in outdoor environments is often degraded by the presence of wind. Mitigating the impact of wind noise on the perceptual quality of single-channel speech remains a significant challenge due to its non-stationary characteristics. Prior work in noise suppression treats wind noise as a general background noise without explicit modeling of its characteristics. In this paper, we leverage ultrasound as an auxiliary modality to explicitly sense the airflow and characterize the wind noise. We propose a multi-modal deep-learning framework to fuse the ultrasonic Doppler features and speech signals for wind noise reduction. Our results show that DeWinder can significantly improve the noise reduction capabilities of state-of-the-art speech enhancement models.
【8】 VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion
标题: VC-ENHANCE:具有集成噪音抑制和语音转换的语音恢复
作者:Kyungguen Byun,Jason Filos,Erik Visser,Sunkuk Moon
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:噪声抑制(NS)算法在许多情况下可以有效地改善语音质量。然而,积极的噪声抑制可能会损害目标语音,降低语音清晰度和质量,尽管消除了噪声。本研究提出了一种显式语音恢复方法,使用语音转换(VC)技术进行噪声抑制后的恢复。我们观察到,高质量的语音可以通过基于扩散的语音转换阶段恢复,条件是从去噪语音中提取的目标说话人嵌入和语音内容信息。这种语音恢复可以实现诸如带宽扩展、去混响和图像修复的增强效果。我们的实验结果表明,这两个阶段的NS+VC框架优于单级增强模型的输出语音质量,客观指标衡量,而得分略低的语音清晰度。为了进一步提高组合系统的可懂度,我们提出了一种内容编码器自适应方法,用于在噪声条件下进行鲁棒的内容提取。摘要:Noise suppression (NS) algorithms are effective in improving speech quality in many cases. However, aggressive noise suppression can damage the target speech, reducing both speech intelligibility and quality despite removing the noise. This study proposes an explicit speech restoration method using a voice conversion (VC) technique for restoration after noise suppression. We observed that high-quality speech can be restored through a diffusion-based voice conversion stage, conditioned on the target speaker embedding and speech content information extracted from the de-noised speech. This speech restoration can achieve enhancement effects such as bandwidth extension, de-reverberation, and in-painting. Our experimental results demonstrate that this two-stage NS+VC framework outperforms single-stage enhancement models in terms of output speech quality, as measured by objective metrics, while scoring slightly lower in speech intelligibility. To further improve the intelligibility of the combined system, we propose a content encoder adaptation method for robust content extraction in noisy conditions.
【9】 Retrieval Augmented Correction of Named Entity Speech Recognition Errors
标题: 命名实体语音识别错误的检索增强纠正
作者:Ernest Pusateri,Anmol Walia,Anirudh Kashi,Bortik Bandyopadhyay,Nadia Hyder,Sayantan Mahinder,Raviteja Anantha,Daben Liu,Sashank Gondala
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:近年来,端到端自动语音识别(ASR)系统已经证明了自己非常准确和高性能,但这些系统仍然有一个显着的错误率为实体名称,很少出现在他们的训练数据。在端到端ASR系统兴起的同时,大型语言模型(LLM)已被证明是各种自然语言处理(NLP)任务的通用工具。在有相关知识数据库的NLP任务中,检索增强生成(RAG)与LLM一起使用时取得了令人印象深刻的结果。在这项工作中,我们提出了一个类似RAG的技术,用于纠正语音识别实体名称错误。我们的方法使用矢量数据库来索引一组相关的实体。在运行时,从可能错误的文本ASR假设生成数据库查询,并且使用这些查询检索的实体与ASR假设一起被馈送到已经被适配为校正ASR错误的LLM。总的来说,我们最好的系统实现了33%-39%的相对字错误率降低合成测试集集中在语音助理查询罕见的音乐实体没有倒退的STOP测试集,一个公开的语音助理测试集,涵盖了许多领域。摘要:In recent years, end-to-end automatic speech recognition (ASR) systems have proven themselves remarkably accurate and performant, but these systems still have a significant error rate for entity names which appear infrequently in their training data. In parallel to the rise of end-to-end ASR systems, large language models (LLMs) have proven to be a versatile tool for various natural language processing (NLP) tasks. In NLP tasks where a database of relevant knowledge is available, retrieval augmented generation (RAG) has achieved impressive results when used with LLMs. In this work, we propose a RAG-like technique for correcting speech recognition entity name errors. Our approach uses a vector database to index a set of relevant entities. At runtime, database queries are generated from possibly errorful textual ASR hypotheses, and the entities retrieved using these queries are fed, along with the ASR hypotheses, to an LLM which has been adapted to correct ASR errors. Overall, our best system achieves 33%-39% relative word error rate reductions on synthetic test sets focused on voice assistant queries of rare music entities without regressing on the STOP test set, a publicly available voice assistant test set covering many domains.
【10】 Property Neurons in Self-Supervised Speech Transformers
标题: 自我监督语音Transformer中的属性神经元
作者:Tzu-Quan Lin,Guan-Ting Lin,Hung-yi Lee,Hao Tang
备注:Accepted by SLT 2024
链接:点击下载PDF文件
摘要:已经有许多关于分析自监督语音Transformers的研究,特别是使用逐层分析。然而,希望有一种方法,可以精确地定位负责语音的特定属性的神经元的子集,能够进行模型修剪和模型编辑。在这项工作中,我们确定了一组属性神经元的前馈层的Transformers研究如何语音相关的属性,如电话,性别和音高,存储。当删除特定属性的神经元(一种简单的模型编辑形式)时,相应的下游性能显着下降,显示了属性神经元的重要性。我们应用这种方法修剪前馈层的Transformers,其中大部分的模型参数。我们发现,在修剪过程中保护属性神经元比基于范数的修剪更有效。摘要:There have been many studies on analyzing self-supervised speech Transformers, in particular, with layer-wise analysis. It is, however, desirable to have an approach that can pinpoint exactly a subset of neurons that is responsible for a particular property of speech, being amenable to model pruning and model editing. In this work, we identify a set of property neurons in the feedforward layers of Transformers to study how speech-related properties, such as phones, gender, and pitch, are stored. When removing neurons of a particular property (a simple form of model editing), the respective downstream performance significantly degrades, showing the importance of the property neurons. We apply this approach to pruning the feedforward layers in Transformers, where most of the model parameters are. We show that protecting property neurons during pruning is significantly more effective than norm-based pruning.
【11】 LLaMA-Omni: Seamless Speech Interaction with Large Language Models
标题: LLaMA-Omni:与大型语言模型的无缝语音交互
作者:Qingkai Fang,Shoutao Guo,Yan Zhou,Zhengrui Ma,Shaolei Zhang,Yang Feng
备注:Preprint. Project: this https URL
链接:点击下载PDF文件
摘要:像GPT-4 o这样的模型可以通过语音与大型语言模型(LLM)进行实时交互,与传统的基于文本的交互相比,显著增强了用户体验。然而,如何构建基于开源LLM的语音交互模型,目前还缺乏探索。为了解决这个问题,我们提出了LLaMA-Omni,一种新的模型架构,设计用于与LLM进行低延迟和高质量的语音交互。LLaMA-Omni集成了一个预训练的语音编码器,一个语音适配器,一个LLM和一个流式语音解码器。它消除了对语音转录的需要,并且可以以极低的延迟直接从语音指令同时生成文本和语音响应。我们基于最新的Llama-3.1-8B-Instruct模型构建我们的模型。为了使模型与语音交互场景相匹配,我们构建了一个名为InstructS 2S-200 K的数据集,其中包括200 K语音指令和相应的语音响应。实验结果表明,与以往的语音语言模型相比,LLaMA-Omni提供了更好的响应在内容和风格,响应延迟低至226毫秒。此外,在4个GPU上训练LLaMA-Omni需要不到3天的时间,为未来高效开发语音语言模型铺平了道路。摘要:Models like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how to build speech interaction models based on open-source LLMs. To address this, we propose LLaMA-Omni, a novel model architecture designed for low-latency and high-quality speech interaction with LLMs. LLaMA-Omni integrates a pretrained speech encoder, a speech adaptor, an LLM, and a streaming speech decoder. It eliminates the need for speech transcription, and can simultaneously generate text and speech responses directly from speech instructions with extremely low latency. We build our model based on the latest Llama-3.1-8B-Instruct model. To align the model with speech interaction scenarios, we construct a dataset named InstructS2S-200K, which includes 200K speech instructions and corresponding speech responses. Experimental results show that compared to previous speech-language models, LLaMA-Omni provides better responses in both content and style, with a response latency as low as 226ms. Additionally, training LLaMA-Omni takes less than 3 days on just 4 GPUs, paving the way for the efficient development of speech-language models in the future.
【12】 MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
标题: MoWE-音频:具有混合弱编码器的多任务音频LLM
作者:Wenyu Zhang,Shuo Sun,Bin Wang,Xunlong Zou,Zhuohan Liu,Yingxu He,Geyu Lin,Nancy F. Chen,Ai Ti Aw
链接:点击下载PDF文件
摘要:大型语言模型(LLM)的快速发展显着增强了自然语言处理能力,促进了AudioLLM的开发,它可以处理和理解语音和音频输入以及文本。现有的AudioLLM通常将预训练的音频编码器与预训练的LLM相结合,随后在特定的音频任务上进行微调。然而,预先训练的音频编码器具有有限的能力来捕获新任务和数据集的特征。为了解决这个问题,我们建议将混合的“弱”编码器(MoWE)到AudioLLM框架。MoWE用相对较轻权重的编码器池补充基本编码器,基于音频输入选择性地激活以增强特征提取而不显著增加模型大小。我们的实证结果表明,MoWE有效地提高了多任务性能,扩大了AudioLLM的适用性,以更多样化的音频任务。摘要:The rapid advancements in large language models (LLMs) have significantly enhanced natural language processing capabilities, facilitating the development of AudioLLMs that process and understand speech and audio inputs alongside text. Existing AudioLLMs typically combine a pre-trained audio encoder with a pre-trained LLM, which are subsequently finetuned on specific audio tasks. However, the pre-trained audio encoder has constrained capacity to capture features for new tasks and datasets. To address this, we propose to incorporate mixtures of weak' encoders (MoWE) into the AudioLLM framework. MoWE supplements a base encoder with a pool of relatively light weight encoders, selectively activated based on the audio input to enhance feature extraction without significantly increasing model size. Our empirical results demonstrate that MoWE effectively improves multi-task performance, broadening the applicability of AudioLLMs to more diverse audio tasks.
【13】 Sine, Transient, Noise Neural Modeling of Piano Notes
标题: 钢琴音符的弦、瞬时、噪音神经建模
作者:Riccardo Simionato,Stefano Fasciani
链接:点击下载PDF文件
摘要:介绍了一种新颖的钢琴音色仿真方法。我们建议利用正弦,瞬态,和噪声分解设计一个微分频谱建模合成器复制钢琴音符。三个子模块从钢琴录音中学习这些分量,并生成相应的谐波、瞬态和噪声信号。将仿真分为三个独立的可训练模型,降低了建模任务的复杂性。准谐波内容是使用由物理导出的公式指导的可微分正弦模型产生的,其参数是从音频记录中自动估计的。噪声子模块使用可学习的时变滤波器,瞬态使用深度卷积网络生成。从单个音符开始,我们使用基于卷积的网络模拟毛键中不同键之间的耦合。结果表明,该模型匹配的部分分布的目标,而预测的能量在较高的部分的频谱提出了更多的挑战。瞬态和噪声分量的频谱中的能量分布总体上是准确的。虽然该模型在计算和内存效率上更高,但感知测试显示,在准确建模音符的攻击阶段方面存在局限性。尽管如此,它通常在模拟单音符和三弦琴时达到感知准确性。摘要:This paper introduces a novel method for emulating piano sounds. We propose to exploit the sine, transient, and noise decomposition to design a differentiable spectral modeling synthesizer replicating piano notes. Three sub-modules learn these components from piano recordings and generate the corresponding harmonic, transient, and noise signals. Splitting the emulation into three independently trainable models reduces the modeling tasks' complexity. The quasi-harmonic content is produced using a differentiable sinusoidal model guided by physics-derived formulas, whose parameters are automatically estimated from audio recordings. The noise sub-module uses a learnable time-varying filter, and the transients are generated using a deep convolutional network. From singular notes, we emulate the coupling between different keys in trichords with a convolutional-based network. Results show the model matches the partial distribution of the target while predicting the energy in the higher part of the spectrum presents more challenges. The energy distribution in the spectra of the transient and noise components is accurate overall. While the model is more computationally and memory efficient, perceptual tests reveal limitations in accurately modeling the attack phase of notes. Despite this, it generally achieves perceptual accuracy in emulating single notes and trichords.
【14】 An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition
标题: 一种有效的长尾语音识别上下文平衡自适应方法
作者:Yi-Cheng Wang,Li-Ting Pai,Bi-Cheng Yan,Hsin-Wei Wang,Chi-Han Lin,Berlin Chen
备注:Accepted by SLT 2024
链接:点击下载PDF文件
摘要:端到端(E2 E)自动语音识别(ASR)模型已经成为各种商业应用的标准实践。然而,在现实世界中,单词分布的长尾性质通常导致E2 E ASR模型在常见单词上表现良好,但在识别不常见单词方面表现不佳。最近,上下文适配器(CA)的概念被提出来注入外部知识表示的上下文单词列表到E2 E ASR模型。虽然CA可以提高对稀有词的识别性能,但仍然存在两个关键的数据不平衡问题。首先,当在训练期间使用低频词作为上下文词时,由于这些词很少出现在话语中,CA在关注令牌摘要:End-to-end (E2E) automatic speech recognition (ASR) models have become standard practice for various commercial applications. However, in real-world scenarios, the long-tailed nature of word distribution often leads E2E ASR models to perform well on common words but fall short in recognizing uncommon ones. Recently, the notion of a contextual adapter (CA) was proposed to infuse external knowledge represented by a context word list into E2E ASR models. Although CA can improve recognition performance on rare words, two crucial data imbalance problems remain. First, when using low-frequency words as context words during training, since these words rarely occur in the utterance, CA becomes prone to overfit on attending to the
【15】 Attention-Based Beamformer For Multi-Channel Speech Enhancement
标题: 用于多通道语音增强的基于注意力的束形成器
作者:Jinglin Bai,Hao Li,Xueliang Zhang,Fei Chen
链接:点击下载PDF文件
【16】 Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
标题: 通过对比学习和扩散模型通过自然语言指导增强情感文本到语音的可控制性
作者:Xin Jing,Kun Zhou,Andreas Triantafyllopoulos,Björn W. Schuller
链接:点击下载PDF文件
【17】 Human-mimetic binaural ear design and sound source direction estimation for task realization of musculoskeletal humanoids
标题: 用于肌肉骨骼人形机器人任务实现的仿人双耳设计和声音方向估计
作者:Yusuke Omura,Kento Kawaharazuka,Yuya Nagamatsu,Yuya Koga,Manabu Nishiura,Yasunori Toshimitsu,Yuki Asano,Kei Okada,Koji Kawasaki,Masayuki Inaba
备注:Accepted at ROBOMECH Journal
链接:点击下载PDF文件
【18】 Soft Acoustic Curvature Sensor: Design and Development
标题: 软声学弯曲传感器:设计与开发
作者:Mohammad Sheikh Sofla,Hanita Golshanian,Vishnu Rajendran S,Amir Ghalamzan E
备注:To appear in Robotics and Automation Letter
链接:点击下载PDF文件
【19】 SpeechTaxi: On Multilingual Semantic Speech Classification
标题: SpeechTaxi:多语言语音语义分类
作者:Lennart Keller,Goran Glavaš
链接:点击下载PDF文件
【20】 VoiceWukong: Benchmarking Deepfake Voice Detection
标题: Voice悟空:Deepfake语音检测基准
作者:Ziwei Yan,Yanjie Zhao,Haoyu Wang
链接:点击下载PDF文件
【21】 An End-to-End Approach for Chord-Conditioned Song Generation
标题: 和弦条件歌曲生成的端到端方法
作者:Shuochen Gao,Shun Lei,Fan Zhuo,Hangyu Liu,Feng Liu,Boshi Tang,Qiaochu Huang,Shiyin Kang,Zhiyong Wu
链接:点击下载PDF文件
【22】 Spectral oversubtraction? An approach for speech enhancement after robot ego speech filtering in semi-real-time
标题: 光谱过度减法?半实时机器人自我语音过滤后的语音增强方法
作者:Yue Li,Koen V. Hindriks,Florian A. Kunneman
备注:6 pages, 2 figures, submitted to 2025 IEEE ICASSP
链接:点击下载PDF文件
【23】 A Two-Stage Band-Split Mamba-2 Network for Music Separation
标题: 用于音乐分离的两级带宽Mamba-2网络
作者:Jinglin Bai,Yuan Fang,Jiajie Wang,Xueliang Zhang
链接:点击下载PDF文件
【24】 RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
标题: RobustCSV:基于HuBERT的旋律提取器和对抗学习,用于稳健的歌唱声音转换
作者:Wei Chen,Xintao Zhao,Jun Chen,Binzhu Sha,Zhiwei Lin,Zhiyong Wu
备注:Accepted by ISCSLP 2024
链接:点击下载PDF文件
【25】 Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models
标题: 增强大型音频语言模型音频问题回答中的时间理解
作者:Arvind Krishna Sridhar,Yinyi Guo,Erik Visser
备注:5 pages, 3 figures
链接:点击下载PDF文件
【26】 Advancing Topic Segmentation of Broadcasted Speech with Multilingual Semantic Embeddings
标题: 利用多语言语义嵌入推进广播语音的主题分割
作者:Sakshi Deo Shukla,Pavel Denisov,Tugtekin Turan
链接:点击下载PDF文件
【27】 MTDA-HSED: Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection
标题: MTDA-HMED:用于异类声音事件检测的互助调谐和双分支聚集
作者:Zehao Wang,Haobo Yue,Zhicheng Zhang,Da Mu,Jin Tang,Jianqin Yin
备注:Submit to Icassp2025
链接:点击下载PDF文件
【28】 DENSE: Dynamic Embedding Causal Target Speech Extraction
标题: DENSE:动态嵌入因果目标语音提取
作者:Yiwen Wang,Zeyu Yuan,Xihong Wu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【29】 Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis
标题: 绘制音频:利用多指令进行视频到音频合成
作者:Qi Yang,Binjie Mao,Zili Wang,Xing Nie,Pengfei Gao,Ying Guo,Cheng Zhen,Pengfei Yan,Shiming Xiang
备注:14 pages, 11 figures
链接:点击下载PDF文件
【30】 Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
标题: 无监督音乐音频音色传输的潜在扩散桥梁
作者:Michele Mancusi,Yurii Halychansky,Kin Wai Cheuk,Chieh-Hsin Lai,Stefan Uhlich,Junghyun Koo,Marco A. Martínez-Ramírez,Wei-Hsiang Liao,Giorgio Fabbro,Yuhki Mitsufuji
链接:点击下载PDF文件
【31】 Investigating Causal Cues: Strengthening Spoofed Audio Detection with Human-Discernible Linguistic Features
标题: 调查因果线索:利用人类可辨别的语言特征加强欺骗音频检测
作者:Zahra Khanjani,Tolulope Ale,Jianwu Wang,Lavon Davis,Christine Mallinson,Vandana P. Janeja
链接:点击下载PDF文件
【32】 SongCreator: Lyrics-based Universal Song Generation
标题: SongCreator:基于歌词的通用歌曲一代
作者:Shun Lei,Yixuan Zhou,Boshi Tang,Max W. Y. Lam,Feng Liu,Hangyu Liu,Jingcheng Wu,Shiyin Kang,Zhiyong Wu,Helen Meng
备注:work in progress
链接:点击下载PDF文件
【33】 Musical Chords: A Novel Java Algorithm and App Utility to Enumerate Chord-Progressions Adhering to Music Theory Guidelines
标题: 音乐和弦:一种新颖的Java算法和应用程序实用程序,用于列举符合音乐理论准则的和弦进行
作者:Aditya Lakshminarasimhan
备注:10 Pages, 5 Figures
链接:点击下载PDF文件
【34】 Continuous Learning of Transformer-based Audio Deepfake Detection
标题: 基于变形器的音频深度伪造检测的持续学习
作者:Tuan Duy Nguyen Le,Kah Kuan Teh,Huy Dat Tran
备注:Submitted to INTERSPEECH 2024
链接:点击下载PDF文件
【15】 Attention-Based Beamformer For Multi-Channel Speech Enhancement
标题: 用于多通道语音增强的基于注意力的束形成器
作者:Jinglin Bai,Hao Li,Xueliang Zhang,Fei Chen
链接:点击下载PDF文件
摘要:最小方差无失真响应(MVDR)是一种经典的自适应波束形成器,理论上可以保证信号在目标方向上的无失真传输。它的降噪性能实际上取决于噪声空间协方差矩阵(SCM)估计的准确性。尽管最近的深度学习在多通道语音增强方面表现出了卓越的性能,但无失真响应的特性仍然使MVDR在实际应用中非常受欢迎。在本文中,我们提出了一种基于注意力的机制来计算语音和噪声SCM,然后应用MVDR来获得增强的语音。此外,使用原地卷积算子和频率无关LSTM的深度学习架构已被证明在促进SCM估计方面是有效的。该模型以端到端的方式进行优化。实验结果表明,该方法是非常有效的跟踪运动或静止的说话人在非因果和因果条件下,优于其他基线。值得一提的是,我们的模型只有35万个参数,很容易部署在边缘设备上。摘要:Minimum Variance Distortionless Response (MVDR) is a classical adaptive beamformer that theoretically ensures the distortionless transmission of signals in the target direction. Its performance in noise reduction actually depends on the accuracy of the noise spatial covariance matrix (SCM) estimate. Although recent deep learning has shown remarkable performance in multi-channel speech enhancement, the property of distortionless response still makes MVDR highly popular in real applications. In this paper, we propose an attention-based mechanism to calculate the speech and noise SCM and then apply MVDR to obtain the enhanced speech. Moreover, a deep learning architecture using the inplace convolution operator and frequency-independent LSTM has proven effective in facilitating SCM estimation. The model is optimized in an end-to-end manner. Experimental results indicate that the proposed method is extremely effective in tracking moving or stationary speakers under non-causal and causal conditions, outperforming other baselines. It is worth mentioning that our model has only 0.35 million parameters, making it easy to be deployed on edge devices.
【16】 Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
标题: 通过对比学习和扩散模型通过自然语言指导增强情感文本到语音的可控制性
作者:Xin Jing,Kun Zhou,Andreas Triantafyllopoulos,Björn W. Schuller
链接:点击下载PDF文件
摘要:虽然目前的情感文本到语音(TTS)系统可以生成高度可理解的情感语音,实现输出语音的情感渲染的精细控制仍然是一个重大的挑战。在本文中,我们介绍了ParaEVITS,这是一种新颖的情感TTS框架,它利用自然语言的组合性来增强对情感渲染的控制。通过结合一个文本音频编码器的启发ParaCLAP,一个对比语言音频预训练(CLAP)模型计算语言学,扩散模型进行训练,以生成基于文本情感风格描述的情感嵌入。我们的框架首先使用音频编码器对参考音频进行训练,然后微调扩散模型以处理来自ParaCLAP文本编码器的文本输入。在推理过程中,语音属性,如音高,抖动和响度操纵只使用文本条件。我们的实验表明,ParaEVITS有效地控制情感渲染,而不影响语音质量。演讲演示是公开的。摘要:While current emotional text-to-speech (TTS) systems can generate highly intelligible emotional speech, achieving fine control over emotion rendering of the output speech still remains a significant challenge. In this paper, we introduce ParaEVITS, a novel emotional TTS framework that leverages the compositionality of natural language to enhance control over emotional rendering. By incorporating a text-audio encoder inspired by ParaCLAP, a contrastive language-audio pretraining (CLAP) model for computational paralinguistics, the diffusion model is trained to generate emotional embeddings based on textual emotional style descriptions. Our framework first trains on reference audio using the audio encoder, then fine-tunes a diffusion model to process textual inputs from ParaCLAP's text encoder. During inference, speech attributes such as pitch, jitter, and loudness are manipulated using only textual conditioning. Our experiments demonstrate that ParaEVITS effectively control emotion rendering without compromising speech quality. Speech demos are publicly available.
【17】 Human-mimetic binaural ear design and sound source direction estimation for task realization of musculoskeletal humanoids
标题: 用于肌肉骨骼人形机器人任务实现的仿人双耳设计和声音方向估计
作者:Yusuke Omura,Kento Kawaharazuka,Yuya Nagamatsu,Yuya Koga,Manabu Nishiura,Yasunori Toshimitsu,Yuki Asano,Kei Okada,Koji Kawasaki,Masayuki Inaba
备注:Accepted at ROBOMECH Journal
链接:点击下载PDF文件
摘要:肌肉骨骼类人机器人的类人环境识别对于在真实复杂环境中实现任务和用作测试对象的假人是重要的。人类整合各种感官信息来感知周围环境,听觉对于识别视线外或触摸不到的物体特别有用。在本研究中,我们的目标是实现类人听觉环境识别和任务实现的肌肉骨骼类人通过配备一个类似人类的听觉处理系统。人类通过估计声源的方向并基于传入声音的时域和频域的变化以及听觉信息在中枢神经系统中的整合来检测环境声音,从而实现基于声音的环境识别。我们提出了一个仿人听觉信息处理系统,它由三个部分组成:仿人双耳耳单元,它模仿人耳的结构和特性,声源方向估计系统,和环境声音检测系统,它模仿在中枢神经系统中的处理。我们将其应用于Musashi,一个人类模仿肌肉骨骼类人动物,并让它执行任务,需要在真实的嘈杂环境中的视图之外的声音信息,以确认所提出的方法的有用性。摘要:Human-like environment recognition by musculoskeletal humanoids is important for task realization in real complex environments and for use as dummies for test subjects. Humans integrate various sensory information to perceive their surroundings, and hearing is particularly useful for recognizing objects out of view or out of touch. In this research, we aim to realize human-like auditory environmental recognition and task realization for musculoskeletal humanoids by equipping them with a human-like auditory processing system. Humans realize sound-based environmental recognition by estimating directions of the sound sources and detecting environmental sounds based on changes in the time and frequency domain of incoming sounds and the integration of auditory information in the central nervous system. We propose a human mimetic auditory information processing system, which consists of three components: the human mimetic binaural ear unit, which mimics human ear structure and characteristics, the sound source direction estimation system, and the environmental sound detection system, which mimics processing in the central nervous system. We apply it to Musashi, a human mimetic musculoskeletal humanoid, and have it perform tasks that require sound information outside of view in real noisy environments to confirm the usefulness of the proposed methods.
【18】 Soft Acoustic Curvature Sensor: Design and Development
标题: 软声学弯曲传感器:设计与开发
作者:Mohammad Sheikh Sofla,Hanita Golshanian,Vishnu Rajendran S,Amir Ghalamzan E
备注:To appear in Robotics and Automation Letter
链接:点击下载PDF文件
摘要:介绍了一种新型的软声曲率传感器。SAC采用集成音频组件,并在灵活的结构中具有声学通道。由通道一端的扬声器产生的参考声波传播并由另一通道端的麦克风接收。我们以前的研究表明,声波能量耗散随声道变形而变化,这使我们设计了一种能够因弯曲而大变形的新型声道。然后,我们使用机器学习(ML)模型来建立通道变形和声音调制之间的复杂映射。各种声音频率和ML模型进行了评估,以提高曲率检测精度。该传感器使用软材料和3D打印制成,经过实验验证,曲率测量误差在0至60 m-1曲率范围内保持在3.5 m-1以内。这些结果证明了所提出的方法估计曲率的有效性。由于其灵活的结构,SAC传感器在软机器人中具有应用潜力,包括连续体机械手,软夹持器和可穿戴设备的形状测量。摘要:This paper introduces a novel Soft Acoustic Curvature (SAC) sensor. SAC incorporates integrated audio components and features an acoustic channel within a flexible structure. A reference acoustic wave, generated by a speaker at one end of the channel, propagates and is received by a microphone at the other channel's end. Our previous study revealed that acoustic wave energy dissipation varies with acoustic channel deformation, leading us to design a novel channel capable of large deformation due to bending. We then use Machine Learning (ML) models to establish a complex mapping between channel deformations and sound modulation. Various sound frequencies and ML models were evaluated to enhance curvature detection accuracy. The sensor, constructed using soft material and 3D printing, was validated experimentally, with curvature measurement errors remaining within 3.5 m-1 for a range of 0 to 60 m-1 curvatures. These results demonstrate the effectiveness of the proposed method for estimating curvatures. With its flexible structure, the SAC sensor holds potential for applications in soft robotics, including shape measurement for continuum manipulators, soft grippers, and wearable devices.
【19】 SpeechTaxi: On Multilingual Semantic Speech Classification
标题: SpeechTaxi:多语言语音语义分类
作者:Lennart Keller,Goran Glavaš
链接:点击下载PDF文件
摘要:多语言语音编码和转录的最新进展提出了最有效的语义语音分类方法的问题。具体而言,(1)通过微调最先进的多语言语音编码器(MSE)获得的端到端(E2 E)分类器是否可以匹配或超过(2)级联(CA)的性能,其中语音首先被转录为文本,并将分类委托给基于文本的分类器。为了回答这个问题,我们首先构建了SpeechTaxi,这是一个80小时的多语言数据集,用于圣经经文的语义语音分类,涵盖28种不同的语言。然后,我们利用SpeechTaxi进行了广泛的实验,比较E2 E和CA在单语语义语音分类以及跨语言迁移。我们发现,基于MSE的E2 E在单语设置中优于CA,即,在语言数据上训练时。然而,MSE似乎具有较差的跨语言迁移能力,E2 E在(1)zero-shot迁移到训练中未见过的语言和(2)多语言训练(即,多语种联合培训。最后,我们设计了一种新的CA方法的基础上转录罗马化文本作为一种语言不可知的中间表示,并表明它代表了一个强大的解决方案,没有本地ASR支持的语言。我们的SpeechTaxi数据集可在以下网址公开获取:https: huggingface.co datasets LennartKeller SpeechTaxi 。摘要:Recent advancements in multilingual speech encoding as well as transcription raise the question of the most effective approach to semantic speech classification. Concretely, can (1) end-to-end (E2E) classifiers obtained by fine-tuning state-of-the-art multilingual speech encoders (MSEs) match or surpass the performance of (2) cascading (CA), where speech is first transcribed into text and classification is delegated to a text-based classifier. To answer this, we first construct SpeechTaxi, an 80-hour multilingual dataset for semantic speech classification of Bible verses, covering 28 diverse languages. We then leverage SpeechTaxi to conduct a wide range of experiments comparing E2E and CA in monolingual semantic speech classification as well as in cross-lingual transfer. We find that E2E based on MSEs outperforms CA in monolingual setups, i.e., when trained on in-language data. However, MSEs seem to have poor cross-lingual transfer abilities, with E2E substantially lagging CA both in (1) zero-shot transfer to languages unseen in training and (2) multilingual training, i.e., joint training on multiple languages. Finally, we devise a novel CA approach based on transcription to Romanized text as a language-agnostic intermediate representation and show that it represents a robust solution for languages without native ASR support. Our SpeechTaxi dataset is publicly available at: https: huggingface.co datasets LennartKeller SpeechTaxi .
【20】 VoiceWukong: Benchmarking Deepfake Voice Detection
标题: Voice悟空:Deepfake语音检测基准
作者:Ziwei Yan,Yanjie Zhao,Haoyu Wang
链接:点击下载PDF文件
摘要:随着文本到语音(TTS)和语音转换(VC)等技术的快速发展,检测深度伪造语音变得越来越重要。然而,学术界和工业界都缺乏一个全面和直观的基准来评估探测器。现有的数据集在语言多样性方面受到限制,并且缺乏在现实世界的生产环境中遇到的许多操作。 为了填补这一空白,我们提出了VoiceWukong,这是一个旨在评估deepfake语音检测器性能的基准测试。为了构建数据集,我们首先收集了由19种先进且广泛认可的商业工具和15种开源工具生成的deepfake语音。然后,我们创建了38个数据变量,涵盖了六种类型的操作,构建了用于deepfake语音检测的评估数据集。因此,VoiceWukong包括265,200个英语和148,200个中文deepfake语音样本。使用VoiceWukong,我们评估了12个最先进的探测器。AASIST 2的最佳等误率(EER)为13.50%,而其他所有算法均超过20%。我们的研究结果表明,这些检测器在现实应用中面临着重大挑战,性能急剧下降。此外,我们还进行了一项有300多名参与者的用户研究。将结果与12个检测器和多模型大语言模型(MLLM)的性能进行比较,即,Qwen 2-Audio,其中不同的检测器和人类在不同的欺骗水平下表现出不同的识别能力,而LALM则完全没有检测能力。此外,我们还提供了deepfake语音检测的排行榜,可在{https: www.example.com}上公开获取。voicewukong.github.io摘要:With the rapid advancement of technologies like text-to-speech (TTS) and voice conversion (VC), detecting deepfake voices has become increasingly crucial. However, both academia and industry lack a comprehensive and intuitive benchmark for evaluating detectors. Existing datasets are limited in language diversity and lack many manipulations encountered in real-world production environments. To fill this gap, we propose VoiceWukong, a benchmark designed to evaluate the performance of deepfake voice detectors. To build the dataset, we first collected deepfake voices generated by 19 advanced and widely recognized commercial tools and 15 open-source tools. We then created 38 data variants covering six types of manipulations, constructing the evaluation dataset for deepfake voice detection. VoiceWukong thus includes 265,200 English and 148,200 Chinese deepfake voice samples. Using VoiceWukong, we evaluated 12 state-of-the-art detectors. AASIST2 achieved the best equal error rate (EER) of 13.50%, while all others exceeded 20%. Our findings reveal that these detectors face significant challenges in real-world applications, with dramatically declining performance. In addition, we conducted a user study with more than 300 participants. The results are compared with the performance of the 12 detectors and a multimodel large language model (MLLM), i.e., Qwen2-Audio, where different detectors and humans exhibit varying identification capabilities for deepfake voices at different deception levels, while the LALM demonstrates no detection ability at all. Furthermore, we provide a leaderboard for deepfake voice detection, publicly available at {https: voicewukong.github.io}.
【21】 An End-to-End Approach for Chord-Conditioned Song Generation
标题: 和弦条件歌曲生成的端到端方法
作者:Shuochen Gao,Shun Lei,Fan Zhuo,Hangyu Liu,Feng Liu,Boshi Tang,Qiaochu Huang,Shiyin Kang,Zhiyong Wu
链接:点击下载PDF文件
摘要:歌曲生成任务的目标是从给定的歌词合成由人声和伴奏组成的音乐。虽然现有的方法,Juventure,已经探索了这一任务,其约束控制的代往往导致音乐表现的不足。为了缓解这个问题,我们引入了一个重要的概念,从音乐创作,即和弦,歌曲生成网络。和弦构成了伴奏的基础,并为声乐旋律提供了相关的和声。针对自动和弦提取器的不准确性,设计了一种基于动态权值序列的交叉注意机制,将提取的和弦信息整合到歌曲生成中,减少了帧级错误,并在此基础上提出了一种新的和弦条件歌曲生成器模型CSG,实验结果表明,该方法在音乐性能和生成歌曲的控制精度方面优于其他方法.摘要:The Song Generation task aims to synthesize music composed of vocals and accompaniment from given lyrics. While the existing method, Jukebox, has explored this task, its constrained control over the generations often leads to deficiency in music performance. To mitigate the issue, we introduce an important concept from music composition, namely chords, to song generation networks. Chords form the foundation of accompaniment and provide vocal melody with associated harmony. Given the inaccuracy of automatic chord extractors, we devise a robust cross-attention mechanism augmented with dynamic weight sequence to integrate extracted chord information into song generations and reduce frame-level flaws, and propose a novel model termed Chord-Conditioned Song Generator (CSG) based on it. Experimental evidence demonstrates our proposed method outperforms other approaches in terms of musical performance and control precision of generated songs.
【22】 Spectral oversubtraction? An approach for speech enhancement after robot ego speech filtering in semi-real-time
标题: 光谱过度减法?半实时机器人自我语音过滤后的语音增强方法
作者:Yue Li,Koen V. Hindriks,Florian A. Kunneman
备注:6 pages, 2 figures, submitted to 2025 IEEE ICASSP
链接:点击下载PDF文件
摘要:谱减法,广泛使用的简单,已被用来解决机器人自我语音过滤(RESF)的问题,从机器人的单通道麦克风录音时,它正在说话的人中断检测语音内容。然而,这种方法遭受在基频范围(FFR)中的过减法,从而导致语音内容识别降级。为了解决这个问题,我们提出了一个基于双掩码一致性的度量生成对抗网络(CMGAN),以增强检测到的语音,提高识别结果。我们的模型用高频信息和长期特征补偿减影过多的血流储备分数值,然后对新的频谱图进行去噪。此外,我们还介绍了一种增量处理方法,该方法允许在长固定长度输入上训练的网络上使用流式输入进行半实时音频处理。两个数据集的评估,包括一个看不见的噪音,证明显着提高识别准确性和有效性的建议的双掩模方法和增量处理,提高鲁棒性的建议RESF管道在现实世界的HRI场景。摘要:Spectral subtraction, widely used for its simplicity, has been employed to address the Robot Ego Speech Filtering (RESF) problem for detecting speech contents of human interruption from robot's single-channel microphone recordings when it is speaking. However, this approach suffers from oversubtraction in the fundamental frequency range (FFR), leading to degraded speech content recognition. To address this, we propose a Two-Mask Conformer-based Metric Generative Adversarial Network (CMGAN) to enhance the detected speech and improve recognition results. Our model compensates for oversubtracted FFR values with high-frequency information and long-term features and then de-noises the new spectrogram. In addition, we introduce an incremental processing method that allows semi-real-time audio processing with streaming input on a network trained on long fixed-length input. Evaluations of two datasets, including one with unseen noise, demonstrate significant improvements in recognition accuracy and the effectiveness of the proposed two-mask approach and incremental processing, enhancing the robustness of the proposed RESF pipeline in real-world HRI scenarios.
【23】 A Two-Stage Band-Split Mamba-2 Network for Music Separation
标题: 用于音乐分离的两级带宽Mamba-2网络
作者:Jinglin Bai,Yuan Fang,Jiajie Wang,Xueliang Zhang
链接:点击下载PDF文件
摘要:音乐源分离(MSS)旨在将混合音乐分离成不同的音轨,如人声,低音,鼓等。由于音乐信号的复杂性,MSS被认为是一项具有挑战性的音频分离任务。尽管RNN和Transformer架构并不完美,但它们通常用于为MSS建模音乐序列。最近,Mamba-2已经在各种顺序建模任务中表现出了高效率,但其优越性尚未在MSS中得到研究。本文应用Mamba-2算法,采用两阶段策略,在掩模方法的基础上引入残差映射,有效地补偿了掩模中缺失的细节,进一步提高了分离性能。实验证明了双向Mamba-2的优越性和两级网络在MSS中的有效性。源代码可在https: github.com baijinglin TS-BSmamba2上公开访问。摘要:Music source separation (MSS) aims to separate mixed music into its distinct tracks, such as vocals, bass, drums, and more. MSS is considered to be a challenging audio separation task due to the complexity of music signals. Although the RNN and Transformer architecture are not perfect, they are commonly used to model the music sequence for MSS. Recently, Mamba-2 has already demonstrated high efficiency in various sequential modeling tasks, but its superiority has not been investigated in MSS. This paper applies Mamba-2 with a two-stage strategy, which introduces residual mapping based on the mask method, effectively compensating for the details absent in the mask and further improving separation performance. Experiments confirm the superiority of bidirectional Mamba-2 and the effectiveness of the two-stage network in MSS. The source code is publicly accessible at https: github.com baijinglin TS-BSmamba2.
【24】 RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
标题: RobustCSV:基于HuBERT的旋律提取器和对抗学习,用于稳健的歌唱声音转换
作者:Wei Chen,Xintao Zhao,Jun Chen,Binzhu Sha,Zhiwei Lin,Zhiyong Wu
备注:Accepted by ISCSLP 2024
链接:点击下载PDF文件
摘要:由于在推断过程中使用非鲁棒的方法来提取音高和能量,因此歌唱声音转换(SVC)受到噪声敏感性的阻碍。由于干净的信号是SVC中源音频的关键,因此音乐源分离预处理为处理嘈杂音频提供了可行的解决方案,例如与背景音乐(BGM)一起唱歌。然而,目前的分离方法难以完全去除噪声或过度抑制信号分量,影响了处理后音频的自然度和相似性。为了解决这个问题,我们的研究引入了RobustSVC,这是一种新颖的任意对一SVC框架,可以将嘈杂的人声转换为目标歌手演唱的干净人声。我们用基于HuBERT的旋律提取器替换非鲁棒特征,并使用具有三个鉴别器的对抗训练机制来减少自监督表示中的信息泄漏。实验结果表明,RobustSVC算法具有较好的抗噪性,在有噪和无噪的情况下都能获得比基线算法更高的相似度和自然度。摘要:Singing voice conversion (SVC) is hindered by noise sensitivity due to the use of non-robust methods for extracting pitch and energy during the inference. As clean signals are key for the source audio in SVC, music source separation preprocessing offers a viable solution for handling noisy audio, like singing with background music (BGM). However, current separating methods struggle to fully remove noise or excessively suppress signal components, affecting the naturalness and similarity of the processed audio. To tackle this, our study introduces RobustSVC, a novel any-to-one SVC framework that converts noisy vocals into clean vocals sung by the target singer. We replace the non-robust feature with a HuBERT-based melody extractor and use adversarial training mechanisms with three discriminators to reduce information leakage in self-supervised representations. Experimental results show that RobustSVC is noise-robust and achieves higher similarity and naturalness than baseline methods in both noisy and clean vocal conditions.
【25】 Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models
标题: 增强大型音频语言模型音频问题回答中的时间理解
作者:Arvind Krishna Sridhar,Yinyi Guo,Erik Visser
备注:5 pages, 3 figures
链接:点击下载PDF文件
摘要:音频问题分类任务包括音频事件分类、音频字幕和开放式推理。最近,由于大型音频语言模型的出现,音频问题分类引起了人们的关注。当前的文献集中于通过投影模块将音频编码器与仅文本的大型语言模型集成来构建LALM。虽然大型音频语言模型在一般音频理解方面表现出色,但它们在时间推理方面受到限制,这可能会阻碍其商业应用和设备部署。本文讨论了这些挑战和限制音频时间推理。首先,我们介绍了一种数据增强技术,用于使用LLM生成可靠的音频时间问题和答案。其次,我们提出了一个持续的微调课程学习策略,专注于时间推理,而不影响微调任务的性能。最后,我们开发了一个可靠和透明的自动化指标,在LLM的帮助下,智能地测量大型音频语言模型响应和地面真实数据之间的相关性。我们证明了我们提出的技术使用SOTA LALM在公共音频基准数据集上的有效性。摘要:The Audio Question Answering task includes audio event classification, audio captioning, and open ended reasoning. Recently, Audio Question Answering has garnered attention due to the advent of Large Audio Language Models. Current literature focuses on constructing LALMs by integrating audio encoders with text only Large Language Models through a projection module. While Large Audio Language Models excel in general audio understanding, they are limited in temporal reasoning which may hinder their commercial applications and on device deployment. This paper addresses these challenges and limitations in audio temporal reasoning. First, we introduce a data augmentation technique for generating reliable audio temporal questions and answers using an LLM. Second, we propose a continued finetuning curriculum learning strategy to specialize in temporal reasoning without compromising performance on finetuned tasks. Finally, we develop a reliable and transparent automated metric, assisted by an LLM, to measure the correlation between Large Audio Language Model responses and ground truth data intelligently. We demonstrate the effectiveness of our proposed techniques using SOTA LALMs on public audio benchmark datasets.
【26】 Advancing Topic Segmentation of Broadcasted Speech with Multilingual Semantic Embeddings
标题: 利用多语言语义嵌入推进广播语音的主题分割
作者:Sakshi Deo Shukla,Pavel Denisov,Tugtekin Turan
链接:点击下载PDF文件
摘要:基于语音的主题分割的最新进展突出了预训练语音编码器直接从语音中捕获语义表示的潜力。传统上,主题分割依赖于流水线方法,其中自动语音识别系统的成绩单被生成,随后是基于文本的分割算法。在本文中,我们介绍了一个端到端的计划,绕过这个传统的两步过程中,直接采用语义语音编码器进行分割。专注于广播新闻领域,这带来了独特的挑战,由于扬声器的多样性和单一录音中的主题,我们解决了以端到端的方式有效地访问话题变化点的挑战。此外,我们提出了一个新的基准口语新闻主题分割,利用数据集具有约1000小时的公开可用的录音在六种欧洲语言,包括在印地语的评估集,以测试模型的跨域性能在跨语言,zero-shot的情况。这种设置反映了现实世界的多样性以及对适应各种语言环境的模型的需求。我们的结果表明,虽然传统的管道方法实现了最先进的$P_k$得分为0.2431的英语,我们的端到端模型提供了一个有竞争力的$P_k$得分为0.2564。当进行多语言训练时,这些分数分别进一步提高到0.1988和0.2370。为了支持进一步的研究,我们发布了我们的模型以及数据准备脚本,促进了对多语言口语新闻主题分割的开放研究。摘要:Recent advancements in speech-based topic segmentation have highlighted the potential of pretrained speech encoders to capture semantic representations directly from speech. Traditionally, topic segmentation has relied on a pipeline approach in which transcripts of the automatic speech recognition systems are generated, followed by text-based segmentation algorithms. In this paper, we introduce an end-to-end scheme that bypasses this conventional two-step process by directly employing semantic speech encoders for segmentation. Focused on the broadcasted news domain, which poses unique challenges due to the diversity of speakers and topics within single recordings, we address the challenge of accessing topic change points efficiently in an end-to-end manner. Furthermore, we propose a new benchmark for spoken news topic segmentation by utilizing a dataset featuring approximately 1000 hours of publicly available recordings across six European languages and including an evaluation set in Hindi to test the model's cross-domain performance in a cross-lingual, zero-shot scenario. This setup reflects real-world diversity and the need for models adapting to various linguistic settings. Our results demonstrate that while the traditional pipeline approach achieves a state-of-the-art $P_k$ score of 0.2431 for English, our end-to-end model delivers a competitive $P_k$ score of 0.2564. When trained multilingually, these scores further improve to 0.1988 and 0.2370, respectively. To support further research, we release our model along with data preparation scripts, facilitating open research on multilingual spoken news topic segmentation.
【27】 MTDA-HSED: Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection
标题: MTDA-HMED:用于异类声音事件检测的互助调谐和双分支聚集
作者:Zehao Wang,Haobo Yue,Zhicheng Zhang,Da Mu,Jin Tang,Jianqin Yin
备注:Submit to Icassp2025
链接:点击下载PDF文件
摘要:声音事件检测(SED)在理解和感知声学场景中起着至关重要的作用。以前的方法已经证明了令人印象深刻的能力。然而,它们在从异构数据集学习复杂场景的特征方面存在不足。本文介绍了一种新型的双分支体系结构,名为用于异构声音事件检测的互助调谐和双分支聚合(MTDA-HSED)。MTDA-HSED架构采用互助音频适配器(M3 A)来有效地解决多场景问题,并使用双分支中间融合(DBMF)模块来解决多粒度问题。具体来说,M3 A作为适配器集成到BEAT块中,通过在多场景数据集上对其进行微调来提高BEAT的性能。DBMF模块连接BEAT和CNN分支,这有助于来自BEAT和CNN分支的信息的深度融合。实验结果表明,在DESED和MAESTRO Real数据集上,该方法比mpAUC的基线提高了 textbf{$5 %$}.代码为 href{https: github.com Visitor-W MTDA}{here}。摘要:Sound Event Detection (SED) plays a vital role in comprehending and perceiving acoustic scenes. Previous methods have demonstrated impressive capabilities. However, they are deficient in learning features of complex scenes from heterogeneous dataset. In this paper, we introduce a novel dual-branch architecture named Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection (MTDA-HSED). The MTDA-HSED architecture employs the Mutual-Assistance Audio Adapter (M3A) to effectively tackle the multi-scenario problem and uses the Dual-Branch Mid-Fusion (DBMF) module to tackle the multi-granularity problem. Specifically, M3A is integrated into the BEATs block as an adapter to improve the BEATs' performance by fine-tuning it on the multi-scenario dataset. The DBMF module connects BEATs and CNN branches, which facilitates the deep fusion of information from the BEATs and the CNN branches. Experimental results show that the proposed methods exceed the baseline of mpAUC by textbf{$5 %$} on the DESED and MAESTRO Real datasets. Code is href{https: github.com Visitor-W MTDA}{here}.
【28】 DENSE: Dynamic Embedding Causal Target Speech Extraction
标题: DENSE:动态嵌入因果目标语音提取
作者:Yiwen Wang,Zeyu Yuan,Xihong Wu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:目标语音提取(TSE)的重点是从混合信号中提取特定目标说话人的语音。现有的TSE模型通常利用静态嵌入作为用于提取目标说话者的语音的条件。然而,静态嵌入往往无法捕捉提取的语音信号的上下文信息,这可能会限制模型的性能。我们提出了一种新的动态嵌入因果目标语音提取模型,以解决这一限制。我们的方法采用了自回归机制,根据提取的语音生成上下文相关的嵌入,从而实现实时的帧级提取。实验结果表明,该模型提高了短时目标可懂度(STOI)和信号失真比(SDR),为复杂场景下的目标语音提取提供了一种有前途的解决方案。摘要:Target speech extraction (TSE) focuses on extracting the speech of a specific target speaker from a mixture of signals. Existing TSE models typically utilize static embeddings as conditions for extracting the target speaker's voice. However, the static embeddings often fail to capture the contextual information of the extracted speech signal, which may limit the model's performance. We propose a novel dynamic embedding causal target speech extraction model to address this limitation. Our approach incorporates an autoregressive mechanism to generate context-dependent embeddings based on the extracted speech, enabling real-time, frame-level extraction. Experimental results demonstrate that the proposed model enhances short-time objective intelligibility (STOI) and signal-to-distortion ratio (SDR), offering a promising solution for target speech extraction in challenging scenarios.
【29】 Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis
标题: 绘制音频:利用多指令进行视频到音频合成
作者:Qi Yang,Binjie Mao,Zili Wang,Xing Nie,Pengfei Gao,Ying Guo,Cheng Zhen,Pengfei Yan,Shiming Xiang
备注:14 pages, 11 figures
链接:点击下载PDF文件
摘要:Foley是电影制作中常用的一个术语,指的是在无声电影或视频中添加日常音效,以增强听觉体验。视频到音频(V2 A),作为一种特殊类型的自动福利任务,提出了与视听同步相关的固有挑战。这些挑战包括保持输入视频和生成的音频之间的内容一致性,以及视频内的时间和响度属性的对齐。为了解决这些问题,我们构建了一个可控的视频到音频的合成模型,称为绘制音频,它支持多个输入指令,通过绘制掩码和响度信号。为了确保合成的音频和目标视频之间的内容一致性,我们引入了掩蔽注意模块(MAM),它采用掩蔽视频指令,使模型能够专注于感兴趣的区域。此外,我们还实现了时间响度模块(TLM),它使用辅助响度信号来确保声音的合成在响度和时间维度上与视频保持一致。此外,我们扩展了一个大规模的V2 A数据集,命名为VGGSound-Caption,通过注释字幕提示。在两个大规模V2 A数据集上对具有挑战性的基准进行了大量实验,验证了Draw an Audio达到了最先进的水平。项目页面:https: yannqi.github.io Draw-an-Audio 。摘要:Foley is a term commonly used in filmmaking, referring to the addition of daily sound effects to silent films or videos to enhance the auditory experience. Video-to-Audio (V2A), as a particular type of automatic foley task, presents inherent challenges related to audio-visual synchronization. These challenges encompass maintaining the content consistency between the input video and the generated audio, as well as the alignment of temporal and loudness properties within the video. To address these issues, we construct a controllable video-to-audio synthesis model, termed Draw an Audio, which supports multiple input instructions through drawn masks and loudness signals. To ensure content consistency between the synthesized audio and target video, we introduce the Mask-Attention Module (MAM), which employs masked video instruction to enable the model to focus on regions of interest. Additionally, we implement the Time-Loudness Module (TLM), which uses an auxiliary loudness signal to ensure the synthesis of sound that aligns with the video in both loudness and temporal dimensions. Furthermore, we have extended a large-scale V2A dataset, named VGGSound-Caption, by annotating caption prompts. Extensive experiments on challenging benchmarks across two large-scale V2A datasets verify Draw an Audio achieves the state-of-the-art. Project page: https: yannqi.github.io Draw-an-Audio .
【30】 Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
标题: 无监督音乐音频音色传输的潜在扩散桥梁
作者:Michele Mancusi,Yurii Halychansky,Kin Wai Cheuk,Chieh-Hsin Lai,Stefan Uhlich,Junghyun Koo,Marco A. Martínez-Ramírez,Wei-Hsiang Liao,Giorgio Fabbro,Yuhki Mitsufuji
链接:点击下载PDF文件
摘要:音乐音色转移是一项具有挑战性的任务,涉及修改音频信号的音色特性,同时保持其旋律结构。在本文中,我们提出了一种基于双扩散桥的新方法,使用CocoChorales数据集进行训练,该数据集由未配对的单声道单乐器音频数据组成。每个扩散模型都是在具有高斯先验的特定仪器上训练的。在推断期间,将一个模型指定为源模型以将输入音频映射到其对应的高斯先验,并且将另一个模型指定为目标模型以从该高斯先验重构目标音频,从而促进音色传递。我们将我们的方法与现有的无监督音色传输模型(如VAEGAN和高斯流桥(GFB))进行比较。实验结果表明,我们的方法实现了更好的Fr 'echet音频距离(FAD)和旋律保存,反映了较低的音高距离(DPD)相比,VAEGAN和GFB。此外,我们发现,从高斯先验,$ sigma$的噪声水平,可以进行调整,以控制旋律保存的程度和转移的音色量。摘要:Music timbre transfer is a challenging task that involves modifying the timbral characteristics of an audio signal while preserving its melodic structure. In this paper, we propose a novel method based on dual diffusion bridges, trained using the CocoChorales Dataset, which consists of unpaired monophonic single-instrument audio data. Each diffusion model is trained on a specific instrument with a Gaussian prior. During inference, a model is designated as the source model to map the input audio to its corresponding Gaussian prior, and another model is designated as the target model to reconstruct the target audio from this Gaussian prior, thereby facilitating timbre transfer. We compare our approach against existing unsupervised timbre transfer models such as VAEGAN and Gaussian Flow Bridges (GFB). Experimental results demonstrate that our method achieves both better Fr 'echet Audio Distance (FAD) and melody preservation, as reflected by lower pitch distances (DPD) compared to VAEGAN and GFB. Additionally, we discover that the noise level from the Gaussian prior, $ sigma$, can be adjusted to control the degree of melody preservation and amount of timbre transferred.
【31】 Investigating Causal Cues: Strengthening Spoofed Audio Detection with Human-Discernible Linguistic Features
标题: 调查因果线索:利用人类可辨别的语言特征加强欺骗音频检测
作者:Zahra Khanjani,Tolulope Ale,Jianwu Wang,Lavon Davis,Christine Mallinson,Vandana P. Janeja
链接:点击下载PDF文件
摘要:几种类型的欺骗音频,如模仿,重放攻击和深度伪造,对信息完整性造成了社会挑战。最近,研究人员与社会语言学专家合作,用人类耳朵可以识别的专家定义语言特征(EDLF)来标记欺骗性音频样本:音调,停顿,辅音停止的单词首字母和单词结尾释放,可听到的吸气或呼气以及整体音频质量。已经确定,当几种深度伪造检测算法使用这些EDLF增强音频数据的传统和常见特征时,它们会有所改进。在本文中,使用由多种类型的欺骗音频增强社会语言学注释组成的混合数据集,我们研究了因果发现和可辨别的语言特征和音频片段中的标签之间的推断,将因果模型的结果与专家地面实况验证标记过程进行比较。我们的研究结果表明,因果模型表明了结合语言特征以帮助识别欺骗音频的实用性,以及将人类知识纳入模型和技术以加强AI模型的整体需求和机会。因果发现和推理可用作训练人类辨别欺骗音频以及自动化EDLF标记的基础,以提高常见的基于人工智能的欺骗音频检测器的性能。摘要:Several types of spoofed audio, such as mimicry, replay attacks, and deepfakes, have created societal challenges to information integrity. Recently, researchers have worked with sociolinguistics experts to label spoofed audio samples with Expert Defined Linguistic Features (EDLFs) that can be discerned by the human ear: pitch, pause, word-initial and word-final release bursts of consonant stops, audible intake or outtake of breath, and overall audio quality. It is established that there is an improvement in several deepfake detection algorithms when they augmented the traditional and common features of audio data with these EDLFs. In this paper, using a hybrid dataset comprised of multiple types of spoofed audio augmented with sociolinguistic annotations, we investigate causal discovery and inferences between the discernible linguistic features and the label in the audio clips, comparing the findings of the causal models with the expert ground truth validation labeling process. Our findings suggest that the causal models indicate the utility of incorporating linguistic features to help discern spoofed audio, as well as the overall need and opportunity to incorporate human knowledge into models and techniques for strengthening AI models. The causal discovery and inference can be used as a foundation of training humans to discern spoofed audio as well as automating EDLFs labeling for the purpose of performance improvement of the common AI-based spoofed audio detectors.
【32】 SongCreator: Lyrics-based Universal Song Generation
标题: SongCreator:基于歌词的通用歌曲一代
作者:Shun Lei,Yixuan Zhou,Boshi Tang,Max W. Y. Lam,Feng Liu,Hangyu Liu,Jingcheng Wu,Shiyin Kang,Zhiyong Wu,Helen Meng
备注:work in progress
链接:点击下载PDF文件
摘要:音乐是人类文化的重要组成部分,是人类智慧和创造力的集中体现,而歌曲是其中不可或缺的一部分。虽然以前的作品已经探讨了歌曲生成的各个方面,如歌唱声音,声乐创作和乐器安排等,根据歌词生成具有人声和伴奏的歌曲仍然是一个重大挑战,阻碍了音乐生成模型在现实世界中的应用。在这种情况下,我们提出了SongCreator,一个歌曲生成系统,旨在应对这一挑战。该模型具有两个新颖的设计:精心设计的双序列语言模型(DSLM),以捕获用于歌曲生成的人声和伴奏信息,以及DSLM的额外注意力掩码策略,该策略允许我们的模型理解,生成和编辑歌曲,使其适合于各种与歌曲相关的生成任务。大量的实验证明了SongCreator的有效性,在所有八个任务上都实现了最先进或有竞争力的性能。值得注意的是,它在歌词到歌曲和歌词到人声方面大大超过了以前的作品。此外,它能够通过不同的提示独立地控制生成的歌曲中的人声和伴奏的声学条件,显示出其潜在的适用性。我们的样品可在https: songcreator.github.io 上获得。摘要:Music is an integral part of human culture, embodying human intelligence and creativity, of which songs compose an essential part. While various aspects of song generation have been explored by previous works, such as singing voice, vocal composition and instrumental arrangement, etc., generating songs with both vocals and accompaniment given lyrics remains a significant challenge, hindering the application of music generation models in the real world. In this light, we propose SongCreator, a song-generation system designed to tackle this challenge. The model features two novel designs: a meticulously designed dual-sequence language model (DSLM) to capture the information of vocals and accompaniment for song generation, and an additional attention mask strategy for DSLM, which allows our model to understand, generate and edit songs, making it suitable for various song-related generation tasks. Extensive experiments demonstrate the effectiveness of SongCreator by achieving state-of-the-art or competitive performances on all eight tasks. Notably, it surpasses previous works by a large margin in lyrics-to-song and lyrics-to-vocals. Additionally, it is able to independently control the acoustic conditions of the vocals and accompaniment in the generated song through different prompts, exhibiting its potential applicability. Our samples are available at https: songcreator.github.io .
【33】 Musical Chords: A Novel Java Algorithm and App Utility to Enumerate Chord-Progressions Adhering to Music Theory Guidelines
标题: 音乐和弦:一种新颖的Java算法和应用程序实用程序,用于列举符合音乐理论准则的和弦进行
作者:Aditya Lakshminarasimhan
备注:10 Pages, 5 Figures
链接:点击下载PDF文件
摘要:一首歌的骨干是它的和弦进行,一系列的和弦,提高和谐和增加到整体组成。对于从初学者到创意艺术家的个人来说,理解和实施音乐理论语法可以扼杀音乐创作过程并导致歌曲作者的障碍。市场上现有的Chord Progression方法仅限于产生预先选择的进行,并且通常无法符合音乐理论指南或为其他音乐家提供API以进行构建。由于四和弦和八和弦进行尚未被枚举,因此在和弦进行上进行训练的机器学习用例有限,并且移动应用程序无法为用户提供独特或未开发的进行。为了解决这些限制,已经开发了一种新颖的Java算法和自动音乐理论和弦进行和变化生成器App。这个应用程序提供了一个钢琴用户界面,应用音乐理论来生成所有可能的四和弦和八和弦进行,并产生三个由用户选择的生成进行的交替变化。该算法阐明了总共3,297个4弦进行式和总共405,216个8弦进行式。在4和弦进行池中,有1,533个主要4和弦进行和1,764个次要4和弦进行。在8和弦进行池中,有182,094个主要进行和223,122个次要进行。这种创新的方法为音乐家提供了一个全面和可定制的音乐创作工具,使他们能够开发自己的签名声音。摘要:A song's backbone is its chord progressions, a series of chords that improve the harmony and add to the overall composition. For individuals ranging from beginners to creative artists, comprehending and implementing music theory grammar for their own compositions can stifle the music creation process and cause song-writer's block. The existing Chord Progression approaches in the marketplace are limited on producing only pre-selected progressions and often fail to conform to music theory guidelines or provide APIs for other musicians to build on. Because four-chord and eight-chord progressions are yet to be enumerated, Machine learning use-cases that train on chord progressions are limited, and mobile applications don't provide users with unique or unexplored progressions. To address these limitations, a novel Java Algorithm and automated music theory chord progression and variations generator App has been developed. This App offers a piano user interface, that applies music theory to generate all possible four-chord and eight-chord progressions and produces three alternate variations of the generated progressions selected by the user. The Algorithm elucidates 3,297 Total 4-Chord Progressions and 405,216 Total 8-Chord Progressions. Within the 4-Chord Progression pool, there are 1,533 Major 4-chord Progressions and 1,764 Minor 4-Chord Progressions. Within the 8-chord Progression pool, there are 182,094 Major Progressions and 223,122 Minor Progressions. This innovative approach provides musicians with a comprehensive and customizable tool for their music creation, allowing them to develop their signature sounds.
【34】 Continuous Learning of Transformer-based Audio Deepfake Detection
标题: 基于变形器的音频深度伪造检测的持续学习
作者:Tuan Duy Nguyen Le,Kah Kuan Teh,Huy Dat Tran
备注:Submitted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:本文提出了一种新的音频deepfake检测框架,其主要目标有两个:i)在可用的假数据上获得尽可能高的准确性,以及ii)以Few-Shot学习方式有效地对新的假数据进行连续学习。具体来说,我们使用各种深度音频生成方法进行大型音频deepfake收集。数据通过额外的增强方法进一步增强,以增加压缩、远场记录、噪声和其他失真中的变化。然后,我们采用音频谱图Transformer来进行音频深度伪造检测模型。因此,所提出的方法在各种基准数据集上实现了有前途的性能。此外,我们提出了一个持续学习插件模块,以最有效地更新训练模型,使用最少的新假类型的标记数据点。所提出的方法优于传统的直接微调方法,标记的数据点少得多。摘要:This paper proposes a novel framework for audio deepfake detection with two main objectives: i) attaining the highest possible accuracy on available fake data, and ii) effectively performing continuous learning on new fake data in a few-shot learning manner. Specifically, we conduct a large audio deepfake collection using various deep audio generation methods. The data is further enhanced with additional augmentation methods to increase variations amidst compressions, far-field recordings, noise, and other distortions. We then adopt the Audio Spectrogram Transformer for the audio deepfake detection model. Accordingly, the proposed method achieves promising performance on various benchmark datasets. Furthermore, we present a continuous learning plugin module to update the trained model most effectively with the fewest possible labeled data points of the new fake type. The proposed method outperforms the conventional direct fine-tuning approach with much fewer labeled data points.
机器翻译,仅供参考
![]()
