本文经arXiv每日学术速递授权转载
【1】 CLAIR-A: Leveraging Large Language Models to Judge Audio Captions
标题: CLAIR-A:利用大型语言模型来判断音频字幕
作者:Tsung-Han Wu,Joseph E. Gonzalez,Trevor Darrell,David M. Chan
备注:Code is publicly available at this https URL
链接:点击下载PDF文件
【2】 Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space
标题: 增强语音命令的合成训练数据:从基于SVR的过滤到SSL潜在空间中的域适应
作者:Sebastião Quintas,Isabelle Ferrané,Thomas Pellegrini
链接:点击下载PDF文件
【3】 $text{M}^text{6}(text{GPT})^text{3}$: Generating Multitrack Modifiable Multi-Minute MIDI Music from Text using Genetic algorithms, Probabilistic methods and GPT Models in any Progression and Time signature
标题: $ ext{M}' ext{6}( ext{GPT}) ext{3}$:在任何进度和时间签名中使用遗传算法、概率方法和GPT模型从文本生成多轨可修改多分钟预设音乐
作者:Jakub Poćwiardowski,Mateusz Modrzejewski,Marek S. Tatara
备注:12 pages, 1 figure
链接:点击下载PDF文件
【4】 Exploring bat song syllable representations in self-supervised audio encoders
标题: 探索自我监督音频编码器中的蝙蝠歌曲音节表示
作者:Marianne de Heer Kloots,Mirjam Knörnschild
Journal-ref:Proceedings of the 4th International Workshop on Vocal Interactivity in-and-between Humans, Animals and Robots (VIHAR), Kos, GR, 6-9 Sept 2024
链接:点击下载PDF文件
【5】 Hidden in Plain Sound: Environmental Backdoor Poisoning Attacks on Whisper, and Mitigations
标题: 隐藏在普通声音中:对Whisper的环境后门中毒攻击和缓解措施
作者:Jonatan Bartolini,Todor Stoyanov,Alberto Giaretta
备注:13 pages, 12 figures, 6 tables
链接:点击下载PDF文件
【6】 FruitsMusic: A Real-World Corpus of Japanese Idol-Group Songs
标题: FruitsMusic:日本偶像组合歌曲的现实世界素材库
作者:Hitoshi Suda,Shunsuke Yoshida,Tomohiko Nakamura,Satoru Fukayama,Jun Ogata
备注:Accepted at the 25th International Society for Music Information Retrieval (ISMIR) Conference 2024, San Francisco, United States
链接:点击下载PDF文件
【7】 SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
标题: SoundBeam遇上M2D:利用音频基础模型提取目标声音
作者:Carlos Hernandez-Olivan,Marc Delcroix,Tsubasa Ochiai,Daisuke Niizumi,Naohiro Tawara,Tomohiro Nakatani,Shoko Araki
链接:点击下载PDF文件
【8】 ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning
标题: ViolinDiff:通过音调弯曲调节增强表达力的小提琴合成
作者:Daewoong Kim,Hao-Wen Dong,Dasaem Jeong
链接:点击下载PDF文件
【9】 AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost
标题: AutoMode-ASB:学习选择ASB系统以获得更好的质量和成本
作者:Ahmet Gündüz,Yunsu Kim,Kamer Ali Yuksel,Mohamed Al-Badrashiny,Thiago Castro Ferreira,Hassan Sawaf
备注:SPECOM 2024 Conference
链接:点击下载PDF文件
【10】 AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework
标题: AudioEditor:免训练的基于扩散的音频编辑框架
作者:Yuhang Jia,Yang Chen,Jinghua Zhao,Shiwan Zhao,Wenjia Zeng,Yong Chen,Yong Qin
链接:点击下载PDF文件
【11】 A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
标题: 具有空间线索保留的轻量级实时双耳语音增强模型
作者:Jingyuan Wang,Jie Zhang,Shihao Chen,Miao Sun
链接:点击下载PDF文件
【12】 Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
标题: 用于鲁棒语音识别的队列感知域自适应生成对抗网络
作者:Chien-Chun Wang,Li-Wei Chen,Cheng-Kang Chou,Hung-Shin Lee,Berlin Chen,Hsin-Min Wang
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【13】 Simultaneous Music Separation and Generation Using Multi-Track Latent Diffusion Models
标题: 使用多轨潜在扩散模型同时音乐分离和生成
作者:Tornike Karchkhadze,Mohammad Rasool Izadi,Shlomo Dubnov
链接:点击下载PDF文件
【14】 Large Language Models Are Strong Audio-Visual Speech Recognition Learners
标题: 大型语言模型是强大的视听语音识别学习者
作者:Umberto Cappellazzo,Minsu Kim,Honglie Chen,Pingchuan Ma,Stavros Petridis,Daniele Falavigna,Alessio Brutti,Maja Pantic
备注:The code will be made available at this link: this https URL
链接:点击下载PDF文件
【15】 Measuring Sound Symbolism in Audio-visual Models
标题: 测量视听模型中的声音象征性
作者:Wei-Cheng Tseng,Yi-Jen Shih,David Harwath,Raymond Mooney
备注:SLT 2024
链接:点击下载PDF文件
【16】 WaveletGPT: Wavelets Meet Large Language Models
标题: WaveletGPT:WaveletGPT满足大型语言模型
作者:Prateek Verma
备注:16 pages, 4 figures
链接:点击下载PDF文件
【17】 NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
标题: NDVQ:具有基于正态分布的载体量化的鲁棒神经音频编解码器
作者:Zhikang Niu,Sanyuan Chen,Long Zhou,Ziyang Ma,Xie Chen,Shujie Liu
链接:点击下载PDF文件
【18】 AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
标题: AudioComposer:利用自然语言描述实现细粒度音频生成
作者:Yuanyuan Wang,Hangting Chen,Dongchao Yang,Zhiyong Wu,Helen Meng,Xixin Wu
链接:点击下载PDF文件
【19】 Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech Enhancement
标题: 用于脑辅助语音增强的几何限制的脑电通道选择
作者:Keying Zuo,Qingtian Xu,Jie Zhang,Zhenhua Ling
链接:点击下载PDF文件
【20】 Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features
标题: 具有复杂频谱图和可学习时间特征的语音分解Transformer
作者:Younghoo Kwon,Jung-Woo Choi
备注:5 pages, 2 figures, submitted to ICASSP 2024
链接:点击下载PDF文件
【21】 Multichannel-to-Multichannel Target Sound Extraction Using Direction and Timestamp Clues
标题: 使用方向和时间戳线索的多通道到多通道目标声音提取
作者:Dayun Choi,Jung-Woo Choi
备注:5 pages, 4 figures
链接:点击下载PDF文件
【22】 DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
标题: DeFT-Mamba:通用多通道声音分离和复音音频分类
作者:Dongheon Lee,Jung-Woo Choi
备注:5 pages, 2 figures
链接:点击下载PDF文件
【23】 Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
标题: 使用说话者感知的CTC在多说话者语音识别中理清说话者
作者:Jiawen Kang,Lingwei Meng,Mingyu Cui,Yuejiao Wang,Xixin Wu,Xunying Liu,Helen Meng
链接:点击下载PDF文件
【24】 Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
标题: 具有混合专家的稳健视听语音识别模型
作者:Yihan Wu,Yifan Peng,Yichen Lu,Xuankai Chang,Ruihua Song,Shinji Watanabe
备注:6 pages, 2 figures, accepted by IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
【25】 META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
标题: META-CAT:通过Meta信息级联用于多说话者ASB的说话者知情的语音嵌入
作者:Jinhan Wang,Weiqing Wang,Kunal Dhawan,Taejin Park,Myungjong Kim,Ivan Medennikov,He Huang,Nithin Koluguri,Jagadeesh Balam,Boris Ginsburg
链接:点击下载PDF文件
标题: WaveletGPT:WaveletGPT满足大型语言模型
作者:Prateek Verma
备注:16 pages, 4 figures
链接:点击下载PDF文件
【2】 NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
标题: NDVQ:具有基于正态分布的载体量化的鲁棒神经音频编解码器
作者:Zhikang Niu,Sanyuan Chen,Long Zhou,Ziyang Ma,Xie Chen,Shujie Liu
链接:点击下载PDF文件
【3】 AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
标题: AudioComposer:利用自然语言描述实现细粒度音频生成
作者:Yuanyuan Wang,Hangting Chen,Dongchao Yang,Zhiyong Wu,Helen Meng,Xixin Wu
链接:点击下载PDF文件
【4】 Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech Enhancement
标题: 用于脑辅助语音增强的几何限制的脑电通道选择
作者:Keying Zuo,Qingtian Xu,Jie Zhang,Zhenhua Ling
链接:点击下载PDF文件
【5】 Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features
标题: 具有复杂频谱图和可学习时间特征的语音分解Transformer
作者:Younghoo Kwon,Jung-Woo Choi
备注:5 pages, 2 figures, submitted to ICASSP 2024
链接:点击下载PDF文件
【6】 Multichannel-to-Multichannel Target Sound Extraction Using Direction and Timestamp Clues
标题: 使用方向和时间戳线索的多通道到多通道目标声音提取
作者:Dayun Choi,Jung-Woo Choi
备注:5 pages, 4 figures
链接:点击下载PDF文件
【7】 DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
标题: DeFT-Mamba:通用多通道声音分离和复音音频分类
作者:Dongheon Lee,Jung-Woo Choi
备注:5 pages, 2 figures
链接:点击下载PDF文件
【8】 Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
标题: 使用说话者感知的CTC在多说话者语音识别中理清说话者
作者:Jiawen Kang,Lingwei Meng,Mingyu Cui,Yuejiao Wang,Xixin Wu,Xunying Liu,Helen Meng
链接:点击下载PDF文件
【9】 Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
标题: 具有混合专家的稳健视听语音识别模型
作者:Yihan Wu,Yifan Peng,Yichen Lu,Xuankai Chang,Ruihua Song,Shinji Watanabe
备注:6 pages, 2 figures, accepted by IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
【10】 META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
标题: META-CAT:通过Meta信息级联用于多说话者ASB的说话者知情的语音嵌入
作者:Jinhan Wang,Weiqing Wang,Kunal Dhawan,Taejin Park,Myungjong Kim,Ivan Medennikov,He Huang,Nithin Koluguri,Jagadeesh Balam,Boris Ginsburg
链接:点击下载PDF文件
【11】 CLAIR-A: Leveraging Large Language Models to Judge Audio Captions
标题: CLAIR-A:利用大型语言模型来判断音频字幕
作者:Tsung-Han Wu,Joseph E. Gonzalez,Trevor Darrell,David M. Chan
备注:Code is publicly available at this https URL
链接:点击下载PDF文件
【12】 Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space
标题: 增强语音命令的合成训练数据:从基于SVR的过滤到SSL潜在空间中的域适应
作者:Sebastião Quintas,Isabelle Ferrané,Thomas Pellegrini
链接:点击下载PDF文件
【13】 $text{M}^text{6}(text{GPT})^text{3}$: Generating Multitrack Modifiable Multi-Minute MIDI Music from Text using Genetic algorithms, Probabilistic methods and GPT Models in any Progression and Time signature
标题: $ ext{M}' ext{6}( ext{GPT}) ext{3}$:在任何进度和时间签名中使用遗传算法、概率方法和GPT模型从文本生成多轨可修改多分钟预设音乐
作者:Jakub Poćwiardowski,Mateusz Modrzejewski,Marek S. Tatara
备注:12 pages, 1 figure
链接:点击下载PDF文件
【14】 Exploring bat song syllable representations in self-supervised audio encoders
标题: 探索自我监督音频编码器中的蝙蝠歌曲音节表示
作者:Marianne de Heer Kloots,Mirjam Knörnschild
Journal-ref:Proceedings of the 4th International Workshop on Vocal Interactivity in-and-between Humans, Animals and Robots (VIHAR), Kos, GR, 6-9 Sept 2024
链接:点击下载PDF文件
【15】 Hidden in Plain Sound: Environmental Backdoor Poisoning Attacks on Whisper, and Mitigations
标题: 隐藏在普通声音中:对Whisper的环境后门中毒攻击和缓解措施
作者:Jonatan Bartolini,Todor Stoyanov,Alberto Giaretta
备注:13 pages, 12 figures, 6 tables
链接:点击下载PDF文件
【16】 FruitsMusic: A Real-World Corpus of Japanese Idol-Group Songs
标题: FruitsMusic:日本偶像组合歌曲的现实世界素材库
作者:Hitoshi Suda,Shunsuke Yoshida,Tomohiko Nakamura,Satoru Fukayama,Jun Ogata
备注:Accepted at the 25th International Society for Music Information Retrieval (ISMIR) Conference 2024, San Francisco, United States
链接:点击下载PDF文件
【17】 SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
标题: SoundBeam遇上M2D:利用音频基础模型提取目标声音
作者:Carlos Hernandez-Olivan,Marc Delcroix,Tsubasa Ochiai,Daisuke Niizumi,Naohiro Tawara,Tomohiro Nakatani,Shoko Araki
链接:点击下载PDF文件
【18】 ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning
标题: ViolinDiff:通过音调弯曲调节增强表达力的小提琴合成
作者:Daewoong Kim,Hao-Wen Dong,Dasaem Jeong
链接:点击下载PDF文件
【19】 AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost
标题: AutoMode-ASB:学习选择ASB系统以获得更好的质量和成本
作者:Ahmet Gündüz,Yunsu Kim,Kamer Ali Yuksel,Mohamed Al-Badrashiny,Thiago Castro Ferreira,Hassan Sawaf
备注:SPECOM 2024 Conference
链接:点击下载PDF文件
【20】 AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework
标题: AudioEditor:免训练的基于扩散的音频编辑框架
作者:Yuhang Jia,Yang Chen,Jinghua Zhao,Shiwan Zhao,Wenjia Zeng,Yong Chen,Yong Qin
链接:点击下载PDF文件
【21】 A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
标题: 具有空间线索保留的轻量级实时双耳语音增强模型
作者:Jingyuan Wang,Jie Zhang,Shihao Chen,Miao Sun
链接:点击下载PDF文件
【22】 Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
标题: 用于鲁棒语音识别的队列感知域自适应生成对抗网络
作者:Chien-Chun Wang,Li-Wei Chen,Cheng-Kang Chou,Hung-Shin Lee,Berlin Chen,Hsin-Min Wang
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【23】 Simultaneous Music Separation and Generation Using Multi-Track Latent Diffusion Models
标题: 使用多轨潜在扩散模型同时音乐分离和生成
作者:Tornike Karchkhadze,Mohammad Rasool Izadi,Shlomo Dubnov
链接:点击下载PDF文件
【24】 Large Language Models Are Strong Audio-Visual Speech Recognition Learners
标题: 大型语言模型是强大的视听语音识别学习者
作者:Umberto Cappellazzo,Minsu Kim,Honglie Chen,Pingchuan Ma,Stavros Petridis,Daniele Falavigna,Alessio Brutti,Maja Pantic
备注:The code will be made available at this link: this https URL
链接:点击下载PDF文件
【25】 Measuring Sound Symbolism in Audio-visual Models
标题: 测量视听模型中的声音象征性
作者:Wei-Cheng Tseng,Yi-Jen Shih,David Harwath,Raymond Mooney
备注:SLT 2024
链接:点击下载PDF文件
标题: CLAIR-A:利用大型语言模型来判断音频字幕
作者:Tsung-Han Wu,Joseph E. Gonzalez,Trevor Darrell,David M. Chan
备注:Code is publicly available at this https URL
链接:点击下载PDF文件
摘要:自动音频字幕(AAC)任务要求模型生成音频输入的自然语言描述。评估这些机器生成的音频字幕是一项复杂的任务,需要考虑各种因素,其中包括听觉场景理解,声音对象推理,时间连贯性和场景的环境背景。虽然目前的方法侧重于特定方面,但它们通常无法提供与人类判断一致的总体评分。在这项工作中,我们提出了CLAIR-A,一个简单而灵活的方法,利用大型语言模型(LLM)的zero-shot能力,直接要求LLM的语义距离分数来评估候选音频字幕。在我们的评估中,与传统指标相比,CLAIR-A更好地预测了人类对质量的判断,与特定领域的FENSE指标相比,相对准确度提高了5.8%,与Clotho-Eval数据集上的最佳通用指标相比,相对准确度提高了11%。此外,CLAIR-A通过允许语言模型解释其分数背后的推理提供了更高的透明度,这些解释比基线方法提供的解释高出30%。CLAIR-A在https: github.com DavidMChan clair-a上公开提供。摘要:The Automated Audio Captioning (AAC) task asks models to generate natural language descriptions of an audio input. Evaluating these machine-generated audio captions is a complex task that requires considering diverse factors, among them, auditory scene understanding, sound-object inference, temporal coherence, and the environmental context of the scene. While current methods focus on specific aspects, they often fail to provide an overall score that aligns well with human judgment. In this work, we propose CLAIR-A, a simple and flexible method that leverages the zero-shot capabilities of large language models (LLMs) to evaluate candidate audio captions by directly asking LLMs for a semantic distance score. In our evaluations, CLAIR-A better predicts human judgements of quality compared to traditional metrics, with a 5.8% relative accuracy improvement compared to the domain-specific FENSE metric and up to 11% over the best general-purpose measure on the Clotho-Eval dataset. Moreover, CLAIR-A offers more transparency by allowing the language model to explain the reasoning behind its scores, with these explanations rated up to 30% better by human evaluators than those provided by baseline methods. CLAIR-A is made publicly available at https: github.com DavidMChan clair-a.
【2】 Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space
标题: 增强语音命令的合成训练数据:从基于SVR的过滤到SSL潜在空间中的域适应
作者:Sebastião Quintas,Isabelle Ferrané,Thomas Pellegrini
链接:点击下载PDF文件
摘要:使用合成语音作为数据增强在自动语音识别和语音分类任务等领域越来越受欢迎。尽管具有语音克隆能力的新颖的文本到语音系统允许使用基于短音频段的更大量的语音,但是已知的是,这些系统倾向于产生幻觉并且经常产生很可能对下游任务具有负面影响的坏数据。在目前的工作中,我们进行了一组实验,围绕zero-shot学习与合成语音数据的语音命令分类的特定任务。我们在Google Speech Commands数据集上的结果表明,一个简单的基于ASR的过滤方法可以对生成的数据质量产生很大的影响,从而转化为更好的性能。此外,尽管生成的语音数据质量很好,但我们还表明,使用自监督(WavLM)特征时,合成语音和真实语音仍然可以轻松区分,这是CycleGAN进一步探索的一个方面,以弥合两种类型的语音材料之间的差距。摘要:The use of synthetic speech as data augmentation is gaining increasing popularity in fields such as automatic speech recognition and speech classification tasks. Despite novel text-to-speech systems with voice cloning capabilities, that allow the usage of a larger amount of voices based on short audio segments, it is known that these systems tend to hallucinate and oftentimes produce bad data that will most likely have a negative impact on the downstream task. In the present work, we conduct a set of experiments around zero-shot learning with synthetic speech data for the specific task of speech commands classification. Our results on the Google Speech Commands dataset show that a simple ASR-based filtering method can have a big impact in the quality of the generated data, translating to a better performance. Furthermore, despite the good quality of the generated speech data, we also show that synthetic and real speech can still be easily distinguishable when using self-supervised (WavLM) features, an aspect further explored with a CycleGAN to bridge the gap between the two types of speech material.
【3】 $text{M}^text{6}(text{GPT})^text{3}$: Generating Multitrack Modifiable Multi-Minute MIDI Music from Text using Genetic algorithms, Probabilistic methods and GPT Models in any Progression and Time signature
标题: $ ext{M}' ext{6}( ext{GPT}) ext{3}$:在任何进度和时间签名中使用遗传算法、概率方法和GPT模型从文本生成多轨可修改多分钟预设音乐
作者:Jakub Poćwiardowski,Mateusz Modrzejewski,Marek S. Tatara
备注:12 pages, 1 figure
链接:点击下载PDF文件
摘要:这项工作介绍了$ text{M}^ text{6}( text{GPT})^ text{3}$ Composer系统,能够生成完整的,多分钟的音乐作品,具有复杂的结构,在任何时间签名,在自然语言的输入描述的领域。该系统利用自回归Transformer语言模型将自然语言提示映射到JSON格式的组合参数。所定义的结构包括时间标记、音阶、和弦进行和效价唤醒值,从这些值创建伴奏、旋律、低音、主题和打击乐音轨。我们提出了一种遗传算法的旋律元素的生成。该算法结合了突变与音乐意义和适应度函数的基础上正态分布和预定义的音乐特征值。价值观自适应地演变,受情绪参数和不同的演奏风格的影响。用于生成任何时间的敲击签名的系统利用概率方法,包括马尔可夫链。通过人为和客观的评估,我们证明了我们的音乐生成方法在特定的、具有音乐意义的指标上优于基线,为纯粹基于神经网络的系统提供了一种有价值的替代方案。摘要:This work introduces the $ text{M}^ text{6}( text{GPT})^ text{3}$ Composer system, capable of generating complete, multi-minute musical compositions with complex structures in any time signature, in the MIDI domain from input descriptions in natural language. The system utilizes an autoregressive transformer language model to map natural language prompts to composition parameters in JSON format. The defined structure includes time signature, scales, chord progressions, and valence-arousal values, from which accompaniment, melody, bass, motif, and percussion tracks are created. We propose a genetic algorithm for the generation of melodic elements. The algorithm incorporates mutations with musical significance and a fitness function based on normal distribution and predefined musical feature values. The values adaptively evolve, influenced by emotional parameters and distinct playing styles. The system for generating percussion in any time signature utilises probabilistic methods, including Markov chains. Through both human and objective evaluations, we demonstrate that our music generation approach outperforms baselines on specific, musically meaningful metrics, offering a valuable alternative to purely neural network-based systems.
【4】 Exploring bat song syllable representations in self-supervised audio encoders
标题: 探索自我监督音频编码器中的蝙蝠歌曲音节表示
作者:Marianne de Heer Kloots,Mirjam Knörnschild
Journal-ref:Proceedings of the 4th International Workshop on Vocal Interactivity in-and-between Humans, Animals and Robots (VIHAR), Kos, GR, 6-9 Sept 2024
链接:点击下载PDF文件
摘要:在人类发出的声音上训练的深度学习模型能在多大程度上区分另一个物种的发声类型?我们分析了几种自监督音频编码器中蝙蝠歌音节的编码,发现在人类语音上预训练的模型生成了不同音节类型的最独特的表示。这些发现形成了跨物种迁移学习在蝙蝠生物声学中应用的第一步,以及对音频编码器模型中分布外信号处理的更好理解。摘要:How well can deep learning models trained on human-generated sounds distinguish between another species' vocalization types? We analyze the encoding of bat song syllables in several self-supervised audio encoders, and find that models pre-trained on human speech generate the most distinctive representations of different syllable types. These findings form first steps towards the application of cross-species transfer learning in bat bioacoustics, as well as an improved understanding of out-of-distribution signal processing in audio encoder models.
【5】 Hidden in Plain Sound: Environmental Backdoor Poisoning Attacks on Whisper, and Mitigations
标题: 隐藏在普通声音中:对Whisper的环境后门中毒攻击和缓解措施
作者:Jonatan Bartolini,Todor Stoyanov,Alberto Giaretta
备注:13 pages, 12 figures, 6 tables
链接:点击下载PDF文件
摘要:由于基于transformer的模型的普及,语音识别(SR)在各种应用领域中越来越受欢迎,例如充满关键任务设备的工业和机器人环境。虽然基于transformer的SR可以为简化人机接口提供各种好处,但对这些模型的网络安全方面的研究却乏善可陈。特别是关于后门中毒攻击。在本文中,我们提出了一种新的中毒方法,映射不同的环境触发声音的目标短语的不同长度,在微调阶段。我们在Whisper上测试了我们的方法,Whisper是最流行的基于transformer的SR模型之一,在几种测试条件下,它非常容易受到我们的攻击。为了减轻本文中提出的攻击,我们研究了使用Silero VAD,一种最先进的语音活动检测(VAD)模型,作为防御机制。我们的实验表明,可以使用VAD模型来过滤恶意触发器并减轻我们的攻击,并根据触发声音的类型和测试条件获得不同程度的成功。摘要:Thanks to the popularisation of transformer-based models, speech recognition (SR) is gaining traction in various application fields, such as industrial and robotics environments populated with mission-critical devices. While transformer-based SR can provide various benefits for simplifying human-machine interfacing, the research on the cybersecurity aspects of these models is lacklustre. In particular, concerning backdoor poisoning attacks. In this paper, we propose a new poisoning approach that maps different environmental trigger sounds to target phrases of different lengths, during the fine-tuning phase. We test our approach on Whisper, one of the most popular transformer-based SR model, showing that it is highly vulnerable to our attack, under several testing conditions. To mitigate the attack proposed in this paper, we investigate the use of Silero VAD, a state-of-the-art voice activity detection (VAD) model, as a defence mechanism. Our experiments show that it is possible to use VAD models to filter out malicious triggers and mitigate our attacks, with a varying degree of success, depending on the type of trigger sound and testing conditions.
【6】 FruitsMusic: A Real-World Corpus of Japanese Idol-Group Songs
标题: FruitsMusic:日本偶像组合歌曲的现实世界素材库
作者:Hitoshi Suda,Shunsuke Yoshida,Tomohiko Nakamura,Satoru Fukayama,Jun Ogata
备注:Accepted at the 25th International Society for Music Information Retrieval (ISMIR) Conference 2024, San Francisco, United States
链接:点击下载PDF文件
摘要:本研究提出了FruitsMusic,一个真实世界中日本偶像团体歌曲的元数据语料库,精确地注释了谁唱什么和什么时候。日本偶像组合歌曲对日本流行文化至关重要,具有独特的声乐编曲风格,歌曲被分成几个片段,每个片段分配一个特定的个人或多个歌手。为了提高识别这种结构的歌手日记化方法,我们使用YouTube上的40个日本偶像团体的音乐视频构建了FruitsMusic作为资源。该语料库包括详细的注释,涵盖了各种流派,划分和分配风格的歌曲,以及4至9名成员的团体。FruitsMusic还促进了各种音乐信息检索技术的开发,例如歌词转录和歌手识别,不仅有利于日本偶像组合歌曲,而且还有利于各种文化的单个或多个歌手的歌曲。本文对FruitsMusic进行了全面的概述,包括其创作方法和与会话语音相比的独特特征。此外,本文还评估了当前使用FruitsMusic在具有挑战性的现实条件下进行歌手嵌入提取和日记化的方法的功效。此外,本文探讨了潜在的改进,通过评估人类的表现,在自动日记的性能。摘要:This study presents FruitsMusic, a metadata corpus of Japanese idol-group songs in the real world, precisely annotated with who sings what and when. Japanese idol-group songs, vital to Japanese pop culture, feature a unique vocal arrangement style, where songs are divided into several segments, and a specific individual or multiple singers are assigned to each segment. To enhance singer diarization methods for recognizing such structures, we constructed FruitsMusic as a resource using 40 music videos of Japanese idol groups from YouTube. The corpus includes detailed annotations, covering songs across various genres, division and assignment styles, and groups ranging from 4 to 9 members. FruitsMusic also facilitates the development of various music information retrieval techniques, such as lyrics transcription and singer identification, benefiting not only Japanese idol-group songs but also a wide range of songs featuring single or multiple singers from various cultures. This paper offers a comprehensive overview of FruitsMusic, including its creation methodology and unique characteristics compared to conversational speech. Additionally, this paper evaluates the efficacy of current methods for singer embedding extraction and diarization in challenging real-world conditions using FruitsMusic. Furthermore, this paper examines potential improvements in automatic diarization performance through evaluating human performance.
【7】 SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
标题: SoundBeam遇上M2D:利用音频基础模型提取目标声音
作者:Carlos Hernandez-Olivan,Marc Delcroix,Tsubasa Ochiai,Daisuke Niizumi,Naohiro Tawara,Tomohiro Nakatani,Shoko Araki
链接:点击下载PDF文件
摘要:目标声提取(TSE)是指利用线索从混合声中分离出所需的声音。TSE系统需要同时解决两个问题:识别目标声源和从混合声中提取目标信号。为了增加实用性,同一系统应与各种类型的声音一起工作。问题的双重性和各种各样的声音使得从头开始训练强大的TSE系统具有挑战性。在本文中,为了解决这个问题,我们探索使用预训练的音频基础模型,该模型可以在TSE系统中提供丰富的声音特征表示。我们选择了M2D基础模型,它似乎特别适合TSE任务,因为它是使用由声音标签预测和改进的掩蔽预测组成的双重目标进行训练的。这些目标涉及声音识别和TSE的信号提取问题。我们提出了一个新的TSE系统,它将M2D的特征表示集成到SoundBeam中,这是一个强大的TSE系统,可以利用目标声音类别标签和预先录制的注册(或音频查询)作为线索。我们的实验表明,使用M2D可以提高提取性能,尤其是在使用注册线索时。摘要:Target sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at once, identifying the target source and extracting the target signal from the mixture. For increased practicability, the same system should work with various types of sound. The duality of the problem and the wide variety of sounds make it challenging to train a powerful TSE system from scratch. In this paper, to tackle this problem, we explore using a pre-trained audio foundation model that can provide rich feature representations of sounds within a TSE system. We chose the masked-modeling duo (M2D) foundation model, which appears especially suited for the TSE task, as it is trained using a dual objective consisting of sound-label predictions and improved masked prediction. These objectives are related to sound identification and the signal extraction problems of TSE. We propose a new TSE system that integrates the feature representation from M2D into SoundBeam, which is a strong TSE system that can exploit both target sound class labels and pre-recorded enrollments (or audio queries) as clues. We show experimentally that using M2D can increase extraction performance, especially when employing enrollment clues.
【8】 ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning
标题: ViolinDiff:通过音调弯曲调节增强表达力的小提琴合成
作者:Daewoong Kim,Hao-Wen Dong,Dasaem Jeong
链接:点击下载PDF文件
摘要:基频(F0)的自然轮廓建模在音乐音频合成中起着至关重要的作用。然而,在复调音乐中转录和管理多个F0轮廓是具有挑战性的,并且尚未探索用于复调乐器合成的显式F0轮廓建模。在本文中,我们提出了ViolinDiff,一个两阶段的基于扩散的合成框架。对于一个给定的小提琴演奏文件,第一阶段估计的F0轮廓作为音高弯曲信息,和第二阶段生成梅尔声谱图结合这些表达的细节。量化指标和听力测试结果表明,该模型产生更真实的小提琴声音比模型没有明确的音高弯曲建模。音频样本可在线获取:daewoung.github.io ViolinDiff-Demo。摘要:Modeling the natural contour of fundamental frequency (F0) plays a critical role in music audio synthesis. However, transcribing and managing multiple F0 contours in polyphonic music is challenging, and explicit F0 contour modeling has not yet been explored for polyphonic instrumental synthesis. In this paper, we present ViolinDiff, a two-stage diffusion-based synthesis framework. For a given violin MIDI file, the first stage estimates the F0 contour as pitch bend information, and the second stage generates mel spectrogram incorporating these expressive details. The quantitative metrics and listening test results show that the proposed model generates more realistic violin sounds than the model without explicit pitch bend modeling. Audio samples are available online: daewoung.github.io ViolinDiff-Demo.
【9】 AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost
标题: AutoMode-ASB:学习选择ASB系统以获得更好的质量和成本
作者:Ahmet Gündüz,Yunsu Kim,Kamer Ali Yuksel,Mohamed Al-Badrashiny,Thiago Castro Ferreira,Hassan Sawaf
备注:SPECOM 2024 Conference
链接:点击下载PDF文件
摘要:我们提出了AutoMode-ASR,一种新的框架,有效地集成了多个ASR系统,以提高整体转录质量,同时优化成本。该想法是训练决策模型,以在运行系统之前仅基于音频输入为每个分段选择最佳ASR系统。我们实现这一目标,通过集成二进制分类器确定两个系统之间的偏好。这些分类器配备了各种功能,如音频嵌入,质量估计和信号属性。此外,我们演示了如何使用质量估计可以进一步提高性能,以最小的成本增加。实验结果表明,WER的相对减少16.2%,成本节省65%,和75%的速度提高,相比使用一个单一的最佳模型的所有部分。我们的框架与商业和开源黑盒ASR系统兼容,因为它不需要更改模型代码。摘要:We present AutoMode-ASR, a novel framework that effectively integrates multiple ASR systems to enhance the overall transcription quality while optimizing cost. The idea is to train a decision model to select the optimal ASR system for each segment based solely on the audio input before running the systems. We achieve this by ensembling binary classifiers determining the preference between two systems. These classifiers are equipped with various features, such as audio embeddings, quality estimation, and signal properties. Additionally, we demonstrate how using a quality estimator can further improve performance with minimal cost increase. Experimental results show a relative reduction in WER of 16.2%, a cost saving of 65%, and a speed improvement of 75%, compared to using a single-best model for all segments. Our framework is compatible with commercial and open-source black-box ASR systems as it does not require changes in model codes.
【10】 AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework
标题: AudioEditor:免训练的基于扩散的音频编辑框架
作者:Yuhang Jia,Yang Chen,Jinghua Zhao,Shiwan Zhao,Wenjia Zeng,Yong Chen,Yong Qin
链接:点击下载PDF文件
摘要:基于扩散的文本到音频(TTA)生成已经取得了实质性的进展,利用潜在扩散模型(LDM)生成高质量,多样化和与预防相关的音频。然而,除了生成,音频编辑的任务仍然同样重要,但受到的关注相对较少。音频编辑任务面临两个主要挑战:执行精确编辑和保留未编辑的部分。虽然基于LDM的工作流已经有效地解决了图像处理领域中的这些挑战,但类似的方法几乎没有应用于音频编辑。在本文中,我们介绍了AudioEditor,这是一个基于预训练的基于扩散的TTA模型的免训练音频编辑框架。AudioEditor集成了空文本反转和EOT抑制方法,使模型能够在执行准确编辑的同时保留原始音频功能。全面的客观和主观实验验证了AudioEditor在提供高质量音频编辑方面的有效性。代码和演示可以在https: github.com NKU-HLT AudioEditor上找到。摘要:Diffusion-based text-to-audio (TTA) generation has made substantial progress, leveraging latent diffusion model (LDM) to produce high-quality, diverse and instruction-relevant audios. However, beyond generation, the task of audio editing remains equally important but has received comparatively little attention. Audio editing tasks face two primary challenges: executing precise edits and preserving the unedited sections. While workflows based on LDMs have effectively addressed these challenges in the field of image processing, similar approaches have been scarcely applied to audio editing. In this paper, we introduce AudioEditor, a training-free audio editing framework built on the pretrained diffusion-based TTA model. AudioEditor incorporates Null-text Inversion and EOT-suppression methods, enabling the model to preserve original audio features while executing accurate edits. Comprehensive objective and subjective experiments validate the effectiveness of AudioEditor in delivering high-quality audio edits. Code and demo can be found at https: github.com NKU-HLT AudioEditor.
【11】 A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
标题: 具有空间线索保留的轻量级实时双耳语音增强模型
作者:Jingyuan Wang,Jie Zhang,Shihao Chen,Miao Sun
链接:点击下载PDF文件
摘要:双耳语音增强(BSE)的目的是联合提高听力设备接收到的含噪信号的语音质量和可懂度,并保留目标的空间线索以用于自然收听。现有的方法往往受到噪声抑制(NR)的能力和空间线索保存(SCP)的准确性和复杂的声学场景中的高计算需求之间的折衷。在这项工作中,我们提出了一个基于学习的轻量级双耳复卷积网络(LBCCN),它通过过滤低频带并保留其余部分而在NR中表现出色。此外,我们的方法显式地结合了通道间相对声学传递函数的估计,以确保空间线索的保真度和语音清晰度。实验结果表明,在各种噪声条件下,该方法都能获得与现有方法相当的NR性能,但计算量更低,SCP更好。可复制的代码和音频示例可在https: github.com jywanng LBCCN上找到。摘要:Binaural speech enhancement (BSE) aims to jointly improve the speech quality and intelligibility of noisy signals received by hearing devices and preserve the spatial cues of the target for natural listening. Existing methods often suffer from the compromise between noise reduction (NR) capacity and spatial cues preservation (SCP) accuracy and a high computational demand in complex acoustic scenes. In this work, we present a learning-based lightweight binaural complex convolutional network (LBCCN), which excels in NR by filtering low-frequency bands and keeping the rest. Additionally, our approach explicitly incorporates the estimation of interchannel relative acoustic transfer function to ensure the spatial cues fidelity and speech clarity. Results show that the proposed LBCCN can achieve a comparable NR performance to state-of-the-art methods under various noise conditions, but with a much lower computational cost and a better SCP. The reproducible code and audio examples are available at https: github.com jywanng LBCCN.
【12】 Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
标题: 用于鲁棒语音识别的队列感知域自适应生成对抗网络
作者:Chien-Chun Wang,Li-Wei Chen,Cheng-Kang Chou,Hung-Shin Lee,Berlin Chen,Hsin-Min Wang
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:虽然预先训练的自动语音识别(ASR)系统在匹配域上表现出令人印象深刻的性能,但当遇到来自不可见的记录环境和条件的通道失配时,它们的性能往往会下降。为了解决这个问题,我们提出了一种新的信道感知数据模拟方法,用于鲁棒的ASR训练。我们的方法利用了通道提取技术和生成对抗网络(GANs)的协同能力。我们首先训练一个能够从任意音频中提取嵌入的通道编码器。最重要的是,使用最少量的目标域数据提取通道嵌入,并用于指导基于GAN的语音合成器。该合成器生成的语音忠实地保留了输入的语音内容,同时模仿目标域的通道特性。我们评估我们的方法具有挑战性的客家横跨台湾(HAT)和台湾横跨台湾(TAT)语料库,实现相对字符错误率(CER)减少20.02%和9.64%,分别与基线相比。这些结果突出了我们的通道感知数据模拟方法的有效性,用于弥合源域和目标域声学之间的差距。摘要:While pre-trained automatic speech recognition (ASR) systems demonstrate impressive performance on matched domains, their performance often degrades when confronted with channel mismatch stemming from unseen recording environments and conditions. To mitigate this issue, we propose a novel channel-aware data simulation method for robust ASR training. Our method harnesses the synergistic power of channel-extractive techniques and generative adversarial networks (GANs). We first train a channel encoder capable of extracting embeddings from arbitrary audio. On top of this, channel embeddings are extracted using a minimal amount of target-domain data and used to guide a GAN-based speech synthesizer. This synthesizer generates speech that faithfully preserves the phonetic content of the input while mimicking the channel characteristics of the target domain. We evaluate our method on the challenging Hakka Across Taiwan (HAT) and Taiwanese Across Taiwan (TAT) corpora, achieving relative character error rate (CER) reductions of 20.02% and 9.64%, respectively, compared to the baselines. These results highlight the efficacy of our channel-aware data simulation method for bridging the gap between source- and target-domain acoustics.
【13】 Simultaneous Music Separation and Generation Using Multi-Track Latent Diffusion Models
标题: 使用多轨潜在扩散模型同时音乐分离和生成
作者:Tornike Karchkhadze,Mohammad Rasool Izadi,Shlomo Dubnov
链接:点击下载PDF文件
摘要:扩散模型最近在音乐生成和音乐源分离任务中显示出强大的潜力。虽然在早期阶段,一种趋势正在出现,将这些任务整合到一个单一的框架中,因为两者都涉及生成音乐对齐的部分,并且可以被视为同一生成过程的方面。在这项工作中,我们引入了一个潜在的基于扩散的多轨道生成模型,能够通过学习共享音乐上下文的轨道的联合概率分布的源分离和多轨道音乐合成。我们的模型还可以通过创建给定其他轨道的任何子集来生成安排。我们在Slakh2100数据集上训练了我们的模型,将其与现有的同步生成和分离模型进行了比较,并观察到源分离,音乐和安排生成任务的客观指标的显着改进。在https: msg-ld.github.io 上可以找到正确的例子。摘要:Diffusion models have recently shown strong potential in both music generation and music source separation tasks. Although in early stages, a trend is emerging towards integrating these tasks into a single framework, as both involve generating musically aligned parts and can be seen as facets of the same generative process. In this work, we introduce a latent diffusion-based multi-track generation model capable of both source separation and multi-track music synthesis by learning the joint probability distribution of tracks sharing a musical context. Our model also enables arrangement generation by creating any subset of tracks given the others. We trained our model on the Slakh2100 dataset, compared it with an existing simultaneous generation and separation model, and observed significant improvements across objective metrics for source separation, music, and arrangement generation tasks. Sound examples are available at https: msg-ld.github.io .
【14】 Large Language Models Are Strong Audio-Visual Speech Recognition Learners
标题: 大型语言模型是强大的视听语音识别学习者
作者:Umberto Cappellazzo,Minsu Kim,Honglie Chen,Pingchuan Ma,Stavros Petridis,Daniele Falavigna,Alessio Brutti,Maja Pantic
备注:The code will be made available at this link: this https URL
链接:点击下载PDF文件
摘要:多模态大型语言模型(MLLM)由于其强大的多模态理解能力而成为近年来的研究热点。例如,在音频和语音域中,LLM可以通过仅连接用音频编码器计算的音频令牌和文本令牌来配备(自动)语音识别(ASR)能力,以实现最先进的结果。相反,像视觉和视听语音识别(VSR AVSR)这样的任务,也利用了噪声不变的嘴唇运动信息,很少或根本没有受到关注。为了弥补这一差距,我们提出了Llama-AVSR,一个新的MLLM具有强大的视听语音识别能力。它利用预先训练的音频和视频编码器来产生特定于模态的令牌,这些令牌与文本令牌一起由预先训练的LLM(例如,图3.1 -8B)以自回归方式产生所得响应。Llama-AVSR需要少量的可训练参数,因为只有特定模态的投影仪和LoRA模块被训练,而多模态编码器和LLM保持冻结。我们在LRS 3(最大的公共AVSR基准)上评估了我们提出的方法,并在ASR和AVSR任务中取得了新的最先进结果,WER分别为0.81%和0.77%。为了支持我们的结果,我们研究了支撑Llama-AVSR有效性的关键因素:预训练编码器和LLM的选择,LoRA模块的有效集成,以及通过模态感知压缩率获得的最佳性能效率权衡。摘要:Multimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and speech domains, an LLM can be equipped with (automatic) speech recognition (ASR) abilities by just concatenating the audio tokens, computed with an audio encoder, and the text tokens to achieve state-of-the-art results. On the contrary, tasks like visual and audio-visual speech recognition (VSR AVSR), which also exploit noise-invariant lip movement information, have received little or no attention. To bridge this gap, we propose Llama-AVSR, a new MLLM with strong audio-visual speech recognition capabilities. It leverages pre-trained audio and video encoders to produce modality-specific tokens which, together with the text tokens, are processed by a pre-trained LLM (e.g., Llama3.1-8B) to yield the resulting response in an auto-regressive fashion. Llama-AVSR requires a small number of trainable parameters as only modality-specific projectors and LoRA modules are trained whereas the multi-modal encoders and LLM are kept frozen. We evaluate our proposed approach on LRS3, the largest public AVSR benchmark, and we achieve new state-of-the-art results for the tasks of ASR and AVSR with a WER of 0.81% and 0.77%, respectively. To bolster our results, we investigate the key factors that underpin the effectiveness of Llama-AVSR: the choice of the pre-trained encoders and LLM, the efficient integration of LoRA modules, and the optimal performance-efficiency trade-off obtained via modality-aware compression rates.
【15】 Measuring Sound Symbolism in Audio-visual Models
标题: 测量视听模型中的声音象征性
作者:Wei-Cheng Tseng,Yi-Jen Shih,David Harwath,Raymond Mooney
备注:SLT 2024
链接:点击下载PDF文件
摘要:视听预训练模型最近得到了广泛的关注,并在各种视听任务上表现出卓越的性能。这项研究调查了预先训练的视听模型是否表现出声音和视觉表示之间的非任意关联,称为声音符号学,这也在人类中观察到。我们开发了一个专门的数据集与合成的图像和音频样本,并评估这些模型使用非参数的方法在zero-shot设置。我们的研究结果揭示了模型的输出和建立的声音符号模式之间的显着相关性,特别是在语音数据训练的模型中。这些结果表明,这些模型可以捕获类似于人类语言处理的声音-意义联系,为认知架构和机器学习策略提供见解。摘要:Audio-visual pre-trained models have gained substantial attention recently and demonstrated superior performance on various audio-visual tasks. This study investigates whether pre-trained audio-visual models demonstrate non-arbitrary associations between sounds and visual representations$ unicode{x2013}$known as sound symbolism$ unicode{x2013}$which is also observed in humans. We developed a specialized dataset with synthesized images and audio samples and assessed these models using a non-parametric approach in a zero-shot setting. Our findings reveal a significant correlation between the models' outputs and established patterns of sound symbolism, particularly in models trained on speech data. These results suggest that such models can capture sound-meaning connections akin to human language processing, providing insights into both cognitive architectures and machine learning strategies.
【16】 WaveletGPT: Wavelets Meet Large Language Models
标题: WaveletGPT:WaveletGPT满足大型语言模型
作者:Prateek Verma
备注:16 pages, 4 figures
链接:点击下载PDF文件
摘要:大型语言模型(LLM)带来了新一轮人工智能进步,影响到每个科学领域和学科。它们的训练目标很简单:根据先前的上下文预测下一个标记。我们生活在一个世界里,我们周围的大多数数据,例如,文本、音频和音乐都具有多尺度结构,本文在LLM的预训练过程中引入了传统的信号处理思想,即小波,以充分利用这种结构。在不向GPT风格的LLM架构添加 textbf{任何额外参数}的情况下,我们在文本、原始音频和符号音乐中实现了几乎两倍的预训练性能。这是通过在中间嵌入上施加结构来实现的。当训练相同数量的训练步骤时,我们在性能上获得了显着的增益,这与预训练更大的神经架构相当。我们的架构允许每个下一个令牌预测访问中间嵌入在不同的时间分辨率在每个Transformer解码器块。这项工作有望为将多速率信号处理思想融入传统LLM预训练铺平道路。此外,我们展示了通过改善内部结构而不仅仅是追求规模来提升模型性能。摘要:Large Language Models (LLMs) have ushered in a new wave of artificial intelligence advancements impacting every scientific field and discipline. They are trained on a simple objective: to predict the next token given the previous context. We live in a world where most of the data around us, e.g., text, audio, and music, has a multi-scale structure associated with it. This paper infuses LLMs with traditional signal processing ideas, namely wavelets, during pre-training to take advantage of the structure. Without adding textbf{any extra parameters} to a GPT-style LLM architecture, we achieve the same pre-training performance almost twice as fast in text, raw audio, and symbolic music. This is achieved by imposing a structure on intermediate embeddings. When trained for the same number of training steps, we achieve significant gains in performance, which is comparable to pre-training a larger neural architecture. Our architecture allows every next token prediction access to intermediate embeddings at different temporal resolutions in every Transformer decoder block. This work will hopefully pave the way for incorporating multi-rate signal processing ideas into traditional LLM pre-training. Further, we showcase pushing model performance by improving internal structure instead of just going after scale.
【17】 NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
标题: NDVQ:具有基于正态分布的载体量化的鲁棒神经音频编解码器
作者:Zhikang Niu,Sanyuan Chen,Long Zhou,Ziyang Ma,Xie Chen,Shujie Liu
链接:点击下载PDF文件
摘要:基于矢量量化(VQ)的离散音频编解码器模型在音频压缩和自回归音频生成方面取得了巨大的成功。然而,现有的模型在感知质量和信号失真方面面临着巨大的挑战,特别是当在极低的带宽中操作时,根源于VQ码本对噪声的敏感性。这种退化对几个下游任务(例如基于编解码器的语音合成)提出了重大挑战。为了解决这个问题,我们提出了一种新的矢量量化方法,基于正态分布的矢量量化(NDVQ),通过引入一个明确的边缘之间的矢量量化码通过学习方差。具体来说,我们的方法涉及到将波形映射到一个潜在的空间,并通过选择最有可能的正态分布对其进行量化,每个码本条目表示由其均值和方差定义的唯一正态分布。使用这些基于分布的VQ编解码器代码,解码器重构输入波形。NDVQ的训练与额外的分布相关的损失,以及重建和歧视损失。实验表明,NDVQ优于现有的音频压缩基线,如EnCodec,在音频质量和zero-shot TTS,特别是在非常低的带宽的情况下。摘要:Built upon vector quantization (VQ), discrete audio codec models have achieved great success in audio compression and auto-regressive audio generation. However, existing models face substantial challenges in perceptual quality and signal distortion, especially when operating in extremely low bandwidth, rooted in the sensitivity of the VQ codebook to noise. This degradation poses significant challenges for several downstream tasks, such as codec-based speech synthesis. To address this issue, we propose a novel VQ method, Normal Distribution-based Vector Quantization (NDVQ), by introducing an explicit margin between the VQ codes via learning a variance. Specifically, our approach involves mapping the waveform to a latent space and quantizing it by selecting the most likely normal distribution, with each codebook entry representing a unique normal distribution defined by its mean and variance. Using these distribution-based VQ codec codes, a decoder reconstructs the input waveform. NDVQ is trained with additional distribution-related losses, alongside reconstruction and discrimination losses. Experiments demonstrate that NDVQ outperforms existing audio compression baselines, such as EnCodec, in terms of audio quality and zero-shot TTS, particularly in very low bandwidth scenarios.
【18】 AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
标题: AudioComposer:利用自然语言描述实现细粒度音频生成
作者:Yuanyuan Wang,Hangting Chen,Dongchao Yang,Zhiyong Wu,Helen Meng,Xixin Wu
链接:点击下载PDF文件
摘要:目前的文本到音频(TTA)模型主要使用粗糙的文本描述作为输入来生成音频,这阻碍了模型生成具有对内容和风格的细粒度控制的音频。一些研究试图通过引入额外的帧级条件或控制网络来提高粒度。然而,这通常会导致复杂的系统设计和困难,由于需要参考帧级的条件。为了解决这些挑战,我们提出了AudioComposer,一种新的TTA生成框架,它完全依赖于自然语言描述(NLDs)来提供内容规范和样式控制信息。为了进一步增强音频生成建模,我们采用具有交叉注意机制的基于流的扩散Transformers,将文本描述有效地融入音频生成过程,不仅可以同时考虑文本输入中的内容和风格信息,而且与其他架构相比还可以加速生成。此外,我们提出了一种新的和全面的自动数据模拟管道来构建数据与细粒度的文本描述,这显着地阐明了该地区的数据稀缺的问题。实验表明,我们的框架使用单独的NLD作为输入的内容规范和风格控制的有效性。发电质量和可控性超过了最先进的TTA模型,即使模型尺寸较小。摘要:Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content and style. Some studies try to improve the granularity by incorporating additional frame-level conditions or control networks. However, this usually leads to complex system design and difficulties due to the requirement for reference frame-level conditions. To address these challenges, we propose AudioComposer, a novel TTA generation framework that relies solely on natural language descriptions (NLDs) to provide both content specification and style control information. To further enhance audio generative modeling, we employ flow-based diffusion transformers with the cross-attention mechanism to incorporate text descriptions effectively into audio generation processes, which can not only simultaneously consider the content and style information in the text inputs, but also accelerate generation compared to other architectures. Furthermore, we propose a novel and comprehensive automatic data simulation pipeline to construct data with fine-grained text descriptions, which significantly alleviates the problem of data scarcity in the area. Experiments demonstrate the effectiveness of our framework using solely NLDs as inputs for content specification and style control. The generation quality and controllability surpass state-of-the-art TTA models, even with a smaller model size.
【19】 Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech Enhancement
标题: 用于脑辅助语音增强的几何限制的脑电通道选择
作者:Keying Zuo,Qingtian Xu,Jie Zhang,Zhenhua Ling
链接:点击下载PDF文件
摘要:脑辅助语音增强(Brain-assisted speech enhancement,BASE)的目的是利用脑电(electroencephalogram,EEG)信号作为辅助手段,在复杂的多说话人场景中提取目标说话人,因为听者的听觉注意力可以从大脑的神经电信号中解码出来。这有利于潜在的整合EEG电极与听力设备,以提高听力受损的听众,这是由最近提出的BASEN模型所示的语音清晰度。由于多通道EEG信号通常是高度相关的,有些甚至与听力无关,盲目地合并所有EEG通道将导致高的经济和计算成本。因此,在这项工作中,我们提出了一个几何约束的EEG通道选择方法的基础。我们设计了一个新的加权多膨胀时间卷积网络(WDTCN)作为骨干,以取代BASEN中的Conv-TasNet。给定一个由电极几何形状定义的用于可行积分的原始通道集,然后我们提出了一个用于WD-TCN的几何约束卷积正则化选择(GC-ConvRS)模块,以找到信息丰富的EEG子集。在公开数据集上的实验结果表明了该方法的优越性。GC-ConvRS可以进一步细化有用的EEG子集的几何约束,从而在性能和集成成本之间实现更好的权衡。摘要:Brain-assisted speech enhancement (BASE) aims to extract the target speaker in complex multi-talker scenarios using electroencephalogram (EEG) signals as an assistive modality, as the auditory attention of the listener can be decoded from electroneurographic signals of the brain. This facilitates a potential integration of EEG electrodes with listening devices to improve the speech intelligibility of hearing-impaired listeners, which was shown by the recently-proposed BASEN model. As in general the multichannel EEG signals are highly correlated and some are even irrelevant to listening, blindly incorporating all EEG channels would lead to a high economic and computational cost. In this work, we therefore propose a geometry-constrained EEG channel selection approach for BASE. We design a new weighted multi-dilation temporal convolutional network (WDTCN) as the backbone to replace the Conv-TasNet in BASEN. Given a raw channel set that is defined by the electrode geometry for feasible integration, we then propose a geometry-constrained convolutional regularization selection (GC-ConvRS) module for WD-TCN to find an informative EEG subset. Experimental results on a public dataset show the superiority of the proposed WD-TCN over BASEN. The GC-ConvRS can further refine the useful EEG subset subject to the geometry constraint, resulting in a better trade-off between performance and integration cost.
【20】 Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features
标题: 具有复杂频谱图和可学习时间特征的语音分解Transformer
作者:Younghoo Kwon,Jung-Woo Choi
备注:5 pages, 2 figures, submitted to ICASSP 2024
链接:点击下载PDF文件
摘要:我们提出了一个基于变换器的语音declipping模型,有效地恢复在广泛的输入信号失真比(SDR)的削波信号。虽然最近基于时域深度神经网络(DNN)的declippers的性能优于传统的手工制作和基于频谱的DNN方法,但它们仍然难以处理低SDR输入。为了解决这个问题,我们引入了一个在时频(TF)域中运行的基于变换器的架构。TF-变换器架构在低SDR信号的语音增强任务中表现出显著的性能,但对于时域伪影(如削波)不能是最佳的。为了克服基于频谱的DNN的局限性,我们设计了一个额外的卷积块,直接从时域波形中提取时间特征。复杂频谱图和学习的时间特征的联合分析允许模型在高和低SDR输入上提高性能。我们的方法还保留了处理过程中的语音信号的未剪辑部分,防止退化时,通常只使用频谱信息。在对VoiceBank-DEMAND和DNS挑战数据集的评估中,所提出的模型在各种指标上始终优于最先进的(SOTA)declipping模型,证明了其鲁棒性和可推广性。摘要:We present a transformer-based speech-declipping model that effectively recovers clipped signals across a wide range of input signal-to-distortion ratios (SDRs). While recent time-domain deep neural network (DNN)-based declippers have outperformed traditional handcrafted and spectrogram-based DNN approaches, they still struggle with low-SDR inputs. To address this, we incorporate a transformer-based architecture that operates in the time-frequency (TF) domain. The TF-transformer architecture has demonstrated remarkable performance in the speech enhancement task for low-SDR signals but cannot be optimal for the time-domain artifact like clipping. To overcome the limitations of spectrogram-based DNNs, we design an extra convolutional block that directly extracts temporal features from time-domain waveforms. The joint analysis of complex spectrogram and learned temporal features allows the model to improve performance on both high- and low-SDR inputs. Our approach also preserves the unclipped portions of the speech signal during processing, preventing degradation typically seen when only spectral information is used. In evaluations on the VoiceBank-DEMAND and DNS challenge datasets, the proposed model consistently outperformed state-of-the-art (SOTA) declipping models across various metrics, demonstrating its robustness and generalizability.
【21】 Multichannel-to-Multichannel Target Sound Extraction Using Direction and Timestamp Clues
标题: 使用方向和时间戳线索的多通道到多通道目标声音提取
作者:Dayun Choi,Jung-Woo Choi
备注:5 pages, 4 figures
链接:点击下载PDF文件
摘要:我们提出了一个多通道到多通道的目标声音提取(M2M-TSE)的框架分离多通道目标信号从多通道混合声源。目标声音提取(TSE)使用用户提供的线索来隔离特定的目标信号,通常专注于具有类别标签或时间激活图的单通道提取。然而,为了保存和利用多声道音频信号中的空间信息,必须提取目标声源的多声道信号。此外,用于提取的线索还可以包括空间或时间线索,如到达方向(DoA)或源激活的时间戳。为了解决这些挑战,我们提出了一个M2M框架,提取多通道的声音信号的基础上的时空线索。我们证明,我们的基于变压器的架构可以成功地完成M2M-TSE任务的多通道信号合成的音频信号在不同的房间环境中的不同类别。此外,我们表明,多通道提取任务在DNN中引入了足够的归纳偏差,使其能够直接处理DoA线索,而无需利用手工制作的空间特征。摘要:We propose a multichannel-to-multichannel target sound extraction (M2M-TSE) framework for separating multichannel target signals from a multichannel mixture of sound sources. Target sound extraction (TSE) isolates a specific target signal using user-provided clues, typically focusing on single-channel extraction with class labels or temporal activation maps. However, to preserve and utilize spatial information in multichannel audio signals, it is essential to extract multichannel signals of a target sound source. Moreover, the clue for extraction can also include spatial or temporal cues like direction-of-arrival (DoA) or timestamps of source activation. To address these challenges, we present an M2M framework that extracts a multichannel sound signal based on spatio-temporal clues. We demonstrate that our transformer-based architecture can successively accomplish the M2M-TSE task for multichannel signals synthesized from audio signals of diverse classes in different room environments. Furthermore, we show that the multichannel extraction task introduces sufficient inductive bias in the DNN, allowing it to directly handle DoA clues without utilizing hand-crafted spatial features.
【22】 DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
标题: DeFT-Mamba:通用多通道声音分离和复音音频分类
作者:Dongheon Lee,Jung-Woo Choi
备注:5 pages, 2 figures
链接:点击下载PDF文件
摘要:本文提出了一个通用的声音分离和复调音频分类的框架,解决了在多声道混合中分离和分类单个声源的挑战。该框架DeFT-Mamba利用密集频率-时间注意网络(DeFTAN)结合Mamba提取声音对象,通过门控卷积块捕获局部时频关系,通过位置混合Mamba捕获全局时频关系。DeFT-Mamba大大超过了现有的分离和分类网络,特别是在涉及类内复调的复杂场景中。此外,基于分类的源计数方法被引入到识别多个源的存在,优于传统的基于阈值的方法。分离细化调整也提出了进一步提高性能。所提出的框架进行了训练和测试,在这项工作中开发的多通道通用声音分离数据集,旨在模仿现实环境中的移动源和不同的开始和复调事件的偏移。摘要:This paper presents a framework for universal sound separation and polyphonic audio classification, addressing the challenges of separating and classifying individual sound sources in a multichannel mixture. The proposed framework, DeFT-Mamba, utilizes the dense frequency-time attentive network (DeFTAN) combined with Mamba to extract sound objects, capturing the local time-frequency relations through gated convolution block and the global time-frequency relations through position-wise Hybrid Mamba. DeFT-Mamba surpasses existing separation and classification networks by a large margin, particularly in complex scenarios involving in-class polyphony. Additionally, a classification-based source counting method is introduced to identify the presence of multiple sources, outperforming conventional threshold-based approaches. Separation refinement tuning is also proposed to improve performance further. The proposed framework is trained and tested on a multichannel universal sound separation dataset developed in this work, designed to mimic realistic environments with moving sources and varying onsets and offsets of polyphonic events.
【23】 Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
标题: 使用说话者感知的CTC在多说话者语音识别中理清说话者
作者:Jiawen Kang,Lingwei Meng,Mingyu Cui,Yuejiao Wang,Xixin Wu,Xunying Liu,Helen Meng
链接:点击下载PDF文件
摘要:多说话者语音识别(MTASR)在解开和转录重叠语音方面面临着独特的挑战。为了解决这些挑战,本文研究了联结主义时间分类(CTC)在与MTASR的序列化输出训练(SOT)结合时在说话者解纠缠中的作用。我们的可视化显示,CTC指导编码器表示不同的扬声器在不同的时间区域的声学嵌入。利用这一洞察力,我们提出了一个新的扬声器感知CTC(SACTC)的训练目标,贝叶斯风险CTC框架的基础上。SACTC是针对多说话者场景定制的CTC变体,其通过约束编码器以在特定时间帧表示不同说话者的令牌来显式地建模说话者解纠缠。当与SOT集成时,SOT-SACTC模型在不同程度的语音重叠上始终优于标准SOT-CTC。具体来说,我们观察到相对字错误率降低10%的整体和15%的低重叠语音。这项工作代表了MTASR任务的CTC为基础的增强的初步探索,提供了一个新的视角说话人解开多说话人语音识别。摘要:Multi-talker speech recognition (MTASR) faces unique challenges in disentangling and transcribing overlapping speech. To address these challenges, this paper investigates the role of Connectionist Temporal Classification (CTC) in speaker disentanglement when incorporated with Serialized Output Training (SOT) for MTASR. Our visualization reveals that CTC guides the encoder to represent different speakers in distinct temporal regions of acoustic embeddings. Leveraging this insight, we propose a novel Speaker-Aware CTC (SACTC) training objective, based on the Bayes risk CTC framework. SACTC is a tailored CTC variant for multi-talker scenarios, it explicitly models speaker disentanglement by constraining the encoder to represent different speakers' tokens at specific time frames. When integrated with SOT, the SOT-SACTC model consistently outperforms standard SOT-CTC across various degrees of speech overlap. Specifically, we observe relative word error rate reductions of 10% overall and 15% on low-overlap speech. This work represents an initial exploration of CTC-based enhancements for MTASR tasks, offering a new perspective on speaker disentanglement in multi-talker speech recognition.
【24】 Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
标题: 具有混合专家的稳健视听语音识别模型
作者:Yihan Wu,Yifan Peng,Yichen Lu,Xuankai Chang,Ruihua Song,Shinji Watanabe
备注:6 pages, 2 figures, accepted by IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
摘要:视觉信号可以通过提供额外的上下文信息来提高视听语音识别的准确性。鉴于视觉信号的复杂性,视听语音识别模型需要在不同的视频场景中具有强大的泛化能力,这是一个重大挑战。在本文中,我们引入EVA,利用视听ASR的专家混合来执行“野外”视频的鲁棒语音识别。具体来说,我们首先将视觉信息编码成视觉符号序列,并通过轻量级投影将其映射到语音空间。然后,我们建立EVA上一个强大的预训练的语音识别模型,确保其泛化能力。此外,为了有效地整合视觉信息,我们通过一个专家混合模块将视觉信息注入到ASR模型中。实验表明,我们的模型在三个基准测试中取得了最先进的结果,这证明了EVA在不同视频领域的泛化能力。摘要:Visual signals can enhance audiovisual speech recognition accuracy by providing additional contextual information. Given the complexity of visual signals, an audiovisual speech recognition model requires robust generalization capabilities across diverse video scenarios, presenting a significant challenge. In this paper, we introduce EVA, leveraging the mixture-of-Experts for audioVisual ASR to perform robust speech recognition for in-the-wild'' videos. Specifically, we first encode visual information into visual tokens sequence and map them into speech space by a lightweight projection. Then, we build EVA upon a robust pretrained speech recognition model, ensuring its generalization ability. Moreover, to incorporate visual information effectively, we inject visual information into the ASR model through a mixture-of-experts module. Experiments show our model achieves state-of-the-art results on three benchmarks, which demonstrates the generalization ability of EVA across diverse video domains.
【25】 META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
标题: META-CAT:通过Meta信息级联用于多说话者ASB的说话者知情的语音嵌入
作者:Jinhan Wang,Weiqing Wang,Kunal Dhawan,Taejin Park,Myungjong Kim,Ivan Medennikov,He Huang,Nithin Koluguri,Jagadeesh Balam,Boris Ginsburg
链接:点击下载PDF文件
摘要:我们提出了一种新的端到端多说话人自动语音识别(ASR)框架,使多说话人(MS)ASR和目标说话人(TS)ASR。我们提出的模型以完全端到端的方式进行训练,并从预先训练的扬声器日记模块中纳入扬声器监督。我们介绍了一种直观而有效的方法,用于掩蔽ASR编码器激活使用从扬声器的监督模块,一种技术,我们术语元猫(元信息级联),可以应用于MS-ASR和TS-ASR的输出。我们的研究结果表明,所提出的架构在MS-ASR和TS-ASR任务中都达到了有竞争力的性能,而不需要传统的方法,如神经掩码估计或在音频或特征级别的掩蔽。此外,我们展示了一个统一的双任务模型,可以有效地处理MS-ASR和TS-ASR任务的一瞥。因此,这项工作表明,一个强大的端到端的多说话者ASR框架可以实现一个精简的架构,避免了需要在以前的研究中采用的复杂的扬声器过滤机制。摘要:We propose a novel end-to-end multi-talker automatic speech recognition (ASR) framework that enables both multi-speaker (MS) ASR and target-speaker (TS) ASR. Our proposed model is trained in a fully end-to-end manner, incorporating speaker supervision from a pre-trained speaker diarization module. We introduce an intuitive yet effective method for masking ASR encoder activations using output from the speaker supervision module, a technique we term Meta-Cat (meta-information concatenation), that can be applied to both MS-ASR and TS-ASR. Our results demonstrate that the proposed architecture achieves competitive performance in both MS-ASR and TS-ASR tasks, without the need for traditional methods, such as neural mask estimation or masking at the audio or feature level. Furthermore, we demonstrate a glimpse of a unified dual-task model which can efficiently handle both MS-ASR and TS-ASR tasks. Thus, this work illustrates that a robust end-to-end multi-talker ASR framework can be implemented with a streamlined architecture, obviating the need for the complex speaker filtering mechanisms employed in previous studies.
eess.AS音频处理
【1】 WaveletGPT: Wavelets Meet Large Language Models标题: WaveletGPT:WaveletGPT满足大型语言模型
作者:Prateek Verma
备注:16 pages, 4 figures
链接:点击下载PDF文件
摘要:大型语言模型(LLM)带来了新一轮人工智能进步,影响到每个科学领域和学科。它们的训练目标很简单:根据先前的上下文预测下一个标记。我们生活在一个世界里,我们周围的大多数数据,例如,文本、音频和音乐都具有多尺度结构,本文在LLM的预训练过程中引入了传统的信号处理思想,即小波,以充分利用这种结构。在不向GPT风格的LLM架构添加 textbf{任何额外参数}的情况下,我们在文本、原始音频和符号音乐中实现了几乎两倍的预训练性能。这是通过在中间嵌入上施加结构来实现的。当训练相同数量的训练步骤时,我们在性能上获得了显着的增益,这与预训练更大的神经架构相当。我们的架构允许每个下一个令牌预测访问中间嵌入在不同的时间分辨率在每个Transformer解码器块。这项工作有望为将多速率信号处理思想融入传统LLM预训练铺平道路。此外,我们展示了通过改善内部结构而不仅仅是追求规模来提升模型性能。摘要:Large Language Models (LLMs) have ushered in a new wave of artificial intelligence advancements impacting every scientific field and discipline. They are trained on a simple objective: to predict the next token given the previous context. We live in a world where most of the data around us, e.g., text, audio, and music, has a multi-scale structure associated with it. This paper infuses LLMs with traditional signal processing ideas, namely wavelets, during pre-training to take advantage of the structure. Without adding textbf{any extra parameters} to a GPT-style LLM architecture, we achieve the same pre-training performance almost twice as fast in text, raw audio, and symbolic music. This is achieved by imposing a structure on intermediate embeddings. When trained for the same number of training steps, we achieve significant gains in performance, which is comparable to pre-training a larger neural architecture. Our architecture allows every next token prediction access to intermediate embeddings at different temporal resolutions in every Transformer decoder block. This work will hopefully pave the way for incorporating multi-rate signal processing ideas into traditional LLM pre-training. Further, we showcase pushing model performance by improving internal structure instead of just going after scale.
【2】 NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
标题: NDVQ:具有基于正态分布的载体量化的鲁棒神经音频编解码器
作者:Zhikang Niu,Sanyuan Chen,Long Zhou,Ziyang Ma,Xie Chen,Shujie Liu
链接:点击下载PDF文件
摘要:基于矢量量化(VQ)的离散音频编解码器模型在音频压缩和自回归音频生成方面取得了巨大的成功。然而,现有的模型在感知质量和信号失真方面面临着巨大的挑战,特别是当在极低的带宽中操作时,根源于VQ码本对噪声的敏感性。这种退化对几个下游任务(例如基于编解码器的语音合成)提出了重大挑战。为了解决这个问题,我们提出了一种新的矢量量化方法,基于正态分布的矢量量化(NDVQ),通过引入一个明确的边缘之间的矢量量化码通过学习方差。具体来说,我们的方法涉及到将波形映射到一个潜在的空间,并通过选择最有可能的正态分布对其进行量化,每个码本条目表示由其均值和方差定义的唯一正态分布。使用这些基于分布的VQ编解码器代码,解码器重构输入波形。NDVQ的训练与额外的分布相关的损失,以及重建和歧视损失。实验表明,NDVQ优于现有的音频压缩基线,如EnCodec,在音频质量和zero-shot TTS,特别是在非常低的带宽的情况下。摘要:Built upon vector quantization (VQ), discrete audio codec models have achieved great success in audio compression and auto-regressive audio generation. However, existing models face substantial challenges in perceptual quality and signal distortion, especially when operating in extremely low bandwidth, rooted in the sensitivity of the VQ codebook to noise. This degradation poses significant challenges for several downstream tasks, such as codec-based speech synthesis. To address this issue, we propose a novel VQ method, Normal Distribution-based Vector Quantization (NDVQ), by introducing an explicit margin between the VQ codes via learning a variance. Specifically, our approach involves mapping the waveform to a latent space and quantizing it by selecting the most likely normal distribution, with each codebook entry representing a unique normal distribution defined by its mean and variance. Using these distribution-based VQ codec codes, a decoder reconstructs the input waveform. NDVQ is trained with additional distribution-related losses, alongside reconstruction and discrimination losses. Experiments demonstrate that NDVQ outperforms existing audio compression baselines, such as EnCodec, in terms of audio quality and zero-shot TTS, particularly in very low bandwidth scenarios.
【3】 AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
标题: AudioComposer:利用自然语言描述实现细粒度音频生成
作者:Yuanyuan Wang,Hangting Chen,Dongchao Yang,Zhiyong Wu,Helen Meng,Xixin Wu
链接:点击下载PDF文件
摘要:目前的文本到音频(TTA)模型主要使用粗糙的文本描述作为输入来生成音频,这阻碍了模型生成具有对内容和风格的细粒度控制的音频。一些研究试图通过引入额外的帧级条件或控制网络来提高粒度。然而,这通常会导致复杂的系统设计和困难,由于需要参考帧级的条件。为了解决这些挑战,我们提出了AudioComposer,一种新的TTA生成框架,它完全依赖于自然语言描述(NLDs)来提供内容规范和样式控制信息。为了进一步增强音频生成建模,我们采用基于流的扩散Transformers与交叉注意机制,有效地将文本描述纳入音频生成过程,不仅可以同时考虑文本输入中的内容和风格信息,而且与其他架构相比,还可以加速生成。此外,我们提出了一种新的和全面的自动数据模拟管道来构建数据与细粒度的文本描述,这显着地阐明了该地区的数据稀缺的问题。实验表明,我们的框架使用单独的NLD作为输入的内容规范和风格控制的有效性。发电质量和可控性超过了最先进的TTA模型,即使模型尺寸较小。摘要:Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content and style. Some studies try to improve the granularity by incorporating additional frame-level conditions or control networks. However, this usually leads to complex system design and difficulties due to the requirement for reference frame-level conditions. To address these challenges, we propose AudioComposer, a novel TTA generation framework that relies solely on natural language descriptions (NLDs) to provide both content specification and style control information. To further enhance audio generative modeling, we employ flow-based diffusion transformers with the cross-attention mechanism to incorporate text descriptions effectively into audio generation processes, which can not only simultaneously consider the content and style information in the text inputs, but also accelerate generation compared to other architectures. Furthermore, we propose a novel and comprehensive automatic data simulation pipeline to construct data with fine-grained text descriptions, which significantly alleviates the problem of data scarcity in the area. Experiments demonstrate the effectiveness of our framework using solely NLDs as inputs for content specification and style control. The generation quality and controllability surpass state-of-the-art TTA models, even with a smaller model size.
【4】 Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech Enhancement
标题: 用于脑辅助语音增强的几何限制的脑电通道选择
作者:Keying Zuo,Qingtian Xu,Jie Zhang,Zhenhua Ling
链接:点击下载PDF文件
摘要:脑辅助语音增强(Brain-assisted speech enhancement,BASE)的目的是利用脑电(electroencephalogram,EEG)信号作为辅助手段,在复杂的多说话人场景中提取目标说话人,因为听者的听觉注意力可以从大脑的神经电信号中解码出来。这有利于潜在的整合EEG电极与听力设备,以提高听力受损的听众,这是由最近提出的BASEN模型所示的语音清晰度。由于多通道EEG信号通常是高度相关的,有些甚至与听力无关,盲目地合并所有EEG通道将导致高的经济和计算成本。因此,在这项工作中,我们提出了一个几何约束的EEG通道选择方法的基础。我们设计了一个新的加权多膨胀时间卷积网络(WDTCN)作为骨干,以取代BASEN中的Conv-TasNet。给定一个由电极几何形状定义的用于可行积分的原始通道集,然后我们提出了一个用于WD-TCN的几何约束卷积正则化选择(GC-ConvRS)模块,以找到信息丰富的EEG子集。在公开数据集上的实验结果表明了该方法的优越性。GC-ConvRS可以进一步细化有用的EEG子集的几何约束,从而在性能和集成成本之间实现更好的权衡。摘要:Brain-assisted speech enhancement (BASE) aims to extract the target speaker in complex multi-talker scenarios using electroencephalogram (EEG) signals as an assistive modality, as the auditory attention of the listener can be decoded from electroneurographic signals of the brain. This facilitates a potential integration of EEG electrodes with listening devices to improve the speech intelligibility of hearing-impaired listeners, which was shown by the recently-proposed BASEN model. As in general the multichannel EEG signals are highly correlated and some are even irrelevant to listening, blindly incorporating all EEG channels would lead to a high economic and computational cost. In this work, we therefore propose a geometry-constrained EEG channel selection approach for BASE. We design a new weighted multi-dilation temporal convolutional network (WDTCN) as the backbone to replace the Conv-TasNet in BASEN. Given a raw channel set that is defined by the electrode geometry for feasible integration, we then propose a geometry-constrained convolutional regularization selection (GC-ConvRS) module for WD-TCN to find an informative EEG subset. Experimental results on a public dataset show the superiority of the proposed WD-TCN over BASEN. The GC-ConvRS can further refine the useful EEG subset subject to the geometry constraint, resulting in a better trade-off between performance and integration cost.
【5】 Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features
标题: 具有复杂频谱图和可学习时间特征的语音分解Transformer
作者:Younghoo Kwon,Jung-Woo Choi
备注:5 pages, 2 figures, submitted to ICASSP 2024
链接:点击下载PDF文件
摘要:我们提出了一个基于变换器的语音declipping模型,有效地恢复在广泛的输入信号失真比(SDR)的削波信号。虽然最近基于时域深度神经网络(DNN)的declippers的性能优于传统的手工制作和基于频谱的DNN方法,但它们仍然难以处理低SDR输入。为了解决这个问题,我们引入了一个在时频(TF)域中运行的基于变换器的架构。TF-变换器架构在低SDR信号的语音增强任务中表现出显著的性能,但对于时域伪影(如削波)不能是最佳的。为了克服基于频谱的DNN的局限性,我们设计了一个额外的卷积块,直接从时域波形中提取时间特征。复杂频谱图和学习的时间特征的联合分析允许模型在高和低SDR输入上提高性能。我们的方法还保留了处理过程中的语音信号的未剪辑部分,防止退化时,通常只使用频谱信息。在对VoiceBank-DEMAND和DNS挑战数据集的评估中,所提出的模型在各种指标上始终优于最先进的(SOTA)declipping模型,证明了其鲁棒性和可推广性。摘要:We present a transformer-based speech-declipping model that effectively recovers clipped signals across a wide range of input signal-to-distortion ratios (SDRs). While recent time-domain deep neural network (DNN)-based declippers have outperformed traditional handcrafted and spectrogram-based DNN approaches, they still struggle with low-SDR inputs. To address this, we incorporate a transformer-based architecture that operates in the time-frequency (TF) domain. The TF-transformer architecture has demonstrated remarkable performance in the speech enhancement task for low-SDR signals but cannot be optimal for the time-domain artifact like clipping. To overcome the limitations of spectrogram-based DNNs, we design an extra convolutional block that directly extracts temporal features from time-domain waveforms. The joint analysis of complex spectrogram and learned temporal features allows the model to improve performance on both high- and low-SDR inputs. Our approach also preserves the unclipped portions of the speech signal during processing, preventing degradation typically seen when only spectral information is used. In evaluations on the VoiceBank-DEMAND and DNS challenge datasets, the proposed model consistently outperformed state-of-the-art (SOTA) declipping models across various metrics, demonstrating its robustness and generalizability.
【6】 Multichannel-to-Multichannel Target Sound Extraction Using Direction and Timestamp Clues
标题: 使用方向和时间戳线索的多通道到多通道目标声音提取
作者:Dayun Choi,Jung-Woo Choi
备注:5 pages, 4 figures
链接:点击下载PDF文件
摘要:我们提出了一个多通道到多通道的目标声音提取(M2M-TSE)的框架分离多通道目标信号从多通道混合声源。目标声音提取(TSE)使用用户提供的线索隔离特定目标信号,通常侧重于使用类别标签或时间激活图的单通道提取。然而,为了保存和利用多声道音频信号中的空间信息,必须提取目标声源的多声道信号。此外,用于提取的线索还可以包括空间或时间线索,如到达方向(DoA)或源激活的时间戳。为了解决这些挑战,我们提出了一个M2M框架,提取多通道的声音信号的基础上的时空线索。我们证明,我们基于变换器的架构可以成功地完成从不同房间环境中的不同类别的音频信号合成的多通道信号的M2M-TSE任务。此外,我们表明,多通道提取任务在DNN中引入了足够的归纳偏差,使其能够直接处理DoA线索,而无需利用手工制作的空间特征。摘要:We propose a multichannel-to-multichannel target sound extraction (M2M-TSE) framework for separating multichannel target signals from a multichannel mixture of sound sources. Target sound extraction (TSE) isolates a specific target signal using user-provided clues, typically focusing on single-channel extraction with class labels or temporal activation maps. However, to preserve and utilize spatial information in multichannel audio signals, it is essential to extract multichannel signals of a target sound source. Moreover, the clue for extraction can also include spatial or temporal cues like direction-of-arrival (DoA) or timestamps of source activation. To address these challenges, we present an M2M framework that extracts a multichannel sound signal based on spatio-temporal clues. We demonstrate that our transformer-based architecture can successively accomplish the M2M-TSE task for multichannel signals synthesized from audio signals of diverse classes in different room environments. Furthermore, we show that the multichannel extraction task introduces sufficient inductive bias in the DNN, allowing it to directly handle DoA clues without utilizing hand-crafted spatial features.
【7】 DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
标题: DeFT-Mamba:通用多通道声音分离和复音音频分类
作者:Dongheon Lee,Jung-Woo Choi
备注:5 pages, 2 figures
链接:点击下载PDF文件
摘要:本文提出了一个通用的声音分离和复调音频分类的框架,解决了在多声道混合中分离和分类单个声源的挑战。该框架DeFT-Mamba利用密集频率-时间注意网络(DeFTAN)结合Mamba提取声音对象,通过门控卷积块捕获局部时频关系,通过位置混合Mamba捕获全局时频关系。DeFT-Mamba大大超过了现有的分离和分类网络,特别是在涉及类内复调的复杂场景中。此外,基于分类的源计数方法被引入到识别多个源的存在,优于传统的基于阈值的方法。分离细化调整也提出了进一步提高性能。所提出的框架进行了训练和测试,在这项工作中开发的多通道通用声音分离数据集,旨在模仿现实环境中的移动源和不同的开始和复调事件的偏移。摘要:This paper presents a framework for universal sound separation and polyphonic audio classification, addressing the challenges of separating and classifying individual sound sources in a multichannel mixture. The proposed framework, DeFT-Mamba, utilizes the dense frequency-time attentive network (DeFTAN) combined with Mamba to extract sound objects, capturing the local time-frequency relations through gated convolution block and the global time-frequency relations through position-wise Hybrid Mamba. DeFT-Mamba surpasses existing separation and classification networks by a large margin, particularly in complex scenarios involving in-class polyphony. Additionally, a classification-based source counting method is introduced to identify the presence of multiple sources, outperforming conventional threshold-based approaches. Separation refinement tuning is also proposed to improve performance further. The proposed framework is trained and tested on a multichannel universal sound separation dataset developed in this work, designed to mimic realistic environments with moving sources and varying onsets and offsets of polyphonic events.
【8】 Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
标题: 使用说话者感知的CTC在多说话者语音识别中理清说话者
作者:Jiawen Kang,Lingwei Meng,Mingyu Cui,Yuejiao Wang,Xixin Wu,Xunying Liu,Helen Meng
链接:点击下载PDF文件
摘要:多说话人语音识别(MTASR)在分离和转录重叠语音方面面临着独特的挑战。为了解决这些挑战,本文研究了连接主义时间分类(CTC)在说话人解纠缠中的作用,结合序列化输出训练(SOT)的MTASR。我们的可视化显示,CTC指导编码器表示不同的扬声器在不同的时间区域的声学嵌入。利用这一洞察力,我们提出了一个新的扬声器感知CTC(SACTC)的训练目标,贝叶斯风险CTC框架的基础上。SACTC是针对多说话者场景定制的CTC变体,其通过约束编码器以在特定时间帧表示不同说话者的令牌来显式地建模说话者解纠缠。当与SOT集成时,SOT-SACTC模型在不同程度的语音重叠上始终优于标准SOT-CTC。具体来说,我们观察到总体相对单词错误率降低了10%,低重叠语音的相对单词错误率降低了15%。这项工作代表了MTASR任务的CTC为基础的增强的初步探索,提供了一个新的视角说话人解开多说话人语音识别。摘要:Multi-talker speech recognition (MTASR) faces unique challenges in disentangling and transcribing overlapping speech. To address these challenges, this paper investigates the role of Connectionist Temporal Classification (CTC) in speaker disentanglement when incorporated with Serialized Output Training (SOT) for MTASR. Our visualization reveals that CTC guides the encoder to represent different speakers in distinct temporal regions of acoustic embeddings. Leveraging this insight, we propose a novel Speaker-Aware CTC (SACTC) training objective, based on the Bayes risk CTC framework. SACTC is a tailored CTC variant for multi-talker scenarios, it explicitly models speaker disentanglement by constraining the encoder to represent different speakers' tokens at specific time frames. When integrated with SOT, the SOT-SACTC model consistently outperforms standard SOT-CTC across various degrees of speech overlap. Specifically, we observe relative word error rate reductions of 10% overall and 15% on low-overlap speech. This work represents an initial exploration of CTC-based enhancements for MTASR tasks, offering a new perspective on speaker disentanglement in multi-talker speech recognition.
【9】 Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
标题: 具有混合专家的稳健视听语音识别模型
作者:Yihan Wu,Yifan Peng,Yichen Lu,Xuankai Chang,Ruihua Song,Shinji Watanabe
备注:6 pages, 2 figures, accepted by IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
摘要:视觉信号可以通过提供额外的上下文信息来提高视听语音识别的准确性。鉴于视觉信号的复杂性,视听语音识别模型需要在不同的视频场景中具有强大的泛化能力,这是一个重大挑战。在本文中,我们引入EVA,利用视听ASR的专家混合来执行“野外”视频的鲁棒语音识别。具体来说,我们首先将视觉信息编码成视觉符号序列,并通过轻量级投影将其映射到语音空间。然后,我们建立EVA上一个强大的预训练的语音识别模型,确保其泛化能力。此外,为了有效地整合视觉信息,我们通过一个专家混合模块将视觉信息注入到ASR模型中。实验表明,我们的模型在三个基准测试中取得了最先进的结果,这证明了EVA在不同视频领域的泛化能力。摘要:Visual signals can enhance audiovisual speech recognition accuracy by providing additional contextual information. Given the complexity of visual signals, an audiovisual speech recognition model requires robust generalization capabilities across diverse video scenarios, presenting a significant challenge. In this paper, we introduce EVA, leveraging the mixture-of-Experts for audioVisual ASR to perform robust speech recognition for in-the-wild'' videos. Specifically, we first encode visual information into visual tokens sequence and map them into speech space by a lightweight projection. Then, we build EVA upon a robust pretrained speech recognition model, ensuring its generalization ability. Moreover, to incorporate visual information effectively, we inject visual information into the ASR model through a mixture-of-experts module. Experiments show our model achieves state-of-the-art results on three benchmarks, which demonstrates the generalization ability of EVA across diverse video domains.
【10】 META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
标题: META-CAT:通过Meta信息级联用于多说话者ASB的说话者知情的语音嵌入
作者:Jinhan Wang,Weiqing Wang,Kunal Dhawan,Taejin Park,Myungjong Kim,Ivan Medennikov,He Huang,Nithin Koluguri,Jagadeesh Balam,Boris Ginsburg
链接:点击下载PDF文件
摘要:我们提出了一种新的端到端多说话人自动语音识别(ASR)框架,使多说话人(MS)ASR和目标说话人(TS)ASR。我们提出的模型以完全端到端的方式进行训练,并从预先训练的扬声器日记模块中纳入扬声器监督。我们介绍了一种直观而有效的方法,用于掩蔽ASR编码器激活使用从扬声器的监督模块,一种技术,我们术语元猫(元信息级联),可以应用于MS-ASR和TS-ASR的输出。我们的研究结果表明,所提出的架构在MS-ASR和TS-ASR任务中都达到了有竞争力的性能,而不需要传统的方法,如神经掩码估计或在音频或特征级别的掩蔽。此外,我们展示了一个统一的双任务模型,可以有效地处理MS-ASR和TS-ASR任务的一瞥。因此,这项工作表明,一个强大的端到端的多说话者ASR框架可以实现一个精简的架构,避免了需要在以前的研究中采用的复杂的扬声器过滤机制。摘要:We propose a novel end-to-end multi-talker automatic speech recognition (ASR) framework that enables both multi-speaker (MS) ASR and target-speaker (TS) ASR. Our proposed model is trained in a fully end-to-end manner, incorporating speaker supervision from a pre-trained speaker diarization module. We introduce an intuitive yet effective method for masking ASR encoder activations using output from the speaker supervision module, a technique we term Meta-Cat (meta-information concatenation), that can be applied to both MS-ASR and TS-ASR. Our results demonstrate that the proposed architecture achieves competitive performance in both MS-ASR and TS-ASR tasks, without the need for traditional methods, such as neural mask estimation or masking at the audio or feature level. Furthermore, we demonstrate a glimpse of a unified dual-task model which can efficiently handle both MS-ASR and TS-ASR tasks. Thus, this work illustrates that a robust end-to-end multi-talker ASR framework can be implemented with a streamlined architecture, obviating the need for the complex speaker filtering mechanisms employed in previous studies.
【11】 CLAIR-A: Leveraging Large Language Models to Judge Audio Captions
标题: CLAIR-A:利用大型语言模型来判断音频字幕
作者:Tsung-Han Wu,Joseph E. Gonzalez,Trevor Darrell,David M. Chan
备注:Code is publicly available at this https URL
链接:点击下载PDF文件
摘要:自动音频字幕(AAC)任务要求模型生成音频输入的自然语言描述。评估这些机器生成的音频字幕是一项复杂的任务,需要考虑各种因素,其中包括听觉场景理解,声音对象推理,时间连贯性和场景的环境背景。虽然目前的方法侧重于特定方面,但它们通常无法提供与人类判断一致的总体评分。在这项工作中,我们提出了CLAIR-A,一个简单而灵活的方法,利用大型语言模型(LLM)的zero-shot能力,直接要求LLM的语义距离分数来评估候选音频字幕。在我们的评估中,与传统指标相比,CLAIR-A更好地预测了人类对质量的判断,与特定领域的FENSE指标相比,相对准确度提高了5.8%,与Clotho-Eval数据集上的最佳通用指标相比,相对准确度提高了11%。此外,CLAIR-A通过允许语言模型解释其分数背后的推理提供了更高的透明度,这些解释比基线方法提供的解释高出30%。CLAIR-A在https: github.com DavidMChan clair-a上公开提供。摘要:The Automated Audio Captioning (AAC) task asks models to generate natural language descriptions of an audio input. Evaluating these machine-generated audio captions is a complex task that requires considering diverse factors, among them, auditory scene understanding, sound-object inference, temporal coherence, and the environmental context of the scene. While current methods focus on specific aspects, they often fail to provide an overall score that aligns well with human judgment. In this work, we propose CLAIR-A, a simple and flexible method that leverages the zero-shot capabilities of large language models (LLMs) to evaluate candidate audio captions by directly asking LLMs for a semantic distance score. In our evaluations, CLAIR-A better predicts human judgements of quality compared to traditional metrics, with a 5.8% relative accuracy improvement compared to the domain-specific FENSE metric and up to 11% over the best general-purpose measure on the Clotho-Eval dataset. Moreover, CLAIR-A offers more transparency by allowing the language model to explain the reasoning behind its scores, with these explanations rated up to 30% better by human evaluators than those provided by baseline methods. CLAIR-A is made publicly available at https: github.com DavidMChan clair-a.
【12】 Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space
标题: 增强语音命令的合成训练数据:从基于SVR的过滤到SSL潜在空间中的域适应
作者:Sebastião Quintas,Isabelle Ferrané,Thomas Pellegrini
链接:点击下载PDF文件
摘要:使用合成语音作为数据增强在自动语音识别和语音分类任务等领域越来越受欢迎。尽管具有语音克隆能力的新颖的文本到语音系统允许使用基于短音频段的更大量的语音,但是已知的是,这些系统倾向于产生幻觉并且经常产生很可能对下游任务具有负面影响的坏数据。在目前的工作中,我们进行了一组实验,围绕zero-shot学习与合成语音数据的语音命令分类的特定任务。我们在Google Speech Commands数据集上的结果表明,一个简单的基于ASR的过滤方法可以对生成的数据质量产生很大的影响,从而转化为更好的性能。此外,尽管生成的语音数据质量很好,但我们还表明,使用自监督(WavLM)特征时,合成语音和真实语音仍然可以轻松区分,这是CycleGAN进一步探索的一个方面,以弥合两种类型的语音材料之间的差距。摘要:The use of synthetic speech as data augmentation is gaining increasing popularity in fields such as automatic speech recognition and speech classification tasks. Despite novel text-to-speech systems with voice cloning capabilities, that allow the usage of a larger amount of voices based on short audio segments, it is known that these systems tend to hallucinate and oftentimes produce bad data that will most likely have a negative impact on the downstream task. In the present work, we conduct a set of experiments around zero-shot learning with synthetic speech data for the specific task of speech commands classification. Our results on the Google Speech Commands dataset show that a simple ASR-based filtering method can have a big impact in the quality of the generated data, translating to a better performance. Furthermore, despite the good quality of the generated speech data, we also show that synthetic and real speech can still be easily distinguishable when using self-supervised (WavLM) features, an aspect further explored with a CycleGAN to bridge the gap between the two types of speech material.
【13】 $text{M}^text{6}(text{GPT})^text{3}$: Generating Multitrack Modifiable Multi-Minute MIDI Music from Text using Genetic algorithms, Probabilistic methods and GPT Models in any Progression and Time signature
标题: $ ext{M}' ext{6}( ext{GPT}) ext{3}$:在任何进度和时间签名中使用遗传算法、概率方法和GPT模型从文本生成多轨可修改多分钟预设音乐
作者:Jakub Poćwiardowski,Mateusz Modrzejewski,Marek S. Tatara
备注:12 pages, 1 figure
链接:点击下载PDF文件
摘要:这项工作介绍了$ text{M}^ text{6}( text{GPT})^ text{3}$ Composer系统,能够生成完整的,多分钟的音乐作品,具有复杂的结构,在任何时间签名,在自然语言的输入描述的领域。该系统利用自回归Transformer语言模型将自然语言提示映射到JSON格式的组合参数。所定义的结构包括时间标记、音阶、和弦进行和效价唤醒值,从这些值创建伴奏、旋律、低音、主题和打击乐音轨。我们提出了一种遗传算法的旋律元素的生成。该算法结合了突变与音乐意义和适应度函数的基础上正态分布和预定义的音乐特征值。价值观自适应地演变,受情绪参数和不同的演奏风格的影响。用于生成任何时间的敲击签名的系统利用概率方法,包括马尔可夫链。通过人为和客观的评估,我们证明了我们的音乐生成方法在特定的、具有音乐意义的指标上优于基线,为纯粹基于神经网络的系统提供了一种有价值的替代方案。摘要:This work introduces the $ text{M}^ text{6}( text{GPT})^ text{3}$ Composer system, capable of generating complete, multi-minute musical compositions with complex structures in any time signature, in the MIDI domain from input descriptions in natural language. The system utilizes an autoregressive transformer language model to map natural language prompts to composition parameters in JSON format. The defined structure includes time signature, scales, chord progressions, and valence-arousal values, from which accompaniment, melody, bass, motif, and percussion tracks are created. We propose a genetic algorithm for the generation of melodic elements. The algorithm incorporates mutations with musical significance and a fitness function based on normal distribution and predefined musical feature values. The values adaptively evolve, influenced by emotional parameters and distinct playing styles. The system for generating percussion in any time signature utilises probabilistic methods, including Markov chains. Through both human and objective evaluations, we demonstrate that our music generation approach outperforms baselines on specific, musically meaningful metrics, offering a valuable alternative to purely neural network-based systems.
【14】 Exploring bat song syllable representations in self-supervised audio encoders
标题: 探索自我监督音频编码器中的蝙蝠歌曲音节表示
作者:Marianne de Heer Kloots,Mirjam Knörnschild
Journal-ref:Proceedings of the 4th International Workshop on Vocal Interactivity in-and-between Humans, Animals and Robots (VIHAR), Kos, GR, 6-9 Sept 2024
链接:点击下载PDF文件
摘要:在人类发出的声音上训练的深度学习模型能在多大程度上区分另一个物种的发声类型?我们分析了几种自监督音频编码器中蝙蝠歌音节的编码,发现在人类语音上预训练的模型生成了不同音节类型的最独特的表示。这些发现形成了跨物种迁移学习在蝙蝠生物声学中应用的第一步,以及对音频编码器模型中分布外信号处理的更好理解。摘要:How well can deep learning models trained on human-generated sounds distinguish between another species' vocalization types? We analyze the encoding of bat song syllables in several self-supervised audio encoders, and find that models pre-trained on human speech generate the most distinctive representations of different syllable types. These findings form first steps towards the application of cross-species transfer learning in bat bioacoustics, as well as an improved understanding of out-of-distribution signal processing in audio encoder models.
【15】 Hidden in Plain Sound: Environmental Backdoor Poisoning Attacks on Whisper, and Mitigations
标题: 隐藏在普通声音中:对Whisper的环境后门中毒攻击和缓解措施
作者:Jonatan Bartolini,Todor Stoyanov,Alberto Giaretta
备注:13 pages, 12 figures, 6 tables
链接:点击下载PDF文件
摘要:由于基于transformer的模型的普及,语音识别(SR)在各种应用领域中越来越受欢迎,例如充满关键任务设备的工业和机器人环境。虽然基于transformer的SR可以为简化人机接口提供各种好处,但对这些模型的网络安全方面的研究却乏善可陈。特别是关于后门中毒攻击。在本文中,我们提出了一种新的中毒方法,映射不同的环境触发声音的目标短语的不同长度,在微调阶段。我们在Whisper上测试了我们的方法,Whisper是最流行的基于transformer的SR模型之一,在几种测试条件下,它非常容易受到我们的攻击。为了减轻本文中提出的攻击,我们研究了使用Silero VAD,一种最先进的语音活动检测(VAD)模型,作为防御机制。我们的实验表明,可以使用VAD模型来过滤恶意触发器并减轻我们的攻击,并根据触发声音的类型和测试条件获得不同程度的成功。摘要:Thanks to the popularisation of transformer-based models, speech recognition (SR) is gaining traction in various application fields, such as industrial and robotics environments populated with mission-critical devices. While transformer-based SR can provide various benefits for simplifying human-machine interfacing, the research on the cybersecurity aspects of these models is lacklustre. In particular, concerning backdoor poisoning attacks. In this paper, we propose a new poisoning approach that maps different environmental trigger sounds to target phrases of different lengths, during the fine-tuning phase. We test our approach on Whisper, one of the most popular transformer-based SR model, showing that it is highly vulnerable to our attack, under several testing conditions. To mitigate the attack proposed in this paper, we investigate the use of Silero VAD, a state-of-the-art voice activity detection (VAD) model, as a defence mechanism. Our experiments show that it is possible to use VAD models to filter out malicious triggers and mitigate our attacks, with a varying degree of success, depending on the type of trigger sound and testing conditions.
【16】 FruitsMusic: A Real-World Corpus of Japanese Idol-Group Songs
标题: FruitsMusic:日本偶像组合歌曲的现实世界素材库
作者:Hitoshi Suda,Shunsuke Yoshida,Tomohiko Nakamura,Satoru Fukayama,Jun Ogata
备注:Accepted at the 25th International Society for Music Information Retrieval (ISMIR) Conference 2024, San Francisco, United States
链接:点击下载PDF文件
摘要:本研究提出了FruitsMusic,一个真实世界中日本偶像团体歌曲的元数据语料库,精确地注释了谁唱什么和什么时候。日本偶像团体歌曲,对日本流行文化至关重要,具有独特的声乐编排风格,歌曲分为几个部分,每个部分都有一个特定的个人或多个歌手。为了提高识别这种结构的歌手日记化方法,我们使用YouTube上的40个日本偶像团体的音乐视频构建了FruitsMusic作为资源。该语料库包括详细的注释,涵盖了各种流派,划分和分配风格的歌曲,以及4至9名成员的团体。FruitsMusic还促进了各种音乐信息检索技术的开发,例如歌词转录和歌手识别,不仅有利于日本偶像组合歌曲,而且还有利于各种文化的单个或多个歌手的歌曲。本文对FruitsMusic进行了全面的概述,包括其创作方法和与会话语音相比的独特特征。此外,本文还评估了当前使用FruitsMusic在具有挑战性的现实条件下进行歌手嵌入提取和日记化的方法的功效。此外,本文探讨了潜在的改进,通过评估人类的表现,在自动日记的性能。摘要:This study presents FruitsMusic, a metadata corpus of Japanese idol-group songs in the real world, precisely annotated with who sings what and when. Japanese idol-group songs, vital to Japanese pop culture, feature a unique vocal arrangement style, where songs are divided into several segments, and a specific individual or multiple singers are assigned to each segment. To enhance singer diarization methods for recognizing such structures, we constructed FruitsMusic as a resource using 40 music videos of Japanese idol groups from YouTube. The corpus includes detailed annotations, covering songs across various genres, division and assignment styles, and groups ranging from 4 to 9 members. FruitsMusic also facilitates the development of various music information retrieval techniques, such as lyrics transcription and singer identification, benefiting not only Japanese idol-group songs but also a wide range of songs featuring single or multiple singers from various cultures. This paper offers a comprehensive overview of FruitsMusic, including its creation methodology and unique characteristics compared to conversational speech. Additionally, this paper evaluates the efficacy of current methods for singer embedding extraction and diarization in challenging real-world conditions using FruitsMusic. Furthermore, this paper examines potential improvements in automatic diarization performance through evaluating human performance.
【17】 SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
标题: SoundBeam遇上M2D:利用音频基础模型提取目标声音
作者:Carlos Hernandez-Olivan,Marc Delcroix,Tsubasa Ochiai,Daisuke Niizumi,Naohiro Tawara,Tomohiro Nakatani,Shoko Araki
链接:点击下载PDF文件
摘要:目标声提取(TSE)是指利用线索从混合声中分离出所需的声音。TSE系统需要同时解决两个问题:识别目标声源和从混合声中提取目标信号。为了增加实用性,同一系统应与各种类型的声音一起工作。问题的双重性和各种各样的声音使得从头开始训练强大的TSE系统具有挑战性。在本文中,为了解决这个问题,我们探索使用预训练的音频基础模型,该模型可以在TSE系统中提供丰富的声音特征表示。我们选择了M2D基础模型,它似乎特别适合TSE任务,因为它是使用由声音标签预测和改进的掩蔽预测组成的双重目标进行训练的。这些目标涉及声音识别和TSE的信号提取问题。我们提出了一个新的TSE系统,它将M2D的特征表示集成到SoundBeam中,这是一个强大的TSE系统,可以利用目标声音类别标签和预先录制的注册(或音频查询)作为线索。我们的实验表明,使用M2D可以提高提取性能,尤其是在使用注册线索时。摘要:Target sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at once, identifying the target source and extracting the target signal from the mixture. For increased practicability, the same system should work with various types of sound. The duality of the problem and the wide variety of sounds make it challenging to train a powerful TSE system from scratch. In this paper, to tackle this problem, we explore using a pre-trained audio foundation model that can provide rich feature representations of sounds within a TSE system. We chose the masked-modeling duo (M2D) foundation model, which appears especially suited for the TSE task, as it is trained using a dual objective consisting of sound-label predictions and improved masked prediction. These objectives are related to sound identification and the signal extraction problems of TSE. We propose a new TSE system that integrates the feature representation from M2D into SoundBeam, which is a strong TSE system that can exploit both target sound class labels and pre-recorded enrollments (or audio queries) as clues. We show experimentally that using M2D can increase extraction performance, especially when employing enrollment clues.
【18】 ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning
标题: ViolinDiff:通过音调弯曲调节增强表达力的小提琴合成
作者:Daewoong Kim,Hao-Wen Dong,Dasaem Jeong
链接:点击下载PDF文件
摘要:基频(F0)的自然轮廓建模在音乐音频合成中起着至关重要的作用。然而,在复调音乐中转录和管理多个F0轮廓是具有挑战性的,并且尚未探索用于复调乐器合成的显式F0轮廓建模。在本文中,我们提出了ViolinDiff,一个两阶段的基于扩散的合成框架。对于一个给定的小提琴演奏文件,第一阶段估计的F0轮廓作为音高弯曲信息,和第二阶段生成梅尔声谱图结合这些表达的细节。量化指标和听力测试结果表明,该模型产生更真实的小提琴声音比模型没有明确的音高弯曲建模。音频样本可在线获取:daewoung.github.io ViolinDiff-Demo。摘要:Modeling the natural contour of fundamental frequency (F0) plays a critical role in music audio synthesis. However, transcribing and managing multiple F0 contours in polyphonic music is challenging, and explicit F0 contour modeling has not yet been explored for polyphonic instrumental synthesis. In this paper, we present ViolinDiff, a two-stage diffusion-based synthesis framework. For a given violin MIDI file, the first stage estimates the F0 contour as pitch bend information, and the second stage generates mel spectrogram incorporating these expressive details. The quantitative metrics and listening test results show that the proposed model generates more realistic violin sounds than the model without explicit pitch bend modeling. Audio samples are available online: daewoung.github.io ViolinDiff-Demo.
【19】 AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost
标题: AutoMode-ASB:学习选择ASB系统以获得更好的质量和成本
作者:Ahmet Gündüz,Yunsu Kim,Kamer Ali Yuksel,Mohamed Al-Badrashiny,Thiago Castro Ferreira,Hassan Sawaf
备注:SPECOM 2024 Conference
链接:点击下载PDF文件
摘要:我们提出了AutoMode-ASR,一种新的框架,有效地集成了多个ASR系统,以提高整体转录质量,同时优化成本。该想法是训练决策模型,以在运行系统之前仅基于音频输入为每个分段选择最佳ASR系统。我们实现这一目标,通过集成二进制分类器确定两个系统之间的偏好。这些分类器配备了各种功能,如音频嵌入,质量估计和信号属性。此外,我们演示了如何使用质量估计可以进一步提高性能,以最小的成本增加。实验结果表明,WER的相对减少16.2%,成本节省65%,和75%的速度提高,相比使用一个单一的最佳模型的所有部分。我们的框架与商业和开源黑盒ASR系统兼容,因为它不需要更改模型代码。摘要:We present AutoMode-ASR, a novel framework that effectively integrates multiple ASR systems to enhance the overall transcription quality while optimizing cost. The idea is to train a decision model to select the optimal ASR system for each segment based solely on the audio input before running the systems. We achieve this by ensembling binary classifiers determining the preference between two systems. These classifiers are equipped with various features, such as audio embeddings, quality estimation, and signal properties. Additionally, we demonstrate how using a quality estimator can further improve performance with minimal cost increase. Experimental results show a relative reduction in WER of 16.2%, a cost saving of 65%, and a speed improvement of 75%, compared to using a single-best model for all segments. Our framework is compatible with commercial and open-source black-box ASR systems as it does not require changes in model codes.
【20】 AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework
标题: AudioEditor:免训练的基于扩散的音频编辑框架
作者:Yuhang Jia,Yang Chen,Jinghua Zhao,Shiwan Zhao,Wenjia Zeng,Yong Chen,Yong Qin
链接:点击下载PDF文件
摘要:基于扩散的文本到音频(TTA)生成已经取得了实质性的进展,利用潜在扩散模型(LDM)生成高质量,多样化和与预防相关的音频。然而,除了生成,音频编辑的任务仍然同样重要,但受到的关注相对较少。音频编辑任务面临两个主要挑战:执行精确编辑和保留未编辑的部分。虽然基于LDM的工作流已经有效地解决了图像处理领域中的这些挑战,但类似的方法几乎没有应用于音频编辑。在本文中,我们介绍了AudioEditor,这是一个基于预训练的基于扩散的TTA模型的免训练音频编辑框架。AudioEditor集成了空文本反转和EOT抑制方法,使模型能够在执行准确编辑的同时保留原始音频功能。全面的客观和主观实验验证了AudioEditor在提供高质量音频编辑方面的有效性。代码和演示可以在https: github.com NKU-HLT AudioEditor上找到。摘要:Diffusion-based text-to-audio (TTA) generation has made substantial progress, leveraging latent diffusion model (LDM) to produce high-quality, diverse and instruction-relevant audios. However, beyond generation, the task of audio editing remains equally important but has received comparatively little attention. Audio editing tasks face two primary challenges: executing precise edits and preserving the unedited sections. While workflows based on LDMs have effectively addressed these challenges in the field of image processing, similar approaches have been scarcely applied to audio editing. In this paper, we introduce AudioEditor, a training-free audio editing framework built on the pretrained diffusion-based TTA model. AudioEditor incorporates Null-text Inversion and EOT-suppression methods, enabling the model to preserve original audio features while executing accurate edits. Comprehensive objective and subjective experiments validate the effectiveness of AudioEditor in delivering high-quality audio edits. Code and demo can be found at https: github.com NKU-HLT AudioEditor.
【21】 A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
标题: 具有空间线索保留的轻量级实时双耳语音增强模型
作者:Jingyuan Wang,Jie Zhang,Shihao Chen,Miao Sun
链接:点击下载PDF文件
摘要:双耳语音增强(BSE)的目的是联合提高听力设备接收到的含噪信号的语音质量和可懂度,并保留目标的空间线索以用于自然收听。现有的方法往往受到噪声抑制(NR)的能力和空间线索保存(SCP)的准确性和复杂的声学场景中的高计算需求之间的折衷。在这项工作中,我们提出了一个基于学习的轻量级双耳复卷积网络(LBCCN),它通过过滤低频带并保留其余部分而在NR中表现出色。此外,我们的方法显式地结合了通道间相对声学传递函数的估计,以确保空间线索的保真度和语音清晰度。实验结果表明,在各种噪声条件下,该方法都能获得与现有方法相当的NR性能,但计算量更低,SCP更好。可复制的代码和音频示例可在https: github.com jywanng LBCCN上找到。摘要:Binaural speech enhancement (BSE) aims to jointly improve the speech quality and intelligibility of noisy signals received by hearing devices and preserve the spatial cues of the target for natural listening. Existing methods often suffer from the compromise between noise reduction (NR) capacity and spatial cues preservation (SCP) accuracy and a high computational demand in complex acoustic scenes. In this work, we present a learning-based lightweight binaural complex convolutional network (LBCCN), which excels in NR by filtering low-frequency bands and keeping the rest. Additionally, our approach explicitly incorporates the estimation of interchannel relative acoustic transfer function to ensure the spatial cues fidelity and speech clarity. Results show that the proposed LBCCN can achieve a comparable NR performance to state-of-the-art methods under various noise conditions, but with a much lower computational cost and a better SCP. The reproducible code and audio examples are available at https: github.com jywanng LBCCN.
【22】 Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
标题: 用于鲁棒语音识别的队列感知域自适应生成对抗网络
作者:Chien-Chun Wang,Li-Wei Chen,Cheng-Kang Chou,Hung-Shin Lee,Berlin Chen,Hsin-Min Wang
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:虽然预先训练的自动语音识别(ASR)系统在匹配域上表现出令人印象深刻的性能,但当遇到来自不可见的记录环境和条件的通道失配时,它们的性能往往会下降。为了解决这个问题,我们提出了一种新的信道感知数据模拟方法,用于鲁棒的ASR训练。我们的方法利用了通道提取技术和生成对抗网络(GANs)的协同能力。我们首先训练一个能够从任意音频中提取嵌入的通道编码器。最重要的是,使用最少量的目标域数据提取通道嵌入,并用于指导基于GAN的语音合成器。该合成器生成的语音忠实地保留了输入的语音内容,同时模仿目标域的通道特性。我们评估我们的方法具有挑战性的客家横跨台湾(HAT)和台湾横跨台湾(TAT)语料库,实现相对字符错误率(CER)减少20.02%和9.64%,分别与基线相比。这些结果突出了我们的通道感知数据模拟方法的有效性,用于弥合源域和目标域声学之间的差距。摘要:While pre-trained automatic speech recognition (ASR) systems demonstrate impressive performance on matched domains, their performance often degrades when confronted with channel mismatch stemming from unseen recording environments and conditions. To mitigate this issue, we propose a novel channel-aware data simulation method for robust ASR training. Our method harnesses the synergistic power of channel-extractive techniques and generative adversarial networks (GANs). We first train a channel encoder capable of extracting embeddings from arbitrary audio. On top of this, channel embeddings are extracted using a minimal amount of target-domain data and used to guide a GAN-based speech synthesizer. This synthesizer generates speech that faithfully preserves the phonetic content of the input while mimicking the channel characteristics of the target domain. We evaluate our method on the challenging Hakka Across Taiwan (HAT) and Taiwanese Across Taiwan (TAT) corpora, achieving relative character error rate (CER) reductions of 20.02% and 9.64%, respectively, compared to the baselines. These results highlight the efficacy of our channel-aware data simulation method for bridging the gap between source- and target-domain acoustics.
【23】 Simultaneous Music Separation and Generation Using Multi-Track Latent Diffusion Models
标题: 使用多轨潜在扩散模型同时音乐分离和生成
作者:Tornike Karchkhadze,Mohammad Rasool Izadi,Shlomo Dubnov
链接:点击下载PDF文件
摘要:扩散模型最近在音乐生成和音乐源分离任务中显示出强大的潜力。虽然在早期阶段,一种趋势正在出现,将这些任务整合到一个单一的框架中,因为两者都涉及生成音乐对齐的部分,并且可以被视为同一生成过程的方面。在这项工作中,我们引入了一个潜在的基于扩散的多轨道生成模型,能够通过学习共享音乐上下文的轨道的联合概率分布的源分离和多轨道音乐合成。我们的模型还可以通过创建给定其他轨道的任何子集来生成安排。我们在Slakh2100数据集上训练了我们的模型,将其与现有的同步生成和分离模型进行了比较,并观察到源分离,音乐和安排生成任务的客观指标的显着改进。在https: msg-ld.github.io 上可以找到正确的例子。摘要:Diffusion models have recently shown strong potential in both music generation and music source separation tasks. Although in early stages, a trend is emerging towards integrating these tasks into a single framework, as both involve generating musically aligned parts and can be seen as facets of the same generative process. In this work, we introduce a latent diffusion-based multi-track generation model capable of both source separation and multi-track music synthesis by learning the joint probability distribution of tracks sharing a musical context. Our model also enables arrangement generation by creating any subset of tracks given the others. We trained our model on the Slakh2100 dataset, compared it with an existing simultaneous generation and separation model, and observed significant improvements across objective metrics for source separation, music, and arrangement generation tasks. Sound examples are available at https: msg-ld.github.io .
【24】 Large Language Models Are Strong Audio-Visual Speech Recognition Learners
标题: 大型语言模型是强大的视听语音识别学习者
作者:Umberto Cappellazzo,Minsu Kim,Honglie Chen,Pingchuan Ma,Stavros Petridis,Daniele Falavigna,Alessio Brutti,Maja Pantic
备注:The code will be made available at this link: this https URL
链接:点击下载PDF文件
摘要:多模态大型语言模型(MLLM)由于其强大的多模态理解能力而成为近年来的研究热点。例如,在音频和语音域中,LLM可以通过仅连接用音频编码器计算的音频令牌和文本令牌来配备(自动)语音识别(ASR)能力,以实现最先进的结果。相反,像视觉和视听语音识别(VSR AVSR)这样的任务,也利用了噪声不变的嘴唇运动信息,很少或根本没有受到关注。为了弥补这一差距,我们提出了Llama-AVSR,一个新的MLLM具有强大的视听语音识别能力。它利用预先训练的音频和视频编码器来产生特定于模态的令牌,这些令牌与文本令牌一起由预先训练的LLM(例如,图3.1 -8B)以自回归方式产生所得响应。Llama-AVSR需要少量的可训练参数,因为只有特定模态的投影仪和LoRA模块被训练,而多模态编码器和LLM保持冻结。我们评估我们提出的方法LRS 3,最大的公共AVSR基准,我们实现了新的国家的最先进的结果,ASR和AVSR的任务,WER分别为0.81%和0.77%。为了支持我们的结果,我们研究了支撑Llama-AVSR有效性的关键因素:预训练编码器和LLM的选择,LoRA模块的有效集成,以及通过模态感知压缩率获得的最佳性能效率权衡。摘要:Multimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and speech domains, an LLM can be equipped with (automatic) speech recognition (ASR) abilities by just concatenating the audio tokens, computed with an audio encoder, and the text tokens to achieve state-of-the-art results. On the contrary, tasks like visual and audio-visual speech recognition (VSR AVSR), which also exploit noise-invariant lip movement information, have received little or no attention. To bridge this gap, we propose Llama-AVSR, a new MLLM with strong audio-visual speech recognition capabilities. It leverages pre-trained audio and video encoders to produce modality-specific tokens which, together with the text tokens, are processed by a pre-trained LLM (e.g., Llama3.1-8B) to yield the resulting response in an auto-regressive fashion. Llama-AVSR requires a small number of trainable parameters as only modality-specific projectors and LoRA modules are trained whereas the multi-modal encoders and LLM are kept frozen. We evaluate our proposed approach on LRS3, the largest public AVSR benchmark, and we achieve new state-of-the-art results for the tasks of ASR and AVSR with a WER of 0.81% and 0.77%, respectively. To bolster our results, we investigate the key factors that underpin the effectiveness of Llama-AVSR: the choice of the pre-trained encoders and LLM, the efficient integration of LoRA modules, and the optimal performance-efficiency trade-off obtained via modality-aware compression rates.
【25】 Measuring Sound Symbolism in Audio-visual Models
标题: 测量视听模型中的声音象征性
作者:Wei-Cheng Tseng,Yi-Jen Shih,David Harwath,Raymond Mooney
备注:SLT 2024
链接:点击下载PDF文件
摘要:视听预训练模型最近得到了广泛的关注,并在各种视听任务上表现出卓越的性能。这项研究调查了预先训练的视听模型是否表现出声音和视觉表示之间的非任意关联,称为声音符号学,这也在人类中观察到。我们开发了一个专门的数据集与合成的图像和音频样本,并评估这些模型使用非参数的方法在zero-shot设置。我们的研究结果揭示了模型的输出和建立的声音符号模式之间的显着相关性,特别是在语音数据训练的模型中。这些结果表明,这些模型可以捕获类似于人类语言处理的声音-意义联系,为认知架构和机器学习策略提供见解。摘要:Audio-visual pre-trained models have gained substantial attention recently and demonstrated superior performance on various audio-visual tasks. This study investigates whether pre-trained audio-visual models demonstrate non-arbitrary associations between sounds and visual representations$ unicode{x2013}$known as sound symbolism$ unicode{x2013}$which is also observed in humans. We developed a specialized dataset with synthesized images and audio samples and assessed these models using a non-parametric approach in a zero-shot setting. Our findings reveal a significant correlation between the models' outputs and established patterns of sound symbolism, particularly in models trained on speech data. These results suggest that such models can capture sound-meaning connections akin to human language processing, providing insights into both cognitive architectures and machine learning strategies.
机器翻译,仅供参考
