本文经arXiv每日学术速递授权转载
【1】 Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
标题: 利用基于LID的协作混合专家模型增强代码转换语音识别
作者:Hukai Huang,Jiayan Lin,Kaidi Wang,Yishuang Li,Wenhao Guan,Qingyang Hong,Lin Li
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
【2】 Activity-Guided Industrial Anomalous Sound Detection against Interferences
标题: 活动引导的工业异常声音抗干扰检测
作者:Yunjoo Lee,Jaechang Kim,Jungseul Ok
备注:Thsis is an extended version of this https URL
链接:点击下载PDF文件
【3】 The Role of Large Language Models in Musicology: Are We Ready to Trust the Machines?
标题: 大型语言模型在音乐学中的作用:我们准备好信任机器了吗?
作者:Pedro Ramoneda,Emilia Parada-Cabaleiro,Benno Weck,Xavier Serra
链接:点击下载PDF文件
【4】 USTC-KXDIGIT System Description for ASVspoof5 Challenge
标题: USTC-KXDIGIT ASVspoof 5挑战赛系统描述
作者:Yihao Chen,Haochen Wu,Nan Jiang,Xiang Xia,Qing Gu,Yunqi Hao,Pengfei Cai,Yu Guan,Jialong Wang,Weilin Xie,Lei Fang,Sian Fang,Yan Song,Wu Guo,Lin Liu,Minqiang Xu
备注:ASVspoof5 workshop paper
链接:点击下载PDF文件
【5】 Pureformer-VC: Non-parallel One-Shot Voice Conversion with Pure Transformer Blocks and Triplet Discriminative Training
标题: Pureformer-VC:使用纯Transformer块和三重区分训练的非并行单次语音转换
作者:Wenhan Yao,Zedong Xing,Xiarun Chen,Jia Liu,Yongqiang He,Weiping Wen
备注:submmited to ICASSP 2025
链接:点击下载PDF文件
【6】 VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka
标题: VoxHakka:台湾客语方言多样化的多说话人文本到语音系统
作者:Li-Wei Chen,Hung-Shin Lee,Chen-Chi Chang
备注:Submitted to O-COCOSDA 2024
链接:点击下载PDF文件
【7】 Effective Noise-aware Data Simulation for Domain-adaptive Speech Enhancement Leveraging Dynamic Stochastic Perturbation
标题: 利用动态随机扰动的域自适应语音增强的有效噪音感知数据模拟
作者:Chien-Chun Wang,Li-Wei Chen,Hung-Shin Lee,Berlin Chen,Hsin-Min Wang
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
【8】 Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement
标题: Spectron:使用具有对抗细化的条件Transformer提取目标说话人
作者:Tathagata Bandyopadhyay
链接:点击下载PDF文件
【9】 A multilingual training strategy for low resource Text to Speech
标题: 低资源文本到语音的多语言训练策略
作者:Asma Amalas,Mounir Ghogho,Mohamed Chetouani,Rachid Oulad Haj Thami
备注:12 pages, 2 figures
链接:点击下载PDF文件
【10】 Interpretable Convolutional SyncNet
标题: 可解释卷积同步网络
作者:Sungjoon Park,Jaesub Yun,Donggeon Lee,Minsik Park
备注:8+5 pages
链接:点击下载PDF文件
【11】 A Framework for Synthetic Audio Conversations Generation using Large Language Models
标题: 使用大型语言模型生成合成音频对话的框架
作者:Kaung Myat Kyaw,Jonathan Hoyin Chan
备注:This work has been submitted for consideration at the WI-IAT'24 to be held in December 2024
链接:点击下载PDF文件
【12】 SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
标题: SoCodec:一种语义有序的多流语音编解码器,用于基于高效语言模型的文本到语音合成
作者:Haohan Guo,Fenglong Xie,Kun Xie,Dongchao Yang,Dake Guo,Xixin Wu,Helen Meng
备注:Accepted by SLT 2024
链接:点击下载PDF文件
【13】 MMT-BERT: Chord-aware Symbolic Music Generation Based on Multitrack Music Transformer and MusicBERT
标题: MMT-BERT:基于多轨音乐Transformer和MusicBERT的和弦感知符号音乐生成
作者:Jinlong Zhu,Keigo Sakurai,Ren Togo,Takahiro Ogawa,Miki Haseyama
备注:Accepted to the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
【14】 Dissecting Temporal Understanding in Text-to-Audio Retrieval
标题: 文本到音频检索中的时态理解剖析
作者:Andreea-Maria Oncescu,João F. Henriques,A. Sophia Koepke
备注:9 pages, 5 figures, ACM Multimedia 2024, this https URL
链接:点击下载PDF文件
【15】 LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
标题: LibriheavightMix:一个20,000小时的数据集,用于单通道回响多说话者语音分离、ASB和说话者拨号
作者:Zengrui Jin,Yifan Yang,Mohan Shi,Wei Kang,Xiaoyu Yang,Zengwei Yao,Fangjun Kuang,Liyong Guo,Lingwei Meng,Long Lin,Yong Xu,Shi-Xiong Zhang,Daniel Povey
备注:InterSpeech 2024
链接:点击下载PDF文件
【16】 Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
标题: 用于多说话人自动语音识别的重叠编码分离序列化语音信息引导
作者:Hao Shi,Yuan Gao,Zhaoheng Ni,Tatsuya Kawahara
链接:点击下载PDF文件
【17】 MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
标题: MaskGCT:带屏蔽生成编解码器的Zero-Shot文本到语音
作者:Yuancheng Wang,Haoyue Zhan,Liwei Liu,Ruihong Zeng,Haotian Guo,Jiachen Zheng,Qiang Zhang,Shunsi Zhang,Zhizheng Wu
链接:点击下载PDF文件
【18】 Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion
标题: 看到你的言语风格:一种新颖的Zero-Shot身份解开基于面部的语音转换
作者:Yan Rong,Li Liu
链接:点击下载PDF文件
【19】 FLUX that Plays Music
标题: 播放音乐的FLOX
作者:Zhengcong Fei,Mingyuan Fan,Changqian Yu,Junshi Huang
链接:点击下载PDF文件
【20】 Multi-scale Multi-instance Visual Sound Localization and Segmentation
标题: 多尺度多实例视觉声音定位与分割
作者:Shentong Mo,Haofan Wang
链接:点击下载PDF文件
【21】 Multi-label Zero-Shot Audio Classification with Temporal Attention
标题: 具有时间注意力的多标签Zero-Shot音频分类
作者:Duygu Dogan,Huang Xie,Toni Heittola,Tuomas Virtanen
备注:Accepted to International Workshop on Acoustic Signal Enhancement (IWAENC) 2024
链接:点击下载PDF文件
【22】 Density Adaptive Attention-based Speech Network: Enhancing Feature Understanding for Mental Health Disorders
标题: 密度自适应基于注意力的语音网络:增强对心理健康障碍的特征理解
作者:Georgios Ioannides,Adrian Kieback,Aman Chadha,Aaron Elkins
链接:点击下载PDF文件
【23】 Contrastive Augmentation: An Unsupervised Learning Approach for Keyword Spotting in Speech Technology
标题: 对比增强:语音技术中关键词发现的无监督学习方法
作者:Weinan Dai,Yifeng Jiang,Yuanjing Liu,Jinkun Chen,Xin Sun,Jinglei Tao
备注:This paper has been accepted by the ICPR2024
链接:点击下载PDF文件
【24】 REFFLY: Melody-Constrained Lyrics Editing Model
标题: REFFLY:旋律约束歌词编辑模型
作者:Songyan Zhao,Bingxuan Li,Yufei Tian,Nanyun Peng
链接:点击下载PDF文件
【25】 Towards a dynamical model of English vowels. Evidence from diphthongisation
标题: 迈向英语元音的动态模型。双元音化的证据
作者:Patrycja Strycharczuk,Sam Kirkham,Emily Gorman,Takayuki Nagamine
链接:点击下载PDF文件
【26】 ProGRes: Prompted Generative Rescoring on ASR n-Best
标题: ProGRes:在ASB n-Best上进行的预定生成重新评分
作者:Ada Defne Tur,Adel Moumen,Mirco Ravanelli
备注:IEEE Spoken Language Technology Workshop
链接:点击下载PDF文件
【27】 Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning
标题: 使用谱-时态图专注池和多任务学习的逐例查询关键词发现
作者:Zhenyu Wang,Shuyu Kong,Li Wan,Biqiao Zhang,Yiteng Huang,Mumin Jin,Ming Sun,Xin Lei,Zhaojun Yang
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
【28】 The USTC-NERCSLIP Systems for the CHiME-8 NOTSOFAR-1 Challenge
标题: CHiME-8诺丁索FAR-1挑战赛的USTC-NERCSLIP系统
作者:Shutong Niu,Ruoyu Wang,Jun Du,Gaobin Yang,Yanhui Tu,Siyuan Wu,Shuangqing Qian,Huaxin Wu,Haitao Xu,Xueyang Zhang,Guolong Zhong,Xindi Yu,Jieru Chen,Mengzhi Wang,Di Cai,Tian Gao,Genshun Wan,Feng Ma,Jia Pan,Jianqing Gao
链接:点击下载PDF文件
【29】 vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
标题: vec 2wav 2.0:通过离散令牌声码器推进语音转换
作者:Yiwei Guo,Zhihan Li,Junjie Li,Chenpeng Du,Hankun Wang,Shuai Wang,Xie Chen,Kai Yu
备注:5 pages, 4 figures
链接:点击下载PDF文件
【30】 Reassessing Noise Augmentation Methods in the Context of Adversarial Speech
标题: 对抗性言语背景下重新评估噪音增强方法
作者:Karla Pizzi,Matías P. Pizarro B,Asja Fischer
链接:点击下载PDF文件
【31】 Steered Response Power-Based Direction-of-Arrival Estimation Exploiting an Auxiliary Microphone
标题: 利用辅助麦克风的定向响应基于功率的到达方向估计
作者:Klaus Brümann,Simon Doclo
备注:5 pages, 3 figures, conference: EUSIPCO 2024 in Lyon
链接:点击下载PDF文件
【32】 Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
标题: 多说话者ASB的语音基础模型的资源高效自适应
作者:Weiqing Wang,Kunal Dhawan,Taejin Park,Krishna C. Puvvada,Ivan Medennikov,Somshubra Majumdar,He Huang,Jagadeesh Balam,Boris Ginsburg
备注:Accepted by SLT 2024
链接:点击下载PDF文件
【33】 Suppressing Noise Disparity in Training Data for Automatic Pathological Speech Detection
标题: 抑制训练数据中的噪音差异以实现自动病理语音检测
作者:Mahdi Amiri,Ina Kodrasi
备注:To appear in IWAENC 2024
链接:点击下载PDF文件
【34】 EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance
标题: EnCLAP++:分析EnCLAP框架以优化自动音频字幕性能
作者:Jaeyeon Kim,Minjeon Jeon,Jaeyoon Jung,Sang Hoon Woo,Jinjoo Lee
备注:Accepted to DCASE2024 Workshop
链接:点击下载PDF文件
【35】 Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
标题: 通过自动音频字幕的辅助检索模型扩展EnCLAP
作者:Jaeyeon Kim,Jaeyoon Jung,Minjeong Jeon,Sang Hoon Woo,Jinjoo Lee
备注:DCASE2024 Challenge Technical Report. Ranked 2nd in Task 6 Automated Audio Captioning
链接:点击下载PDF文件
【36】 BUET Multi-disease Heart Sound Dataset: A Comprehensive Auscultation Dataset for Developing Computer-Aided Diagnostic Systems
标题: BUET多疾病心脏声音数据集:用于开发计算机辅助诊断系统的全面听诊数据集
作者:Shams Nafisa Ali,Afia Zahin,Samiul Based Shuvo,Nusrat Binta Nizam,Shoyad Ibn Sabur Khan Nuhash,Sayeed Sajjad Razin,S. M. Sakeef Sani,Farihin Rahman,Nawshad Binta Nizam,Farhat Binte Azam,Rakib Hossen,Sumaiya Ohab,Nawsabah Noor,Taufiq Hasan
备注:14 pages, 13 figures
链接:点击下载PDF文件
【37】 Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
标题: 视听人员识别与验证的情态融合方法比较分析
作者:Aref Farhadipour,Masoumeh Chapariniya,Teodora Vukovic,Volker Dellwo
备注:This paper has been submitted to a conference
链接:点击下载PDF文件
【38】 Digit Recognition using Multimodal Spiking Neural Networks
标题: 使用多峰尖峰神经网络的数字识别
作者:William Bjorndahl,Jack Easton,Austin Modoff,Eric C. Larson,Joseph Camp,Prasanna Rangarajan
备注:4 pages, 2 figures, submitted to 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing
链接:点击下载PDF文件
【39】 DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
标题: DCIM-AVSR:通过双适形器交互模块实现高效的视听语音识别
作者:Xinyu Wang,Qian Wang
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【40】 Progressive Residual Extraction based Pre-training for Speech Representation Learning
标题: 基于渐进剩余提取的预训练的语音表示学习
作者:Tianrui Wang,Jin Li,Ziyang Ma,Rui Cao,Xie Chen,Longbiao Wang,Meng Ge,Xiaobao Wang,Yuguang Wang,Jianwu Dang,Nyima Tashi
链接:点击下载PDF文件
标题: CHiME-8诺丁索FAR-1挑战赛的USTC-NERCSLIP系统
作者:Shutong Niu,Ruoyu Wang,Jun Du,Gaobin Yang,Yanhui Tu,Siyuan Wu,Shuangqing Qian,Huaxin Wu,Haitao Xu,Xueyang Zhang,Guolong Zhong,Xindi Yu,Jieru Chen,Mengzhi Wang,Di Cai,Tian Gao,Genshun Wan,Feng Ma,Jia Pan,Jianqing Gao
链接:点击下载PDF文件
【2】 vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
标题: vec 2wav 2.0:通过离散令牌声码器推进语音转换
作者:Yiwei Guo,Zhihan Li,Junjie Li,Chenpeng Du,Hankun Wang,Shuai Wang,Xie Chen,Kai Yu
备注:5 pages, 4 figures
链接:点击下载PDF文件
【3】 Reassessing Noise Augmentation Methods in the Context of Adversarial Speech
标题: 对抗性言语背景下重新评估噪音增强方法
作者:Karla Pizzi,Matías P. Pizarro B,Asja Fischer
链接:点击下载PDF文件
【4】 Steered Response Power-Based Direction-of-Arrival Estimation Exploiting an Auxiliary Microphone
标题: 利用辅助麦克风的定向响应基于功率的到达方向估计
作者:Klaus Brümann,Simon Doclo
备注:5 pages, 3 figures, conference: EUSIPCO 2024 in Lyon
链接:点击下载PDF文件
【5】 Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
标题: 多说话者ASB的语音基础模型的资源高效自适应
作者:Weiqing Wang,Kunal Dhawan,Taejin Park,Krishna C. Puvvada,Ivan Medennikov,Somshubra Majumdar,He Huang,Jagadeesh Balam,Boris Ginsburg
备注:Accepted by SLT 2024
链接:点击下载PDF文件
【6】 Suppressing Noise Disparity in Training Data for Automatic Pathological Speech Detection
标题: 抑制训练数据中的噪音差异以实现自动病理语音检测
作者:Mahdi Amiri,Ina Kodrasi
备注:To appear in IWAENC 2024
链接:点击下载PDF文件
【7】 EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance
标题: EnCLAP++:分析EnCLAP框架以优化自动音频字幕性能
作者:Jaeyeon Kim,Minjeon Jeon,Jaeyoon Jung,Sang Hoon Woo,Jinjoo Lee
备注:Accepted to DCASE2024 Workshop
链接:点击下载PDF文件
【8】 Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
标题: 通过自动音频字幕的辅助检索模型扩展EnCLAP
作者:Jaeyeon Kim,Jaeyoon Jung,Minjeong Jeon,Sang Hoon Woo,Jinjoo Lee
备注:DCASE2024 Challenge Technical Report. Ranked 2nd in Task 6 Automated Audio Captioning
链接:点击下载PDF文件
【9】 BUET Multi-disease Heart Sound Dataset: A Comprehensive Auscultation Dataset for Developing Computer-Aided Diagnostic Systems
标题: BUET多疾病心脏声音数据集:用于开发计算机辅助诊断系统的全面听诊数据集
作者:Shams Nafisa Ali,Afia Zahin,Samiul Based Shuvo,Nusrat Binta Nizam,Shoyad Ibn Sabur Khan Nuhash,Sayeed Sajjad Razin,S. M. Sakeef Sani,Farihin Rahman,Nawshad Binta Nizam,Farhat Binte Azam,Rakib Hossen,Sumaiya Ohab,Nawsabah Noor,Taufiq Hasan
备注:14 pages, 13 figures
链接:点击下载PDF文件
【10】 Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
标题: 视听人员识别与验证的情态融合方法比较分析
作者:Aref Farhadipour,Masoumeh Chapariniya,Teodora Vukovic,Volker Dellwo
备注:This paper has been submitted to a conference
链接:点击下载PDF文件
【11】 Digit Recognition using Multimodal Spiking Neural Networks
标题: 使用多峰尖峰神经网络的数字识别
作者:William Bjorndahl,Jack Easton,Austin Modoff,Eric C. Larson,Joseph Camp,Prasanna Rangarajan
备注:4 pages, 2 figures, submitted to 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing
链接:点击下载PDF文件
【12】 DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
标题: DCIM-AVSR:通过双适形器交互模块实现高效的视听语音识别
作者:Xinyu Wang,Qian Wang
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【13】 Progressive Residual Extraction based Pre-training for Speech Representation Learning
标题: 基于渐进剩余提取的预训练的语音表示学习
作者:Tianrui Wang,Jin Li,Ziyang Ma,Rui Cao,Xie Chen,Longbiao Wang,Meng Ge,Xiaobao Wang,Yuguang Wang,Jianwu Dang,Nyima Tashi
链接:点击下载PDF文件
【14】 BELT-2: Bootstrapping EEG-to-Language representation alignment for multi-task brain decoding
标题: BELT-2:引导脑电与语言表示对齐,用于多任务大脑解码
作者:Jinzhao Zhou,Yiqun Duan,Fred Chang,Thomas Do,Yu-Kai Wang,Chin-Teng Lin
链接:点击下载PDF文件
【15】 Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
标题: 利用基于LID的协作混合专家模型增强代码转换语音识别
作者:Hukai Huang,Jiayan Lin,Kaidi Wang,Yishuang Li,Wenhao Guan,Qingyang Hong,Lin Li
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
【16】 Activity-Guided Industrial Anomalous Sound Detection against Interferences
标题: 活动引导的工业异常声音抗干扰检测
作者:Yunjoo Lee,Jaechang Kim,Jungseul Ok
备注:Thsis is an extended version of this https URL
链接:点击下载PDF文件
【17】 The Role of Large Language Models in Musicology: Are We Ready to Trust the Machines?
标题: 大型语言模型在音乐学中的作用:我们准备好信任机器了吗?
作者:Pedro Ramoneda,Emilia Parada-Cabaleiro,Benno Weck,Xavier Serra
链接:点击下载PDF文件
【18】 USTC-KXDIGIT System Description for ASVspoof5 Challenge
标题: USTC-KXDIGIT ASVspoof 5挑战赛系统描述
作者:Yihao Chen,Haochen Wu,Nan Jiang,Xiang Xia,Qing Gu,Yunqi Hao,Pengfei Cai,Yu Guan,Jialong Wang,Weilin Xie,Lei Fang,Sian Fang,Yan Song,Wu Guo,Lin Liu,Minqiang Xu
备注:ASVspoof5 workshop paper
链接:点击下载PDF文件
【19】 Pureformer-VC: Non-parallel One-Shot Voice Conversion with Pure Transformer Blocks and Triplet Discriminative Training
标题: Pureformer-VC:使用纯Transformer块和三重区分训练的非并行单次语音转换
作者:Wenhan Yao,Zedong Xing,Xiarun Chen,Jia Liu,Yongqiang He,Weiping Wen
备注:submmited to ICASSP 2025
链接:点击下载PDF文件
【20】 VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka
标题: VoxHakka:台湾客语方言多样化的多说话人文本到语音系统
作者:Li-Wei Chen,Hung-Shin Lee,Chen-Chi Chang
备注:Submitted to O-COCOSDA 2024
链接:点击下载PDF文件
【21】 Effective Noise-aware Data Simulation for Domain-adaptive Speech Enhancement Leveraging Dynamic Stochastic Perturbation
标题: 利用动态随机扰动的域自适应语音增强的有效噪音感知数据模拟
作者:Chien-Chun Wang,Li-Wei Chen,Hung-Shin Lee,Berlin Chen,Hsin-Min Wang
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
【22】 Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement
标题: Spectron:使用具有对抗细化的条件Transformer提取目标说话人
作者:Tathagata Bandyopadhyay
链接:点击下载PDF文件
【23】 A multilingual training strategy for low resource Text to Speech
标题: 低资源文本到语音的多语言训练策略
作者:Asma Amalas,Mounir Ghogho,Mohamed Chetouani,Rachid Oulad Haj Thami
备注:12 pages, 2 figures
链接:点击下载PDF文件
【24】 Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language
标题: 个性化唇读:用视觉和语言适应独特的唇动
作者:Jeong Hun Yeo,Chae Won Kim,Hyunjun Kim,Hyeongseop Rha,Seunghee Han,Wen-Huang Cheng,Yong Man Ro
备注:Code available: this https URL
链接:点击下载PDF文件
【25】 Interpretable Convolutional SyncNet
标题: 可解释卷积同步网络
作者:Sungjoon Park,Jaesub Yun,Donggeon Lee,Minsik Park
备注:8+5 pages
链接:点击下载PDF文件
【26】 A Framework for Synthetic Audio Conversations Generation using Large Language Models
标题: 使用大型语言模型生成合成音频对话的框架
作者:Kaung Myat Kyaw,Jonathan Hoyin Chan
备注:This work has been submitted for consideration at the WI-IAT'24 to be held in December 2024
链接:点击下载PDF文件
【27】 SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
标题: SoCodec:一种语义有序的多流语音编解码器,用于基于高效语言模型的文本到语音合成
作者:Haohan Guo,Fenglong Xie,Kun Xie,Dongchao Yang,Dake Guo,Xixin Wu,Helen Meng
备注:Accepted by SLT 2024
链接:点击下载PDF文件
【28】 MMT-BERT: Chord-aware Symbolic Music Generation Based on Multitrack Music Transformer and MusicBERT
标题: MMT-BERT:基于多轨音乐Transformer和MusicBERT的和弦感知符号音乐生成
作者:Jinlong Zhu,Keigo Sakurai,Ren Togo,Takahiro Ogawa,Miki Haseyama
备注:Accepted to the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
【29】 Dissecting Temporal Understanding in Text-to-Audio Retrieval
标题: 文本到音频检索中的时态理解剖析
作者:Andreea-Maria Oncescu,João F. Henriques,A. Sophia Koepke
备注:9 pages, 5 figures, ACM Multimedia 2024, this https URL
链接:点击下载PDF文件
【30】 LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
标题: LibriheavightMix:一个20,000小时的数据集,用于单通道回响多说话者语音分离、ASB和说话者拨号
作者:Zengrui Jin,Yifan Yang,Mohan Shi,Wei Kang,Xiaoyu Yang,Zengwei Yao,Fangjun Kuang,Liyong Guo,Lingwei Meng,Long Lin,Yong Xu,Shi-Xiong Zhang,Daniel Povey
备注:InterSpeech 2024
链接:点击下载PDF文件
【31】 Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
标题: 用于多说话人自动语音识别的重叠编码分离序列化语音信息引导
作者:Hao Shi,Yuan Gao,Zhaoheng Ni,Tatsuya Kawahara
链接:点击下载PDF文件
【32】 MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
标题: MaskGCT:带屏蔽生成编解码器的Zero-Shot文本到语音
作者:Yuancheng Wang,Haoyue Zhan,Liwei Liu,Ruihong Zeng,Haotian Guo,Jiachen Zheng,Qiang Zhang,Shunsi Zhang,Zhizheng Wu
链接:点击下载PDF文件
【33】 Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion
标题: 看到你的言语风格:一种新颖的Zero-Shot身份解开基于面部的语音转换
作者:Yan Rong,Li Liu
链接:点击下载PDF文件
【34】 FLUX that Plays Music
标题: 播放音乐的FLOX
作者:Zhengcong Fei,Mingyuan Fan,Changqian Yu,Junshi Huang
链接:点击下载PDF文件
【35】 Multi-scale Multi-instance Visual Sound Localization and Segmentation
标题: 多尺度多实例视觉声音定位与分割
作者:Shentong Mo,Haofan Wang
链接:点击下载PDF文件
【36】 Multi-label Zero-Shot Audio Classification with Temporal Attention
标题: 具有时间注意力的多标签Zero-Shot音频分类
作者:Duygu Dogan,Huang Xie,Toni Heittola,Tuomas Virtanen
备注:Accepted to International Workshop on Acoustic Signal Enhancement (IWAENC) 2024
链接:点击下载PDF文件
【37】 Density Adaptive Attention-based Speech Network: Enhancing Feature Understanding for Mental Health Disorders
标题: 密度自适应基于注意力的语音网络:增强对心理健康障碍的特征理解
作者:Georgios Ioannides,Adrian Kieback,Aman Chadha,Aaron Elkins
链接:点击下载PDF文件
【38】 Contrastive Augmentation: An Unsupervised Learning Approach for Keyword Spotting in Speech Technology
标题: 对比增强:语音技术中关键词发现的无监督学习方法
作者:Weinan Dai,Yifeng Jiang,Yuanjing Liu,Jinkun Chen,Xin Sun,Jinglei Tao
备注:This paper has been accepted by the ICPR2024
链接:点击下载PDF文件
【39】 REFFLY: Melody-Constrained Lyrics Editing Model
标题: REFFLY:旋律约束歌词编辑模型
作者:Songyan Zhao,Bingxuan Li,Yufei Tian,Nanyun Peng
链接:点击下载PDF文件
【40】 Towards a dynamical model of English vowels. Evidence from diphthongisation
标题: 迈向英语元音的动态模型。双元音化的证据
作者:Patrycja Strycharczuk,Sam Kirkham,Emily Gorman,Takayuki Nagamine
链接:点击下载PDF文件
【41】 ProGRes: Prompted Generative Rescoring on ASR n-Best
标题: ProGRes:在ASB n-Best上进行的预定生成重新评分
作者:Ada Defne Tur,Adel Moumen,Mirco Ravanelli
备注:IEEE Spoken Language Technology Workshop
链接:点击下载PDF文件
【42】 Speaker Tagging Correction With Non-Autoregressive Language Models
标题: 使用非自回归语言模型进行说话者标记纠正
作者:Grigor Kirakosyan,Davit Karamyan
备注:6 pages, 7 tables
链接:点击下载PDF文件
【43】 Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning
标题: 使用谱-时态图专注池和多任务学习的逐例查询关键词发现
作者:Zhenyu Wang,Shuyu Kong,Li Wan,Biqiao Zhang,Yiteng Huang,Mumin Jin,Ming Sun,Xin Lei,Zhaojun Yang
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
标题: 利用基于LID的协作混合专家模型增强代码转换语音识别
作者:Hukai Huang,Jiayan Lin,Kaidi Wang,Yishuang Li,Wenhao Guan,Qingyang Hong,Lin Li
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
摘要:由于不同语言之间语音相似性建模的固有困难,码转换语音识别提出了一个艰巨的挑战。本研究提出了一个合作的MoE,一个混合的专家(MoE)模型,利用专家组之间的协作机制。最初,前面的路由网络显式学习语言识别(LID)任务,并根据获得的LID权重选择专家。该过程确保了到MoE层的鲁棒路由信息,减轻了来自不同语言域对专家网络参数更新的干扰。LID权重还用于促进组间协作,从而实现特定语言表示的集成。此外,在每个语言专家组内,门控网络在无监督的情况下运行,以促进语言以外属性的合作。大量的实验证明了我们的方法的有效性,实现了显着的性能增强相比,替代方法。重要的是,我们的方法保留了MoE模型的高效推理能力,而不需要额外的预训练。摘要:Due to the inherent difficulty in modeling phonetic similarities across different languages, code-switching speech recognition presents a formidable challenge. This study proposes a Collaborative-MoE, a Mixture of Experts (MoE) model that leverages a collaborative mechanism among expert groups. Initially, a preceding routing network explicitly learns Language Identification (LID) tasks and selects experts based on acquired LID weights. This process ensures robust routing information to the MoE layer, mitigating interference from diverse language domains on expert network parameter updates. The LID weights are also employed to facilitate inter-group collaboration, enabling the integration of language-specific representations. Furthermore, within each language expert group, a gating network operates unsupervised to foster collaboration on attributes beyond language. Extensive experiments demonstrate the efficacy of our approach, achieving significant performance enhancements compared to alternative methods. Importantly, our method preserves the efficient inference capabilities characteristic of MoE models without necessitating additional pre-training.
【2】 Activity-Guided Industrial Anomalous Sound Detection against Interferences
标题: 活动引导的工业异常声音抗干扰检测
作者:Yunjoo Lee,Jaechang Kim,Jungseul Ok
备注:Thsis is an extended version of this https URL
链接:点击下载PDF文件
摘要:我们解决了工业声音数据异常检测的实际情况,其中目标机器的声音被背景噪声和相邻机器的干扰所破坏。克服这一挑战是困难的,因为在没有额外信息的情况下,干扰通常与目标机器几乎无法区分。为了解决这个问题,我们提出了SSAD,一个框架的源分离(SS),其次是异常检测(AD),它利用机器的活动信息,往往很容易在实际环境中。SSAD由两个部分组成:(i)活动通知SS,即使在具有相似音色的干扰下也能实现有效的源分离,以及(ii)两步掩蔽,通过强调与机器活动一致的异常来增强异常检测。我们的实验表明,SSAD实现了相当的精度与基线完全访问干净的信号,而SSAD只提供了一个损坏的信号和活动信息。此外,由于具有两步掩蔽的活动通知SS和AD,SSAD优于标准方法,特别是在干扰的情况下。它突出了SSAD在解决工业声音数据异常检测的复杂性方面的实际功效。摘要:We address a practical scenario of anomaly detection for industrial sound data, where the sound of a target machine is corrupted by background noise and interference from neighboring machines. Overcoming this challenge is difficult since the interference is often virtually indistinguishable from the target machine without additional information. To address the issue, we propose SSAD, a framework of source separation (SS) followed by anomaly detection (AD), which leverages machine activity information, often readily available in practical settings. SSAD consists of two components: (i) activity-informed SS, enabling effective source separation even given interference with similar timbre, and (ii) two-step masking, robustifying anomaly detection by emphasizing anomalies aligned with the machine activity. Our experiments demonstrate that SSAD achieves comparable accuracy to a baseline with full access to clean signals, while SSAD is provided only a corrupted signal and activity information. In addition, thanks to the activity-informed SS and AD with the two-step masking, SSAD outperforms standard approaches, particularly in cases with interference. It highlights the practical efficacy of SSAD in addressing the complexities of anomaly detection in industrial sound data.
【3】 The Role of Large Language Models in Musicology: Are We Ready to Trust the Machines?
标题: 大型语言模型在音乐学中的作用:我们准备好信任机器了吗?
作者:Pedro Ramoneda,Emilia Parada-Cabaleiro,Benno Weck,Xavier Serra
链接:点击下载PDF文件
摘要:在这项工作中,我们将探讨大型语言模型(LLM)在音乐学中的使用和可靠性。通过与专家和学生的讨论,我们评估了目前对这种如今无处不在的技术的接受程度和担忧。我们的目标是更进一步,提出一种半自动的方法来创建一个初始的基准使用检索增强生成模型和多项选择题生成,由人类专家验证。我们对400个人类验证问题的评估表明,目前的香草LLM比从音乐词典中检索增强生成的可靠性低。本文认为,在音乐学的LLM的潜力,需要音乐学驱动的研究,可以专门的LLM包括准确和可靠的领域知识。摘要:In this work, we explore the use and reliability of Large Language Models (LLMs) in musicology. From a discussion with experts and students, we assess the current acceptance and concerns regarding this, nowadays ubiquitous, technology. We aim to go one step further, proposing a semi-automatic method to create an initial benchmark using retrieval-augmented generation models and multiple-choice question generation, validated by human experts. Our evaluation on 400 human-validated questions shows that current vanilla LLMs are less reliable than retrieval augmented generation from music dictionaries. This paper suggests that the potential of LLMs in musicology requires musicology driven research that can specialized LLMs by including accurate and reliable domain knowledge.
【4】 USTC-KXDIGIT System Description for ASVspoof5 Challenge
标题: USTC-KXDIGIT ASVspoof 5挑战赛系统描述
作者:Yihao Chen,Haochen Wu,Nan Jiang,Xiang Xia,Qing Gu,Yunqi Hao,Pengfei Cai,Yu Guan,Jialong Wang,Weilin Xie,Lei Fang,Sian Fang,Yan Song,Wu Guo,Lin Liu,Minqiang Xu
备注:ASVspoof5 workshop paper
链接:点击下载PDF文件
摘要:本文描述了USTC-KXDIGIT系统提交给ASVspoof 5挑战赛的Track 1(语音深度伪造检测)和Track 2(欺骗鲁棒自动说话人验证,SASV)。轨道1展示了来自潜在处理算法的各种技术质量,包括开放和封闭条件。对于这些条件,我们的系统由前端特征提取器和后端分类器的级联组成。我们专注于广泛的嵌入工程和增强后端分类器模型的泛化。具体来说,嵌入工程是基于手工制作的功能和语音表示从一个自我监督的模型,分别用于封闭和开放的条件。为了在各种对抗条件下检测欺骗攻击,我们在增强训练集上训练了多个系统。此外,我们使用语音转换技术从训练集中的真实音频合成假音频,以丰富合成算法。为了利用不同模型架构学习到的互补信息,我们采用了来自不同系统的激活集成和融合分数,以获得欺骗检测的最终决策分数。在评估阶段,所提出的方法在封闭条件下实现了0.3948 minDCF和14.33%EER,在开放条件下实现了0.0750 minDCF和2.59%EER,证明了我们提交的系统在对抗条件下的鲁棒性。在Track 2中,我们继续使用Track 1中的CM系统,并将其与基于CNN的ASV系统融合。该方法在封闭条件下实现了0.2814 min-aDCF,在开放条件下实现了0.0756 min-aDCF,显示了SASV系统的优异性能。摘要:This paper describes the USTC-KXDIGIT system submitted to the ASVspoof5 Challenge for Track 1 (speech deepfake detection) and Track 2 (spoofing-robust automatic speaker verification, SASV). Track 1 showcases a diverse range of technical qualities from potential processing algorithms and includes both open and closed conditions. For these conditions, our system consists of a cascade of a frontend feature extractor and a back-end classifier. We focus on extensive embedding engineering and enhancing the generalization of the back-end classifier model. Specifically, the embedding engineering is based on hand-crafted features and speech representations from a self-supervised model, used for closed and open conditions, respectively. To detect spoof attacks under various adversarial conditions, we trained multiple systems on an augmented training set. Additionally, we used voice conversion technology to synthesize fake audio from genuine audio in the training set to enrich the synthesis algorithms. To leverage the complementary information learned by different model architectures, we employed activation ensemble and fused scores from different systems to obtain the final decision score for spoof detection. During the evaluation phase, the proposed methods achieved 0.3948 minDCF and 14.33% EER in the close condition, and 0.0750 minDCF and 2.59% EER in the open condition, demonstrating the robustness of our submitted systems under adversarial conditions. In Track 2, we continued using the CM system from Track 1 and fused it with a CNN-based ASV system. This approach achieved 0.2814 min-aDCF in the closed condition and 0.0756 min-aDCF in the open condition, showcasing superior performance in the SASV system.
【5】 Pureformer-VC: Non-parallel One-Shot Voice Conversion with Pure Transformer Blocks and Triplet Discriminative Training
标题: Pureformer-VC:使用纯Transformer块和三重区分训练的非并行单次语音转换
作者:Wenhan Yao,Zedong Xing,Xiarun Chen,Jia Liu,Yongqiang He,Weiping Wen
备注:submmited to ICASSP 2025
链接:点击下载PDF文件
摘要:单次语音转换(VC)的目的是改变任何源语音的音色,以匹配的目标说话人与只有一个语音样本。现有的基于风格转换的VC方法依赖于语音表示的解纠缠,难以准确独立地对每个语音分量进行编码,并有效地重新组合成转换后的语音。为了解决这个问题,我们提出了Pureformer-VC,它利用Conformer块来构建一个解纠缠的编码器,并利用Zipformer块来构建一个风格转换解码器作为生成器。在解码器中,我们使用了有效的styleformer块来有效地将说话人特征整合到生成的语音中。该模型使用生成VAE损失编码组件和三重损失的无监督判别训练。我们将styleformer方法应用于Zipformer的共享权重以进行样式传输。实验结果表明,该模型实现了可比的主观分数,并表现出改善的客观指标相比,现有的方法在一个一次性的语音转换场景。摘要:One-shot voice conversion(VC) aims to change the timbre of any source speech to match that of the unseen target speaker with only one speech sample. Existing style transfer-based VC methods relied on speech representation disentanglement and suffered from accurately and independently encoding each speech component and recomposing back to converted speech effectively. To tackle this, we proposed Pureformer-VC, which utilizes Conformer blocks to build a disentangled encoder, and Zipformer blocks to build a style transfer decoder as the generator. In the decoder, we used effective styleformer blocks to integrate speaker characteristics into the generated speech effectively. The models used the generative VAE loss for encoding components and triplet loss for unsupervised discriminative training. We applied the styleformer method to Zipformer's shared weights for style transfer. The experimental results show that the proposed model achieves comparable subjective scores and exhibits improvements in objective metrics compared to existing methods in a one-shot voice conversion scenario.
【6】 VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka
标题: VoxHakka:台湾客语方言多样化的多说话人文本到语音系统
作者:Li-Wei Chen,Hung-Shin Lee,Chen-Chi Chang
备注:Submitted to O-COCOSDA 2024
链接:点击下载PDF文件
摘要:本文介绍了VoxHakka,一个文本到语音(TTS)系统设计的台湾客家话,一个严重不足的资源在台湾所说的语言。利用YourTTS框架,VoxHakka在语音合成中实现了高自然度和准确度以及低实时性,同时支持六种不同的客家方言。这是通过用方言特定数据训练模型来实现的,允许生成说话者感知的客家话语音。为了解决公共可用的客家语音语料库的稀缺性,我们采用了一种具有成本效益的方法,利用网络抓取管道加上自动语音识别(ASR)为基础的数据清洗技术。该过程确保了获得适合TTS训练的高质量、多说话者、多方言数据集。使用比较平均意见分数(CMOS)进行的主观听力测试表明,VoxHakka显着优于现有的公开可用的客家文语转换系统的发音准确性,音调的正确性,和整体自然。这项工作代表了客家语言技术的重大进步,并为语言保护和振兴工作提供了宝贵的资源。摘要:This paper introduces VoxHakka, a text-to-speech (TTS) system designed for Taiwanese Hakka, a critically under-resourced language spoken in Taiwan. Leveraging the YourTTS framework, VoxHakka achieves high naturalness and accuracy and low real-time factor in speech synthesis while supporting six distinct Hakka dialects. This is achieved by training the model with dialect-specific data, allowing for the generation of speaker-aware Hakka speech. To address the scarcity of publicly available Hakka speech corpora, we employed a cost-effective approach utilizing a web scraping pipeline coupled with automatic speech recognition (ASR)-based data cleaning techniques. This process ensured the acquisition of a high-quality, multi-speaker, multi-dialect dataset suitable for TTS training. Subjective listening tests conducted using comparative mean opinion scores (CMOS) demonstrate that VoxHakka significantly outperforms existing publicly available Hakka TTS systems in terms of pronunciation accuracy, tone correctness, and overall naturalness. This work represents a significant advancement in Hakka language technology and provides a valuable resource for language preservation and revitalization efforts.
【7】 Effective Noise-aware Data Simulation for Domain-adaptive Speech Enhancement Leveraging Dynamic Stochastic Perturbation
标题: 利用动态随机扰动的域自适应语音增强的有效噪音感知数据模拟
作者:Chien-Chun Wang,Li-Wei Chen,Hung-Shin Lee,Berlin Chen,Hsin-Min Wang
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
摘要:跨域语音增强(SE)通常面临着严峻的挑战,因为在未知的目标域中噪声和背景信息的缺乏,导致训练和测试条件之间的不匹配。这项研究提出了一种新的数据模拟方法来解决这个问题,利用噪声提取技术和生成对抗网络(GANs),只有有限的目标噪声语音数据。值得注意的是,我们的方法采用噪声编码器从目标域数据中提取噪声嵌入。这些嵌入适当地引导生成器合成声学上适合于目标域的话语,同时真实地保留输入干净语音的语音内容。此外,我们引入了动态随机扰动的概念,它可以在推理过程中将受控扰动注入噪声嵌入,从而使模型能够很好地推广到看不见的噪声条件。在VoiceBank-DEMAND基准数据集上的实验表明,我们的域自适应SE方法优于现有的基于数据模拟的强基线。摘要:Cross-domain speech enhancement (SE) is often faced with severe challenges due to the scarcity of noise and background information in an unseen target domain, leading to a mismatch between training and test conditions. This study puts forward a novel data simulation method to address this issue, leveraging noise-extractive techniques and generative adversarial networks (GANs) with only limited target noisy speech data. Notably, our method employs a noise encoder to extract noise embeddings from target-domain data. These embeddings aptly guide the generator to synthesize utterances acoustically fitted to the target domain while authentically preserving the phonetic content of the input clean speech. Furthermore, we introduce the notion of dynamic stochastic perturbation, which can inject controlled perturbations into the noise embeddings during inference, thereby enabling the model to generalize well to unseen noise conditions. Experiments on the VoiceBank-DEMAND benchmark dataset demonstrate that our domain-adaptive SE method outperforms an existing strong baseline based on data simulation.
【8】 Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement
标题: Spectron:使用具有对抗细化的条件Transformer提取目标说话人
作者:Tathagata Bandyopadhyay
链接:点击下载PDF文件
摘要:最近,基于注意力的Transformers已经成为许多深度学习应用的事实标准,包括自然语言处理,计算机视觉,信号处理等。在本文中,我们提出了一个基于变换器的端到端模型,从单声道多说话人混合音频信号中提取目标说话人的语音。与现有的说话人提取方法不同,我们引入了两个额外的目标,施加说话人嵌入的一致性和波形编码器的可逆性,并联合训练说话人编码器和语音分离器,以更好地捕捉说话人的条件嵌入。此外,我们利用一个多尺度的降噪来改善所提取的语音的感知质量。我们的实验表明,在分离器主干中使用双路径Transformer以及所提出的训练范例将CNN基线提高了3.12 $ dB点。最后,我们将我们的方法与最新的最先进的方法进行比较,结果表明,我们的模型平均比现有方法高出4.1 $ dB,而不会产生额外的数据依赖性。摘要:Recently, attention-based transformers have become a de facto standard in many deep learning applications including natural language processing, computer vision, signal processing, etc.. In this paper, we propose a transformer-based end-to-end model to extract a target speaker's speech from a monaural multi-speaker mixed audio signal. Unlike existing speaker extraction methods, we introduce two additional objectives to impose speaker embedding consistency and waveform encoder invertibility and jointly train both speaker encoder and speech separator to better capture the speaker conditional embedding. Furthermore, we leverage a multi-scale discriminator to refine the perceptual quality of the extracted speech. Our experiments show that the use of a dual path transformer in the separator backbone along with proposed training paradigm improves the CNN baseline by $3.12$ dB points. Finally, we compare our approach with recent state-of-the-arts and show that our model outperforms existing methods by $4.1$ dB points on an average without creating additional data dependency.
【9】 A multilingual training strategy for low resource Text to Speech
标题: 低资源文本到语音的多语言训练策略
作者:Asma Amalas,Mounir Ghogho,Mohamed Chetouani,Rachid Oulad Haj Thami
备注:12 pages, 2 figures
链接:点击下载PDF文件
摘要:由于神经文本到语音(TTS)的最新进展,最近的语音技术已经导致产生高质量的合成语音。然而,这种TTS模型依赖于大量的数据,这些数据的产生成本很高,并且很难扩展到所有现有的语言,特别是很少关注低资源语言。通过知识转移等技术,可以减轻创建数据集的负担。因此,在本文中,我们研究了两个方面;首先,来自社交媒体的数据是否可以用于小型TTS数据集的构建,其次,低资源语言的跨语言迁移学习(TL)是否可以使用这种类型的数据。在这方面,我们具体评估了多语言建模在多大程度上可以作为单语语料库训练的替代方案。要做到这一点,我们探讨如何从外语的数据可以被选择和汇集到训练一个目标低资源语言的TTS模型。我们的研究结果表明,多语种预训练比单语预训练更好地提高了生成语音的可懂度和自然度。摘要:Recent speech technologies have led to produce high quality synthesised speech due to recent advances in neural Text to Speech (TTS). However, such TTS models depend on extensive amounts of data that can be costly to produce and is hardly scalable to all existing languages, especially that seldom attention is given to low resource languages. With techniques such as knowledge transfer, the burden of creating datasets can be alleviated. In this paper, we therefore investigate two aspects; firstly, whether data from social media can be used for a small TTS dataset construction, and secondly whether cross lingual transfer learning (TL) for a low resource language can work with this type of data. In this aspect, we specifically assess to what extent multilingual modeling can be leveraged as an alternative to training on monolingual corporas. To do so, we explore how data from foreign languages may be selected and pooled to train a TTS model for a target low resource language. Our findings show that multilingual pre-training is better than monolingual pre-training at increasing the intelligibility and naturalness of the generated speech.
【10】 Interpretable Convolutional SyncNet
标题: 可解释卷积同步网络
作者:Sungjoon Park,Jaesub Yun,Donggeon Lee,Minsik Park
备注:8+5 pages
链接:点击下载PDF文件
摘要:由于各种原因,视频可能会不同步,因此使用同步网络将视频恢复同步,以执行需要同步视频的任务。以前的最先进的(SOTA)同步网络使用InfoNCE损耗,依赖于Transformer架构,或两者兼而有之。不幸的是,前者使模型的输出难以解释,后者对大图像不友好,从而限制了同步网络的有用性。在这项工作中,我们训练一个卷积同步网络使用平衡BCE损失(BBCE),损失的启发二进制交叉熵(BCE)和InfoNCE损失。与InfoNCE损失相比,BBCE损失不需要复杂的采样方案。我们的模型可以更好地处理更大的图像,其输出可以给出概率解释。概率解释允许我们定义诸如偏移概率和屏幕外比率的度量来评估视听(AV)语音数据集的同步质量。此外,我们的模型在LRS 2数据集上达到了96.5 %$的SOTA准确度,在LRS 3数据集上达到了93.8 %$。摘要:Because videos in the wild can be out of sync for various reasons, a sync-net is used to bring the video back into sync for tasks that require synchronized videos. Previous state-of-the-art (SOTA) sync-nets use InfoNCE loss, rely on the transformer architecture, or both. Unfortunately, the former makes the model's output difficult to interpret, and the latter is unfriendly with large images, thus limiting the usefulness of sync-nets. In this work, we train a convolutional sync-net using the balanced BCE loss (BBCE), a loss inspired by the binary cross entropy (BCE) and the InfoNCE losses. In contrast to the InfoNCE loss, the BBCE loss does not require complicated sampling schemes. Our model can better handle larger images, and its output can be given a probabilistic interpretation. The probabilistic interpretation allows us to define metrics such as probability at offset and offscreen ratio to evaluate the sync quality of audio-visual (AV) speech datasets. Furthermore, our model achieves SOTA accuracy of $96.5 %$ on the LRS2 dataset and $93.8 %$ on the LRS3 dataset.
【11】 A Framework for Synthetic Audio Conversations Generation using Large Language Models
标题: 使用大型语言模型生成合成音频对话的框架
作者:Kaung Myat Kyaw,Jonathan Hoyin Chan
备注:This work has been submitted for consideration at the WI-IAT'24 to be held in December 2024
链接:点击下载PDF文件
摘要:在本文中,我们介绍ConversaSynth,一个框架,旨在生成合成会话音频使用大型语言模型(LLM)与多个人物设置。该框架首先创建各种主题的基于文本的对话,然后使用文本到语音(TTS)系统将其转换为音频。我们的实验表明,ConversaSynth可以有效地生成高质量的合成音频数据集,这可以显着增强音频标记,音频分类和多说话人语音识别模型的训练和评估。结果表明,由ConversaSynth生成的合成数据集表现出极大的多样性和真实性,使其适合开发强大的,适应性强的基于音频的AI系统。摘要:In this paper, we introduce ConversaSynth, a framework designed to generate synthetic conversation audio using large language models (LLMs) with multiple persona settings. The framework first creates diverse and coherent text-based dialogues across various topics, which are then converted into audio using text-to-speech (TTS) systems. Our experiments demonstrate that ConversaSynth effectively generates highquality synthetic audio datasets, which can significantly enhance the training and evaluation of models for audio tagging, audio classification, and multi-speaker speech recognition. The results indicate that the synthetic datasets generated by ConversaSynth exhibit substantial diversity and realism, making them suitable for developing robust, adaptable audio-based AI systems.
【12】 SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
标题: SoCodec:一种语义有序的多流语音编解码器,用于基于高效语言模型的文本到语音合成
作者:Haohan Guo,Fenglong Xie,Kun Xie,Dongchao Yang,Dake Guo,Xixin Wu,Helen Meng
备注:Accepted by SLT 2024
链接:点击下载PDF文件
摘要:长语音序列一直困扰着基于语言模型(LM)的TTS方法的建模复杂度和效率。这项工作提出了SoCodec,语义有序的多流语音编解码器,以解决这个问题。它将语音压缩成一个较短的、多流的离散语义序列,每个帧有多个标记。同时,提出了有序积量化的方法,将该序列约束为有序表示。它可以与多流延迟LM一起应用,以在TTS中沿时间和流轴实现更好的自回归生成。实验结果有力地证明了所提出的方法的有效性,即使将语音的帧移从20ms压缩到240ms(12倍),也能获得优于基线系统的性能。消融研究进一步验证了学习所提出的有序多流语义表示在追求更短的语音序列的有效LM为基础的TTS的重要性。摘要:The long speech sequence has been troubling language models (LM) based TTS approaches in terms of modeling complexity and efficiency. This work proposes SoCodec, a semantic-ordered multi-stream speech codec, to address this issue. It compresses speech into a shorter, multi-stream discrete semantic sequence with multiple tokens at each frame. Meanwhile, the ordered product quantization is proposed to constrain this sequence into an ordered representation. It can be applied with a multi-stream delayed LM to achieve better autoregressive generation along both time and stream axes in TTS. The experimental result strongly demonstrates the effectiveness of the proposed approach, achieving superior performance over baseline systems even if compressing the frameshift of speech from 20ms to 240ms (12x). The ablation studies further validate the importance of learning the proposed ordered multi-stream semantic representation in pursuing shorter speech sequences for efficient LM-based TTS.
【13】 MMT-BERT: Chord-aware Symbolic Music Generation Based on Multitrack Music Transformer and MusicBERT
标题: MMT-BERT:基于多轨音乐Transformer和MusicBERT的和弦感知符号音乐生成
作者:Jinlong Zhu,Keigo Sakurai,Ren Togo,Takahiro Ogawa,Miki Haseyama
备注:Accepted to the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
摘要:我们提出了一种新的符号音乐表示和生成对抗网络(GAN)框架,专门设计用于符号多轨音乐生成。符号音乐生成的主题主要包括音乐数据的预处理和深度学习框架的实现。目前的符号音乐生成技术通常面临两个重大挑战:训练数据缺乏和弦和音阶的信息,以及需要专门设计的模型架构来适应符号音乐表示的独特格式。在本文中,我们解决了上述问题,通过引入新的符号音乐表示与MusicLang和弦分析模型。我们提出了我们的MMT-BERT架构适应的表示。为了构建一个强大的多轨音乐生成器,我们微调了一个预先训练好的MusicBERT模型来作为训练器,并结合了相对论标准损失。这种方法得到了对MusicBERT中编码的符号音乐的深入理解的支持,增强了我们的方法生成的音乐的和谐和人性。实验结果表明,我们的方法,严格遵循国家的最先进的方法的有效性。摘要:We propose a novel symbolic music representation and Generative Adversarial Network (GAN) framework specially designed for symbolic multitrack music generation. The main theme of symbolic music generation primarily encompasses the preprocessing of music data and the implementation of a deep learning framework. Current techniques dedicated to symbolic music generation generally encounter two significant challenges: training data's lack of information about chords and scales and the requirement of specially designed model architecture adapted to the unique format of symbolic music representation. In this paper, we solve the above problems by introducing new symbolic music representation with MusicLang chord analysis model. We propose our MMT-BERT architecture adapting to the representation. To build a robust multitrack music generator, we fine-tune a pre-trained MusicBERT model to serve as the discriminator, and incorporate relativistic standard loss. This approach, supported by the in-depth understanding of symbolic music encoded within MusicBERT, fortifies the consonance and humanity of music generated by our method. Experimental results demonstrate the effectiveness of our approach which strictly follows the state-of-the-art methods.
【14】 Dissecting Temporal Understanding in Text-to-Audio Retrieval
标题: 文本到音频检索中的时态理解剖析
作者:Andreea-Maria Oncescu,João F. Henriques,A. Sophia Koepke
备注:9 pages, 5 figures, ACM Multimedia 2024, this https URL
链接:点击下载PDF文件
摘要:机器学习的最新进展推动了对多模态任务的研究,例如文本到视频和文本到音频检索。这些任务需要模型来理解视频和音频数据的语义内容,包括对象和字符。模型还需要学习空间安排和时间关系。在这项工作中,我们分析的时间顺序的声音,这是一个未充分研究的问题,在文本到音频检索的背景下。特别是,我们剖析了AudioCaps和Clotho数据集上的文本到音频检索的最先进模型的时间理解能力。此外,我们还介绍了一个合成的文本音频数据集,它提供了一个受控的设置,用于评估最近的模型的时间能力。最后,我们提出了一个损失函数,鼓励文本音频模型专注于事件的时间顺序。代码和数据可在https: www.robots.ox.ac.uk ~vgg research audio-retrieval dtu 上获得。摘要:Recent advancements in machine learning have fueled research on multimodal tasks, such as for instance text-to-video and text-to-audio retrieval. These tasks require models to understand the semantic content of video and audio data, including objects, and characters. The models also need to learn spatial arrangements and temporal relationships. In this work, we analyse the temporal ordering of sounds, which is an understudied problem in the context of text-to-audio retrieval. In particular, we dissect the temporal understanding capabilities of a state-of-the-art model for text-to-audio retrieval on the AudioCaps and Clotho datasets. Additionally, we introduce a synthetic text-audio dataset that provides a controlled setting for evaluating temporal capabilities of recent models. Lastly, we present a loss function that encourages text-audio models to focus on the temporal ordering of events. Code and data are available at https: www.robots.ox.ac.uk ~vgg research audio-retrieval dtu .
【15】 LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
标题: LibriheavightMix:一个20,000小时的数据集,用于单通道回响多说话者语音分离、ASB和说话者拨号
作者:Zengrui Jin,Yifan Yang,Mohan Shi,Wei Kang,Xiaoyu Yang,Zengwei Yao,Fangjun Kuang,Liyong Guo,Lingwei Meng,Long Lin,Yong Xu,Shi-Xiong Zhang,Daniel Povey
备注:InterSpeech 2024
链接:点击下载PDF文件
摘要:不断发展的语音处理领域越来越关注复杂的场景,如具有多个同时发言者和远场条件的会议或鸡尾酒会。解决这些挑战的现有方法分为两类:多渠道和单渠道解决方案。单通道方法以其通用性和方便性而闻名,不需要关于麦克风阵列的特定信息。 本文提出了一个大规模的远场重叠语音数据集,制作语音分离,识别和说话人日记的研究。该数据集是在多人、混响环境中解码“谁在什么时候说了什么”的关键资源,这是该领域的一个艰巨挑战。此外,我们介绍了一个管道系统,包括语音分离,识别和日记作为一个基本的基准。关于WHAMR!数据集验证了拟议数据的广泛适用性。摘要:The evolving speech processing landscape is increasingly focused on complex scenarios like meetings or cocktail parties with multiple simultaneous speakers and far-field conditions. Existing methodologies for addressing these challenges fall into two categories: multi-channel and single-channel solutions. Single-channel approaches, notable for their generality and convenience, do not require specific information about microphone arrays. This paper presents a large-scale far-field overlapping speech dataset, crafted to advance research in speech separation, recognition, and speaker diarization. This dataset is a critical resource for decoding Who said What and When'' in multi-talker, reverberant environments, a daunting challenge in the field. Additionally, we introduce a pipeline system encompassing speech separation, recognition, and diarization as a foundational benchmark. Evaluations on the WHAMR! dataset validate the broad applicability of the proposed data.
【16】 Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
标题: 用于多说话人自动语音识别的重叠编码分离序列化语音信息引导
作者:Hao Shi,Yuan Gao,Zhaoheng Ni,Tatsuya Kawahara
链接:点击下载PDF文件
摘要:串行输出训练(SOT)由于其方便灵活的多说话人自动语音识别(ASR)方法而受到越来越多的关注。然而,仅仅在注意力损失的情况下进行训练并不容易。在本文中,我们提出了重叠编码分离(EncSep),以充分利用连接主义的时间分类(CTC)和注意力混合损失的好处。该附加分离器被插入在编码器之后以提取具有CTC损失的多说话者信息。此外,我们提出了序列化的语音信息指导SOT(GEncSep),以进一步利用分离的编码。分离的流被连接以提供单个说话者信息,从而在解码期间引导注意力。在LibriMix上的实验结果表明,该方法可以有效地将单说话人编码与重叠编码分离。CTC损失有助于改善复杂场景下的编码器表示。GEncSep进一步提高了性能。摘要:Serialized output training (SOT) attracts increasing attention due to its convenience and flexibility for multi-speaker automatic speech recognition (ASR). However, it is not easy to train with attention loss only. In this paper, we propose the overlapped encoding separation (EncSep) to fully utilize the benefits of the connectionist temporal classification (CTC) and attention hybrid loss. This additional separator is inserted after the encoder to extract the multi-speaker information with CTC losses. Furthermore, we propose the serialized speech information guidance SOT (GEncSep) to further utilize the separated encodings. The separated streams are concatenated to provide single-speaker information to guide attention during decoding. The experimental results on LibriMix show that the single-speaker encoding can be separated from the overlapped encoding. The CTC loss helps to improve the encoder representation under complex scenarios. GEncSep further improved performance.
【17】 MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
标题: MaskGCT:带屏蔽生成编解码器的Zero-Shot文本到语音
作者:Yuancheng Wang,Haoyue Zhan,Liwei Liu,Ruihong Zeng,Haotian Guo,Jiachen Zheng,Qiang Zhang,Shunsi Zhang,Zhizheng Wu
链接:点击下载PDF文件
摘要:目前,大规模的文语转换系统主要分为两类:自回归和非自回归。自回归系统在鲁棒性方面存在一定的不足,并且不能控制语音持续时间。相比之下,非自回归系统需要明确的预测的语音水平的持续时间,这可能会损害他们的自然性。我们介绍了掩蔽生成编解码器Transformer(MaskGCT),一个完全非自回归模型的TTS,不需要精确的对齐信息之间的文本和语音。MaskGCT是一个两阶段模型:在第一阶段,该模型使用文本来预测从语音自监督学习(SSL)模型中提取的语义令牌,在第二阶段,该模型预测以这些语义令牌为条件的声学令牌。MaskGCT遵循 textit{mask-and-predict}学习范式。在训练过程中,MaskGCT根据给定的条件和提示学习预测掩蔽的语义或声学标记。在推理过程中,模型以并行方式生成指定长度的令牌。我们将MaskGCT扩展到一个大规模的多语言数据集,其中包含10万小时的野外语音。我们的实验表明,MaskGCT实现卓越的或有竞争力的性能相比,国家的最先进的zero-shot TTS系统的质量,相似性和可理解性,同时提供更高的生成效率比基于扩散或自回归TTS模型。音频样本可在https: maskgct.github.io上获得。摘要:Nowadays, large-scale text-to-speech (TTS) systems are primarily divided into two types: autoregressive and non-autoregressive. The autoregressive systems have certain deficiencies in robustness and cannot control speech duration. In contrast, non-autoregressive systems require explicit prediction of phone-level duration, which may compromise their naturalness. We introduce the Masked Generative Codec Transformer (MaskGCT), a fully non-autoregressive model for TTS that does not require precise alignment information between text and speech. MaskGCT is a two-stage model: in the first stage, the model uses text to predict semantic tokens extracted from a speech self-supervised learning (SSL) model, and in the second stage, the model predicts acoustic tokens conditioned on these semantic tokens. MaskGCT follows the textit{mask-and-predict} learning paradigm. During training, MaskGCT learns to predict masked semantic or acoustic tokens based on given conditions and prompts. During inference, the model generates tokens of a specified length in a parallel manner. We scale MaskGCT to a large-scale multilingual dataset with 100K hours of in-the-wild speech. Our experiments demonstrate that MaskGCT achieves superior or competitive performance compared to state-of-the-art zero-shot TTS systems in terms of quality, similarity, and intelligibility while offering higher generation efficiency than diffusion-based or autoregressive TTS models. Audio samples are available at https: maskgct.github.io.
【18】 Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion
标题: 看到你的言语风格:一种新颖的Zero-Shot身份解开基于面部的语音转换
作者:Yan Rong,Li Liu
链接:点击下载PDF文件
摘要:基于人脸的语音转换(FVC)是一种利用人脸图像来生成目标说话人的语音风格的新任务。先前的工作有两个缺点:(1)难以获得与说话者的语音身份信息良好对齐的面部嵌入,以及(2)在从音频输入解耦内容和说话者身份信息方面的不足。为了解决这些问题,我们提出了一种新的FVC方法,身份解开基于人脸的语音转换(ID-FaceVC),它克服了上述两个限制。更确切地说,我们提出了一个基于查询的身份感知对比学习(IAQ-CL)模块来提取特定于说话者的面部特征,以及一个基于互信息的双重解耦(MIDD)模块来从音频中纯化内容特征,确保清晰和高质量的语音转换。此外,与以前的作品不同,我们的方法可以接受音频或文本输入,提供可控的语音生成与可调的情感基调和速度。大量的实验表明,ID-FaceVC在各种指标上都达到了最先进的性能,定性和用户研究结果证实了其在自然性、相似性和多样性方面的有效性。包含音频示例和代码的项目网站可以在https: id-facevc.github.io上找到。摘要:Face-based Voice Conversion (FVC) is a novel task that leverages facial images to generate the target speaker's voice style. Previous work has two shortcomings: (1) suffering from obtaining facial embeddings that are well-aligned with the speaker's voice identity information, and (2) inadequacy in decoupling content and speaker identity information from the audio input. To address these issues, we present a novel FVC method, Identity-Disentanglement Face-based Voice Conversion (ID-FaceVC), which overcomes the above two limitations. More precisely, we propose an Identity-Aware Query-based Contrastive Learning (IAQ-CL) module to extract speaker-specific facial features, and a Mutual Information-based Dual Decoupling (MIDD) module to purify content features from audio, ensuring clear and high-quality voice conversion. Besides, unlike prior works, our method can accept either audio or text inputs, offering controllable speech generation with adjustable emotional tone and speed. Extensive experiments demonstrate that ID-FaceVC achieves state-of-the-art performance across various metrics, with qualitative and user study results confirming its effectiveness in naturalness, similarity, and diversity. Project website with audio samples and code can be found at https: id-facevc.github.io.
【19】 FLUX that Plays Music
标题: 播放音乐的FLOX
作者:Zhengcong Fei,Mingyuan Fan,Changqian Yu,Junshi Huang
链接:点击下载PDF文件
摘要:本文探讨了一个简单的扩展,基于扩散整流Transformers的文本到音乐的生成,称为FluxMusic。通常,随着高级Flux footnote{https: github.com black-forest-labs flux}模型的设计,我们将其转移到mel-谱的潜在VAE空间中。它涉及首先将一系列独立注意力应用于双文本音乐流,然后是堆叠的单音乐流用于去噪补丁预测。我们采用多个预先训练的文本编码器来充分捕获字幕语义信息以及推理灵活性。在这两者之间,粗文本信息结合时间步嵌入被用于调制机制,而细粒度的文本细节与作为输入的音乐补丁序列连接。通过深入的研究,我们证明了具有优化架构的整流训练显著优于文本到音乐任务的既定扩散方法,正如各种自动指标和人类偏好评估所证明的那样。我们的实验数据、代码和模型权重在以下网址公开: url{https: github.com feizc FluxMusic}。摘要:This paper explores a simple extension of diffusion-based rectified flow Transformers for text-to-music generation, termed as FluxMusic. Generally, along with design in advanced Flux footnote{https: github.com black-forest-labs flux} model, we transfers it into a latent VAE space of mel-spectrum. It involves first applying a sequence of independent attention to the double text-music stream, followed by a stacked single music stream for denoised patch prediction. We employ multiple pre-trained text encoders to sufficiently capture caption semantic information as well as inference flexibility. In between, coarse textual information, in conjunction with time step embeddings, is utilized in a modulation mechanism, while fine-grained textual details are concatenated with the music patch sequence as inputs. Through an in-depth study, we demonstrate that rectified flow training with an optimized architecture significantly outperforms established diffusion methods for the text-to-music task, as evidenced by various automatic metrics and human preference evaluations. Our experimental data, code, and model weights are made publicly available at: url{https: github.com feizc FluxMusic}.
【20】 Multi-scale Multi-instance Visual Sound Localization and Segmentation
标题: 多尺度多实例视觉声音定位与分割
作者:Shentong Mo,Haofan Wang
链接:点击下载PDF文件
摘要:视觉声音定位是一个典型的和具有挑战性的问题,预测对应的声源在视频中的对象的位置。以往的方法主要是使用全局音频和单尺度视觉特征之间的视听关联来定位每个图像中的发声对象。尽管它们的性能很有希望,但它们忽略了相应图像的多尺度视觉特征,并且与地面事实相比,它们无法学习区分区域。为了解决这个问题,我们提出了一种新的多尺度多实例视觉声音定位框架,即M2 VSL,它可以直接从输入图像中学习与声源相关的多尺度语义特征来定位发声对象。具体来说,我们的M2 VSL利用可学习的多尺度视觉特征,在相应图像的多层次位置对齐视听表示。我们还介绍了一种新的多尺度多实例Transformer动态聚合多尺度跨模态表示的视觉声音定位。我们对VGGSound-Instruments,VGG-Sound Sources和AVSBench基准进行了广泛的实验。实验结果表明,本文提出的M2 VSL算法在探测目标定位和分割方面具有较好的性能。摘要:Visual sound localization is a typical and challenging problem that predicts the location of objects corresponding to the sound source in a video. Previous methods mainly used the audio-visual association between global audio and one-scale visual features to localize sounding objects in each image. Despite their promising performance, they omitted multi-scale visual features of the corresponding image, and they cannot learn discriminative regions compared to ground truths. To address this issue, we propose a novel multi-scale multi-instance visual sound localization framework, namely M2VSL, that can directly learn multi-scale semantic features associated with sound sources from the input image to localize sounding objects. Specifically, our M2VSL leverages learnable multi-scale visual features to align audio-visual representations at multi-level locations of the corresponding image. We also introduce a novel multi-scale multi-instance transformer to dynamically aggregate multi-scale cross-modal representations for visual sound localization. We conduct extensive experiments on VGGSound-Instruments, VGG-Sound Sources, and AVSBench benchmarks. The results demonstrate that the proposed M2VSL can achieve state-of-the-art performance on sounding object localization and segmentation.
【21】 Multi-label Zero-Shot Audio Classification with Temporal Attention
标题: 具有时间注意力的多标签Zero-Shot音频分类
作者:Duygu Dogan,Huang Xie,Toni Heittola,Tuomas Virtanen
备注:Accepted to International Workshop on Acoustic Signal Enhancement (IWAENC) 2024
链接:点击下载PDF文件
摘要:Zero-shot学习模型能够通过使用辅助信息从所看到的类中转移知识来对新类进行分类。现有的zero-shot学习方法大多集中在单标签分类任务上,本文提出了一种多标签zero-shot音频分类方法。为了解决多标签声音分类的挑战,同时推广到看不见的类,我们适应时间注意。时间注意力机制基于音频片段的声学和语义兼容性将重要性权重分配给不同的音频片段,从而使模型能够通过关注与每个类最相关的片段来捕获音频样本内不同声音类的变化主导性。这导致比采用时间上聚合的声学特征而不加权的方法更准确的多标签zero-shot分类,所述方法同等地对待所有音频段。我们评估我们的方法对一个zero-shot模型的AudioSet子集上使用统一聚合的声学特征,零规则基线,并在监督的情况下提出的方法。我们的研究结果表明,在多标签的情况下,时间注意增强了zero-shot音频分类性能。摘要:Zero-shot learning models are capable of classifying new classes by transferring knowledge from the seen classes using auxiliary information. While most of the existing zero-shot learning methods focused on single-label classification tasks, the present study introduces a method to perform multi-label zero-shot audio classification. To address the challenge of classifying multi-label sounds while generalizing to unseen classes, we adapt temporal attention. The temporal attention mechanism assigns importance weights to different audio segments based on their acoustic and semantic compatibility, thus enabling the model to capture the varying dominance of different sound classes within an audio sample by focusing on the segments most relevant for each class. This leads to more accurate multi-label zero-shot classification than methods employing temporally aggregated acoustic features without weighting, which treat all audio segments equally. We evaluate our approach on a subset of AudioSet against a zero-shot model using uniformly aggregated acoustic features, a zero-rule baseline, and the proposed method in the supervised scenario. Our results show that temporal attention enhances the zero-shot audio classification performance in multi-label scenario.
【22】 Density Adaptive Attention-based Speech Network: Enhancing Feature Understanding for Mental Health Disorders
标题: 密度自适应基于注意力的语音网络:增强对心理健康障碍的特征理解
作者:Georgios Ioannides,Adrian Kieback,Aman Chadha,Aaron Elkins
链接:点击下载PDF文件
摘要:基于语音的抑郁症检测由于其在个体中的独特表现和数据稀缺性,对自动检测提出了重大挑战。为了应对这些挑战,我们引入了DAAMAudioCNNLSTM和DAAMAudioTransformer,这是两个用于音频特征提取和抑郁检测的参数有效且可解释的模型。DAAMAudioCNNLSTM是一个新的CNN-LSTM框架,具有多头密度自适应注意力机制(DAAM),动态关注信息语音段。DAAMAudioTransformer利用Transformer编码器代替CNN-LSTM架构,结合了相同的DAAM模块,以增强注意力和可解释性。这些方法不仅增强了检测的鲁棒性和可解释性,而且还实现了最先进的性能:DAAMAudioCNNLSTM在DAIC-WOZ数据集上的F1宏得分为0.702,DAAMAudioTransformer的F1宏得分为0.72,而不像以前的方法那样在训练 验证过程中依赖于补充信息,如元音位置和说话人信息。这两种模型在利用语音信号进行抑郁症检测方面的显著可解释性和效率代表了向更可靠,临床有用的诊断工具的飞跃,在语音和心理健康护理方面取得了可喜的进步。为了促进这一领域的进一步研究,我们公开了我们的代码。摘要:Speech-based depression detection poses significant challenges for automated detection due to its unique manifestation across individuals and data scarcity. Addressing these challenges, we introduce DAAMAudioCNNLSTM and DAAMAudioTransformer, two parameter efficient and explainable models for audio feature extraction and depression detection. DAAMAudioCNNLSTM features a novel CNN-LSTM framework with multi-head Density Adaptive Attention Mechanism (DAAM), focusing dynamically on informative speech segments. DAAMAudioTransformer, leveraging a transformer encoder in place of the CNN-LSTM architecture, incorporates the same DAAM module for enhanced attention and interpretability. These approaches not only enhance detection robustness and interpretability but also achieve state-of-the-art performance: DAAMAudioCNNLSTM with an F1 macro score of 0.702 and DAAMAudioTransformer with an F1 macro score of 0.72 on the DAIC-WOZ dataset, without reliance on supplementary information such as vowel positions and speaker information during training validation as in previous approaches. Both models' significant explainability and efficiency in leveraging speech signals for depression detection represent a leap towards more reliable, clinically useful diagnostic tools, promising advancements in speech and mental health care. To foster further research in this domain, we make our code publicly available.
【23】 Contrastive Augmentation: An Unsupervised Learning Approach for Keyword Spotting in Speech Technology
标题: 对比增强:语音技术中关键词发现的无监督学习方法
作者:Weinan Dai,Yifeng Jiang,Yuanjing Liu,Jinkun Chen,Xin Sun,Jinglei Tao
备注:This paper has been accepted by the ICPR2024
链接:点击下载PDF文件
摘要:本文讨论了持续存在的挑战,在关键字定位(KWS),语音技术的基本组成部分,关于大量的标记数据的采集进行训练。鉴于难以获得大量的阳性样本和收集新的目标样本时,关键字变化的费力的过程中,我们介绍了一种新的方法相结合的无监督对比学习和一个独特的增强为基础的技术。我们的方法允许神经网络在未标记的数据集上进行训练,从而可能提高具有有限标记数据集的下游任务的性能。我们还建议,类似的高层次的功能表示应采用相同的关键字的语音话语,尽管速度或音量的变化。为了实现这一目标,我们提出了一种基于语音增强的无监督学习方法,利用瓶颈层特征和音频重构信息之间的相似性进行辅助训练。此外,我们提出了一种压缩卷积架构,以解决KWS任务中潜在的冗余和非信息性信息,使模型能够同时学习本地特征并专注于长期信息。此方法在Google Speech Commands V2数据集上实现了强大的性能。受最近在符号识别和口语术语检测方面取得的进展的启发,我们的方法强调了我们在KWS中的对比学习方法的潜力和逐例查询口语术语检测策略的优势。提出的CAB-KWS提供了新的视角在该领域的KWS,展示了有效的方法来减少数据收集的努力,提高系统的鲁棒性。摘要:This paper addresses the persistent challenge in Keyword Spotting (KWS), a fundamental component in speech technology, regarding the acquisition of substantial labeled data for training. Given the difficulty in obtaining large quantities of positive samples and the laborious process of collecting new target samples when the keyword changes, we introduce a novel approach combining unsupervised contrastive learning and a unique augmentation-based technique. Our method allows the neural network to train on unlabeled data sets, potentially improving performance in downstream tasks with limited labeled data sets. We also propose that similar high-level feature representations should be employed for speech utterances with the same keyword despite variations in speed or volume. To achieve this, we present a speech augmentation-based unsupervised learning method that utilizes the similarity between the bottleneck layer feature and the audio reconstructing information for auxiliary training. Furthermore, we propose a compressed convolutional architecture to address potential redundancy and non-informative information in KWS tasks, enabling the model to simultaneously learn local features and focus on long-term information. This method achieves strong performance on the Google Speech Commands V2 Dataset. Inspired by recent advancements in sign spotting and spoken term detection, our method underlines the potential of our contrastive learning approach in KWS and the advantages of Query-by-Example Spoken Term Detection strategies. The presented CAB-KWS provide new perspectives in the field of KWS, demonstrating effective ways to reduce data collection efforts and increase the system's robustness.
【24】 REFFLY: Melody-Constrained Lyrics Editing Model
标题: REFFLY:旋律约束歌词编辑模型
作者:Songyan Zhao,Bingxuan Li,Yufei Tian,Nanyun Peng
链接:点击下载PDF文件
摘要:自动旋律到歌词生成旨在生成与给定旋律一致的歌词。虽然以前的作品可以基于高级控制信号(如关键字或流派)生成歌词,但它们通常面临三个挑战:(1)缺乏可控性,因为以前的作品只能从头开始生成歌词,很少或根本无法控制内容;(2)无法生成具有所需格式的完全结构化的歌曲;以及(3)未能将歌词中的突出词与旋律中的突出音符对齐,从而导致不良的歌词-旋律对齐。在这项工作中,我们引入了REFFLY(歌词修订框架),这是第一个修订框架,旨在将任意形式的纯文本草稿编辑为高质量、完整的歌曲歌词。我们的方法可以确保生成的歌词保留了草稿的原始含义,与旋律保持一致,并坚持所需的歌曲结构。我们证明,REFFLY表现良好,在不同的任务设置,如歌词修改和歌曲翻译。实验结果表明,我们的模型在音乐性和文本质量方面都比Lyra(Tian et al. 2023)和GPT-4等强基线高出25%。摘要:Automatic melody-to-lyric generation aims to produce lyrics that align with a given melody. Although previous work can generate lyrics based on high-level control signals, such as keywords or genre, they often struggle with three challenges: (1) lack of controllability, as prior works are only able to produce lyrics from scratch, with little or no control over the content; (2) inability to generate fully structured songs with the desired format; and (3) failure to align prominent words in the lyrics with prominent notes in the melody, resulting in poor lyrics-melody alignment. In this work, we introduce REFFLY (REvision Framework For Lyrics), the first revision framework designed to edit arbitrary forms of plain text draft into high-quality, full-fledged song lyrics. Our approach ensures that the generated lyrics retain the original meaning of the draft, align with the melody, and adhere to the desired song structures. We demonstrate that REFFLY performs well in diverse task settings, such as lyrics revision and song translation. Experimental results show that our model outperforms strong baselines, such as Lyra (Tian et al. 2023) and GPT-4, by 25% in both musicality and text quality.
【25】 Towards a dynamical model of English vowels. Evidence from diphthongisation
标题: 迈向英语元音的动态模型。双元音化的证据
作者:Patrycja Strycharczuk,Sam Kirkham,Emily Gorman,Takayuki Nagamine
链接:点击下载PDF文件
摘要:双母音元音表现出一定程度的内在动态变化,其程度可以在共时和历时上变化,例如双母音元音可以变成单母音,反之亦然。模拟这种类型的变化需要定义与单元音相对的双元音。然而,制定一个明确的定义已被证明是难以捉摸的声学和清晰度,因为双元音化往往是梯度在这些领域。本研究从发音的角度来探讨双母音元音是否构成一个连贯的语音范畴。我们目前的发音和声学数据从六个扬声器的北部盎格鲁英语生产一套完整的语音长元音。我们分析了几个措施的双元音化,所有这些都表明,双元音是不是绝对不同的长monophthongs。我们考虑到这一观察与发音语音 任务动态模型,其中双元音和长单元音有一个共同的手势表示,包括两个发音目标在每种情况下,但他们根据手势收缩和位置的组件手势不同。我们认为,所有长元音的双目标表示是独立的语音重量的支持,以及在英国英语的历史双元音化和现今的动态元音变化的性质。摘要:Diphthong vowels exhibit a degree of inherent dynamic change, the extent of which can vary synchronically and diachronically, such that diphthong vowels can become monophthongs and vice versa. Modelling this type of change requires defining diphthongs in opposition to monophthongs. However, formulating an explicit definition has proven elusive in acoustics and articulation, as diphthongisation is often gradient in these domains. In this study, we consider whether diphthong vowels form a coherent phonetic category from the articulatory point of view. We present articulometry and acoustic data from six speakers of Northern Anglo-English producing a full set of phonologically long vowels. We analyse several measures of diphthongisation, all of which suggest that diphthongs are not categorically distinct from long monophthongs. We account for this observation with an Articulatory Phonology Task Dynamic model in which diphthongs and long monophthongs have a common gestural representation, comprising two articulatory targets in each case, but they differ according to gestural constriction and location of the component gestures. We argue that a two-target representation for all long vowels is independently supported by phonological weight, as well as by the nature of historical diphthongisation and present-day dynamic vowel variation in British English.
【26】 ProGRes: Prompted Generative Rescoring on ASR n-Best
标题: ProGRes:在ASB n-Best上进行的预定生成重新评分
作者:Ada Defne Tur,Adel Moumen,Mirco Ravanelli
备注:IEEE Spoken Language Technology Workshop
链接:点击下载PDF文件
摘要:大型语言模型(LLM)已经显示出它们通过有效地对波束搜索过程中生成的n个最佳假设进行重新评分来提高语音识别器性能的能力。然而,最好的方式来利用最新的生成抑制调整LLM的假设重新评分仍然不清楚。本文提出了一种新的方法,使用预调LLM动态扩展的n-最好的语音识别假设与新的假设产生通过适当的提示LLM。具体而言,我们引入了一种新的zero-shot方法,用于ASR n-最佳重新评分,该方法结合了置信度评分、LLM序列评分和基于序列的假设生成。我们比较了Llama-3-Instruct,GPT-3.5 Turbo和GPT-4 Turbo作为基于序列的生成器,Llama-3作为序列评分器LLM。我们评估了我们的方法,使用不同的语音识别器,并观察到显着的相对改善的字错误率(WER)从5%到25%不等。摘要:Large Language Models (LLMs) have shown their ability to improve the performance of speech recognizers by effectively rescoring the n-best hypotheses generated during the beam search process. However, the best way to exploit recent generative instruction-tuned LLMs for hypothesis rescoring is still unclear. This paper proposes a novel method that uses instruction-tuned LLMs to dynamically expand the n-best speech recognition hypotheses with new hypotheses generated through appropriately-prompted LLMs. Specifically, we introduce a new zero-shot method for ASR n-best rescoring, which combines confidence scores, LLM sequence scoring, and prompt-based hypothesis generation. We compare Llama-3-Instruct, GPT-3.5 Turbo, and GPT-4 Turbo as prompt-based generators with Llama-3 as sequence scorer LLM. We evaluated our approach using different speech recognizers and observed significant relative improvement in the word error rate (WER) ranging from 5% to 25%.
【27】 Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning
标题: 使用谱-时态图专注池和多任务学习的逐例查询关键词发现
作者:Zhenyu Wang,Shuyu Kong,Li Wan,Biqiao Zhang,Yiteng Huang,Mumin Jin,Ming Sun,Xin Lei,Zhaojun Yang
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:现有的关键词识别(KWS)系统主要依赖于预定义的关键词短语。然而,识别定制关键字的能力对于定制与智能设备的交互至关重要。在本文中,我们提出了一种新的查询的例子(QbyE)KWS系统,采用频谱-时间图注意池和多任务学习。该框架旨在有效地学习QbyE KWS任务的说话人不变和语言信息嵌入。在这个框架内,我们研究了三种不同的编码器建模网络架构:LiCoNet,Conformer和ECAPA_TDNN。大量的内部数据集上的实验结果$629$扬声器已经证明了所提出的QbyE框架的有效性,在最大限度地发挥潜力的更简单的模型,如LiCoNet。特别是,LiCoNet的效率提高了13倍,实现了与计算密集型Conformer模型相当的性能(在0.3 FA Hr时,FRR为1.98% vs. 1.63 %)。摘要:Existing keyword spotting (KWS) systems primarily rely on predefined keyword phrases. However, the ability to recognize customized keywords is crucial for tailoring interactions with intelligent devices. In this paper, we present a novel Query-by-Example (QbyE) KWS system that employs spectral-temporal graph attentive pooling and multi-task learning. This framework aims to effectively learn speaker-invariant and linguistic-informative embeddings for QbyE KWS tasks. Within this framework, we investigate three distinct network architectures for encoder modeling: LiCoNet, Conformer and ECAPA_TDNN. The experimental results on a substantial internal dataset of $629$ speakers have demonstrated the effectiveness of the proposed QbyE framework in maximizing the potential of simpler models such as LiCoNet. Particularly, LiCoNet, which is 13x more efficient, achieves comparable performance to the computationally intensive Conformer model (1.98% vs. 1.63 % FRR at 0.3 FAs Hr).
【28】 The USTC-NERCSLIP Systems for the CHiME-8 NOTSOFAR-1 Challenge
标题: CHiME-8诺丁索FAR-1挑战赛的USTC-NERCSLIP系统
作者:Shutong Niu,Ruoyu Wang,Jun Du,Gaobin Yang,Yanhui Tu,Siyuan Wu,Shuangqing Qian,Huaxin Wu,Haitao Xu,Xueyang Zhang,Guolong Zhong,Xindi Yu,Jieru Chen,Mengzhi Wang,Di Cai,Tian Gao,Genshun Wan,Feng Ma,Jia Pan,Jianqing Gao
链接:点击下载PDF文件
摘要:本技术报告概述了CHiME-8 NOTSOFAR-1挑战赛的提交系统。这个挑战的主要困难是在各个会议室记录的数据集,它捕捉了现实世界的复杂性,例如高重叠率,背景噪音,可变数量的发言者和自然的对话风格。为了解决这些问题,我们在几个方面对系统进行了优化:对于前端语音信号处理,我们引入了数据驱动的日志化和分离联合训练方法(JDS)来提高音频质量。此外,我们还集成了传统的引导源分离(GSS)的多通道跟踪,为JDS提供补充信息。对于后端语音识别,我们使用WavLM、ConvNeXt和Transformer创新技术增强了Whisper,并应用了多任务训练和Noise KLD增强技术,从而显著提高了ASR的鲁棒性和准确性。我们的系统在CHiME-8 NOTSOFAR-1 Dev-set-2多通道和单通道轨道上分别获得了14.265%和22.989%的时间约束最小排列字错误率(tcpWER)。摘要:This technical report outlines our submission system for the CHiME-8 NOTSOFAR-1 Challenge. The primary difficulty of this challenge is the dataset recorded across various conference rooms, which captures real-world complexities such as high overlap rates, background noises, a variable number of speakers, and natural conversation styles. To address these issues, we optimized the system in several aspects: For front-end speech signal processing, we introduced a data-driven joint training method for diarization and separation (JDS) to enhance audio quality. Additionally, we also integrated traditional guided source separation (GSS) for multi-channel track to provide complementary information for the JDS. For back-end speech recognition, we enhanced Whisper with WavLM, ConvNeXt, and Transformer innovations, applying multi-task training and Noise KLD augmentation, to significantly advance ASR robustness and accuracy. Our system attained a Time-Constrained minimum Permutation Word Error Rate (tcpWER) of 14.265% and 22.989% on the CHiME-8 NOTSOFAR-1 Dev-set-2 multi-channel and single-channel tracks, respectively.
【29】 vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
标题: vec 2wav 2.0:通过离散令牌声码器推进语音转换
作者:Yiwei Guo,Zhihan Li,Junjie Li,Chenpeng Du,Hankun Wang,Shuai Wang,Xie Chen,Kai Yu
备注:5 pages, 4 figures
链接:点击下载PDF文件
摘要:我们提出了一种新的语音离散令牌声码器,vec2wav 2.0,它推进语音转换(VC)。我们使用来自语音自监督模型的离散标记作为源语音的内容特征,并将VC视为提示声码任务。为了修正内容令牌中扬声器音色的损失,vec2wav 2.0利用WavLM功能提供强音色相关信息。提出了一种新的自适应Snake激活函数,以更好地将音色纳入波形重建过程。通过这种方式,vec2wav 2.0学会在不同的参考提示下适当地改变扬声器音色。此外,vec2wav 2.0不需要监督数据来有效训练。实验结果表明,vec2wav 2.0在任何对任何VC的音频质量和扬声器相似性方面都优于所有其他基线。消融研究验证了所提出的技术的效果。此外,vec2wav 2.0即使只在单语语料库上训练,也实现了具有竞争力的跨语言VC。因此,vec2wav 2.0显示音色可能只能由语音令牌声码器操纵,推动了VC和语音合成的前沿。摘要:We propose a new speech discrete token vocoder, vec2wav 2.0, which advances voice conversion (VC). We use discrete tokens from speech self-supervised models as the content features of source speech, and treat VC as a prompted vocoding task. To amend the loss of speaker timbre in the content tokens, vec2wav 2.0 utilizes the WavLM features to provide strong timbre-dependent information. A novel adaptive Snake activation function is proposed to better incorporate timbre into the waveform reconstruction process. In this way, vec2wav 2.0 learns to alter the speaker timbre appropriately given different reference prompts. Also, no supervised data is required for vec2wav 2.0 to be effectively trained. Experimental results demonstrate that vec2wav 2.0 outperforms all other baselines to a considerable margin in terms of audio quality and speaker similarity in any-to-any VC. Ablation studies verify the effects made by the proposed techniques. Moreover, vec2wav 2.0 achieves competitive cross-lingual VC even only trained on monolingual corpus. Thus, vec2wav 2.0 shows timbre can potentially be manipulated only by speech token vocoders, pushing the frontiers of VC and speech synthesis.
【30】 Reassessing Noise Augmentation Methods in the Context of Adversarial Speech
标题: 对抗性言语背景下重新评估噪音增强方法
作者:Karla Pizzi,Matías P. Pizarro B,Asja Fischer
链接:点击下载PDF文件
摘要:在这项研究中,我们研究了噪声增强训练是否可以同时提高自动语音识别(ASR)系统的对抗鲁棒性。我们对四种不同的最先进的ASR架构的对抗鲁棒性进行了比较分析,其中每种ASR架构都在三种不同的增强条件下进行训练:一种受到背景噪声,速度变化和混响的影响,另一种只受到速度变化的影响,第三种没有任何形式的数据增强。结果表明,噪声增强不仅提高了模型对含噪语音的性能,而且提高了模型对对抗性攻击的鲁棒性。摘要:In this study, we investigate if noise-augmented training can concurrently improve adversarial robustness in automatic speech recognition (ASR) systems. We conduct a comparative analysis of the adversarial robustness of four different state-of-the-art ASR architectures, where each of the ASR architectures is trained under three different augmentation conditions: one subject to background noise, speed variations, and reverberations, another subject to speed variations only, and a third without any form of data augmentation. The results demonstrate that noise augmentation not only improves model performance on noisy speech but also the model's robustness to adversarial attacks.
【31】 Steered Response Power-Based Direction-of-Arrival Estimation Exploiting an Auxiliary Microphone
标题: 利用辅助麦克风的定向响应基于功率的到达方向估计
作者:Klaus Brümann,Simon Doclo
备注:5 pages, 3 figures, conference: EUSIPCO 2024 in Lyon
链接:点击下载PDF文件
摘要:利用紧凑型麦克风阵列(CMA)精确估计语音源的到达方向(DOA)通常会受到背景噪声和混响的影响。一种常用的DOA估计方法是转向响应功率相位变换(SRP-PHAT)函数,它已被证明是可靠的工作在中等水平的噪声和混响。由于对于紧密间隔的麦克风,噪声和混响的空间相干性在扩展的频率范围内可能很高,这可能对SRP-PHAT频谱产生负面影响,从而导致DOA估计误差。假设辅助麦克风在空间上与CMA分离的未知位置处可用,在本文中,我们提出基于辅助麦克风和CMA麦克风之间的SRP-PHAT谱来计算CMA麦克风之间的SRP-PHAT谱。对于不同级别的噪声和混响,我们展示了辅助麦克风需要在空间上与CMA分离多远,以使基于辅助麦克风的SRP-PHAT谱比没有辅助麦克风的SRP-PHAT谱更可靠。这些研究结果进行了验证的基础上模拟麦克风信号的几个辅助麦克风的位置和两个不同的噪声和混响条件。摘要:Accurately estimating the direction-of-arrival (DOA) of a speech source using a compact microphone array (CMA) is often complicated by background noise and reverberation. A commonly used DOA estimation method is the steered response power with phase transform (SRP-PHAT) function, which has been shown to work reliably in moderate levels of noise and reverberation. Since for closely spaced microphones the spatial coherence of noise and reverberation may be high over an extended frequency range, this may negatively affect the SRP-PHAT spectra, resulting in DOA estimation errors. Assuming the availability of an auxiliary microphone at an unknown position which is spatially separated from the CMA, in this paper we propose to compute the SRP-PHAT spectra between the microphones of the CMA based on the SRP-PHAT spectra between the auxiliary microphone and the microphones of the CMA. For different levels of noise and reverberation, we show how far the auxiliary microphone needs to be spatially separated from the CMA for the auxiliary microphone-based SRP-PHAT spectra to be more reliable than the SRP-PHAT spectra without the auxiliary microphone. These findings are validated based on simulated microphone signals for several auxiliary microphone positions and two different noise and reverberation conditions.
【32】 Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
标题: 多说话者ASB的语音基础模型的资源高效自适应
作者:Weiqing Wang,Kunal Dhawan,Taejin Park,Krishna C. Puvvada,Ivan Medennikov,Somshubra Majumdar,He Huang,Jagadeesh Balam,Boris Ginsburg
备注:Accepted by SLT 2024
链接:点击下载PDF文件
摘要:语音基础模型已经在各种任务中实现了最先进的(SoTA)性能,例如数百种语言的自动语音识别(ASR)。然而,由于数据稀缺性和稀疏性,多说话者ASR对于这些模型来说仍然是一项具有挑战性的任务。在本文中,我们提出的方法,使语音基础模型处理和理解有限的训练数据的多扬声器语音。具体来说,我们适应语音基础模型的多扬声器ASR任务,只使用电话数据。值得注意的是,适应模型在会议数据上也表现良好,无需任何微调,证明了我们方法的泛化能力。我们进行了几项消融研究,以分析不同参数和策略对模型性能的影响。我们的研究结果突出了我们的方法的有效性。结果表明,更少的参数得到更好的整体cpWER,这虽然违反直觉,提供了见解,以适应语音基础模型的多扬声器ASR任务与最小的注释数据。摘要:Speech foundation models have achieved state-of-the-art (SoTA) performance across various tasks, such as automatic speech recognition (ASR) in hundreds of languages. However, multi-speaker ASR remains a challenging task for these models due to data scarcity and sparsity. In this paper, we present approaches to enable speech foundation models to process and understand multi-speaker speech with limited training data. Specifically, we adapt a speech foundation model for the multi-speaker ASR task using only telephonic data. Remarkably, the adapted model also performs well on meeting data without any fine-tuning, demonstrating the generalization ability of our approach. We conduct several ablation studies to analyze the impact of different parameters and strategies on model performance. Our findings highlight the effectiveness of our methods. Results show that less parameters give better overall cpWER, which, although counter-intuitive, provides insights into adapting speech foundation models for multi-speaker ASR tasks with minimal annotated data.
【33】 Suppressing Noise Disparity in Training Data for Automatic Pathological Speech Detection
标题: 抑制训练数据中的噪音差异以实现自动病理语音检测
作者:Mahdi Amiri,Ina Kodrasi
备注:To appear in IWAENC 2024
链接:点击下载PDF文件
摘要:虽然自动病理语音检测方法显示出良好的效果时,干净的录音,他们是容易受到加性噪声。最近已经表明,通常用于开发和评估这种方法的数据库是嘈杂的,健康和病理记录之间的噪声特性是不同的。因此,在这些数据库上训练的自动方法通常学习区分噪声而不是语音病理。本文介绍了一种方法,以减轻这种噪声差异的训练数据。使用来自一组说话者的录音的噪声估计来增强来自另一组的录音,噪声特性在所有录音中变得一致。实验结果表明,这种方法在减轻训练数据中的噪声差异的有效性,从而使自动病理语音检测专注于病理判别线索,而不是噪声判别的。摘要:Although automatic pathological speech detection approaches show promising results when clean recordings are available, they are vulnerable to additive noise. Recently it has been shown that databases commonly used to develop and evaluate such approaches are noisy, with the noise characteristics between healthy and pathological recordings being different. Consequently, automatic approaches trained on these databases often learn to discriminate noise rather than speech pathology. This paper introduces a method to mitigate this noise disparity in training data. Using noise estimates from recordings from one group of speakers to augment recordings from the other group, the noise characteristics become consistent across all recordings. Experimental results demonstrate the efficacy of this approach in mitigating noise disparity in training data, thereby enabling automatic pathological speech detection to focus on pathology-discriminant cues rather than noise-discriminant ones.
【34】 EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance
标题: EnCLAP++:分析EnCLAP框架以优化自动音频字幕性能
作者:Jaeyeon Kim,Minjeon Jeon,Jaeyoon Jung,Sang Hoon Woo,Jinjoo Lee
备注:Accepted to DCASE2024 Workshop
链接:点击下载PDF文件
摘要:在这项工作中,我们的目标是分析和优化的EnCLAP框架,一个国家的最先进的模型在自动音频字幕。我们研究了修改声学编码器组件的影响,探索了不同数据集尺度的预训练,并研究了重新排序方案的有效性。通过对生成的字幕进行大量的实验和定量分析,我们开发了EnCLAP++,这是一个大大超过原始版本的增强版本。摘要:In this work, we aim to analyze and optimize the EnCLAP framework, a state-of-the-art model in automated audio captioning. We investigate the impact of modifying the acoustic encoder components, explore pretraining with different dataset scales, and study the effectiveness of a reranking scheme. Through extensive experimentation and quantitative analysis of generated captions, we develop EnCLAP++, an enhanced version that significantly surpasses the original.
【35】 Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
标题: 通过自动音频字幕的辅助检索模型扩展EnCLAP
作者:Jaeyeon Kim,Jaeyoon Jung,Minjeong Jeon,Sang Hoon Woo,Jinjoo Lee
备注:DCASE2024 Challenge Technical Report. Ranked 2nd in Task 6 Automated Audio Captioning
链接:点击下载PDF文件
摘要:在本技术报告中,我们描述了我们提交给DCASE2024挑战任务6(自动音频字幕)和任务8(基于文档的音频检索)。我们开发我们的方法建立在EnCLAP音频字幕框架和优化它的任务6的挑战。值得注意的是,我们概述了基本组件的变化和重新排序过程的纳入。此外,我们提交了一个补充检索模型,我们修改后的框架的副产品,任务8。我们提出的系统在Task6上实现了0.542的FENSE分数,在Task8上实现了0.386的mAP@10分数,显著优于基线模型。摘要:In this technical report, we describe our submission to DCASE2024 Challenge Task6 (Automated Audio Captioning) and Task8 (Language-based Audio Retrieval). We develop our approach building upon the EnCLAP audio captioning framework and optimizing it for Task6 of the challenge. Notably, we outline the changes in the underlying components and the incorporation of the reranking process. Additionally, we submit a supplementary retriever model, a byproduct of our modified framework, to Task8. Our proposed systems achieve FENSE score of 0.542 on Task6 and mAP@10 score of 0.386 on Task8, significantly outperforming the baseline models.
【36】 BUET Multi-disease Heart Sound Dataset: A Comprehensive Auscultation Dataset for Developing Computer-Aided Diagnostic Systems
标题: BUET多疾病心脏声音数据集:用于开发计算机辅助诊断系统的全面听诊数据集
作者:Shams Nafisa Ali,Afia Zahin,Samiul Based Shuvo,Nusrat Binta Nizam,Shoyad Ibn Sabur Khan Nuhash,Sayeed Sajjad Razin,S. M. Sakeef Sani,Farihin Rahman,Nawshad Binta Nizam,Farhat Binte Azam,Rakib Hossen,Sumaiya Ohab,Nawsabah Noor,Taufiq Hasan
备注:14 pages, 13 figures
链接:点击下载PDF文件
摘要:心脏听诊是诊断心血管疾病(CVD)的重要工具,通常依赖于临床医生的主观解释,在一致性和准确性方面存在局限性。为了解决这个问题,我们介绍了BUET多疾病心音(BMD-HS)数据集-一个全面而精心策划的心音记录集合。该数据集涵盖了五种不同类别的常见心音的864个记录,代表了广泛的心脏瓣膜疾病,重点关注诊断上具有挑战性的病例。BMD-HS数据集的突出特点是其创新的多标签注释系统,该系统可以捕获各种疾病和独特的疾病状态。该系统显著增强了数据集在自动心音分类和诊断中开发高级机器学习模型的实用性。通过弥合传统听诊实践与当代数据驱动诊断方法之间的差距,BMD-HS数据集有望彻底改变CVD诊断和管理,为心脏健康研究的发展提供宝贵的资源。该数据集可在此链接上公开获取:https: github.com mHealthBuet BMD-HS-Dataset。摘要:Cardiac auscultation, an integral tool in diagnosing cardiovascular diseases (CVDs), often relies on the subjective interpretation of clinicians, presenting a limitation in consistency and accuracy. Addressing this, we introduce the BUET Multi-disease Heart Sound (BMD-HS) dataset - a comprehensive and meticulously curated collection of heart sound recordings. This dataset, encompassing 864 recordings across five distinct classes of common heart sounds, represents a broad spectrum of valvular heart diseases, with a focus on diagnostically challenging cases. The standout feature of the BMD-HS dataset is its innovative multi-label annotation system, which captures a diverse range of diseases and unique disease states. This system significantly enhances the dataset's utility for developing advanced machine learning models in automated heart sound classification and diagnosis. By bridging the gap between traditional auscultation practices and contemporary data-driven diagnostic methods, the BMD-HS dataset is poised to revolutionize CVD diagnosis and management, providing an invaluable resource for the advancement of cardiac health research. The dataset is publicly available at this link: https: github.com mHealthBuet BMD-HS-Dataset.
【37】 Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
标题: 视听人员识别与验证的情态融合方法比较分析
作者:Aref Farhadipour,Masoumeh Chapariniya,Teodora Vukovic,Volker Dellwo
备注:This paper has been submitted to a conference
链接:点击下载PDF文件
摘要:多模式学习包括整合来自各种模式的信息,以加强学习和理解。通过对语音和人脸两种模态的处理,比较了三种模态融合策略在身份识别和验证中的应用。在本文中,一维卷积神经网络用于从语音中提取x向量,而预训练的VGGFace2网络和迁移学习用于人脸模态。此外,gammatonegram被用作与Darknet 19预训练网络进行交互的语音表示。建议的系统进行评估,使用K折交叉验证技术的118扬声器的测试集的VoxCeleb2数据集。在同等条件下,对单模态和三种拟议的多模态策略进行了比较评估。结果表明,伽玛射线图和面部特征的特征融合策略在人物识别任务中性能最高,准确率达到98.37%。然而,在验证任务中,将面部特征与x向量连接起来的EER达到0.62%。摘要:Multimodal learning involves integrating information from various modalities to enhance learning and comprehension. We compare three modality fusion strategies in person identification and verification by processing two modalities: voice and face. In this paper, a one-dimensional convolutional neural network is employed for x-vector extraction from voice, while the pre-trained VGGFace2 network and transfer learning are utilized for face modality. In addition, gammatonegram is used as speech representation in engagement with the Darknet19 pre-trained network. The proposed systems are evaluated using the K-fold cross-validation technique on the 118 speakers of the test set of the VoxCeleb2 dataset. The comparative evaluations are done for single-modality and three proposed multimodal strategies in equal situations. Results demonstrate that the feature fusion strategy of gammatonegram and facial features achieves the highest performance, with an accuracy of 98.37% in the person identification task. However, concatenating facial features with the x-vector reaches 0.62% for EER in verification tasks.
【38】 Digit Recognition using Multimodal Spiking Neural Networks
标题: 使用多峰尖峰神经网络的数字识别
作者:William Bjorndahl,Jack Easton,Austin Modoff,Eric C. Larson,Joseph Camp,Prasanna Rangarajan
备注:4 pages, 2 figures, submitted to 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing
链接:点击下载PDF文件
摘要:尖峰神经网络(SNN)是第三代神经网络,它受到生物学的启发,以模拟大脑中信号交换的方式处理数据。在计算机视觉领域,SNN已经获得了极大的关注,这在很大程度上是由于基于事件的传感器的可用性,该传感器响应于场景辐射的变化而产生空间分辨的尖峰训练。SNN由于其神经形态性质而用于处理基于事件的数据。拟议的工作探讨神经形态的优势,融合多个感官输入分类任务。具体来说,我们研究了SNN在数字分类中的性能,通过从使用基于事件的传感器创建的数据集传递视觉模态分支(Neuromorphic-MNIST [N-MNIST])和听觉模态分支(Spiking Heidelberg Digits [SHD])来生成一系列时间依赖事件。据观察,多模态SNN优于单峰视觉和单峰听觉SNN。此外,据观察,感觉融合的过程是不敏感的视觉和听觉分支相结合的深度。这项工作在N-MNIST和SHD数据集上实现了98.43%的准确性,使用多模态SNN连接后期深度的视觉和听觉分支。摘要:Spiking neural networks (SNNs) are the third generation of neural networks that are biologically inspired to process data in a fashion that emulates the exchange of signals in the brain. Within the Computer Vision community SNNs have garnered significant attention due in large part to the availability of event-based sensors that produce a spatially resolved spike train in response to changes in scene radiance. SNNs are used to process event-based data due to their neuromorphic nature. The proposed work examines the neuromorphic advantage of fusing multiple sensory inputs in classification tasks. Specifically we study the performance of a SNN in digit classification by passing in a visual modality branch (Neuromorphic-MNIST [N-MNIST]) and an auditory modality branch (Spiking Heidelberg Digits [SHD]) from datasets that were created using event-based sensors to generate a series of time-dependent events. It is observed that multi-modal SNNs outperform unimodal visual and unimodal auditory SNNs. Furthermore, it is observed that the process of sensory fusion is insensitive to the depth at which the visual and auditory branches are combined. This work achieves a 98.43% accuracy on the combined N-MNIST and SHD dataset using a multimodal SNN that concatenates the visual and auditory branches at a late depth.
【39】 DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
标题: DCIM-AVSR:通过双适形器交互模块实现高效的视听语音识别
作者:Xinyu Wang,Qian Wang
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:语音识别是使机器能够解释和处理人类语音的技术,将口头语言转换为文本或命令。这项技术对于虚拟助理、转录服务和通信工具等应用程序至关重要。视听语音识别(AVSR)模型通过结合嘴唇运动和面部表情等视觉模态,增强了传统的语音识别,特别是在嘈杂的环境中。虽然在具有众多参数的大规模数据集上训练的传统AVSR模型可以实现卓越的准确性,通常超过人类的表现,但它们也带来了高昂的训练成本和部署挑战。为了解决这些问题,我们引入了一个有效的AVSR模型,减少了参数的数量,通过集成的双构象相互作用模块(DCIM)。此外,我们提出了一种预训练方法,通过选择性地更新参数来进一步优化模型性能,从而显著提高效率。与传统的模型,需要系统独立学习音频和视觉模态之间的层次关系,我们的方法将这种区别直接纳入模型架构。这种设计提高了效率和性能,从而为AVSR任务提供了更实用和有效的解决方案。摘要:Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription services, and communication tools. The Audio-Visual Speech Recognition (AVSR) model enhances traditional speech recognition, particularly in noisy environments, by incorporating visual modalities like lip movements and facial expressions. While traditional AVSR models trained on large-scale datasets with numerous parameters can achieve remarkable accuracy, often surpassing human performance, they also come with high training costs and deployment challenges. To address these issues, we introduce an efficient AVSR model that reduces the number of parameters through the integration of a Dual Conformer Interaction Module (DCIM). In addition, we propose a pre-training method that further optimizes model performance by selectively updating parameters, leading to significant improvements in efficiency. Unlike conventional models that require the system to independently learn the hierarchical relationship between audio and visual modalities, our approach incorporates this distinction directly into the model architecture. This design enhances both efficiency and performance, resulting in a more practical and effective solution for AVSR tasks.
【40】 Progressive Residual Extraction based Pre-training for Speech Representation Learning
标题: 基于渐进剩余提取的预训练的语音表示学习
作者:Tianrui Wang,Jin Li,Ziyang Ma,Rui Cao,Xie Chen,Longbiao Wang,Meng Ge,Xiaobao Wang,Yuguang Wang,Jianwu Dang,Nyima Tashi
链接:点击下载PDF文件
摘要:自监督学习(SSL)在语音处理中获得了极大的关注,在语音识别等语言任务中表现出色。然而,联合提高预训练模型在各种下游任务上的性能,每个任务都需要不同的语音信息,这带来了巨大的挑战。为此,我们提出了一个渐进的残差提取为基础的自监督学习方法,命名为ProgRE。具体来说,我们引入了两个轻量级的和专门的任务模块到编码器风格的SSL骨干,以提高其能力,从语音中提取音高变化和扬声器信息。此外,为了防止增强音高变化和说话人信息与无关内容信息学习的干扰,我们从主分支中剩余地移除由这两个模块提取的信息。然后使用HuBERT的语音掩蔽预测来训练主分支,以确保Transformer的深层功能在内容任务上的性能。通过这种方式,我们可以逐步从输入语音中提取音高变化、说话人和内容表示。最后,我们可以使用不同的层权重将多种表示与不同的语音信息相结合,以获得各种下游任务的特定于任务的表示。实验结果表明,我们提出的方法实现了联合性能的改善,在各种任务,如说话人识别,语音识别,情感识别,语音增强,语音转换,相比优秀的SSL方法,如wav2vec2.0,HuBERT,WavLM。摘要:Self-supervised learning (SSL) has garnered significant attention in speech processing, excelling in linguistic tasks such as speech recognition. However, jointly improving the performance of pre-trained models on various downstream tasks, each requiring different speech information, poses significant challenges. To this purpose, we propose a progressive residual extraction based self-supervised learning method, named ProgRE. Specifically, we introduce two lightweight and specialized task modules into an encoder-style SSL backbone to enhance its ability to extract pitch variation and speaker information from speech. Furthermore, to prevent the interference of reinforced pitch variation and speaker information with irrelevant content information learning, we residually remove the information extracted by these two modules from the main branch. The main branch is then trained using HuBERT's speech masking prediction to ensure the performance of the Transformer's deep-layer features on content tasks. In this way, we can progressively extract pitch variation, speaker, and content representations from the input speech. Finally, we can combine multiple representations with diverse speech information using different layer weights to obtain task-specific representations for various downstream tasks. Experimental results indicate that our proposed method achieves joint performance improvements on various tasks, such as speaker identification, speech recognition, emotion recognition, speech enhancement, and voice conversion, compared to excellent SSL methods such as wav2vec2.0, HuBERT, and WavLM.
eess.AS音频处理
【1】 The USTC-NERCSLIP Systems for the CHiME-8 NOTSOFAR-1 Challenge标题: CHiME-8诺丁索FAR-1挑战赛的USTC-NERCSLIP系统
作者:Shutong Niu,Ruoyu Wang,Jun Du,Gaobin Yang,Yanhui Tu,Siyuan Wu,Shuangqing Qian,Huaxin Wu,Haitao Xu,Xueyang Zhang,Guolong Zhong,Xindi Yu,Jieru Chen,Mengzhi Wang,Di Cai,Tian Gao,Genshun Wan,Feng Ma,Jia Pan,Jianqing Gao
链接:点击下载PDF文件
摘要:本技术报告概述了CHiME-8 NOTSOFAR-1挑战赛的提交系统。这个挑战的主要困难是在各个会议室记录的数据集,它捕捉了现实世界的复杂性,例如高重叠率,背景噪音,可变数量的发言者和自然的对话风格。为了解决这些问题,我们在几个方面对系统进行了优化:对于前端语音信号处理,我们引入了数据驱动的日志化和分离联合训练方法(JDS)来提高音频质量。此外,我们还集成了传统的引导源分离(GSS)的多通道跟踪,为JDS提供补充信息。对于后端语音识别,我们使用WavLM、ConvNeXt和Transformer创新技术增强了Whisper,并应用了多任务训练和Noise KLD增强技术,从而显著提高了ASR的鲁棒性和准确性。我们的系统在CHiME-8 NOTSOFAR-1 Dev-set-2多通道和单通道轨道上分别获得了14.265%和22.989%的时间约束最小排列字错误率(tcpWER)。摘要:This technical report outlines our submission system for the CHiME-8 NOTSOFAR-1 Challenge. The primary difficulty of this challenge is the dataset recorded across various conference rooms, which captures real-world complexities such as high overlap rates, background noises, a variable number of speakers, and natural conversation styles. To address these issues, we optimized the system in several aspects: For front-end speech signal processing, we introduced a data-driven joint training method for diarization and separation (JDS) to enhance audio quality. Additionally, we also integrated traditional guided source separation (GSS) for multi-channel track to provide complementary information for the JDS. For back-end speech recognition, we enhanced Whisper with WavLM, ConvNeXt, and Transformer innovations, applying multi-task training and Noise KLD augmentation, to significantly advance ASR robustness and accuracy. Our system attained a Time-Constrained minimum Permutation Word Error Rate (tcpWER) of 14.265% and 22.989% on the CHiME-8 NOTSOFAR-1 Dev-set-2 multi-channel and single-channel tracks, respectively.
【2】 vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
标题: vec 2wav 2.0:通过离散令牌声码器推进语音转换
作者:Yiwei Guo,Zhihan Li,Junjie Li,Chenpeng Du,Hankun Wang,Shuai Wang,Xie Chen,Kai Yu
备注:5 pages, 4 figures
链接:点击下载PDF文件
摘要:我们提出了一种新的语音离散令牌声码器,vec2wav 2.0,它推进语音转换(VC)。我们使用来自语音自监督模型的离散标记作为源语音的内容特征,并将VC视为提示声码任务。为了修正内容令牌中扬声器音色的损失,vec2wav 2.0利用WavLM功能提供强音色相关信息。提出了一种新的自适应Snake激活函数,以更好地将音色纳入波形重建过程。通过这种方式,vec2wav 2.0学会在不同的参考提示下适当地改变扬声器音色。此外,vec2wav 2.0不需要监督数据来有效训练。实验结果表明,vec2wav 2.0在任何对任何VC的音频质量和扬声器相似性方面都优于所有其他基线。消融研究验证了所提出的技术的效果。此外,vec2wav 2.0即使只在单语语料库上训练,也实现了具有竞争力的跨语言VC。因此,vec2wav 2.0显示音色可能只能由语音令牌声码器操纵,推动了VC和语音合成的前沿。摘要:We propose a new speech discrete token vocoder, vec2wav 2.0, which advances voice conversion (VC). We use discrete tokens from speech self-supervised models as the content features of source speech, and treat VC as a prompted vocoding task. To amend the loss of speaker timbre in the content tokens, vec2wav 2.0 utilizes the WavLM features to provide strong timbre-dependent information. A novel adaptive Snake activation function is proposed to better incorporate timbre into the waveform reconstruction process. In this way, vec2wav 2.0 learns to alter the speaker timbre appropriately given different reference prompts. Also, no supervised data is required for vec2wav 2.0 to be effectively trained. Experimental results demonstrate that vec2wav 2.0 outperforms all other baselines to a considerable margin in terms of audio quality and speaker similarity in any-to-any VC. Ablation studies verify the effects made by the proposed techniques. Moreover, vec2wav 2.0 achieves competitive cross-lingual VC even only trained on monolingual corpus. Thus, vec2wav 2.0 shows timbre can potentially be manipulated only by speech token vocoders, pushing the frontiers of VC and speech synthesis.
【3】 Reassessing Noise Augmentation Methods in the Context of Adversarial Speech
标题: 对抗性言语背景下重新评估噪音增强方法
作者:Karla Pizzi,Matías P. Pizarro B,Asja Fischer
链接:点击下载PDF文件
摘要:在这项研究中,我们研究了噪声增强训练是否可以同时提高自动语音识别(ASR)系统的对抗鲁棒性。我们对四种不同的最先进的ASR架构的对抗鲁棒性进行了比较分析,其中每种ASR架构都在三种不同的增强条件下进行训练:一种受到背景噪声,速度变化和混响的影响,另一种只受到速度变化的影响,第三种没有任何形式的数据增强。结果表明,噪声增强不仅提高了模型对含噪语音的性能,而且提高了模型对对抗性攻击的鲁棒性。摘要:In this study, we investigate if noise-augmented training can concurrently improve adversarial robustness in automatic speech recognition (ASR) systems. We conduct a comparative analysis of the adversarial robustness of four different state-of-the-art ASR architectures, where each of the ASR architectures is trained under three different augmentation conditions: one subject to background noise, speed variations, and reverberations, another subject to speed variations only, and a third without any form of data augmentation. The results demonstrate that noise augmentation not only improves model performance on noisy speech but also the model's robustness to adversarial attacks.
【4】 Steered Response Power-Based Direction-of-Arrival Estimation Exploiting an Auxiliary Microphone
标题: 利用辅助麦克风的定向响应基于功率的到达方向估计
作者:Klaus Brümann,Simon Doclo
备注:5 pages, 3 figures, conference: EUSIPCO 2024 in Lyon
链接:点击下载PDF文件
摘要:利用紧凑型麦克风阵列(CMA)精确估计语音源的到达方向(DOA)通常会受到背景噪声和混响的影响。一种常用的DOA估计方法是转向响应功率相位变换(SRP-PHAT)函数,它已被证明是可靠的工作在中等水平的噪声和混响。由于对于紧密间隔的麦克风,噪声和混响的空间相干性在扩展的频率范围内可能很高,这可能对SRP-PHAT频谱产生负面影响,从而导致DOA估计误差。假设辅助麦克风在空间上与CMA分离的未知位置处可用,在本文中,我们提出基于辅助麦克风和CMA麦克风之间的SRP-PHAT谱来计算CMA麦克风之间的SRP-PHAT谱。对于不同级别的噪声和混响,我们展示了辅助麦克风需要在空间上与CMA分离多远,以使基于辅助麦克风的SRP-PHAT谱比没有辅助麦克风的SRP-PHAT谱更可靠。这些研究结果进行了验证的基础上模拟麦克风信号的几个辅助麦克风的位置和两个不同的噪声和混响条件。摘要:Accurately estimating the direction-of-arrival (DOA) of a speech source using a compact microphone array (CMA) is often complicated by background noise and reverberation. A commonly used DOA estimation method is the steered response power with phase transform (SRP-PHAT) function, which has been shown to work reliably in moderate levels of noise and reverberation. Since for closely spaced microphones the spatial coherence of noise and reverberation may be high over an extended frequency range, this may negatively affect the SRP-PHAT spectra, resulting in DOA estimation errors. Assuming the availability of an auxiliary microphone at an unknown position which is spatially separated from the CMA, in this paper we propose to compute the SRP-PHAT spectra between the microphones of the CMA based on the SRP-PHAT spectra between the auxiliary microphone and the microphones of the CMA. For different levels of noise and reverberation, we show how far the auxiliary microphone needs to be spatially separated from the CMA for the auxiliary microphone-based SRP-PHAT spectra to be more reliable than the SRP-PHAT spectra without the auxiliary microphone. These findings are validated based on simulated microphone signals for several auxiliary microphone positions and two different noise and reverberation conditions.
【5】 Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
标题: 多说话者ASB的语音基础模型的资源高效自适应
作者:Weiqing Wang,Kunal Dhawan,Taejin Park,Krishna C. Puvvada,Ivan Medennikov,Somshubra Majumdar,He Huang,Jagadeesh Balam,Boris Ginsburg
备注:Accepted by SLT 2024
链接:点击下载PDF文件
摘要:语音基础模型已经在各种任务中实现了最先进的(SoTA)性能,例如数百种语言的自动语音识别(ASR)。然而,由于数据稀缺性和稀疏性,多说话者ASR对于这些模型来说仍然是一项具有挑战性的任务。在本文中,我们提出的方法,使语音基础模型处理和理解有限的训练数据的多扬声器语音。具体来说,我们适应语音基础模型的多扬声器ASR任务,只使用电话数据。值得注意的是,适应模型在会议数据上也表现良好,无需任何微调,证明了我们方法的泛化能力。我们进行了几项消融研究,以分析不同参数和策略对模型性能的影响。我们的研究结果突出了我们的方法的有效性。结果表明,更少的参数得到更好的整体cpWER,这虽然违反直觉,提供了见解,以适应语音基础模型的多扬声器ASR任务与最小的注释数据。摘要:Speech foundation models have achieved state-of-the-art (SoTA) performance across various tasks, such as automatic speech recognition (ASR) in hundreds of languages. However, multi-speaker ASR remains a challenging task for these models due to data scarcity and sparsity. In this paper, we present approaches to enable speech foundation models to process and understand multi-speaker speech with limited training data. Specifically, we adapt a speech foundation model for the multi-speaker ASR task using only telephonic data. Remarkably, the adapted model also performs well on meeting data without any fine-tuning, demonstrating the generalization ability of our approach. We conduct several ablation studies to analyze the impact of different parameters and strategies on model performance. Our findings highlight the effectiveness of our methods. Results show that less parameters give better overall cpWER, which, although counter-intuitive, provides insights into adapting speech foundation models for multi-speaker ASR tasks with minimal annotated data.
【6】 Suppressing Noise Disparity in Training Data for Automatic Pathological Speech Detection
标题: 抑制训练数据中的噪音差异以实现自动病理语音检测
作者:Mahdi Amiri,Ina Kodrasi
备注:To appear in IWAENC 2024
链接:点击下载PDF文件
摘要:虽然自动病理语音检测方法显示出良好的效果时,干净的录音,他们是容易受到加性噪声。最近已经表明,通常用于开发和评估这种方法的数据库是嘈杂的,健康和病理记录之间的噪声特性是不同的。因此,在这些数据库上训练的自动方法通常学习区分噪声而不是语音病理。本文介绍了一种方法,以减轻这种噪声差异的训练数据。使用来自一组说话者的录音的噪声估计来增强来自另一组的录音,噪声特性在所有录音中变得一致。实验结果表明,这种方法在减轻训练数据中的噪声差异的有效性,从而使自动病理语音检测专注于病理判别线索,而不是噪声判别的。摘要:Although automatic pathological speech detection approaches show promising results when clean recordings are available, they are vulnerable to additive noise. Recently it has been shown that databases commonly used to develop and evaluate such approaches are noisy, with the noise characteristics between healthy and pathological recordings being different. Consequently, automatic approaches trained on these databases often learn to discriminate noise rather than speech pathology. This paper introduces a method to mitigate this noise disparity in training data. Using noise estimates from recordings from one group of speakers to augment recordings from the other group, the noise characteristics become consistent across all recordings. Experimental results demonstrate the efficacy of this approach in mitigating noise disparity in training data, thereby enabling automatic pathological speech detection to focus on pathology-discriminant cues rather than noise-discriminant ones.
【7】 EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance
标题: EnCLAP++:分析EnCLAP框架以优化自动音频字幕性能
作者:Jaeyeon Kim,Minjeon Jeon,Jaeyoon Jung,Sang Hoon Woo,Jinjoo Lee
备注:Accepted to DCASE2024 Workshop
链接:点击下载PDF文件
摘要:在这项工作中,我们的目标是分析和优化的EnCLAP框架,一个国家的最先进的模型在自动音频字幕。我们研究了修改声学编码器组件的影响,探索了不同数据集尺度的预训练,并研究了重新排序方案的有效性。通过对生成的字幕进行大量的实验和定量分析,我们开发了EnCLAP++,这是一个大大超过原始版本的增强版本。摘要:In this work, we aim to analyze and optimize the EnCLAP framework, a state-of-the-art model in automated audio captioning. We investigate the impact of modifying the acoustic encoder components, explore pretraining with different dataset scales, and study the effectiveness of a reranking scheme. Through extensive experimentation and quantitative analysis of generated captions, we develop EnCLAP++, an enhanced version that significantly surpasses the original.
【8】 Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
标题: 通过自动音频字幕的辅助检索模型扩展EnCLAP
作者:Jaeyeon Kim,Jaeyoon Jung,Minjeong Jeon,Sang Hoon Woo,Jinjoo Lee
备注:DCASE2024 Challenge Technical Report. Ranked 2nd in Task 6 Automated Audio Captioning
链接:点击下载PDF文件
摘要:在本技术报告中,我们描述了我们提交给DCASE2024挑战任务6(自动音频字幕)和任务8(基于文档的音频检索)。我们开发我们的方法建立在EnCLAP音频字幕框架和优化它的任务6的挑战。值得注意的是,我们概述了基本组件的变化和重新排序过程的纳入。此外,我们提交了一个补充检索模型,我们修改后的框架的副产品,任务8。我们提出的系统在Task6上实现了0.542的FENSE分数,在Task8上实现了0.386的mAP@10分数,显著优于基线模型。摘要:In this technical report, we describe our submission to DCASE2024 Challenge Task6 (Automated Audio Captioning) and Task8 (Language-based Audio Retrieval). We develop our approach building upon the EnCLAP audio captioning framework and optimizing it for Task6 of the challenge. Notably, we outline the changes in the underlying components and the incorporation of the reranking process. Additionally, we submit a supplementary retriever model, a byproduct of our modified framework, to Task8. Our proposed systems achieve FENSE score of 0.542 on Task6 and mAP@10 score of 0.386 on Task8, significantly outperforming the baseline models.
【9】 BUET Multi-disease Heart Sound Dataset: A Comprehensive Auscultation Dataset for Developing Computer-Aided Diagnostic Systems
标题: BUET多疾病心脏声音数据集:用于开发计算机辅助诊断系统的全面听诊数据集
作者:Shams Nafisa Ali,Afia Zahin,Samiul Based Shuvo,Nusrat Binta Nizam,Shoyad Ibn Sabur Khan Nuhash,Sayeed Sajjad Razin,S. M. Sakeef Sani,Farihin Rahman,Nawshad Binta Nizam,Farhat Binte Azam,Rakib Hossen,Sumaiya Ohab,Nawsabah Noor,Taufiq Hasan
备注:14 pages, 13 figures
链接:点击下载PDF文件
摘要:心脏听诊是诊断心血管疾病(CVD)的重要工具,通常依赖于临床医生的主观解释,在一致性和准确性方面存在局限性。为了解决这个问题,我们介绍了BUET多疾病心音(BMD-HS)数据集-一个全面而精心策划的心音记录集合。该数据集涵盖了五种不同类别的常见心音的864个记录,代表了广泛的心脏瓣膜疾病,重点关注诊断上具有挑战性的病例。BMD-HS数据集的突出特点是其创新的多标签注释系统,该系统可以捕获各种疾病和独特的疾病状态。该系统显著增强了数据集在自动心音分类和诊断中开发高级机器学习模型的实用性。通过弥合传统听诊实践与当代数据驱动诊断方法之间的差距,BMD-HS数据集有望彻底改变CVD诊断和管理,为心脏健康研究的发展提供宝贵的资源。该数据集可在此链接上公开获取:https: github.com mHealthBuet BMD-HS-Dataset。摘要:Cardiac auscultation, an integral tool in diagnosing cardiovascular diseases (CVDs), often relies on the subjective interpretation of clinicians, presenting a limitation in consistency and accuracy. Addressing this, we introduce the BUET Multi-disease Heart Sound (BMD-HS) dataset - a comprehensive and meticulously curated collection of heart sound recordings. This dataset, encompassing 864 recordings across five distinct classes of common heart sounds, represents a broad spectrum of valvular heart diseases, with a focus on diagnostically challenging cases. The standout feature of the BMD-HS dataset is its innovative multi-label annotation system, which captures a diverse range of diseases and unique disease states. This system significantly enhances the dataset's utility for developing advanced machine learning models in automated heart sound classification and diagnosis. By bridging the gap between traditional auscultation practices and contemporary data-driven diagnostic methods, the BMD-HS dataset is poised to revolutionize CVD diagnosis and management, providing an invaluable resource for the advancement of cardiac health research. The dataset is publicly available at this link: https: github.com mHealthBuet BMD-HS-Dataset.
【10】 Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
标题: 视听人员识别与验证的情态融合方法比较分析
作者:Aref Farhadipour,Masoumeh Chapariniya,Teodora Vukovic,Volker Dellwo
备注:This paper has been submitted to a conference
链接:点击下载PDF文件
摘要:多模式学习包括整合来自各种模式的信息,以加强学习和理解。通过对语音和人脸两种模态的处理,比较了三种模态融合策略在身份识别和验证中的应用。在本文中,一维卷积神经网络用于从语音中提取x向量,而预训练的VGGFace2网络和迁移学习用于人脸模态。此外,gammatonegram被用作与Darknet 19预训练网络进行交互的语音表示。建议的系统进行评估,使用K折交叉验证技术的118扬声器的测试集的VoxCeleb2数据集。在同等条件下,对单模态和三种拟议的多模态策略进行了比较评估。结果表明,伽玛射线图和面部特征的特征融合策略在人物识别任务中性能最高,准确率达到98.37%。然而,在验证任务中,将面部特征与x向量连接起来的EER达到0.62%。摘要:Multimodal learning involves integrating information from various modalities to enhance learning and comprehension. We compare three modality fusion strategies in person identification and verification by processing two modalities: voice and face. In this paper, a one-dimensional convolutional neural network is employed for x-vector extraction from voice, while the pre-trained VGGFace2 network and transfer learning are utilized for face modality. In addition, gammatonegram is used as speech representation in engagement with the Darknet19 pre-trained network. The proposed systems are evaluated using the K-fold cross-validation technique on the 118 speakers of the test set of the VoxCeleb2 dataset. The comparative evaluations are done for single-modality and three proposed multimodal strategies in equal situations. Results demonstrate that the feature fusion strategy of gammatonegram and facial features achieves the highest performance, with an accuracy of 98.37% in the person identification task. However, concatenating facial features with the x-vector reaches 0.62% for EER in verification tasks.
【11】 Digit Recognition using Multimodal Spiking Neural Networks
标题: 使用多峰尖峰神经网络的数字识别
作者:William Bjorndahl,Jack Easton,Austin Modoff,Eric C. Larson,Joseph Camp,Prasanna Rangarajan
备注:4 pages, 2 figures, submitted to 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing
链接:点击下载PDF文件
摘要:尖峰神经网络(SNN)是第三代神经网络,它受到生物学的启发,以模拟大脑中信号交换的方式处理数据。在计算机视觉领域,SNN已经获得了极大的关注,这在很大程度上是由于基于事件的传感器的可用性,该传感器响应于场景辐射的变化而产生空间分辨的尖峰训练。SNN由于其神经形态性质而用于处理基于事件的数据。拟议的工作探讨神经形态的优势,融合多个感官输入分类任务。具体来说,我们研究了SNN在数字分类中的性能,通过从使用基于事件的传感器创建的数据集传递视觉模态分支(Neuromorphic-MNIST [N-MNIST])和听觉模态分支(Spiking Heidelberg Digits [SHD])来生成一系列时间依赖事件。据观察,多模态SNN优于单峰视觉和单峰听觉SNN。此外,据观察,感觉融合的过程是不敏感的视觉和听觉分支相结合的深度。这项工作在N-MNIST和SHD数据集上实现了98.43%的准确性,使用多模态SNN连接后期深度的视觉和听觉分支。摘要:Spiking neural networks (SNNs) are the third generation of neural networks that are biologically inspired to process data in a fashion that emulates the exchange of signals in the brain. Within the Computer Vision community SNNs have garnered significant attention due in large part to the availability of event-based sensors that produce a spatially resolved spike train in response to changes in scene radiance. SNNs are used to process event-based data due to their neuromorphic nature. The proposed work examines the neuromorphic advantage of fusing multiple sensory inputs in classification tasks. Specifically we study the performance of a SNN in digit classification by passing in a visual modality branch (Neuromorphic-MNIST [N-MNIST]) and an auditory modality branch (Spiking Heidelberg Digits [SHD]) from datasets that were created using event-based sensors to generate a series of time-dependent events. It is observed that multi-modal SNNs outperform unimodal visual and unimodal auditory SNNs. Furthermore, it is observed that the process of sensory fusion is insensitive to the depth at which the visual and auditory branches are combined. This work achieves a 98.43% accuracy on the combined N-MNIST and SHD dataset using a multimodal SNN that concatenates the visual and auditory branches at a late depth.
【12】 DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
标题: DCIM-AVSR:通过双适形器交互模块实现高效的视听语音识别
作者:Xinyu Wang,Qian Wang
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:语音识别是使机器能够解释和处理人类语音的技术,将口头语言转换为文本或命令。这项技术对于虚拟助理、转录服务和通信工具等应用程序至关重要。视听语音识别(AVSR)模型通过结合嘴唇运动和面部表情等视觉模态,增强了传统的语音识别,特别是在嘈杂的环境中。虽然在具有众多参数的大规模数据集上训练的传统AVSR模型可以实现卓越的准确性,通常超过人类的表现,但它们也带来了高昂的训练成本和部署挑战。为了解决这些问题,我们引入了一个有效的AVSR模型,减少了参数的数量,通过集成的双构象相互作用模块(DCIM)。此外,我们提出了一种预训练方法,通过选择性地更新参数来进一步优化模型性能,从而显著提高效率。与传统的模型,需要系统独立学习音频和视觉模态之间的层次关系,我们的方法将这种区别直接纳入模型架构。这种设计提高了效率和性能,从而为AVSR任务提供了更实用和有效的解决方案。摘要:Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription services, and communication tools. The Audio-Visual Speech Recognition (AVSR) model enhances traditional speech recognition, particularly in noisy environments, by incorporating visual modalities like lip movements and facial expressions. While traditional AVSR models trained on large-scale datasets with numerous parameters can achieve remarkable accuracy, often surpassing human performance, they also come with high training costs and deployment challenges. To address these issues, we introduce an efficient AVSR model that reduces the number of parameters through the integration of a Dual Conformer Interaction Module (DCIM). In addition, we propose a pre-training method that further optimizes model performance by selectively updating parameters, leading to significant improvements in efficiency. Unlike conventional models that require the system to independently learn the hierarchical relationship between audio and visual modalities, our approach incorporates this distinction directly into the model architecture. This design enhances both efficiency and performance, resulting in a more practical and effective solution for AVSR tasks.
【13】 Progressive Residual Extraction based Pre-training for Speech Representation Learning
标题: 基于渐进剩余提取的预训练的语音表示学习
作者:Tianrui Wang,Jin Li,Ziyang Ma,Rui Cao,Xie Chen,Longbiao Wang,Meng Ge,Xiaobao Wang,Yuguang Wang,Jianwu Dang,Nyima Tashi
链接:点击下载PDF文件
摘要:自监督学习(SSL)在语音处理中获得了极大的关注,在语音识别等语言任务中表现出色。然而,联合提高预训练模型在各种下游任务上的性能,每个任务都需要不同的语音信息,这带来了巨大的挑战。为此,我们提出了一个渐进的残差提取为基础的自监督学习方法,命名为ProgRE。具体来说,我们引入了两个轻量级的和专门的任务模块到编码器风格的SSL骨干,以提高其能力,从语音中提取音高变化和扬声器信息。此外,为了防止增强音高变化和说话人信息与无关内容信息学习的干扰,我们从主分支中剩余地移除由这两个模块提取的信息。然后使用HuBERT的语音掩蔽预测来训练主分支,以确保Transformer的深层功能在内容任务上的性能。通过这种方式,我们可以逐步从输入语音中提取音高变化、说话人和内容表示。最后,我们可以使用不同的层权重将多种表示与不同的语音信息相结合,以获得各种下游任务的特定于任务的表示。实验结果表明,我们提出的方法实现了联合性能的改善,在各种任务,如说话人识别,语音识别,情感识别,语音增强,语音转换,相比优秀的SSL方法,如wav2vec2.0,HuBERT,WavLM。摘要:Self-supervised learning (SSL) has garnered significant attention in speech processing, excelling in linguistic tasks such as speech recognition. However, jointly improving the performance of pre-trained models on various downstream tasks, each requiring different speech information, poses significant challenges. To this purpose, we propose a progressive residual extraction based self-supervised learning method, named ProgRE. Specifically, we introduce two lightweight and specialized task modules into an encoder-style SSL backbone to enhance its ability to extract pitch variation and speaker information from speech. Furthermore, to prevent the interference of reinforced pitch variation and speaker information with irrelevant content information learning, we residually remove the information extracted by these two modules from the main branch. The main branch is then trained using HuBERT's speech masking prediction to ensure the performance of the Transformer's deep-layer features on content tasks. In this way, we can progressively extract pitch variation, speaker, and content representations from the input speech. Finally, we can combine multiple representations with diverse speech information using different layer weights to obtain task-specific representations for various downstream tasks. Experimental results indicate that our proposed method achieves joint performance improvements on various tasks, such as speaker identification, speech recognition, emotion recognition, speech enhancement, and voice conversion, compared to excellent SSL methods such as wav2vec2.0, HuBERT, and WavLM.
【14】 BELT-2: Bootstrapping EEG-to-Language representation alignment for multi-task brain decoding
标题: BELT-2:引导脑电与语言表示对齐,用于多任务大脑解码
作者:Jinzhao Zhou,Yiqun Duan,Fred Chang,Thomas Do,Yu-Kai Wang,Chin-Teng Lin
链接:点击下载PDF文件
摘要:大型语言模型(LLM)在各种多模态应用中取得了显着的成功。然而,将大型语言模型与人类或大脑动力学相结合仍然相对未被探索。在本文中,我们介绍了BELT-2,一个开创性的多任务模型,旨在提高编码和解码性能的EEG信号。为了提高EEG编码器的质量,BELT-2是第一个创新性的工作:1)采用字节对编码(BPE)级别的EEG语言对齐; 2)在EEG领域集成多任务训练和解码。受 textbf{ texttit {Bridging the Brain with GPT}}思想的启发,我们通过对EEG编码器的中间输出进行前缀调谐,进一步将多任务EEG编码器与LLM连接。这些创新努力使BELT-2成为一项开创性的突破,使其成为该领域第一个能够从非侵入性大脑信号中解码连贯且可读的句子的工作。我们的实验突出了在定量和定性测量方面优于现有技术的显著进步,在ZuCo数据集上实现了BLEU-1得分为52.2%的解码性能。此外,BELT-2在其他翻译基准上显示出从31%到162%的显着改进。代码可通过提供的匿名链接访问~ footnote{https: anonymous.4open.science r BELT-2-0048}。摘要:The remarkable success of large language models (LLMs) across various multi-modality applications is well established. However, integrating large language models with humans, or brain dynamics, remains relatively unexplored. In this paper, we introduce BELT-2, a pioneering multi-task model designed to enhance both encoding and decoding performance from EEG signals. To bolster the quality of the EEG encoder, BELT-2 is the first work to innovatively 1) adopt byte-pair encoding (BPE)-level EEG-language alignment and 2) integrate multi-task training and decoding in the EEG domain. Inspired by the idea of textbf{ textit{Bridging the Brain with GPT}}, we further connect the multi-task EEG encoder with LLMs by utilizing prefix-tuning on intermediary output from the EEG encoder. These innovative efforts make BELT-2 a pioneering breakthrough, making it the first work in the field capable of decoding coherent and readable sentences from non-invasive brain signals. Our experiments highlight significant advancements over prior techniques in both quantitative and qualitative measures, achieving a decoding performance with a BLEU-1 score of 52.2 % on the ZuCo dataset. Furthermore, BELT-2 shows a remarkable improvement ranging from 31 % to 162 % on other translation benchmarks. Codes can be accessed via the provided anonymous link~ footnote{https: anonymous.4open.science r BELT-2-0048}.
【15】 Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
标题: 利用基于LID的协作混合专家模型增强代码转换语音识别
作者:Hukai Huang,Jiayan Lin,Kaidi Wang,Yishuang Li,Wenhao Guan,Qingyang Hong,Lin Li
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
摘要:由于不同语言之间语音相似性建模的固有困难,码转换语音识别提出了一个艰巨的挑战。本研究提出了一个合作的MoE,一个混合的专家(MoE)模型,利用专家组之间的协作机制。最初,前面的路由网络显式学习语言识别(LID)任务,并根据获得的LID权重选择专家。该过程确保了到MoE层的鲁棒路由信息,减轻了来自不同语言域对专家网络参数更新的干扰。LID权重还用于促进组间协作,从而实现特定语言表示的集成。此外,在每个语言专家组内,门控网络在无监督的情况下运行,以促进语言以外属性的合作。大量的实验证明了我们的方法的有效性,实现了显着的性能增强相比,替代方法。重要的是,我们的方法保留了MoE模型的高效推理能力,而不需要额外的预训练。摘要:Due to the inherent difficulty in modeling phonetic similarities across different languages, code-switching speech recognition presents a formidable challenge. This study proposes a Collaborative-MoE, a Mixture of Experts (MoE) model that leverages a collaborative mechanism among expert groups. Initially, a preceding routing network explicitly learns Language Identification (LID) tasks and selects experts based on acquired LID weights. This process ensures robust routing information to the MoE layer, mitigating interference from diverse language domains on expert network parameter updates. The LID weights are also employed to facilitate inter-group collaboration, enabling the integration of language-specific representations. Furthermore, within each language expert group, a gating network operates unsupervised to foster collaboration on attributes beyond language. Extensive experiments demonstrate the efficacy of our approach, achieving significant performance enhancements compared to alternative methods. Importantly, our method preserves the efficient inference capabilities characteristic of MoE models without necessitating additional pre-training.
【16】 Activity-Guided Industrial Anomalous Sound Detection against Interferences
标题: 活动引导的工业异常声音抗干扰检测
作者:Yunjoo Lee,Jaechang Kim,Jungseul Ok
备注:Thsis is an extended version of this https URL
链接:点击下载PDF文件
摘要:我们解决了工业声音数据异常检测的实际情况,其中目标机器的声音被背景噪声和相邻机器的干扰所破坏。克服这一挑战是困难的,因为在没有额外信息的情况下,干扰通常与目标机器几乎无法区分。为了解决这个问题,我们提出了SSAD,一个框架的源分离(SS),其次是异常检测(AD),它利用机器的活动信息,往往很容易在实际环境中。SSAD由两个部分组成:(i)活动通知SS,即使在具有相似音色的干扰下也能实现有效的源分离,以及(ii)两步掩蔽,通过强调与机器活动一致的异常来增强异常检测。我们的实验表明,SSAD实现了相当的精度与基线完全访问干净的信号,而SSAD只提供了一个损坏的信号和活动信息。此外,由于具有两步掩蔽的活动通知SS和AD,SSAD优于标准方法,特别是在干扰的情况下。它突出了SSAD在解决工业声音数据异常检测的复杂性方面的实际功效。摘要:We address a practical scenario of anomaly detection for industrial sound data, where the sound of a target machine is corrupted by background noise and interference from neighboring machines. Overcoming this challenge is difficult since the interference is often virtually indistinguishable from the target machine without additional information. To address the issue, we propose SSAD, a framework of source separation (SS) followed by anomaly detection (AD), which leverages machine activity information, often readily available in practical settings. SSAD consists of two components: (i) activity-informed SS, enabling effective source separation even given interference with similar timbre, and (ii) two-step masking, robustifying anomaly detection by emphasizing anomalies aligned with the machine activity. Our experiments demonstrate that SSAD achieves comparable accuracy to a baseline with full access to clean signals, while SSAD is provided only a corrupted signal and activity information. In addition, thanks to the activity-informed SS and AD with the two-step masking, SSAD outperforms standard approaches, particularly in cases with interference. It highlights the practical efficacy of SSAD in addressing the complexities of anomaly detection in industrial sound data.
【17】 The Role of Large Language Models in Musicology: Are We Ready to Trust the Machines?
标题: 大型语言模型在音乐学中的作用:我们准备好信任机器了吗?
作者:Pedro Ramoneda,Emilia Parada-Cabaleiro,Benno Weck,Xavier Serra
链接:点击下载PDF文件
摘要:在这项工作中,我们将探讨大型语言模型(LLM)在音乐学中的使用和可靠性。通过与专家和学生的讨论,我们评估了目前对这种如今无处不在的技术的接受程度和担忧。我们的目标是更进一步,提出一种半自动的方法来创建一个初始的基准使用检索增强生成模型和多项选择题生成,由人类专家验证。我们对400个人类验证问题的评估表明,目前的香草LLM比从音乐词典中检索增强生成的可靠性低。本文认为,在音乐学的LLM的潜力,需要音乐学驱动的研究,可以专门的LLM包括准确和可靠的领域知识。摘要:In this work, we explore the use and reliability of Large Language Models (LLMs) in musicology. From a discussion with experts and students, we assess the current acceptance and concerns regarding this, nowadays ubiquitous, technology. We aim to go one step further, proposing a semi-automatic method to create an initial benchmark using retrieval-augmented generation models and multiple-choice question generation, validated by human experts. Our evaluation on 400 human-validated questions shows that current vanilla LLMs are less reliable than retrieval augmented generation from music dictionaries. This paper suggests that the potential of LLMs in musicology requires musicology driven research that can specialized LLMs by including accurate and reliable domain knowledge.
【18】 USTC-KXDIGIT System Description for ASVspoof5 Challenge
标题: USTC-KXDIGIT ASVspoof 5挑战赛系统描述
作者:Yihao Chen,Haochen Wu,Nan Jiang,Xiang Xia,Qing Gu,Yunqi Hao,Pengfei Cai,Yu Guan,Jialong Wang,Weilin Xie,Lei Fang,Sian Fang,Yan Song,Wu Guo,Lin Liu,Minqiang Xu
备注:ASVspoof5 workshop paper
链接:点击下载PDF文件
摘要:本文描述了USTC-KXDIGIT系统提交给ASVspoof 5挑战赛的Track 1(语音深度伪造检测)和Track 2(欺骗鲁棒自动说话人验证,SASV)。轨道1展示了来自潜在处理算法的各种技术品质,包括开放和封闭条件。对于这些条件,我们的系统由前端特征提取器和后端分类器的级联组成。我们专注于广泛的嵌入工程和增强后端分类器模型的泛化。具体来说,嵌入工程是基于手工制作的功能和语音表示从一个自我监督的模型,分别用于封闭和开放的条件。为了在各种对抗条件下检测欺骗攻击,我们在增强训练集上训练了多个系统。此外,我们使用语音转换技术从训练集中的真实音频合成假音频,以丰富合成算法。为了利用不同模型架构学习到的互补信息,我们采用了来自不同系统的激活集成和融合分数,以获得欺骗检测的最终决策分数。在评估阶段,所提出的方法在封闭条件下实现了0.3948 minDCF和14.33%EER,在开放条件下实现了0.0750 minDCF和2.59%EER,证明了我们提交的系统在对抗条件下的鲁棒性。在Track 2中,我们继续使用Track 1中的CM系统,并将其与基于CNN的ASV系统融合。该方法在封闭条件下实现了0.2814 min-aDCF,在开放条件下实现了0.0756 min-aDCF,显示了SASV系统的优异性能。摘要:This paper describes the USTC-KXDIGIT system submitted to the ASVspoof5 Challenge for Track 1 (speech deepfake detection) and Track 2 (spoofing-robust automatic speaker verification, SASV). Track 1 showcases a diverse range of technical qualities from potential processing algorithms and includes both open and closed conditions. For these conditions, our system consists of a cascade of a frontend feature extractor and a back-end classifier. We focus on extensive embedding engineering and enhancing the generalization of the back-end classifier model. Specifically, the embedding engineering is based on hand-crafted features and speech representations from a self-supervised model, used for closed and open conditions, respectively. To detect spoof attacks under various adversarial conditions, we trained multiple systems on an augmented training set. Additionally, we used voice conversion technology to synthesize fake audio from genuine audio in the training set to enrich the synthesis algorithms. To leverage the complementary information learned by different model architectures, we employed activation ensemble and fused scores from different systems to obtain the final decision score for spoof detection. During the evaluation phase, the proposed methods achieved 0.3948 minDCF and 14.33% EER in the close condition, and 0.0750 minDCF and 2.59% EER in the open condition, demonstrating the robustness of our submitted systems under adversarial conditions. In Track 2, we continued using the CM system from Track 1 and fused it with a CNN-based ASV system. This approach achieved 0.2814 min-aDCF in the closed condition and 0.0756 min-aDCF in the open condition, showcasing superior performance in the SASV system.
【19】 Pureformer-VC: Non-parallel One-Shot Voice Conversion with Pure Transformer Blocks and Triplet Discriminative Training
标题: Pureformer-VC:使用纯Transformer块和三重区分训练的非并行单次语音转换
作者:Wenhan Yao,Zedong Xing,Xiarun Chen,Jia Liu,Yongqiang He,Weiping Wen
备注:submmited to ICASSP 2025
链接:点击下载PDF文件
摘要:单次语音转换(VC)的目的是改变任何源语音的音色,以匹配的目标说话人与只有一个语音样本。现有的基于风格转换的VC方法依赖于语音表示的解纠缠,难以准确独立地对每个语音分量进行编码,并有效地重新组合成转换后的语音。为了解决这个问题,我们提出了Pureformer-VC,它利用Conformer块来构建一个解纠缠的编码器,并利用Zipformer块来构建一个风格转换解码器作为生成器。在解码器中,我们使用了有效的styleformer块来有效地将说话人特征整合到生成的语音中。该模型使用生成VAE损失编码组件和三重损失的无监督判别训练。我们将styleformer方法应用于Zipformer的共享权重以进行样式传输。实验结果表明,该模型实现了可比的主观分数,并表现出改善的客观指标相比,现有的方法在一个一次性的语音转换场景。摘要:One-shot voice conversion(VC) aims to change the timbre of any source speech to match that of the unseen target speaker with only one speech sample. Existing style transfer-based VC methods relied on speech representation disentanglement and suffered from accurately and independently encoding each speech component and recomposing back to converted speech effectively. To tackle this, we proposed Pureformer-VC, which utilizes Conformer blocks to build a disentangled encoder, and Zipformer blocks to build a style transfer decoder as the generator. In the decoder, we used effective styleformer blocks to integrate speaker characteristics into the generated speech effectively. The models used the generative VAE loss for encoding components and triplet loss for unsupervised discriminative training. We applied the styleformer method to Zipformer's shared weights for style transfer. The experimental results show that the proposed model achieves comparable subjective scores and exhibits improvements in objective metrics compared to existing methods in a one-shot voice conversion scenario.
【20】 VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka
标题: VoxHakka:台湾客语方言多样化的多说话人文本到语音系统
作者:Li-Wei Chen,Hung-Shin Lee,Chen-Chi Chang
备注:Submitted to O-COCOSDA 2024
链接:点击下载PDF文件
摘要:本文介绍VoxHakka,这是一个专为台湾客语(台湾资源严重不足的语言)设计的文本到语音(TTS)系统。利用YourTTS框架,VoxHakka在语音合成中实现了高自然度和准确度以及低实时性,同时支持六种不同的客家方言。这是通过用方言特定数据训练模型来实现的,允许生成说话者感知的客家话语音。为了解决公共可用的客家语音语料库的稀缺性,我们采用了一种具有成本效益的方法,利用网络抓取管道加上自动语音识别(ASR)为基础的数据清洗技术。该过程确保了获得适合TTS训练的高质量、多说话者、多方言数据集。使用比较平均意见分数(CMOS)进行的主观听力测试表明,VoxHakka显着优于现有的公开可用的客家文语转换系统的发音准确性,音调的正确性,和整体自然。这项工作代表了客家语言技术的重大进步,并为语言保护和振兴工作提供了宝贵的资源。摘要:This paper introduces VoxHakka, a text-to-speech (TTS) system designed for Taiwanese Hakka, a critically under-resourced language spoken in Taiwan. Leveraging the YourTTS framework, VoxHakka achieves high naturalness and accuracy and low real-time factor in speech synthesis while supporting six distinct Hakka dialects. This is achieved by training the model with dialect-specific data, allowing for the generation of speaker-aware Hakka speech. To address the scarcity of publicly available Hakka speech corpora, we employed a cost-effective approach utilizing a web scraping pipeline coupled with automatic speech recognition (ASR)-based data cleaning techniques. This process ensured the acquisition of a high-quality, multi-speaker, multi-dialect dataset suitable for TTS training. Subjective listening tests conducted using comparative mean opinion scores (CMOS) demonstrate that VoxHakka significantly outperforms existing publicly available Hakka TTS systems in terms of pronunciation accuracy, tone correctness, and overall naturalness. This work represents a significant advancement in Hakka language technology and provides a valuable resource for language preservation and revitalization efforts.
【21】 Effective Noise-aware Data Simulation for Domain-adaptive Speech Enhancement Leveraging Dynamic Stochastic Perturbation
标题: 利用动态随机扰动的域自适应语音增强的有效噪音感知数据模拟
作者:Chien-Chun Wang,Li-Wei Chen,Hung-Shin Lee,Berlin Chen,Hsin-Min Wang
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
摘要:跨域语音增强(SE)通常面临着严峻的挑战,因为在未知的目标域中噪声和背景信息的缺乏,导致训练和测试条件之间的不匹配。这项研究提出了一种新的数据模拟方法来解决这个问题,利用噪声提取技术和生成对抗网络(GANs),只有有限的目标噪声语音数据。值得注意的是,我们的方法采用噪声编码器从目标域数据中提取噪声嵌入。这些嵌入适当地引导生成器合成声学上适合于目标域的话语,同时真实地保留输入干净语音的语音内容。此外,我们引入了动态随机扰动的概念,它可以在推理过程中将受控扰动注入噪声嵌入,从而使模型能够很好地推广到看不见的噪声条件。在VoiceBank-DEMAND基准数据集上的实验表明,我们的域自适应SE方法优于现有的基于数据模拟的强基线。摘要:Cross-domain speech enhancement (SE) is often faced with severe challenges due to the scarcity of noise and background information in an unseen target domain, leading to a mismatch between training and test conditions. This study puts forward a novel data simulation method to address this issue, leveraging noise-extractive techniques and generative adversarial networks (GANs) with only limited target noisy speech data. Notably, our method employs a noise encoder to extract noise embeddings from target-domain data. These embeddings aptly guide the generator to synthesize utterances acoustically fitted to the target domain while authentically preserving the phonetic content of the input clean speech. Furthermore, we introduce the notion of dynamic stochastic perturbation, which can inject controlled perturbations into the noise embeddings during inference, thereby enabling the model to generalize well to unseen noise conditions. Experiments on the VoiceBank-DEMAND benchmark dataset demonstrate that our domain-adaptive SE method outperforms an existing strong baseline based on data simulation.
【22】 Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement
标题: Spectron:使用具有对抗细化的条件Transformer提取目标说话人
作者:Tathagata Bandyopadhyay
链接:点击下载PDF文件
摘要:最近,基于注意力的Transformers已经成为许多深度学习应用的事实标准,包括自然语言处理,计算机视觉,信号处理等。在本文中,我们提出了一个基于变换器的端到端模型,从单声道多说话人混合音频信号中提取目标说话人的语音。与现有的说话人提取方法不同,我们引入了两个额外的目标,施加说话人嵌入的一致性和波形编码器的可逆性,并联合训练说话人编码器和语音分离器,以更好地捕捉说话人的条件嵌入。此外,我们利用一个多尺度的降噪来改善所提取的语音的感知质量。我们的实验表明,在分离器主干中使用双路径Transformer以及所提出的训练范例将CNN基线提高了3.12 $ dB点。最后,我们将我们的方法与最新的最先进的方法进行比较,结果表明,我们的模型平均比现有方法高出4.1 $ dB,而不会产生额外的数据依赖性。摘要:Recently, attention-based transformers have become a de facto standard in many deep learning applications including natural language processing, computer vision, signal processing, etc.. In this paper, we propose a transformer-based end-to-end model to extract a target speaker's speech from a monaural multi-speaker mixed audio signal. Unlike existing speaker extraction methods, we introduce two additional objectives to impose speaker embedding consistency and waveform encoder invertibility and jointly train both speaker encoder and speech separator to better capture the speaker conditional embedding. Furthermore, we leverage a multi-scale discriminator to refine the perceptual quality of the extracted speech. Our experiments show that the use of a dual path transformer in the separator backbone along with proposed training paradigm improves the CNN baseline by $3.12$ dB points. Finally, we compare our approach with recent state-of-the-arts and show that our model outperforms existing methods by $4.1$ dB points on an average without creating additional data dependency.
【23】 A multilingual training strategy for low resource Text to Speech
标题: 低资源文本到语音的多语言训练策略
作者:Asma Amalas,Mounir Ghogho,Mohamed Chetouani,Rachid Oulad Haj Thami
备注:12 pages, 2 figures
链接:点击下载PDF文件
摘要:由于神经文本到语音(TTS)的最新进展,最近的语音技术已经导致产生高质量的合成语音。然而,这种TTS模型依赖于大量的数据,这些数据的产生成本很高,并且很难扩展到所有现有的语言,特别是很少关注低资源语言。通过知识转移等技术,可以减轻创建数据集的负担。因此,在本文中,我们研究了两个方面;首先,来自社交媒体的数据是否可以用于小型TTS数据集的构建,其次,低资源语言的跨语言迁移学习(TL)是否可以使用这种类型的数据。在这方面,我们具体评估了多语言建模在多大程度上可以作为单语语料库训练的替代方案。要做到这一点,我们探讨如何从外语的数据可以被选择和汇集到训练一个目标低资源语言的TTS模型。我们的研究结果表明,多语种预训练比单语预训练更好地提高了生成语音的可懂度和自然度。摘要:Recent speech technologies have led to produce high quality synthesised speech due to recent advances in neural Text to Speech (TTS). However, such TTS models depend on extensive amounts of data that can be costly to produce and is hardly scalable to all existing languages, especially that seldom attention is given to low resource languages. With techniques such as knowledge transfer, the burden of creating datasets can be alleviated. In this paper, we therefore investigate two aspects; firstly, whether data from social media can be used for a small TTS dataset construction, and secondly whether cross lingual transfer learning (TL) for a low resource language can work with this type of data. In this aspect, we specifically assess to what extent multilingual modeling can be leveraged as an alternative to training on monolingual corporas. To do so, we explore how data from foreign languages may be selected and pooled to train a TTS model for a target low resource language. Our findings show that multilingual pre-training is better than monolingual pre-training at increasing the intelligibility and naturalness of the generated speech.
【24】 Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language
标题: 个性化唇读:用视觉和语言适应独特的唇动
作者:Jeong Hun Yeo,Chae Won Kim,Hyunjun Kim,Hyeongseop Rha,Seunghee Han,Wen-Huang Cheng,Yong Man Ro
备注:Code available: this https URL
链接:点击下载PDF文件
摘要:唇读的目的是通过分析嘴唇的运动来预测口语。尽管唇读技术取得了进步,但当模型应用于看不见的扬声器时,性能会下降,因为它们对视觉信息(如嘴唇外观)的变化很敏感。为了解决这一挑战,说话人自适应唇读技术已经通过专注于有效地使唇读模型适应视觉模态中的目标说话人而得到了发展。对目标语者的语言信息(如词汇选择)进行顺应的有效性在以前的研究中还没有被探讨过。此外,现有的说话人自适应数据集的词汇量和姿势变化有限,限制了以前的说话人自适应方法在现实世界中的验证。为了解决这些问题,我们提出了一种新的说话者自适应唇读方法,该方法在视觉和语言水平上将预先训练的模型适应于目标说话者。具体来说,我们整合了即时调谐和LoRA方法,将它们应用于预先训练的唇读模型,以有效地使模型适应目标扬声器。此外,为了验证其在现实世界中的有效性,我们引入了一个新的数据集,VoxLRS-SA,来自VoxCeleb 2和LRS 3。它包含大约10万个单词的词汇表,提供各种姿势变化,并首次在野生,高级唇读中验证适应方法。通过各种实验,我们证明了现有的说话人自适应方法也提高了在句子级别上的性能。此外,与所提出的自适应方法,我们表明,所提出的方法实现更大的改善时,应用到目标扬声器,相比以前的作品。摘要:Lip reading aims to predict spoken language by analyzing lip movements. Despite advancements in lip reading technologies, performance degrades when models are applied to unseen speakers due to their sensitivity to variations in visual information such as lip appearances. To address this challenge, speaker adaptive lip reading technologies have advanced by focusing on effectively adapting a lip reading model to target speakers in the visual modality. The effectiveness of adapting language information, such as vocabulary choice, of the target speaker has not been explored in the previous works. Moreover, existing datasets for speaker adaptation have limited vocabulary size and pose variations, limiting the validation of previous speaker-adaptive methods in real-world scenarios. To address these issues, we propose a novel speaker-adaptive lip reading method that adapts a pre-trained model to target speakers at both vision and language levels. Specifically, we integrate prompt tuning and the LoRA approach, applying them to a pre-trained lip reading model to effectively adapt the model to target speakers. In addition, to validate its effectiveness in real-world scenarios, we introduce a new dataset, VoxLRS-SA, derived from VoxCeleb2 and LRS3. It contains a vocabulary of approximately 100K words, offers diverse pose variations, and enables the validation of adaptation methods in wild, sentence-level lip reading for the first time. Through various experiments, we demonstrate that the existing speaker-adaptive method also improves performance in the wild at the sentence level. Moreover, with the proposed adaptation method, we show that the proposed method achieves larger improvements when applied to the target speaker, compared to the previous works.
【25】 Interpretable Convolutional SyncNet
标题: 可解释卷积同步网络
作者:Sungjoon Park,Jaesub Yun,Donggeon Lee,Minsik Park
备注:8+5 pages
链接:点击下载PDF文件
摘要:由于各种原因,视频可能会不同步,因此使用同步网络将视频恢复同步,以执行需要同步视频的任务。以前的最先进的(SOTA)同步网络使用InfoNCE损耗,依赖于Transformer架构,或两者兼而有之。不幸的是,前者使模型的输出难以解释,后者对大图像不友好,从而限制了同步网络的有用性。在这项工作中,我们训练一个卷积同步网络使用平衡BCE损失(BBCE),损失的启发二进制交叉熵(BCE)和InfoNCE损失。与InfoNCE损失相比,BBCE损失不需要复杂的采样方案。我们的模型可以更好地处理更大的图像,其输出可以给出概率解释。概率解释允许我们定义诸如偏移概率和屏幕外比率的度量来评估视听(AV)语音数据集的同步质量。此外,我们的模型在LRS 2数据集上达到了96.5 %$的SOTA准确度,在LRS 3数据集上达到了93.8 %$。摘要:Because videos in the wild can be out of sync for various reasons, a sync-net is used to bring the video back into sync for tasks that require synchronized videos. Previous state-of-the-art (SOTA) sync-nets use InfoNCE loss, rely on the transformer architecture, or both. Unfortunately, the former makes the model's output difficult to interpret, and the latter is unfriendly with large images, thus limiting the usefulness of sync-nets. In this work, we train a convolutional sync-net using the balanced BCE loss (BBCE), a loss inspired by the binary cross entropy (BCE) and the InfoNCE losses. In contrast to the InfoNCE loss, the BBCE loss does not require complicated sampling schemes. Our model can better handle larger images, and its output can be given a probabilistic interpretation. The probabilistic interpretation allows us to define metrics such as probability at offset and offscreen ratio to evaluate the sync quality of audio-visual (AV) speech datasets. Furthermore, our model achieves SOTA accuracy of $96.5 %$ on the LRS2 dataset and $93.8 %$ on the LRS3 dataset.
【26】 A Framework for Synthetic Audio Conversations Generation using Large Language Models
标题: 使用大型语言模型生成合成音频对话的框架
作者:Kaung Myat Kyaw,Jonathan Hoyin Chan
备注:This work has been submitted for consideration at the WI-IAT'24 to be held in December 2024
链接:点击下载PDF文件
摘要:在本文中,我们介绍ConversaSynth,一个框架,旨在生成合成会话音频使用大型语言模型(LLM)与多个人物设置。该框架首先创建各种主题的基于文本的对话,然后使用文本到语音(TTS)系统将其转换为音频。我们的实验表明,ConversaSynth可以有效地生成高质量的合成音频数据集,这可以显着增强音频标记,音频分类和多说话人语音识别模型的训练和评估。结果表明,由ConversaSynth生成的合成数据集表现出极大的多样性和真实性,使其适合开发强大的,适应性强的基于音频的AI系统。摘要:In this paper, we introduce ConversaSynth, a framework designed to generate synthetic conversation audio using large language models (LLMs) with multiple persona settings. The framework first creates diverse and coherent text-based dialogues across various topics, which are then converted into audio using text-to-speech (TTS) systems. Our experiments demonstrate that ConversaSynth effectively generates highquality synthetic audio datasets, which can significantly enhance the training and evaluation of models for audio tagging, audio classification, and multi-speaker speech recognition. The results indicate that the synthetic datasets generated by ConversaSynth exhibit substantial diversity and realism, making them suitable for developing robust, adaptable audio-based AI systems.
【27】 SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
标题: SoCodec:一种语义有序的多流语音编解码器,用于基于高效语言模型的文本到语音合成
作者:Haohan Guo,Fenglong Xie,Kun Xie,Dongchao Yang,Dake Guo,Xixin Wu,Helen Meng
备注:Accepted by SLT 2024
链接:点击下载PDF文件
摘要:长语音序列一直困扰着基于语言模型(LM)的TTS方法的建模复杂度和效率。这项工作提出了SoCodec,语义有序的多流语音编解码器,以解决这个问题。它将语音压缩成一个较短的、多流的离散语义序列,每个帧有多个标记。同时,提出了有序积量化的方法,将该序列约束为有序表示。它可以与多流延迟LM一起应用,以在TTS中沿时间和流轴实现更好的自回归生成。实验结果有力地证明了所提出的方法的有效性,即使将语音的帧移从20ms压缩到240ms(12倍),也能获得优于基线系统的性能。消融研究进一步验证了学习所提出的有序多流语义表示在追求更短的语音序列的有效LM为基础的TTS的重要性。摘要:The long speech sequence has been troubling language models (LM) based TTS approaches in terms of modeling complexity and efficiency. This work proposes SoCodec, a semantic-ordered multi-stream speech codec, to address this issue. It compresses speech into a shorter, multi-stream discrete semantic sequence with multiple tokens at each frame. Meanwhile, the ordered product quantization is proposed to constrain this sequence into an ordered representation. It can be applied with a multi-stream delayed LM to achieve better autoregressive generation along both time and stream axes in TTS. The experimental result strongly demonstrates the effectiveness of the proposed approach, achieving superior performance over baseline systems even if compressing the frameshift of speech from 20ms to 240ms (12x). The ablation studies further validate the importance of learning the proposed ordered multi-stream semantic representation in pursuing shorter speech sequences for efficient LM-based TTS.
【28】 MMT-BERT: Chord-aware Symbolic Music Generation Based on Multitrack Music Transformer and MusicBERT
标题: MMT-BERT:基于多轨音乐Transformer和MusicBERT的和弦感知符号音乐生成
作者:Jinlong Zhu,Keigo Sakurai,Ren Togo,Takahiro Ogawa,Miki Haseyama
备注:Accepted to the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
摘要:我们提出了一种新的符号音乐表示和生成对抗网络(GAN)框架,专门设计用于符号多轨音乐生成。符号音乐生成的主题主要包括音乐数据的预处理和深度学习框架的实现。目前的符号音乐生成技术通常面临两个重大挑战:训练数据缺乏和弦和音阶的信息,以及需要专门设计的模型架构来适应符号音乐表示的独特格式。在本文中,我们解决了上述问题,通过引入新的符号音乐表示与MusicLang和弦分析模型。我们提出了我们的MMT-BERT架构适应的表示。为了构建一个强大的多轨音乐生成器,我们微调了一个预先训练好的MusicBERT模型来作为训练器,并结合了相对论标准损失。这种方法得到了对MusicBERT中编码的符号音乐的深入理解的支持,增强了我们的方法生成的音乐的和谐和人性。实验结果表明,我们的方法,严格遵循国家的最先进的方法的有效性。摘要:We propose a novel symbolic music representation and Generative Adversarial Network (GAN) framework specially designed for symbolic multitrack music generation. The main theme of symbolic music generation primarily encompasses the preprocessing of music data and the implementation of a deep learning framework. Current techniques dedicated to symbolic music generation generally encounter two significant challenges: training data's lack of information about chords and scales and the requirement of specially designed model architecture adapted to the unique format of symbolic music representation. In this paper, we solve the above problems by introducing new symbolic music representation with MusicLang chord analysis model. We propose our MMT-BERT architecture adapting to the representation. To build a robust multitrack music generator, we fine-tune a pre-trained MusicBERT model to serve as the discriminator, and incorporate relativistic standard loss. This approach, supported by the in-depth understanding of symbolic music encoded within MusicBERT, fortifies the consonance and humanity of music generated by our method. Experimental results demonstrate the effectiveness of our approach which strictly follows the state-of-the-art methods.
【29】 Dissecting Temporal Understanding in Text-to-Audio Retrieval
标题: 文本到音频检索中的时态理解剖析
作者:Andreea-Maria Oncescu,João F. Henriques,A. Sophia Koepke
备注:9 pages, 5 figures, ACM Multimedia 2024, this https URL
链接:点击下载PDF文件
摘要:机器学习的最新进展推动了对多模态任务的研究,例如文本到视频和文本到音频检索。这些任务需要模型来理解视频和音频数据的语义内容,包括对象和字符。模型还需要学习空间安排和时间关系。在这项工作中,我们分析的时间顺序的声音,这是一个未充分研究的问题,在文本到音频检索的背景下。特别是,我们剖析了AudioCaps和Clotho数据集上的文本到音频检索的最先进模型的时间理解能力。此外,我们还介绍了一个合成的文本音频数据集,它提供了一个受控的设置,用于评估最近的模型的时间能力。最后,我们提出了一个损失函数,鼓励文本音频模型专注于事件的时间顺序。代码和数据可在https: www.robots.ox.ac.uk ~vgg research audio-retrieval dtu 上获得。摘要:Recent advancements in machine learning have fueled research on multimodal tasks, such as for instance text-to-video and text-to-audio retrieval. These tasks require models to understand the semantic content of video and audio data, including objects, and characters. The models also need to learn spatial arrangements and temporal relationships. In this work, we analyse the temporal ordering of sounds, which is an understudied problem in the context of text-to-audio retrieval. In particular, we dissect the temporal understanding capabilities of a state-of-the-art model for text-to-audio retrieval on the AudioCaps and Clotho datasets. Additionally, we introduce a synthetic text-audio dataset that provides a controlled setting for evaluating temporal capabilities of recent models. Lastly, we present a loss function that encourages text-audio models to focus on the temporal ordering of events. Code and data are available at https: www.robots.ox.ac.uk ~vgg research audio-retrieval dtu .
【30】 LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
标题: LibriheavightMix:一个20,000小时的数据集,用于单通道回响多说话者语音分离、ASB和说话者拨号
作者:Zengrui Jin,Yifan Yang,Mohan Shi,Wei Kang,Xiaoyu Yang,Zengwei Yao,Fangjun Kuang,Liyong Guo,Lingwei Meng,Long Lin,Yong Xu,Shi-Xiong Zhang,Daniel Povey
备注:InterSpeech 2024
链接:点击下载PDF文件
摘要:不断发展的语音处理领域越来越关注复杂的场景,如具有多个同时发言者和远场条件的会议或鸡尾酒会。解决这些挑战的现有方法分为两类:多渠道和单渠道解决方案。单通道方法以其通用性和方便性而闻名,不需要关于麦克风阵列的特定信息。 本文提出了一个大规模的远场重叠语音数据集,制作语音分离,识别和说话人日记的研究。该数据集是在多人、混响环境中解码“谁在什么时候说了什么”的关键资源,这是该领域的一个艰巨挑战。此外,我们介绍了一个管道系统,包括语音分离,识别和日记作为一个基本的基准。关于WHAMR!数据集验证了拟议数据的广泛适用性。摘要:The evolving speech processing landscape is increasingly focused on complex scenarios like meetings or cocktail parties with multiple simultaneous speakers and far-field conditions. Existing methodologies for addressing these challenges fall into two categories: multi-channel and single-channel solutions. Single-channel approaches, notable for their generality and convenience, do not require specific information about microphone arrays. This paper presents a large-scale far-field overlapping speech dataset, crafted to advance research in speech separation, recognition, and speaker diarization. This dataset is a critical resource for decoding Who said What and When'' in multi-talker, reverberant environments, a daunting challenge in the field. Additionally, we introduce a pipeline system encompassing speech separation, recognition, and diarization as a foundational benchmark. Evaluations on the WHAMR! dataset validate the broad applicability of the proposed data.
【31】 Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
标题: 用于多说话人自动语音识别的重叠编码分离序列化语音信息引导
作者:Hao Shi,Yuan Gao,Zhaoheng Ni,Tatsuya Kawahara
链接:点击下载PDF文件
摘要:串行输出训练(SOT)由于其方便灵活的多说话人自动语音识别(ASR)方法而受到越来越多的关注。然而,仅仅在注意力损失的情况下进行训练并不容易。在本文中,我们提出了重叠编码分离(EncSep),以充分利用连接主义的时间分类(CTC)和注意力混合损失的好处。该附加分离器被插入在编码器之后以提取具有CTC损失的多说话者信息。此外,我们提出了序列化的语音信息指导SOT(GEncSep),以进一步利用分离的编码。分离的流被连接以提供单个说话者信息,从而在解码期间引导注意力。在LibriMix上的实验结果表明,该方法可以有效地将单说话人编码与重叠编码分离。CTC损失有助于改善复杂场景下的编码器表示。GEncSep进一步提高了性能。摘要:Serialized output training (SOT) attracts increasing attention due to its convenience and flexibility for multi-speaker automatic speech recognition (ASR). However, it is not easy to train with attention loss only. In this paper, we propose the overlapped encoding separation (EncSep) to fully utilize the benefits of the connectionist temporal classification (CTC) and attention hybrid loss. This additional separator is inserted after the encoder to extract the multi-speaker information with CTC losses. Furthermore, we propose the serialized speech information guidance SOT (GEncSep) to further utilize the separated encodings. The separated streams are concatenated to provide single-speaker information to guide attention during decoding. The experimental results on LibriMix show that the single-speaker encoding can be separated from the overlapped encoding. The CTC loss helps to improve the encoder representation under complex scenarios. GEncSep further improved performance.
【32】 MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
标题: MaskGCT:带屏蔽生成编解码器的Zero-Shot文本到语音
作者:Yuancheng Wang,Haoyue Zhan,Liwei Liu,Ruihong Zeng,Haotian Guo,Jiachen Zheng,Qiang Zhang,Shunsi Zhang,Zhizheng Wu
链接:点击下载PDF文件
摘要:目前,大规模的文语转换系统主要分为两类:自回归和非自回归。自回归系统在鲁棒性方面存在一定的不足,并且不能控制语音持续时间。相比之下,非自回归系统需要明确的预测的语音水平的持续时间,这可能会损害他们的自然性。我们介绍了掩蔽生成编解码器Transformer(MaskGCT),一个完全非自回归模型的TTS,不需要精确的对齐信息之间的文本和语音。MaskGCT是一个两阶段模型:在第一阶段,该模型使用文本来预测从语音自监督学习(SSL)模型中提取的语义令牌,在第二阶段,该模型预测以这些语义令牌为条件的声学令牌。MaskGCT遵循 textit{mask-and-predict}学习范式。在训练过程中,MaskGCT根据给定的条件和提示学习预测掩蔽的语义或声学标记。在推理过程中,模型以并行方式生成指定长度的令牌。我们将MaskGCT扩展到一个大规模的多语言数据集,其中包含10万小时的野外语音。我们的实验表明,MaskGCT实现卓越的或有竞争力的性能相比,国家的最先进的zero-shot TTS系统的质量,相似性和可理解性,同时提供更高的生成效率比基于扩散或自回归TTS模型。音频样本可在https: maskgct.github.io上获得。摘要:Nowadays, large-scale text-to-speech (TTS) systems are primarily divided into two types: autoregressive and non-autoregressive. The autoregressive systems have certain deficiencies in robustness and cannot control speech duration. In contrast, non-autoregressive systems require explicit prediction of phone-level duration, which may compromise their naturalness. We introduce the Masked Generative Codec Transformer (MaskGCT), a fully non-autoregressive model for TTS that does not require precise alignment information between text and speech. MaskGCT is a two-stage model: in the first stage, the model uses text to predict semantic tokens extracted from a speech self-supervised learning (SSL) model, and in the second stage, the model predicts acoustic tokens conditioned on these semantic tokens. MaskGCT follows the textit{mask-and-predict} learning paradigm. During training, MaskGCT learns to predict masked semantic or acoustic tokens based on given conditions and prompts. During inference, the model generates tokens of a specified length in a parallel manner. We scale MaskGCT to a large-scale multilingual dataset with 100K hours of in-the-wild speech. Our experiments demonstrate that MaskGCT achieves superior or competitive performance compared to state-of-the-art zero-shot TTS systems in terms of quality, similarity, and intelligibility while offering higher generation efficiency than diffusion-based or autoregressive TTS models. Audio samples are available at https: maskgct.github.io.
【33】 Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion
标题: 看到你的言语风格:一种新颖的Zero-Shot身份解开基于面部的语音转换
作者:Yan Rong,Li Liu
链接:点击下载PDF文件
摘要:基于人脸的语音转换(FVC)是一种利用人脸图像来生成目标说话人的语音风格的新任务。先前的工作有两个缺点:(1)难以获得与说话者的语音身份信息良好对齐的面部嵌入,以及(2)在从音频输入解耦内容和说话者身份信息方面的不足。为了解决这些问题,我们提出了一种新的FVC方法,身份解开基于人脸的语音转换(ID-FaceVC),它克服了上述两个限制。更确切地说,我们提出了一个基于查询的身份感知对比学习(IAQ-CL)模块来提取特定于说话者的面部特征,以及一个基于互信息的双重解耦(MIDD)模块来从音频中纯化内容特征,确保清晰和高质量的语音转换。此外,与以前的作品不同,我们的方法可以接受音频或文本输入,提供可控的语音生成与可调的情感基调和速度。大量的实验表明,ID-FaceVC在各种指标上都达到了最先进的性能,定性和用户研究结果证实了其在自然性、相似性和多样性方面的有效性。包含音频示例和代码的项目网站可以在https: id-facevc.github.io上找到。摘要:Face-based Voice Conversion (FVC) is a novel task that leverages facial images to generate the target speaker's voice style. Previous work has two shortcomings: (1) suffering from obtaining facial embeddings that are well-aligned with the speaker's voice identity information, and (2) inadequacy in decoupling content and speaker identity information from the audio input. To address these issues, we present a novel FVC method, Identity-Disentanglement Face-based Voice Conversion (ID-FaceVC), which overcomes the above two limitations. More precisely, we propose an Identity-Aware Query-based Contrastive Learning (IAQ-CL) module to extract speaker-specific facial features, and a Mutual Information-based Dual Decoupling (MIDD) module to purify content features from audio, ensuring clear and high-quality voice conversion. Besides, unlike prior works, our method can accept either audio or text inputs, offering controllable speech generation with adjustable emotional tone and speed. Extensive experiments demonstrate that ID-FaceVC achieves state-of-the-art performance across various metrics, with qualitative and user study results confirming its effectiveness in naturalness, similarity, and diversity. Project website with audio samples and code can be found at https: id-facevc.github.io.
【34】 FLUX that Plays Music
标题: 播放音乐的FLOX
作者:Zhengcong Fei,Mingyuan Fan,Changqian Yu,Junshi Huang
链接:点击下载PDF文件
摘要:本文探讨了一个简单的扩展,基于扩散整流Transformers的文本到音乐的生成,称为FluxMusic。通常,随着高级Flux footnote{https: github.com black-forest-labs flux}模型的设计,我们将其转移到mel-谱的潜在VAE空间中。它涉及首先将一系列独立注意力应用于双文本音乐流,然后是堆叠的单音乐流用于去噪补丁预测。我们采用多个预先训练的文本编码器来充分捕获字幕语义信息以及推理灵活性。在这两者之间,粗文本信息结合时间步嵌入被用于调制机制,而细粒度的文本细节与作为输入的音乐补丁序列连接。通过深入的研究,我们证明了具有优化架构的整流训练显著优于文本到音乐任务的既定扩散方法,正如各种自动指标和人类偏好评估所证明的那样。我们的实验数据、代码和模型权重在以下网址公开: url{https: github.com feizc FluxMusic}。摘要:This paper explores a simple extension of diffusion-based rectified flow Transformers for text-to-music generation, termed as FluxMusic. Generally, along with design in advanced Flux footnote{https: github.com black-forest-labs flux} model, we transfers it into a latent VAE space of mel-spectrum. It involves first applying a sequence of independent attention to the double text-music stream, followed by a stacked single music stream for denoised patch prediction. We employ multiple pre-trained text encoders to sufficiently capture caption semantic information as well as inference flexibility. In between, coarse textual information, in conjunction with time step embeddings, is utilized in a modulation mechanism, while fine-grained textual details are concatenated with the music patch sequence as inputs. Through an in-depth study, we demonstrate that rectified flow training with an optimized architecture significantly outperforms established diffusion methods for the text-to-music task, as evidenced by various automatic metrics and human preference evaluations. Our experimental data, code, and model weights are made publicly available at: url{https: github.com feizc FluxMusic}.
【35】 Multi-scale Multi-instance Visual Sound Localization and Segmentation
标题: 多尺度多实例视觉声音定位与分割
作者:Shentong Mo,Haofan Wang
链接:点击下载PDF文件
摘要:视觉声音定位是一个典型的和具有挑战性的问题,预测对应的声源在视频中的对象的位置。以往的方法主要是使用全局音频和单尺度视觉特征之间的视听关联来定位每个图像中的发声对象。尽管它们的性能很有希望,但它们忽略了相应图像的多尺度视觉特征,并且与地面事实相比,它们无法学习区分区域。为了解决这个问题,我们提出了一种新的多尺度多实例视觉声音定位框架,即M2 VSL,它可以直接从输入图像中学习与声源相关的多尺度语义特征来定位发声对象。具体来说,我们的M2 VSL利用可学习的多尺度视觉特征,在相应图像的多层次位置对齐视听表示。我们还介绍了一种新的多尺度多实例Transformer动态聚合多尺度跨模态表示的视觉声音定位。我们对VGGSound-Instruments,VGG-Sound Sources和AVSBench基准进行了广泛的实验。实验结果表明,本文提出的M2 VSL算法在探测目标定位和分割方面具有较好的性能。摘要:Visual sound localization is a typical and challenging problem that predicts the location of objects corresponding to the sound source in a video. Previous methods mainly used the audio-visual association between global audio and one-scale visual features to localize sounding objects in each image. Despite their promising performance, they omitted multi-scale visual features of the corresponding image, and they cannot learn discriminative regions compared to ground truths. To address this issue, we propose a novel multi-scale multi-instance visual sound localization framework, namely M2VSL, that can directly learn multi-scale semantic features associated with sound sources from the input image to localize sounding objects. Specifically, our M2VSL leverages learnable multi-scale visual features to align audio-visual representations at multi-level locations of the corresponding image. We also introduce a novel multi-scale multi-instance transformer to dynamically aggregate multi-scale cross-modal representations for visual sound localization. We conduct extensive experiments on VGGSound-Instruments, VGG-Sound Sources, and AVSBench benchmarks. The results demonstrate that the proposed M2VSL can achieve state-of-the-art performance on sounding object localization and segmentation.
【36】 Multi-label Zero-Shot Audio Classification with Temporal Attention
标题: 具有时间注意力的多标签Zero-Shot音频分类
作者:Duygu Dogan,Huang Xie,Toni Heittola,Tuomas Virtanen
备注:Accepted to International Workshop on Acoustic Signal Enhancement (IWAENC) 2024
链接:点击下载PDF文件
摘要:Zero-shot学习模型能够通过使用辅助信息从所看到的类中转移知识来对新类进行分类。现有的zero-shot学习方法大多集中在单标签分类任务上,本文提出了一种多标签zero-shot音频分类方法。为了解决多标签声音分类的挑战,同时推广到看不见的类,我们适应时间注意。时间注意力机制基于音频片段的声学和语义兼容性将重要性权重分配给不同的音频片段,从而使模型能够通过关注与每个类最相关的片段来捕获音频样本内不同声音类的变化主导性。这导致比采用时间上聚合的声学特征而没有加权的方法更准确的多标签zero-shot分类,所述方法同等地对待所有音频段。我们评估我们的方法对一个zero-shot模型的AudioSet子集上使用统一聚合的声学特征,零规则基线,并在监督的情况下提出的方法。我们的研究结果表明,在多标签的情况下,时间注意增强了zero-shot音频分类性能。摘要:Zero-shot learning models are capable of classifying new classes by transferring knowledge from the seen classes using auxiliary information. While most of the existing zero-shot learning methods focused on single-label classification tasks, the present study introduces a method to perform multi-label zero-shot audio classification. To address the challenge of classifying multi-label sounds while generalizing to unseen classes, we adapt temporal attention. The temporal attention mechanism assigns importance weights to different audio segments based on their acoustic and semantic compatibility, thus enabling the model to capture the varying dominance of different sound classes within an audio sample by focusing on the segments most relevant for each class. This leads to more accurate multi-label zero-shot classification than methods employing temporally aggregated acoustic features without weighting, which treat all audio segments equally. We evaluate our approach on a subset of AudioSet against a zero-shot model using uniformly aggregated acoustic features, a zero-rule baseline, and the proposed method in the supervised scenario. Our results show that temporal attention enhances the zero-shot audio classification performance in multi-label scenario.
【37】 Density Adaptive Attention-based Speech Network: Enhancing Feature Understanding for Mental Health Disorders
标题: 密度自适应基于注意力的语音网络:增强对心理健康障碍的特征理解
作者:Georgios Ioannides,Adrian Kieback,Aman Chadha,Aaron Elkins
链接:点击下载PDF文件
摘要:基于语音的抑郁症检测由于其在个体中的独特表现和数据稀缺性,对自动检测提出了重大挑战。为了应对这些挑战,我们引入了DAAMAudioCNNLSTM和DAAMAudioTransformer,这是两个用于音频特征提取和抑郁检测的参数有效且可解释的模型。DAAMAudioCNNLSTM是一个新的CNN-LSTM框架,具有多头密度自适应注意力机制(DAAM),动态关注信息语音段。DAAMAudioTransformer利用Transformer编码器代替CNN-LSTM架构,结合了相同的DAAM模块,以增强注意力和可解释性。这些方法不仅增强了检测的鲁棒性和可解释性,而且还实现了最先进的性能:DAAMAudioCNNLSTM在DAIC-WOZ数据集上的F1宏得分为0.702,DAAMAudioTransformer的F1宏得分为0.72,而不像以前的方法那样在训练 验证过程中依赖于补充信息,如元音位置和说话人信息。这两种模型在利用语音信号进行抑郁症检测方面的显著可解释性和效率代表了向更可靠,临床有用的诊断工具的飞跃,在语音和心理健康护理方面取得了可喜的进步。为了促进这一领域的进一步研究,我们公开了我们的代码。摘要:Speech-based depression detection poses significant challenges for automated detection due to its unique manifestation across individuals and data scarcity. Addressing these challenges, we introduce DAAMAudioCNNLSTM and DAAMAudioTransformer, two parameter efficient and explainable models for audio feature extraction and depression detection. DAAMAudioCNNLSTM features a novel CNN-LSTM framework with multi-head Density Adaptive Attention Mechanism (DAAM), focusing dynamically on informative speech segments. DAAMAudioTransformer, leveraging a transformer encoder in place of the CNN-LSTM architecture, incorporates the same DAAM module for enhanced attention and interpretability. These approaches not only enhance detection robustness and interpretability but also achieve state-of-the-art performance: DAAMAudioCNNLSTM with an F1 macro score of 0.702 and DAAMAudioTransformer with an F1 macro score of 0.72 on the DAIC-WOZ dataset, without reliance on supplementary information such as vowel positions and speaker information during training validation as in previous approaches. Both models' significant explainability and efficiency in leveraging speech signals for depression detection represent a leap towards more reliable, clinically useful diagnostic tools, promising advancements in speech and mental health care. To foster further research in this domain, we make our code publicly available.
【38】 Contrastive Augmentation: An Unsupervised Learning Approach for Keyword Spotting in Speech Technology
标题: 对比增强:语音技术中关键词发现的无监督学习方法
作者:Weinan Dai,Yifeng Jiang,Yuanjing Liu,Jinkun Chen,Xin Sun,Jinglei Tao
备注:This paper has been accepted by the ICPR2024
链接:点击下载PDF文件
摘要:本文讨论了持续存在的挑战,在关键字定位(KWS),语音技术的基本组成部分,关于大量的标记数据的采集进行训练。鉴于难以获得大量的阳性样本和收集新的目标样本时,关键字变化的费力的过程中,我们介绍了一种新的方法相结合的无监督对比学习和一个独特的增强为基础的技术。我们的方法允许神经网络在未标记的数据集上进行训练,从而可能提高具有有限标记数据集的下游任务的性能。我们还建议,类似的高层次的功能表示应采用相同的关键字的语音话语,尽管速度或音量的变化。为了实现这一目标,我们提出了一种基于语音增强的无监督学习方法,利用瓶颈层特征和音频重构信息之间的相似性进行辅助训练。此外,我们提出了一种压缩卷积架构,以解决KWS任务中潜在的冗余和非信息性信息,使模型能够同时学习本地特征并专注于长期信息。此方法在Google Speech Commands V2数据集上实现了强大的性能。受最近在符号识别和口语术语检测方面取得的进展的启发,我们的方法强调了我们在KWS中的对比学习方法的潜力和逐例查询口语术语检测策略的优势。提出的CAB-KWS提供了新的视角在该领域的KWS,展示了有效的方法来减少数据收集的努力,提高系统的鲁棒性。摘要:This paper addresses the persistent challenge in Keyword Spotting (KWS), a fundamental component in speech technology, regarding the acquisition of substantial labeled data for training. Given the difficulty in obtaining large quantities of positive samples and the laborious process of collecting new target samples when the keyword changes, we introduce a novel approach combining unsupervised contrastive learning and a unique augmentation-based technique. Our method allows the neural network to train on unlabeled data sets, potentially improving performance in downstream tasks with limited labeled data sets. We also propose that similar high-level feature representations should be employed for speech utterances with the same keyword despite variations in speed or volume. To achieve this, we present a speech augmentation-based unsupervised learning method that utilizes the similarity between the bottleneck layer feature and the audio reconstructing information for auxiliary training. Furthermore, we propose a compressed convolutional architecture to address potential redundancy and non-informative information in KWS tasks, enabling the model to simultaneously learn local features and focus on long-term information. This method achieves strong performance on the Google Speech Commands V2 Dataset. Inspired by recent advancements in sign spotting and spoken term detection, our method underlines the potential of our contrastive learning approach in KWS and the advantages of Query-by-Example Spoken Term Detection strategies. The presented CAB-KWS provide new perspectives in the field of KWS, demonstrating effective ways to reduce data collection efforts and increase the system's robustness.
【39】 REFFLY: Melody-Constrained Lyrics Editing Model
标题: REFFLY:旋律约束歌词编辑模型
作者:Songyan Zhao,Bingxuan Li,Yufei Tian,Nanyun Peng
链接:点击下载PDF文件
摘要:自动旋律到歌词生成旨在产生与给定旋律一致的歌词。虽然以前的作品可以基于高级控制信号(如关键字或流派)生成歌词,但它们通常面临三个挑战:(1)缺乏可控性,因为以前的作品只能从头开始生成歌词,很少或根本无法控制内容;(2)无法生成具有所需格式的完全结构化的歌曲;以及(3)未能将歌词中的突出词与旋律中的突出音符对齐,从而导致不良的歌词-旋律对齐。在这项工作中,我们介绍REFFLY(歌词修订框架),第一个修订框架,旨在将任意形式的纯文本草稿编辑成高质量,完整的歌词。我们的方法可以确保生成的歌词保留了草稿的原始含义,与旋律保持一致,并坚持所需的歌曲结构。我们证明,REFFLY表现良好,在不同的任务设置,如歌词修改和歌曲翻译。实验结果表明,我们的模型在音乐性和文本质量方面都比Lyra(Tian et al. 2023)和GPT-4等强基线高出25%。摘要:Automatic melody-to-lyric generation aims to produce lyrics that align with a given melody. Although previous work can generate lyrics based on high-level control signals, such as keywords or genre, they often struggle with three challenges: (1) lack of controllability, as prior works are only able to produce lyrics from scratch, with little or no control over the content; (2) inability to generate fully structured songs with the desired format; and (3) failure to align prominent words in the lyrics with prominent notes in the melody, resulting in poor lyrics-melody alignment. In this work, we introduce REFFLY (REvision Framework For Lyrics), the first revision framework designed to edit arbitrary forms of plain text draft into high-quality, full-fledged song lyrics. Our approach ensures that the generated lyrics retain the original meaning of the draft, align with the melody, and adhere to the desired song structures. We demonstrate that REFFLY performs well in diverse task settings, such as lyrics revision and song translation. Experimental results show that our model outperforms strong baselines, such as Lyra (Tian et al. 2023) and GPT-4, by 25% in both musicality and text quality.
【40】 Towards a dynamical model of English vowels. Evidence from diphthongisation
标题: 迈向英语元音的动态模型。双元音化的证据
作者:Patrycja Strycharczuk,Sam Kirkham,Emily Gorman,Takayuki Nagamine
链接:点击下载PDF文件
摘要:双母音元音表现出一定程度的内在动态变化,其程度可以在共时和历时上变化,例如双母音元音可以变成单母音,反之亦然。模拟这种类型的变化需要定义与单元音相对的双元音。然而,制定一个明确的定义已被证明是难以捉摸的声学和清晰度,因为双元音化往往是梯度在这些领域。本研究从发音的角度来探讨双母音元音是否构成一个连贯的语音范畴。我们目前的发音和声学数据从六个扬声器的北部盎格鲁英语生产一套完整的语音长元音。我们分析了双元音化的几种措施,所有这些措施都表明双元音与长元音并没有绝对不同。我们考虑到这一观察与发音语音 任务动态模型,其中双元音和长单元音有一个共同的手势表示,包括两个发音目标在每种情况下,但他们根据手势收缩和位置的组件手势不同。我们认为,所有长元音的双目标表示是独立的语音重量的支持,以及在英国英语的历史双元音化和现今的动态元音变化的性质。摘要:Diphthong vowels exhibit a degree of inherent dynamic change, the extent of which can vary synchronically and diachronically, such that diphthong vowels can become monophthongs and vice versa. Modelling this type of change requires defining diphthongs in opposition to monophthongs. However, formulating an explicit definition has proven elusive in acoustics and articulation, as diphthongisation is often gradient in these domains. In this study, we consider whether diphthong vowels form a coherent phonetic category from the articulatory point of view. We present articulometry and acoustic data from six speakers of Northern Anglo-English producing a full set of phonologically long vowels. We analyse several measures of diphthongisation, all of which suggest that diphthongs are not categorically distinct from long monophthongs. We account for this observation with an Articulatory Phonology Task Dynamic model in which diphthongs and long monophthongs have a common gestural representation, comprising two articulatory targets in each case, but they differ according to gestural constriction and location of the component gestures. We argue that a two-target representation for all long vowels is independently supported by phonological weight, as well as by the nature of historical diphthongisation and present-day dynamic vowel variation in British English.
【41】 ProGRes: Prompted Generative Rescoring on ASR n-Best
标题: ProGRes:在ASB n-Best上进行的预定生成重新评分
作者:Ada Defne Tur,Adel Moumen,Mirco Ravanelli
备注:IEEE Spoken Language Technology Workshop
链接:点击下载PDF文件
摘要:大型语言模型(LLM)已经显示出它们通过有效地对波束搜索过程中生成的n个最佳假设进行重新评分来提高语音识别器性能的能力。然而,最好的方式来利用最新的生成抑制调整LLM的假设重新评分仍然不清楚。本文提出了一种新的方法,使用预调LLM动态扩展的n-最好的语音识别假设与新的假设产生通过适当的提示LLM。具体而言,我们引入了一种新的zero-shot方法,用于ASR n-最佳重新评分,该方法结合了置信度评分、LLM序列评分和基于序列的假设生成。我们比较了Llama-3-Instruct,GPT-3.5 Turbo和GPT-4 Turbo作为基于序列的生成器,Llama-3作为序列评分器LLM。我们评估了我们的方法,使用不同的语音识别器,并观察到显着的相对改善的字错误率(WER)从5%到25%不等。摘要:Large Language Models (LLMs) have shown their ability to improve the performance of speech recognizers by effectively rescoring the n-best hypotheses generated during the beam search process. However, the best way to exploit recent generative instruction-tuned LLMs for hypothesis rescoring is still unclear. This paper proposes a novel method that uses instruction-tuned LLMs to dynamically expand the n-best speech recognition hypotheses with new hypotheses generated through appropriately-prompted LLMs. Specifically, we introduce a new zero-shot method for ASR n-best rescoring, which combines confidence scores, LLM sequence scoring, and prompt-based hypothesis generation. We compare Llama-3-Instruct, GPT-3.5 Turbo, and GPT-4 Turbo as prompt-based generators with Llama-3 as sequence scorer LLM. We evaluated our approach using different speech recognizers and observed significant relative improvement in the word error rate (WER) ranging from 5% to 25%.
【42】 Speaker Tagging Correction With Non-Autoregressive Language Models
标题: 使用非自回归语言模型进行说话者标记纠正
作者:Grigor Kirakosyan,Davit Karamyan
备注:6 pages, 7 tables
链接:点击下载PDF文件
摘要:处理对话的语音应用程序不仅需要识别所说的话,还需要确定谁在什么时候说话。将单词分配给说话者的任务通常通过合并两个独立系统(即,自动语音识别(ASR)系统和说话者日志化(SD)系统)的输出来解决。在实际环境中,由于各种因素,包括具有高时间分辨率的均匀分割、不准确的单词时间戳、不正确的聚类和对说话人数量的估计以及背景噪声,说话人日志化系统可能会经历性能的显著下降。 因此,自动检测错误并在可能的情况下进行更正非常重要。我们使用了一个基于非自回归语言模型的第二遍说话者标记校正系统来校正不同说话者所说的句子边界上的单词中的错误。我们首先表明,所采用的纠错方法导致减少字日记错误率(WDER)的两个数据集:TAL和测试集的Fisher。此外,我们评估了我们的系统在后ASR扬声器标记校正的挑战,并观察到显着改善cpWER相比,基线方法。摘要:Speech applications dealing with conversations require not only recognizing the spoken words but also determining who spoke when. The task of assigning words to speakers is typically addressed by merging the outputs of two separate systems, namely, an automatic speech recognition (ASR) system and a speaker diarization (SD) system. In practical settings, speaker diarization systems can experience significant degradation in performance due to a variety of factors, including uniform segmentation with a high temporal resolution, inaccurate word timestamps, incorrect clustering and estimation of speaker numbers, as well as background noise. Therefore, it is important to automatically detect errors and make corrections if possible. We used a second-pass speaker tagging correction system based on a non-autoregressive language model to correct mistakes in words placed at the borders of sentences spoken by different speakers. We first show that the employed error correction approach leads to reductions in word diarization error rate (WDER) on two datasets: TAL and test set of Fisher. Additionally, we evaluated our system in the Post-ASR Speaker Tagging Correction challenge and observed significant improvements in cpWER compared to baseline methods.
【43】 Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning
标题: 使用谱-时态图专注池和多任务学习的逐例查询关键词发现
作者:Zhenyu Wang,Shuyu Kong,Li Wan,Biqiao Zhang,Yiteng Huang,Mumin Jin,Ming Sun,Xin Lei,Zhaojun Yang
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:现有的关键词识别(KWS)系统主要依赖于预定义的关键词短语。然而,识别定制关键词的能力对于定制与智能设备的交互至关重要。在本文中,我们提出了一种新的查询的例子(QbyE)KWS系统,采用频谱-时间图注意池和多任务学习。该框架旨在有效地学习QbyE KWS任务的说话人不变和语言信息嵌入。在这个框架内,我们研究了三种不同的编码器建模网络架构:LiCoNet,Conformer和ECAPA_TDNN。大量的内部数据集上的实验结果$629$扬声器已经证明了所提出的QbyE框架的有效性,在最大限度地发挥潜力的更简单的模型,如LiCoNet。特别是,LiCoNet的效率提高了13倍,实现了与计算密集型Conformer模型相当的性能(在0.3 FA Hr时,FRR为1.98% vs. 1.63 %)。摘要:Existing keyword spotting (KWS) systems primarily rely on predefined keyword phrases. However, the ability to recognize customized keywords is crucial for tailoring interactions with intelligent devices. In this paper, we present a novel Query-by-Example (QbyE) KWS system that employs spectral-temporal graph attentive pooling and multi-task learning. This framework aims to effectively learn speaker-invariant and linguistic-informative embeddings for QbyE KWS tasks. Within this framework, we investigate three distinct network architectures for encoder modeling: LiCoNet, Conformer and ECAPA_TDNN. The experimental results on a substantial internal dataset of $629$ speakers have demonstrated the effectiveness of the proposed QbyE framework in maximizing the potential of simpler models such as LiCoNet. Particularly, LiCoNet, which is 13x more efficient, achieves comparable performance to the computationally intensive Conformer model (1.98% vs. 1.63 % FRR at 0.3 FAs Hr).
机器翻译,仅供参考
