微信公众号:arXiv_Daily
cs.SD语音
标题: 基于集体学习机制的非并行语音转换最优传输生成对抗网络
链接:https://arxiv.org/abs/2504.13791
备注:7 pages, 2 figures, 3 tables
摘要:在图像合成方面取得了巨大成功之后,生成对抗网络(GAN)模型在语音合成领域也取得了重大进展,利用其通过对抗学习过程适应目标数据精确分布的能力。值得注意的是,在最先进的(SOTA)基于GAN的语音转换(VC)模型领域,真实语音样本和GAN生成的语音样本之间的自然度存在很大差异。此外,虽然许多GAN模型目前在单发电机并行学习方法上操作,但是通过单发电机多并行学习方案可以更有效地实现目标数据分布的优化。因此,本研究引入了一种新的GAN模型,称为基于集体学习机制的最优传输GAN(CLOT-GAN)模型,该模型包含多个鉴别器,包括深度卷积神经网络(DCNN)模型,Vision Transformer(ViT)和conformer。整合各种鉴别器的目的在于它们能够理解梅尔频谱图的共振峰分布,这是由集体学习机制促成的。同时,包含最优传输(OT)损失旨在采用OT理论的原则,精确地弥合源和目标数据分布之间的差距。在VCC 2018、VCTK和CMU-Arctic数据集上的实验验证证实,CLOT-GAN-VC模型在客观和主观评估方面优于现有的VC模型。
摘要:After demonstrating significant success in image synthesis, Generative Adversarial Network (GAN) models have likewise made significant progress in the field of speech synthesis, leveraging their capacity to adapt the precise distribution of target data through adversarial learning processes. Notably, in the realm of State-Of-The-Art (SOTA) GAN-based Voice Conversion (VC) models, there exists a substantial disparity in naturalness between real and GAN-generated speech samples. Furthermore, while many GAN models currently operate on a single generator discriminator learning approach, optimizing target data distribution is more effectively achievable through a single generator multi-discriminator learning scheme. Hence, this study introduces a novel GAN model named Collective Learning Mechanism-based Optimal Transport GAN (CLOT-GAN) model, incorporating multiple discriminators, including the Deep Convolutional Neural Network (DCNN) model, Vision Transformer (ViT), and conformer. The objective of integrating various discriminators lies in their ability to comprehend the formant distribution of mel-spectrograms, facilitated by a collective learning mechanism. Simultaneously, the inclusion of Optimal Transport (OT) loss aims to precisely bridge the gap between the source and target data distribution, employing the principles of OT theory. The experimental validation on VCC 2018, VCTK, and CMU-Arctic datasets confirms that the CLOT-GAN-VC model outperforms existing VC models in objective and subjective assessments.
【2】 MusFlow: Multimodal Music Generation via Conditional Flow Matching
标题: MusFlow:通过条件流匹配的多模式音乐生成链接:https://arxiv.org/abs/2504.13535
摘要:音乐生成的目标是根据不同的条件信息创建符合人类审美的音乐片段。尽管在从特定文本描述(例如,风格、流派、乐器),但实际应用仍然受到普通用户有限的专业知识或编写准确提示的时间的阻碍。为了弥合这一应用差距,本文介绍了MusFlow,一种新的多模态音乐生成模型,使用条件流匹配。我们采用多个多层感知器(MLP)将多模态条件信息对齐到音频的CLAP嵌入空间中。在预训练的VAE特征空间中,通过对齐特征嵌入指导条件流匹配来重建压缩的Mel谱图。MusFlow可以从图像、故事文本和音乐字幕生成音乐。为了收集用于模型训练的数据,受多代理协作的启发,我们构建了一个以微调的Qwen 2-VL模型为中心的智能数据注释工作流。使用此工作流程,我们构建了一个新的多模态音乐数据集MMusSet,每个样本包含图像,故事文本,音乐标题和音乐片段的四重内容。我们进行了四组实验:图像到音乐,故事到音乐,字幕到音乐,和多模态音乐生成。实验结果表明,无论输入条件是单峰还是多峰,MusFlow都能生成高质量的音乐作品。希望本文的工作能够推动音乐生成技术在多媒体领域的应用,使音乐创作更容易实现。我们生成的示例、代码和数据集可以在musflow.github.io上找到。
摘要:Music generation aims to create music segments that align with human aesthetics based on diverse conditional information. Despite advancements in generating music from specific textual descriptions (e.g., style, genre, instruments), the practical application is still hindered by ordinary users' limited expertise or time to write accurate prompts. To bridge this application gap, this paper introduces MusFlow, a novel multimodal music generation model using Conditional Flow Matching. We employ multiple Multi-Layer Perceptrons (MLPs) to align multimodal conditional information into the audio's CLAP embedding space. Conditional flow matching is trained to reconstruct the compressed Mel-spectrogram in the pretrained VAE latent space guided by aligned feature embedding. MusFlow can generate music from images, story texts, and music captions. To collect data for model training, inspired by multi-agent collaboration, we construct an intelligent data annotation workflow centered around a fine-tuned Qwen2-VL model. Using this workflow, we build a new multimodal music dataset, MMusSet, with each sample containing a quadruple of image, story text, music caption, and music piece. We conduct four sets of experiments: image-to-music, story-to-music, caption-to-music, and multimodal music generation. Experimental results demonstrate that MusFlow can generate high-quality music pieces whether the input conditions are unimodal or multimodal. We hope this work can advance the application of music generation in multimedia field, making music creation more accessible. Our generated samples, code and dataset are available at musflow.github.io.
【3】 Acoustic to Articulatory Inversion of Speech; Data Driven Approaches, Challenges, Applications, and Future Scope
标题: 语音的声学到关节翻转;数据驱动的方法、挑战、应用和未来范围链接:https://arxiv.org/abs/2504.13308
备注:This is a review paper about Acoustic to Articulatory inversion of speech, presented in an international conference. This paper has 8 pages and 2 figures
摘要:本文综述了语音声学-发音反转(AAI)的数据驱动方法。本综述文件考虑了过去十年(2011-2021年)出版的相关著作。选择标准包括(a)AAI的类型-说话人相关和说话人无关AAI,(b)工作目标-发音近似,发音特征空间选择和自动语音识别(ASR),探索声学和发音特征之间的相关性,以及计算机辅助语言训练的框架,(c)语料库-同时记录的语音(wav)和医学成像模型,如电磁关节描记术(EMA)、腭电描记术(EPG)、喉描记术、电声门描记术(EGG)、X射线电影放射成像术、超声和实时磁共振成像(rtMRI),(e)评估-由于AAI是一个非线性回归问题,性能评估主要通过相关系数(CC)、均方根误差(RMSE)进行,也考虑了均方误差(MSE)和平均格式误差(MFE)。AAI模型的实际应用可以提供一个更好的和用户友好的发音位置,特别是舌运动的图像反馈系统。这种轨迹反馈系统可用于为病理受试者提供语音、语言和言语治疗。
摘要:This review is focused on the data-driven approaches applied in different applications of Acoustic-to-Articulatory Inversion (AAI) of speech. This review paper considered the relevant works published in the last ten years (2011-2021). The selection criteria includes (a) type of AAI - Speaker Dependent and Speaker Independent AAI, (b) objectives of the work - Articulatory approximation, Articulatory Feature space selection and Automatic Speech Recognition (ASR), explore the correlation between acoustic and articulatory features, and framework for Computer-assisted language training, (c) Corpus - Simultaneously recorded speech (wav) and medical imaging models such as ElectroMagnetic Articulography (EMA), Electropalatography (EPG), Laryngography, Electroglottography (EGG), X-ray Cineradiography, Ultrasound, and real-time Magnetic Resonance Imaging (rtMRI), (d) Methods or models - recent works are considered, and therefore all the works are based on machine learning, (e) Evaluation - as AAI is a non-linear regression problem, the performance evaluation is mostly done by Correlation Coefficient (CC), Root Mean Square Error (RMSE), and also considered Mean Square Error (MSE), and Mean Format Error (MFE). The practical application of the AAI model can provide a better and user-friendly interpretable image feedback system of articulatory positions, especially tongue movement. Such trajectory feedback system can be used to provide phonetic, language, and speech therapy for pathological subjects.
【4】 Modeling L1 Influence on L2 Pronunciation: An MFCC-Based Framework for Explainable Machine Learning and Pedagogical Feedback
标题: 建模L1对L2发音的影响:基于MFCC的可解释机器学习和教学反馈框架链接:https://arxiv.org/abs/2504.13765
备注:27 pages (including references), 4 figures, 1 table. Combines statistical inference and explainable machine learning to model L1 influence in L2 pronunciation using MFCC features. Methodology and code are openly available via Zenodo and OSF: Zenodo: this https URL OSF: this https URL
摘要:本研究旨在探讨梅尔频率倒谱系数(MFCC)在扩展的第二语言(L2)英语语音中捕获第一语言(L1)迁移的程度。从GMU语音口音档案中提取来自普通话和美国英语L1扬声器的语音样本,转换为WAV格式,并处理以获得每个扬声器13个MFCC。结合推理统计(t检验,MANOVA,典型判别分析)和机器学习(随机森林分类)的多方法分析框架确定MFCC-1(宽带能量),MFCC-2(第一共振峰区域)和MFCC-5(清音和摩擦音能量)作为区分L1背景的最具鉴别力的特征。减少功能模型使用这些MFCC显着优于全功能模型,证实了McNemar的测试和非重叠的置信区间。研究结果实证支持第二语言的感知同化模型(PAM-L2)和语音学习模型(SLM),表明L1的条件变化,在L2语音的感知接地和声学量化。在方法上,该研究通过为L2发音建模提出一个透明的,数据高效的管道,为应用语言学和可解释的AI做出了贡献。研究结果还为ESL/EFL教学提供了教学启示,突出了L1特定的功能,可以为以理解力为导向的教学,课程设计和语音评估工具。
摘要:This study investigates the extent to which Mel-Frequency Cepstral Coefficients (MFCCs) capture first language (L1) transfer in extended second language (L2) English speech. Speech samples from Mandarin and American English L1 speakers were extracted from the GMU Speech Accent Archive, converted to WAV format, and processed to obtain thirteen MFCCs per speaker. A multi-method analytic framework combining inferential statistics (t-tests, MANOVA, Canonical Discriminant Analysis) and machine learning (Random Forest classification) identified MFCC-1 (broadband energy), MFCC-2 (first formant region), and MFCC-5 (voicing and fricative energy) as the most discriminative features for distinguishing L1 backgrounds. A reduced-feature model using these MFCCs significantly outperformed the full-feature model, as confirmed by McNemar's test and non-overlapping confidence intervals. The findings empirically support the Perceptual Assimilation Model for L2 (PAM-L2) and the Speech Learning Model (SLM), demonstrating that L1-conditioned variation in L2 speech is both perceptually grounded and acoustically quantifiable. Methodologically, the study contributes to applied linguistics and explainable AI by proposing a transparent, data-efficient pipeline for L2 pronunciation modeling. The results also offer pedagogical implications for ESL/EFL instruction by highlighting L1-specific features that can inform intelligibility-oriented instruction, curriculum design, and speech assessment tools.
【1】 Modeling L1 Influence on L2 Pronunciation: An MFCC-Based Framework for Explainable Machine Learning and Pedagogical Feedback
标题: 建模L1对L2发音的影响:基于MFCC的可解释机器学习和教学反馈框架链接:https://arxiv.org/abs/2504.13765
备注:27 pages (including references), 4 figures, 1 table. Combines statistical inference and explainable machine learning to model L1 influence in L2 pronunciation using MFCC features. Methodology and code are openly available via Zenodo and OSF: Zenodo: this https URL OSF: this https URL
摘要:本研究旨在探讨梅尔频率倒谱系数(MFCC)在扩展的第二语言(L2)英语语音中捕获第一语言(L1)迁移的程度。从GMU语音口音档案中提取来自普通话和美国英语L1扬声器的语音样本,转换为WAV格式,并处理以获得每个扬声器13个MFCC。结合推理统计(t检验,MANOVA,典型判别分析)和机器学习(随机森林分类)的多方法分析框架确定MFCC-1(宽带能量),MFCC-2(第一共振峰区域)和MFCC-5(清音和摩擦音能量)作为区分L1背景的最具鉴别力的特征。减少功能模型使用这些MFCC显着优于全功能模型,证实了McNemar的测试和非重叠的置信区间。这些发现从经验上支持了第二语言感知同化模型(PAM-L2)和言语学习模型(SLM),表明第二语言言语中的第一语言条件变化既有感知基础,又可以在声学上量化。在方法上,该研究通过为L2发音建模提出一个透明的,数据高效的管道,为应用语言学和可解释的AI做出了贡献。研究结果还为ESL/EFL教学提供了教学启示,突出了L1特定的功能,可以为以理解力为导向的教学,课程设计和语音评估工具。
摘要:This study investigates the extent to which Mel-Frequency Cepstral Coefficients (MFCCs) capture first language (L1) transfer in extended second language (L2) English speech. Speech samples from Mandarin and American English L1 speakers were extracted from the GMU Speech Accent Archive, converted to WAV format, and processed to obtain thirteen MFCCs per speaker. A multi-method analytic framework combining inferential statistics (t-tests, MANOVA, Canonical Discriminant Analysis) and machine learning (Random Forest classification) identified MFCC-1 (broadband energy), MFCC-2 (first formant region), and MFCC-5 (voicing and fricative energy) as the most discriminative features for distinguishing L1 backgrounds. A reduced-feature model using these MFCCs significantly outperformed the full-feature model, as confirmed by McNemar's test and non-overlapping confidence intervals. The findings empirically support the Perceptual Assimilation Model for L2 (PAM-L2) and the Speech Learning Model (SLM), demonstrating that L1-conditioned variation in L2 speech is both perceptually grounded and acoustically quantifiable. Methodologically, the study contributes to applied linguistics and explainable AI by proposing a transparent, data-efficient pipeline for L2 pronunciation modeling. The results also offer pedagogical implications for ESL/EFL instruction by highlighting L1-specific features that can inform intelligibility-oriented instruction, curriculum design, and speech assessment tools.
【2】 Collective Learning Mechanism based Optimal Transport Generative Adversarial Network for Non-parallel Voice Conversion
标题: 基于集体学习机制的非并行语音转换最优传输生成对抗网络链接:https://arxiv.org/abs/2504.13791
备注:7 pages, 2 figures, 3 tables
摘要:在图像合成方面取得了巨大成功之后,生成对抗网络(GAN)模型在语音合成领域也取得了重大进展,利用其通过对抗学习过程适应目标数据精确分布的能力。值得注意的是,在最先进的(SOTA)基于GAN的语音转换(VC)模型领域,真实语音样本和GAN生成的语音样本之间的自然度存在很大差异。此外,虽然许多GAN模型目前在单发电机并行学习方法上操作,但是通过单发电机多并行学习方案可以更有效地实现目标数据分布的优化。因此,本研究引入了一种新的GAN模型,称为基于集体学习机制的最优传输GAN(CLOT-GAN)模型,该模型包含多个鉴别器,包括深度卷积神经网络(DCNN)模型,Vision Transformer(ViT)和conformer。整合各种鉴别器的目的在于它们能够理解梅尔频谱图的共振峰分布,这是由集体学习机制促成的。同时,包含最优传输(OT)损失旨在采用OT理论的原则,精确地弥合源和目标数据分布之间的差距。在VCC 2018、VCTK和CMU-Arctic数据集上的实验验证证实,CLOT-GAN-VC模型在客观和主观评估方面优于现有的VC模型。
摘要:After demonstrating significant success in image synthesis, Generative Adversarial Network (GAN) models have likewise made significant progress in the field of speech synthesis, leveraging their capacity to adapt the precise distribution of target data through adversarial learning processes. Notably, in the realm of State-Of-The-Art (SOTA) GAN-based Voice Conversion (VC) models, there exists a substantial disparity in naturalness between real and GAN-generated speech samples. Furthermore, while many GAN models currently operate on a single generator discriminator learning approach, optimizing target data distribution is more effectively achievable through a single generator multi-discriminator learning scheme. Hence, this study introduces a novel GAN model named Collective Learning Mechanism-based Optimal Transport GAN (CLOT-GAN) model, incorporating multiple discriminators, including the Deep Convolutional Neural Network (DCNN) model, Vision Transformer (ViT), and conformer. The objective of integrating various discriminators lies in their ability to comprehend the formant distribution of mel-spectrograms, facilitated by a collective learning mechanism. Simultaneously, the inclusion of Optimal Transport (OT) loss aims to precisely bridge the gap between the source and target data distribution, employing the principles of OT theory. The experimental validation on VCC 2018, VCTK, and CMU-Arctic datasets confirms that the CLOT-GAN-VC model outperforms existing VC models in objective and subjective assessments.
【3】 MusFlow: Multimodal Music Generation via Conditional Flow Matching
标题: MusFlow:通过条件流匹配的多模式音乐生成链接:https://arxiv.org/abs/2504.13535
摘要:音乐生成的目标是根据不同的条件信息创建符合人类审美的音乐片段。尽管在从特定文本描述(例如,风格、流派、乐器),但实际应用仍然受到普通用户有限的专业知识或编写准确提示的时间的阻碍。为了弥合这一应用差距,本文介绍了MusFlow,一种新的多模态音乐生成模型,使用条件流匹配。我们采用多个多层感知器(MLP)将多模态条件信息对齐到音频的CLAP嵌入空间中。在预训练的VAE特征空间中,通过对齐特征嵌入指导条件流匹配来重建压缩的Mel谱图。MusFlow可以从图像、故事文本和音乐字幕生成音乐。为了收集用于模型训练的数据,受多代理协作的启发,我们构建了一个以微调的Qwen 2-VL模型为中心的智能数据注释工作流。使用此工作流程,我们构建了一个新的多模态音乐数据集MMusSet,每个样本包含图像,故事文本,音乐标题和音乐片段的四重内容。我们进行了四组实验:图像到音乐,故事到音乐,字幕到音乐,和多模态音乐生成。实验结果表明,无论输入条件是单峰还是多峰,MusFlow都能生成高质量的音乐作品。希望本文的工作能够推动音乐生成技术在多媒体领域的应用,使音乐创作更容易实现。我们生成的示例、代码和数据集可以在musflow.github.io上找到。
摘要:Music generation aims to create music segments that align with human aesthetics based on diverse conditional information. Despite advancements in generating music from specific textual descriptions (e.g., style, genre, instruments), the practical application is still hindered by ordinary users' limited expertise or time to write accurate prompts. To bridge this application gap, this paper introduces MusFlow, a novel multimodal music generation model using Conditional Flow Matching. We employ multiple Multi-Layer Perceptrons (MLPs) to align multimodal conditional information into the audio's CLAP embedding space. Conditional flow matching is trained to reconstruct the compressed Mel-spectrogram in the pretrained VAE latent space guided by aligned feature embedding. MusFlow can generate music from images, story texts, and music captions. To collect data for model training, inspired by multi-agent collaboration, we construct an intelligent data annotation workflow centered around a fine-tuned Qwen2-VL model. Using this workflow, we build a new multimodal music dataset, MMusSet, with each sample containing a quadruple of image, story text, music caption, and music piece. We conduct four sets of experiments: image-to-music, story-to-music, caption-to-music, and multimodal music generation. Experimental results demonstrate that MusFlow can generate high-quality music pieces whether the input conditions are unimodal or multimodal. We hope this work can advance the application of music generation in multimedia field, making music creation more accessible. Our generated samples, code and dataset are available at musflow.github.io.
【4】 Acoustic to Articulatory Inversion of Speech; Data Driven Approaches, Challenges, Applications, and Future Scope
标题: 语音的声学到关节翻转;数据驱动的方法、挑战、应用和未来范围链接:https://arxiv.org/abs/2504.13308
备注:This is a review paper about Acoustic to Articulatory inversion of speech, presented in an international conference. This paper has 8 pages and 2 figures
摘要:本文综述了语音声学-发音反转(AAI)的数据驱动方法。本综述文件考虑了过去十年(2011-2021年)出版的相关著作。选择标准包括(a)AAI的类型-说话人相关和说话人无关AAI,(b)工作目标-发音近似,发音特征空间选择和自动语音识别(ASR),探索声学和发音特征之间的相关性,以及计算机辅助语言训练的框架,(c)语料库-同时记录的语音(wav)和医学成像模型,如电磁关节描记术(EMA)、腭电描记术(EPG)、喉描记术、电声门描记术(EGG)、X射线电影放射成像术、超声和实时磁共振成像(rtMRI),(e)评估-由于AAI是一个非线性回归问题,性能评估主要通过相关系数(CC)、均方根误差(RMSE)进行,也考虑了均方误差(MSE)和平均格式误差(MFE)。AAI模型的实际应用可以提供一个更好的和用户友好的发音位置,特别是舌运动的图像反馈系统。这种轨迹反馈系统可用于为病理受试者提供语音、语言和言语治疗。
摘要:This review is focused on the data-driven approaches applied in different applications of Acoustic-to-Articulatory Inversion (AAI) of speech. This review paper considered the relevant works published in the last ten years (2011-2021). The selection criteria includes (a) type of AAI - Speaker Dependent and Speaker Independent AAI, (b) objectives of the work - Articulatory approximation, Articulatory Feature space selection and Automatic Speech Recognition (ASR), explore the correlation between acoustic and articulatory features, and framework for Computer-assisted language training, (c) Corpus - Simultaneously recorded speech (wav) and medical imaging models such as ElectroMagnetic Articulography (EMA), Electropalatography (EPG), Laryngography, Electroglottography (EGG), X-ray Cineradiography, Ultrasound, and real-time Magnetic Resonance Imaging (rtMRI), (d) Methods or models - recent works are considered, and therefore all the works are based on machine learning, (e) Evaluation - as AAI is a non-linear regression problem, the performance evaluation is mostly done by Correlation Coefficient (CC), Root Mean Square Error (RMSE), and also considered Mean Square Error (MSE), and Mean Format Error (MFE). The practical application of the AAI model can provide a better and user-friendly interpretable image feedback system of articulatory positions, especially tongue movement. Such trajectory feedback system can be used to provide phonetic, language, and speech therapy for pathological subjects.
机器翻译由腾讯交互翻译提供,仅供参考
