今日论文合集:cs.SD语音6篇,eess.AS音频处理6篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】VoxGenesis: Unsupervised Discovery of Latent Speaker Manifold for Speech  Synthesis
标题:VoxGenesis:语音合成中潜在说话人流形的无监督发现
链接:https://arxiv.org/abs/2403.00529
作者:Weiwei Lin,Chenhang He,Man-Wai Mak,Jiachen Lian,Kong Aik Lee
备注:preprint
摘要:实现对人类声音的细微而准确的模仿一直是人工智能的长期目标。尽管近年来语音合成技术取得了很大的进步,但语音合成模型的主流仍然依赖于有监督的说话人建模和显式参考话语。然而,人类语音有许多方面,如情感,语调和说话风格,很难获得准确的标签。在本文中,我们提出了VoxGenesis,一种新的无监督语音合成框架,可以发现一个潜在的扬声器流形和有意义的语音编辑方向,而无需监督。VoxGenesis在概念上很简单。VoxGenesis不是将语音特征确定性地映射到波形,而是将高斯分布转换为由语义标记调节和对齐的语音分布。这迫使模型学习从语义内容中分离出来的说话者分布。在推理过程中,从高斯分布中采样,使得能够创建具有不同特征的新颖扬声器。更重要的是,对潜在空间的探索揭示了与特定说话者特征(例如性别属性、音高、音调和情感)相关联的人类可解释的方向,从而允许通过沿着这些识别的方向操纵潜在代码来进行语音编辑。我们进行了广泛的实验,以评估所提出的VoxGenesis使用主观和客观的指标,发现它产生显着更多样化和现实的扬声器具有鲜明的特点比以前的方法。我们还表明,潜在的空间操作产生一致的和人类可识别的效果,是不会损害语音质量,这是不可能与以前的方法。VoxGenesis的音频样本可以在以下位置找到:\url{https://bit.ly/VoxGenesis}。
摘要:Achieving nuanced and accurate emulation of human voice has been a longstanding goal in artificial intelligence. Although significant progress has been made in recent years, the mainstream of speech synthesis models still relies on supervised speaker modeling and explicit reference utterances. However, there are many aspects of human voice, such as emotion, intonation, and speaking style, for which it is hard to obtain accurate labels. In this paper, we propose VoxGenesis, a novel unsupervised speech synthesis framework that can discover a latent speaker manifold and meaningful voice editing directions without supervision. VoxGenesis is conceptually simple. Instead of mapping speech features to waveforms deterministically, VoxGenesis transforms a Gaussian distribution into speech distributions conditioned and aligned by semantic tokens. This forces the model to learn a speaker distribution disentangled from the semantic content. During the inference, sampling from the Gaussian distribution enables the creation of novel speakers with distinct characteristics. More importantly, the exploration of latent space uncovers human-interpretable directions associated with specific speaker characteristics such as gender attributes, pitch, tone, and emotion, allowing for voice editing by manipulating the latent codes along these identified directions. We conduct extensive experiments to evaluate the proposed VoxGenesis using both subjective and objective metrics, finding that it produces significantly more diverse and realistic speakers with distinct characteristics than the previous approaches. We also show that latent space manipulation produces consistent and human-identifiable effects that are not detrimental to the speech quality, which was not possible with previous approaches. Audio samples of VoxGenesis can be found at: \url{https://bit.ly/VoxGenesis}.


【2】 Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn  Medical Interview
标题:多轮医学会诊端到端语音识别的后解码级偏
链接:https://arxiv.org/abs/2403.00370
作者:Heyang Liu,Yu Wang,Yanfeng Wang
摘要:端到端(E2 E)方法正在逐步取代自动语音识别(ASR)任务的混合模型。然而,E2 E模型的优化缺乏一种直观的方法来处理解码移位,特别是在具有大量具有特定重要含义的特定领域稀有词的场景中。此外,学术界缺乏知识密集型语音数据集一直是一个重要的限制因素,常用的语音语料库与现实对话表现出显着的差异。为了解决这些挑战,我们提出了医疗面试(MED-IT),一个多轮咨询语音数据集,包含大量的知识密集型命名实体。我们还探索了提高E2 E模型中稀有词识别性能的方法。我们提出了一种新的方法,解码器后偏置,它构造了一个基于训练transmittance的分布的变换概率矩阵。这引导模型优先识别偏置列表中的单词。在我们的实验中,对于训练语音中出现10到20次和1到5次的稀有词子集,所提出的方法分别实现了9.3%和5.1%的相对改善。
摘要:End-to-end (E2E) approach is gradually replacing hybrid models for automatic speech recognition (ASR) tasks. However, the optimization of E2E models lacks an intuitive method for handling decoding shifts, especially in scenarios with a large number of domain-specific rare words that hold specific important meanings. Furthermore, the absence of knowledge-intensive speech datasets in academia has been a significant limiting factor, and the commonly used speech corpora exhibit significant disparities with realistic conversation. To address these challenges, we present Medical Interview (MED-IT), a multi-turn consultation speech dataset that contains a substantial number of knowledge-intensive named entities. We also explore methods to enhance the recognition performance of rare words for E2E models. We propose a novel approach, post-decoder biasing, which constructs a transform probability matrix based on the distribution of training transcriptions. This guides the model to prioritize recognizing words in the biasing list. In our experiments, for subsets of rare words appearing in the training speech between 10 and 20 times, and between 1 and 5 times, the proposed method achieves a relative improvement of 9.3% and 5.1%, respectively.

【3】 CustomListener: Text-guided Responsive Interaction for User-friendly  Listening Head Generation标题:自定义:文本引导的响应式交互,用于用户友好的听力头部生成
链接:https://arxiv.org/abs/2403.00274
作者:Xi Liu,Ying Guo,Cheng Zhen,Tong Li,Yingying Ao,Pengfei Yan
备注:Accepted by CVPR 2024
摘要:听觉头部生成的目的是通过对说话人和听者之间的动态转换关系进行建模,合成一个非言语响应的听觉头部,听觉主体生成技术在虚拟交互中的应用促进了许多研究工作的开展,实现了多样化和细粒度的运动生成。然而,他们只能通过简单的情感标签来操纵动作,而不能自由地控制听者的动作。由于监听器代理应该具有类似人类的属性(例如身份,个性),可以由用户自由定制,这限制了它们的真实性。在本文中,我们提出了一个用户友好的框架,称为定制的实现自由形式的文本之前引导的听众生成。为了实现说者-听者的协调,我们设计了一个静态到动态肖像模块(SDP),它与说话人信息交互,将静态文本转换为具有完成节奏和幅度信息的动态肖像令牌。为了实现片段之间的一致性,设计了过去引导生成模块(Past Guided Generation Module,PGG),通过运动先验来保持自定义听者属性的一致性,并利用以肖像标记和运动先验为条件的基于扩散的结构来实现可控生成。为了训练和评估我们的模型,我们构建了两个基于ViCo和RealTalk的文本注释的听力头数据集,它们提供了文本-视频配对标签。大量的实验验证了该模型的有效性。
摘要:Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion.The applications of listener agent generation in virtual interaction have promoted many works achieving the diverse and fine-grained motion generation. However, they can only manipulate motions through simple emotional labels, but cannot freely control the listener's motions. Since listener agents should have human-like attributes (e.g. identity, personality) which can be freely customized by users, this limits their realism. In this paper, we propose a user-friendly framework called CustomListener to realize the free-form text prior guided listener generation. To achieve speaker-listener coordination, we design a Static to Dynamic Portrait module (SDP), which interacts with speaker information to transform static text into dynamic portrait token with completion rhythm and amplitude information. To achieve coherence between segments, we design a Past Guided Generation Module (PGG) to maintain the consistency of customized listener attributes through the motion prior, and utilize a diffusion-based structure conditioned on the portrait token and the motion prior to realize the controllable generation. To train and evaluate our model, we have constructed two text-annotated listening head datasets based on ViCo and RealTalk, which provide text-video paired labels. Extensive experiments have verified the effectiveness of our model.


【4】 Transcription and translation of videos using fine-tuned XLSR Wav2Vec2  on custom dataset and mBART
标题:在定制数据集和mBART上使用微调的XLSR Wav2Vec2转录和翻译视频
链接:https://arxiv.org/abs/2403.00212
作者:Aniket Tathe,Anand Kamble,Suyash Kumbharkar,Atharva Bhandare,Anirban C. Mitra
摘要:这项研究解决了用最少的数据训练个性化语音的ASR模型的挑战。利用YouTube视频中仅14分钟的自定义音频,我们采用基于检索的语音转换(RVC)来创建自定义Common Voice 16.0语料库。随后,跨语言自监督表示(XLSR)Wav2Vec2模型在此数据集上进行了微调。开发的基于Web的GUI有效地转录和翻译输入印地语视频。通过集成XLSR Wav2Vec2和mBART,该系统将翻译文本与视频时间轴对齐,为多语言视频内容转录和个性化语音翻译提供了一个可访问的解决方案。
摘要:This research addresses the challenge of training an ASR model for personalized voices with minimal data. Utilizing just 14 minutes of custom audio from a YouTube video, we employ Retrieval-Based Voice Conversion (RVC) to create a custom Common Voice 16.0 corpus. Subsequently, a Cross-lingual Self-supervised Representations (XLSR) Wav2Vec2 model is fine-tuned on this dataset. The developed web-based GUI efficiently transcribes and translates input Hindi videos. By integrating XLSR Wav2Vec2 and mBART, the system aligns the translated text with the video timeline, delivering an accessible solution for multilingual video content transcription and translation for personalized voice.

【5】 The Impact of Frequency Bands on Acoustic Anomaly Detection of Machines  using Deep Learning Based Model
标题:频带对基于深度学习模型的机器声学异常检测的影响
链接:https://arxiv.org/abs/2403.00379
作者:Tin Nguyen,Lam Pham,Phat Lam,Dat Ngo,Hieu Tang,Alexander Schindler
摘要:在本文中,我们提出了一种基于深度学习的机器声学异常检测模型,该模型通过分析机器声音来检测异常机器。通过大量的实验,我们表明,多种技术的伪音频,音频段,数据增强,马氏距离,窄频带,主要集中在特征工程,是有效的,以提高系统的性能。在评估技术中,窄频带呈现出显著的影响。事实上,我们提出的模型,重点是窄频带,优于DCASE基准数据集的DCASE 2022任务2开发集。本文中指出的窄频带的重要作用启发了机器声学异常检测任务的研究界进一步研究并提出专注于频带的新型网络架构。
摘要:In this paper, we propose a deep learning based model for Acoustic Anomaly Detection of Machines, the task for detecting abnormal machines by analysing the machine sound. By conducting extensive experiments, we indicate that multiple techniques of pseudo audios, audio segment, data augmentation, Mahalanobis distance, and narrow frequency bands, which mainly focus on feature engineering, are effective to enhance the system performance. Among the evaluating techniques, the narrow frequency bands presents a significant impact. Indeed, our proposed model, which focuses on the narrow frequency bands, outperforms the DCASE baseline on the benchmark dataset of DCASE 2022 Task 2 Development set. The important role of the narrow frequency bands indicated in this paper inspires the research community on the task of Acoustic Anomaly Detection of Machines to further investigate and propose novel network architectures focusing on the frequency bands.


【6】 Efficient Adapter Tuning of Pre-trained Speech Models for Automatic  Speaker Verification
标题:用于自动说话人确认的预训练语音模型的高效自适应调整
链接:https://arxiv.org/abs/2403.00293
作者:Mufan Sang,John H. L. Hansen
备注:Accepted to ICASSP 2024
摘要:自监督语音模型具有良好的泛化能力,在预训练和微调范式下,在各种下游语音任务中表现出令人印象深刻的性能。然而,随着预训练模型的规模不断增长,由于计算和存储开销以及过拟合的风险,微调变得实际上不可行。适配器是插入到预训练模型中的轻量级模块,以促进参数有效的适应。在本文中,我们提出了一个有效的适配器框架,旨在适应自监督语音模型的说话人验证任务。通过并行适配器设计,我们提出的框架将两种类型的适配器插入到预先训练的模型中,允许在中间Transformer层内调整潜在特征,并从所有Transformer层输出嵌入。我们进行了全面的实验,以验证所提出的框架的效率和有效性。在VoxCeleb1数据集上的实验结果表明,所提出的适配器优于微调和其他参数有效的迁移学习方法,在仅更新5%的参数的情况下实现了卓越的性能。
摘要:With excellent generalization ability, self-supervised speech models have shown impressive performance on various downstream speech tasks in the pre-training and fine-tuning paradigm. However, as the growing size of pre-trained models, fine-tuning becomes practically unfeasible due to heavy computation and storage overhead, as well as the risk of overfitting. Adapters are lightweight modules inserted into pre-trained models to facilitate parameter-efficient adaptation. In this paper, we propose an effective adapter framework designed for adapting self-supervised speech models to the speaker verification task. With a parallel adapter design, our proposed framework inserts two types of adapters into the pre-trained model, allowing the adaptation of latent features within intermediate Transformer layers and output embeddings from all Transformer layers. We conduct comprehensive experiments to validate the efficiency and effectiveness of the proposed framework. Experimental results on the VoxCeleb1 dataset demonstrate that the proposed adapters surpass fine-tuning and other parameter-efficient transfer learning methods, achieving superior performance while updating only 5% of the parameters.

eess.AS音频处理
【1】 The Impact of Frequency Bands on Acoustic Anomaly Detection of Machines  using Deep Learning Based Model
标题:频带对基于深度学习模型的机器声学异常检测的影响
链接:https://arxiv.org/abs/2403.00379
作者:Tin Nguyen,Lam Pham,Phat Lam,Dat Ngo,Hieu Tang,Alexander Schindler
摘要:在本文中,我们提出了一种基于深度学习的机器声学异常检测模型,该模型通过分析机器声音来检测异常机器。通过大量的实验,我们表明,多种技术的伪音频,音频段,数据增强,马氏距离,窄频带,主要集中在特征工程,是有效的,以提高系统的性能。在评估技术中,窄频带呈现出显著的影响。事实上,我们提出的模型,重点是窄频带,优于DCASE基准数据集的DCASE 2022任务2开发集。本文中指出的窄频带的重要作用启发了机器声学异常检测任务的研究界进一步研究并提出专注于频带的新型网络架构。
摘要:In this paper, we propose a deep learning based model for Acoustic Anomaly Detection of Machines, the task for detecting abnormal machines by analysing the machine sound. By conducting extensive experiments, we indicate that multiple techniques of pseudo audios, audio segment, data augmentation, Mahalanobis distance, and narrow frequency bands, which mainly focus on feature engineering, are effective to enhance the system performance. Among the evaluating techniques, the narrow frequency bands presents a significant impact. Indeed, our proposed model, which focuses on the narrow frequency bands, outperforms the DCASE baseline on the benchmark dataset of DCASE 2022 Task 2 Development set. The important role of the narrow frequency bands indicated in this paper inspires the research community on the task of Acoustic Anomaly Detection of Machines to further investigate and propose novel network architectures focusing on the frequency bands.

【2】 Efficient Adapter Tuning of Pre-trained Speech Models for Automatic  Speaker Verification
标题:用于自动说话人确认的预训练语音模型的高效自适应调整
链接:https://arxiv.org/abs/2403.00293
作者:Mufan Sang,John H. L. Hansen
备注:Accepted to ICASSP 2024
摘要:自监督语音模型具有良好的泛化能力,在预训练和微调范式下,在各种下游语音任务中表现出令人印象深刻的性能。然而,随着预训练模型的规模不断增长,由于计算和存储开销以及过拟合的风险,微调变得实际上不可行。适配器是插入到预训练模型中的轻量级模块,以促进参数有效的适应。在本文中,我们提出了一个有效的适配器框架,旨在适应自监督语音模型的说话人验证任务。通过并行适配器设计,我们提出的框架将两种类型的适配器插入到预先训练的模型中,允许在中间Transformer层内调整潜在特征,并从所有Transformer层输出嵌入。我们进行了全面的实验,以验证所提出的框架的效率和有效性。在VoxCeleb1数据集上的实验结果表明,所提出的适配器优于微调和其他参数有效的迁移学习方法,在仅更新5%的参数的情况下实现了卓越的性能。
摘要:With excellent generalization ability, self-supervised speech models have shown impressive performance on various downstream speech tasks in the pre-training and fine-tuning paradigm. However, as the growing size of pre-trained models, fine-tuning becomes practically unfeasible due to heavy computation and storage overhead, as well as the risk of overfitting. Adapters are lightweight modules inserted into pre-trained models to facilitate parameter-efficient adaptation. In this paper, we propose an effective adapter framework designed for adapting self-supervised speech models to the speaker verification task. With a parallel adapter design, our proposed framework inserts two types of adapters into the pre-trained model, allowing the adaptation of latent features within intermediate Transformer layers and output embeddings from all Transformer layers. We conduct comprehensive experiments to validate the efficiency and effectiveness of the proposed framework. Experimental results on the VoxCeleb1 dataset demonstrate that the proposed adapters surpass fine-tuning and other parameter-efficient transfer learning methods, achieving superior performance while updating only 5% of the parameters.


【3】 VoxGenesis: Unsupervised Discovery of Latent Speaker Manifold for Speech  Synthesis
标题:VoxGenesis:语音合成中潜在说话人流形的无监督发现
链接:https://arxiv.org/abs/2403.00529
作者:Weiwei Lin,Chenhang He,Man-Wai Mak,Jiachen Lian,Kong Aik Lee
备注:preprint
摘要:实现对人类声音的细微而准确的模仿一直是人工智能的长期目标。尽管近年来语音合成技术取得了很大的进步,但语音合成模型的主流仍然依赖于有监督的说话人建模和显式参考话语。然而,人类语音有许多方面,如情感,语调和说话风格,很难获得准确的标签。在本文中,我们提出了VoxGenesis,一种新的无监督语音合成框架,可以发现一个潜在的扬声器流形和有意义的语音编辑方向,而无需监督。VoxGenesis在概念上很简单。VoxGenesis不是将语音特征确定性地映射到波形,而是将高斯分布转换为由语义标记调节和对齐的语音分布。这迫使模型学习从语义内容中分离出来的说话者分布。在推理过程中,从高斯分布中采样,使得能够创建具有不同特征的新颖扬声器。更重要的是,对潜在空间的探索揭示了与特定说话者特征(例如性别属性、音高、音调和情感)相关联的人类可解释的方向,从而允许通过沿着这些识别的方向操纵潜在代码来进行语音编辑。我们进行了广泛的实验,以评估所提出的VoxGenesis使用主观和客观的指标,发现它产生显着更多样化和现实的扬声器具有鲜明的特点比以前的方法。我们还表明,潜在的空间操作产生一致的和人类可识别的效果,是不会损害语音质量,这是不可能与以前的方法。VoxGenesis的音频样本可以在以下位置找到:\url{https://bit.ly/VoxGenesis}。
摘要:Achieving nuanced and accurate emulation of human voice has been a longstanding goal in artificial intelligence. Although significant progress has been made in recent years, the mainstream of speech synthesis models still relies on supervised speaker modeling and explicit reference utterances. However, there are many aspects of human voice, such as emotion, intonation, and speaking style, for which it is hard to obtain accurate labels. In this paper, we propose VoxGenesis, a novel unsupervised speech synthesis framework that can discover a latent speaker manifold and meaningful voice editing directions without supervision. VoxGenesis is conceptually simple. Instead of mapping speech features to waveforms deterministically, VoxGenesis transforms a Gaussian distribution into speech distributions conditioned and aligned by semantic tokens. This forces the model to learn a speaker distribution disentangled from the semantic content. During the inference, sampling from the Gaussian distribution enables the creation of novel speakers with distinct characteristics. More importantly, the exploration of latent space uncovers human-interpretable directions associated with specific speaker characteristics such as gender attributes, pitch, tone, and emotion, allowing for voice editing by manipulating the latent codes along these identified directions. We conduct extensive experiments to evaluate the proposed VoxGenesis using both subjective and objective metrics, finding that it produces significantly more diverse and realistic speakers with distinct characteristics than the previous approaches. We also show that latent space manipulation produces consistent and human-identifiable effects that are not detrimental to the speech quality, which was not possible with previous approaches. Audio samples of VoxGenesis can be found at: \url{https://bit.ly/VoxGenesis}.

【4】 Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn  Medical Interview
标题:多轮医学会诊端到端语音识别的后解码级偏
链接:https://arxiv.org/abs/2403.00370
作者:Heyang Liu,Yu Wang,Yanfeng Wang
摘要:端到端(E2 E)方法正在逐步取代自动语音识别(ASR)任务的混合模型。然而,E2 E模型的优化缺乏一种直观的方法来处理解码移位,特别是在具有大量具有特定重要含义的特定领域稀有词的场景中。此外,学术界缺乏知识密集型语音数据集一直是一个重要的限制因素,常用的语音语料库与现实对话表现出显着的差异。为了解决这些挑战,我们提出了医疗面试(MED-IT),一个多轮咨询语音数据集,包含大量的知识密集型命名实体。我们还探索了提高E2 E模型中稀有词识别性能的方法。我们提出了一种新的方法,解码器后偏置,它构造了一个基于训练transmittance的分布的变换概率矩阵。这引导模型优先识别偏置列表中的单词。在我们的实验中,对于训练语音中出现10到20次和1到5次的稀有词子集,所提出的方法分别实现了9.3%和5.1%的相对改善。
摘要:End-to-end (E2E) approach is gradually replacing hybrid models for automatic speech recognition (ASR) tasks. However, the optimization of E2E models lacks an intuitive method for handling decoding shifts, especially in scenarios with a large number of domain-specific rare words that hold specific important meanings. Furthermore, the absence of knowledge-intensive speech datasets in academia has been a significant limiting factor, and the commonly used speech corpora exhibit significant disparities with realistic conversation. To address these challenges, we present Medical Interview (MED-IT), a multi-turn consultation speech dataset that contains a substantial number of knowledge-intensive named entities. We also explore methods to enhance the recognition performance of rare words for E2E models. We propose a novel approach, post-decoder biasing, which constructs a transform probability matrix based on the distribution of training transcriptions. This guides the model to prioritize recognizing words in the biasing list. In our experiments, for subsets of rare words appearing in the training speech between 10 and 20 times, and between 1 and 5 times, the proposed method achieves a relative improvement of 9.3% and 5.1%, respectively.

【5】 CustomListener: Text-guided Responsive Interaction for User-friendly  Listening Head Generation标题:CustomListener:文本引导的响应式交互,可生成用户友好的听头
链接:https://arxiv.org/abs/2403.00274
作者:Xi Liu,Ying Guo,Cheng Zhen,Tong Li,Yingying Ao,Pengfei Yan
备注:Accepted by CVPR 2024
摘要:听觉头部生成的目的是通过对说话人和听者之间的动态转换关系进行建模,合成一个非言语响应的听觉头部,听觉主体生成技术在虚拟交互中的应用促进了许多研究工作的开展,实现了多样化和细粒度的运动生成。然而,他们只能通过简单的情感标签来操纵动作,而不能自由地控制听者的动作。由于监听器代理应该具有类似人类的属性(例如身份,个性),可以由用户自由定制,这限制了它们的真实性。在本文中,我们提出了一个用户友好的框架,称为定制的实现自由形式的文本之前引导的听众生成。为了实现说者-听者的协调,我们设计了一个静态到动态肖像模块(SDP),它与说话人信息交互,将静态文本转换为具有完成节奏和幅度信息的动态肖像令牌。为了实现片段之间的一致性,设计了过去引导生成模块(Past Guided Generation Module,PGG),通过运动先验来保持自定义听者属性的一致性,并利用以肖像标记和运动先验为条件的基于扩散的结构来实现可控生成。为了训练和评估我们的模型,我们构建了两个基于ViCo和RealTalk的文本注释的听力头数据集,它们提供了文本-视频配对标签。大量的实验验证了该模型的有效性。
摘要:Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion.The applications of listener agent generation in virtual interaction have promoted many works achieving the diverse and fine-grained motion generation. However, they can only manipulate motions through simple emotional labels, but cannot freely control the listener's motions. Since listener agents should have human-like attributes (e.g. identity, personality) which can be freely customized by users, this limits their realism. In this paper, we propose a user-friendly framework called CustomListener to realize the free-form text prior guided listener generation. To achieve speaker-listener coordination, we design a Static to Dynamic Portrait module (SDP), which interacts with speaker information to transform static text into dynamic portrait token with completion rhythm and amplitude information. To achieve coherence between segments, we design a Past Guided Generation Module (PGG) to maintain the consistency of customized listener attributes through the motion prior, and utilize a diffusion-based structure conditioned on the portrait token and the motion prior to realize the controllable generation. To train and evaluate our model, we have constructed two text-annotated listening head datasets based on ViCo and RealTalk, which provide text-video paired labels. Extensive experiments have verified the effectiveness of our model.


【6】 Transcription and translation of videos using fine-tuned XLSR Wav2Vec2  on custom dataset and mBART
标题:在定制数据集和mBART上使用微调的XLSR Wav2Vec2转录和翻译视频
链接:https://arxiv.org/abs/2403.00212
作者:Aniket Tathe,Anand Kamble,Suyash Kumbharkar,Atharva Bhandare,Anirban C. Mitra
摘要:这项研究解决了用最少的数据训练个性化语音的ASR模型的挑战。利用YouTube视频中仅14分钟的自定义音频,我们采用基于检索的语音转换(RVC)来创建自定义Common Voice 16.0语料库。随后,跨语言自监督表示(XLSR)Wav2Vec2模型在此数据集上进行了微调。开发的基于Web的GUI有效地转录和翻译输入印地语视频。通过集成XLSR Wav2Vec2和mBART,该系统将翻译文本与视频时间轴对齐,为多语言视频内容转录和个性化语音翻译提供了一个可访问的解决方案。
摘要:This research addresses the challenge of training an ASR model for personalized voices with minimal data. Utilizing just 14 minutes of custom audio from a YouTube video, we employ Retrieval-Based Voice Conversion (RVC) to create a custom Common Voice 16.0 corpus. Subsequently, a Cross-lingual Self-supervised Representations (XLSR) Wav2Vec2 model is fine-tuned on this dataset. The developed web-based GUI efficiently transcribes and translates input Hindi videos. By integrating XLSR Wav2Vec2 and mBART, the system aligns the translated text with the video timeline, delivering an accessible solution for multilingual video content transcription and translation for personalized voice.


机器翻译由腾讯交互翻译提供,仅供参考