本文将分享浙江大学计算机学院赵洲老师团队在AAAI-2022 论文中提出的DiffSpeech (用于语音合成)与DiffSinger (用于歌声合成)的官方Pytorch实现。


开源地址

https://github.com/MoonInTheRiver/DiffSinger

论文地址

DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism

https://arxiv.org/abs/2105.02446


安装依赖

DiffSpeech (语音合成的版本)

1.准备工作

数据准备

a) 下载并解压 LJ Speech dataset, 创建软链接: ln -s /xxx/LJSpeech-1.1/ data/raw/ 

b) 下载并解压 我们用MFA预处理好的对齐: tar -xvf mfa_outputs.tar; mv mfa_outputs data/processed/ljspeech/ 

c) 按照如下脚本给数据集打包,打包后的二进制文件用于后续的训练和推理.

export PYTHONPATH=.
CUDA_VISIBLE_DEVICES=0 python data_gen/tts/bin/binarize.py --config configs/tts/lj/fs2.yaml

# `data/binary/ljspeech` will be generated.

声码器准备

我们提供了HifiGAN声码器的预训练模型. 请在训练声学模型前,先把声码器文件解压到checkpoints里。
2.训练样例
首先你需要一个预训练好的FastSpeech2存档点. 你可以用我们预训练好的模型, 或者跑下面这个指令从零开始训练FastSpeech2:
CUDA_VISIBLE_DEVICES=0 python tasks/run.py --config configs/tts/lj/fs2.yaml --exp_name fs2_lj_1 --reset

然后为了训练DiffSpeech, 运行:

CUDA_VISIBLE_DEVICES=0 python tasks/run.py --config usr/configs/lj_ds_beta6.yaml --exp_name lj_exp1 --reset

记得针对你的路径修改usr/configs/lj_ds_beta6.yaml里"fs2_ckpt"这个参数。


3.推理样例
CUDA_VISIBLE_DEVICES=0 python tasks/run.py --config usr/configs/lj_ds_beta6.yaml --exp_name lj_exp1 --reset --infer


我们也提供了:
  • DiffSpeech的预训练模型;
  • FastSpeech 2的预训练模型, 这是为了DiffSpeech里的浅扩散机制;
记得把预训练模型放在checkpoints 目录。
DiffSinger (歌声合成的版本)
0.数据获取
  • 申请表:
    https://github.com/MoonInTheRiver/DiffSinger/blob/master/resources/apply_form.md
  • 数据集预览:
    https://github.com/MoonInTheRiver/DiffSinger/releases/download/pretrain-model/popcs_preview.zip
1.Preparation
数据准备
a) 下载并解压PopCSDownload and extract PopCS, 创建软链接:ln -s /xxx/popcs/ data/processed/popcs
b) 按照如下脚本给数据集打包,打包后的二进制文件用于后续的训练和推理.
export PYTHONPATH=.
CUDA_VISIBLE_DEVICES=0 python data_gen/tts/bin/binarize.py --config usr/configs/popcs_ds_beta6.yaml
# `data/binary/popcs-pmf0` 会生成出来.


声码器准备

我们提供了HifiGAN-Singing的预训练模型, 它专门为了歌声合成系统设计, 采用了NSF的技术。请在训练声学模型前,先把声码器文件解压到checkpoints里。
(更新: 你也可以将训练更多步数的存档点放到声码器的文件夹里)
这个声码器是在大约70小时的较大数据集上训练的, 可以被认为是一个通用声码器。
2.训练样例
首先你需要一个预训练好的FFT-Singer. 你可以用我们预训练好的模型, 或者用如下脚本从零训练FFT-Singer:
# First, train fft-singer;
CUDA_VISIBLE_DEVICES=0 python tasks/run.py --config usr/configs/popcs_fs2.yaml --exp_name popcs_fs2_pmf0_1230 --reset
# Then, infer fft-singer;
CUDA_VISIBLE_DEVICES=0 python tasks/run.py --config usr/configs/popcs_fs2.yaml --exp_name popcs_fs2_pmf0_1230 --reset --infer

然后, 为了训练DiffSinger, 运行:

CUDA_VISIBLE_DEVICES=0 python tasks/run.py --config usr/configs/popcs_ds_beta6_offline.yaml --exp_name popcs_exp2 --reset

记得针对你的路径修改

usr/configs/popcs_ds_beta6_offline.yaml里"fs2_ckpt"这个参数。

3.推理样例

CUDA_VISIBLE_DEVICES=0 python tasks/run.py --config usr/configs/popcs_ds_beta6_offline.yaml --exp_name popcs_exp2 --reset --infer
我们也提供了:
  • DiffSinger的预训练模型;
  • FFT-Singer的预训练模型, 这是为了DiffSinger里的浅扩散机制;

记得把预训练模型放在 checkpoints 目录.

请注意:
-我们原始论文中的PWG版本声码器已投入商业使用,因此我们提供此HifiGAN版本声码器作为替代品
-我们假设提供真实的F0来进行实验,如[1][2][3]等前作所做的那样,重点在频谱建模上,而非F0曲线的预测。如果你想对MIDI数据进行实验,从MIDI和歌词预测F0曲线,请查看文档:
https://github.com/MoonInTheRiver/DiffSinger/blob/master/usr/configs/midi/readme.md
目前已经支持的MIDI数据集有: Opencpop
[1] Adversarially trained multi-singer sequence-to-sequence singing synthesizer. Interspeech 2020.
[2] SEQUENCE-TO-SEQUENCE SINGING SYNTHESIS USING THE FEED-FORWARD TRANSFORMER. ICASSP 2020.
[3] DeepSinger : Singing Voice Synthesis with Data Mined From the Web. KDD 2020.
Tensorboard


Mel 可视化
沿着纵轴, DiffSpeech: [0-80]; FastSpeech2: [80-160].

DiffSpeech vs. FastSpeech 2




Audio Demos

音频样本可以看我们的样例页(https://diffsinger.github.io/)

我们也放了部分由DiffSpeech+HifiGAN (标记为[P]) 和 GTmel+HifiGAN (标记为[G]) 生成的测试集音频样例在:resources/demos_1213.

(对应这个预训练参数:

https://github.com/MoonInTheRiver/DiffSinger/releases/download/pretrain-model/lj_ds_beta6_1213.zip)

更新:新生成的歌声样例在:

resources/demos_0112.

Citation

如果本仓库对你的研究和工作有用,请引用以下论文:

@article{liu2021diffsinger,
title={Diffsinger: Singing voice synthesis via shallow diffusion mechanism},
author={Liu, Jinglin and Li, Chengxi and Ren, Yi and Chen, Feiyang and Liu, Peng and Zhao, Zhou},
journal={arXiv preprint arXiv:2105.02446},
volume={2},
year={2021}}