Hugging Face 新开源的 TTS 模型:Parler-TTS,完全开源免费的一款 TTS 工具。一行命令即可安装!可自主训练定制声音!

项目链接:https://github.com/huggingface/parler-tts

试用链接:https://huggingface.co/spaces/parler-tts/parler_tts_mini


Parler-TTS 是一个轻量级的文本转语音(TTS)模型,能够以特定发音人(性别、音调、说话风格等)的风格生成高质量、自然听起来的语音。
该项目基于Dan Lyth和Simon King的研究,包含了数据集、预处理、训练代码和权重的完整发布。用户可以通过简单的代码安装和运行模型,也可以参与到进一步的模型训练和优化中去。
Parler-TTS 发布 Mini(880M)和 Large(2.3B)版本模型,在 45,000 小时的有声读物数据上训练得到,与v0.1 版本相比,生成速度提高了 4 倍。此外,支持 SDPA 和 Flash Attention 2,以进一步提高速度。
  • Parler-TTS Mini,880M参数模型

  • Parler-TTS Large,2.3B 参数模型


实 操

下面开始安装


# 创建全新python环境,使用3.9版本
conda create -n tts python=3.9

# 激活环境
conda activate tts

# 安装parler-tts
pip install git+https://github.com/huggingface/parler-tts.git

# 或者通过源码来安装
git clone https://github.com/huggingface/parler-tts.git
cd parler-tts
python setup.py install

# 安装特定版本的
numpypip install numpy==1.26.4


安装完毕后,看个示例

import torch
from parler_tts import ParlerTTSForConditionalGeneration
from transformers import AutoTokenizer
import soundfile as sf

device = "cuda:0" if torch.cuda.is_available() else "cpu"

model = ParlerTTSForConditionalGeneration.from_pretrained("parler-tts/parler-tts-mini-v1").to(device)
tokenizer = AutoTokenizer.from_pretrained("parler-tts/parler-tts-mini-v1")

prompt = "Hey, how are you doing today?"
description = "A female speaker delivers a slightly expressive and animated speech with a moderate speed and pitch. The recording is of very high quality, with the speaker's voice sounding clear and very close up."

input_ids = tokenizer(description, return_tensors="pt").input_ids.to(device)
prompt_input_ids = tokenizer(prompt, return_tensors="pt").input_ids.to(device)

generation = model.generate(input_ids=input_ids, prompt_input_ids=prompt_input_ids)
audio_arr = generation.cpu().numpy().squeeze()
sf.write("parler_tts_out.wav", audio_arr, model.config.sampling_rate)
不过比较遗憾的是,目前放出的模型都是基于英文数据来训练,中文的效果不是特别好,需要自己做训练。

官方也提供了训练方法,训练文档地址:

https://github.com/huggingface/parler-tts/blob/main/training/README.md