← 返回资源分享
pyannote.audio: neural building blocks for speaker diarization
论文
论文
发布时间2019-11-04
发表arXiv:1911.01255
作者:Ruiqing Yin,Hervé Bredin,Pavel Korshunov,Juan Manuel Coria,Gregory Gelly,Marvin Lavechin,Diego Fustes,Hadrien Titeux,Wassim Bouaziz,Marie-Philippe Gill
详细介绍
We introduce pyannote.audio, an open-source toolkit written in Python for speaker diarization. Based on PyTorch machine learning framework, it provides a set of trainable end-to-end neural building blocks that can be combined and jointly optimized to build speaker diarization pipelines. pyannote.audio also comes with pre-trained models covering a wide range of domains for voice activity detection, speaker change detection, overlapped speech detection, and speaker embedding -- reaching state-of-the-art performance for most of them.
代码仓库 (3)
pyannote/pyannote-audio官方PyTorch
MarvinLvn/voice-type-classifier
muskang48/Speaker-DiarizationTensorFlow
