MaryTTS is an open-source, multilingual Text-to-Speech Synthesis platform written in Java.It was originally developed as a collaborative project of DFKI’s Language Technology Lab and the Institute of Phonetics at Saarland University. It is now maintained by the Multimodal Speech Processing Group in the Cluster of Excellence MMCI and DFKI.
As of version 5.2, MaryTTS supports German, British and American English, French, Italian, Luxembourgish, Russian, Swedish, Telugu, and Turkish; more languages are in preparation.MaryTTS comes with toolkits for quickly adding support for new languages and for building unit selection and HMM-based synthesis voices.
New to MaryTTS?
First, check out the online demo:
- Speech synthesis interface directly in your web browser
Want to know more about this Modular Architecture for Research on speech sYnthesis (MARY)?
- Browse the Documentation section
- Subscribe to the MARY users mailing list and ask questions there
Documentation
Start here if you want to understand what MARY is doing and how you can use it
- Overview
- History
- Publications
- Architecture Walkthrough
- MaryXML
- Development Wiki
The place to find documentation on how to compile, develop, contribute to MARY
- Javadoc
API details for the MARY system Java implementation
- Tibetan
Some background on the Tibetan language synthesis within MARY
Overview
Processing architecture and modules
the preprocessing or text normalisation; the natural language processing , doing linguistic analysis and annotation; the calculation of acoustic parameters , which translates the linguistically annotated symbolic structure into a table containing only physically relevant parameters; and the synthesis , transforming the parameter table into an audio file.
1. The preprocessing
2. The natural language processing
2.1 Components
part of speech tagger; chunker (a partial syntactic analysis); grapheme to phoneme conversion using a lexicon for the known tokens; grapheme to phoneme rules for the unknown tokens, using a morphological analysis; syllabification, word stress and phonologic rules;
intonation annotation using GToBI; postlexical phonological rules.
2.2 Output
Eine
echte
Herausforderung
.
3. The Calculation of Acoustic Parameters
3.1 Output
_ 10
aI 130 (0,209)
n 62
@ 52 (0,187)
_ 55
E 84 (50,232)
C 71
t 57
@ 52
h 61
E 71 (0,224)
R 63
aU 148 (50,174)
s 86
f 71
O 68
6 31
d 42
6 60
R 60
U 139
N 78 (100,160)
_ 400
# 4. The synthesiser
Technical architecture
multi-threaded: each request is processed in a thread of its own, which allows the server to process multiple requests “in parallel”; flexible: both pure Java modules and “external” modules (external programs reading from stdin and writing to stdout) are supported and can easily be integrated into the system; XML-based: state-of-the-art technologies such as DOM (for internal manipulation of the MaryXML structures) and XSLT (for input markup parsing) are used to make the system as transparent and understandable as possible.
