### This package was made to helps to discovery....
--- Acoustic Model:
The Acoustic Model where built using HTK. To train we use about 15 hours of speech for Brazilian Portuguese. The model contains 14 gaussians mixtures with cross-word triphones. The parameter type was MFCC_E_D_A_Z. The sample rate used was 22050 Hz.
--- Language Model:
We built 2 n-gram language models using the SRILM toolkit. The HDecode LM is a 3-gram and Julius is a 2-gram combine with a reverse 3-gram model (I use the mkbingram to join then). To train the models we use more the 2 milions of sentences. The vocabulary was limited to 65503 words. The model has perplexity of 166 and we use Kneser-Ney discounting to smoothing.
--- Test Files:
The test files are from LapsBenchmark corpus available on: www.laps.ufpa.br/falabrasil/The corpus contain 700 files, but this package contains only 100 files to speed-up the test, the files are storage on "database" directory.We create 3 list of test files:5mfc_test.list - 5 test files25mfc_test.list - 25 test files100mfc_test.list - 100 test files
