This tar file contains perl scripts designed to manipulate speechdatabase transcriptions and word lattice files.

Demonstration Scripts_________________________

demo_lattices.sh - Automatically generate landmark-based pronunciationmodels from the pronlex dictionary. Use these to augment the edgesin a set of Switchboard lattices. Find minimum-distance alignmentof pronunciation models of words on coterminous edges.

demo_praat.sh - Convert TIMIT transcriptions to Praat TextGrid format,by way of HTK-format MLF.

Tool Scripts__________________________________

align_edges.pl - Compute the alignment between pronunciation tagson parallel edges in a word lattice.

augment_lattice.pl - Augment a word lattice with pronunciationsfrom a dictionary.

cfg - contains configuration files for phn2lm.pl.

consecutivelabels.pl - Separate a string of labels into consecutivelines, e.g., in order to divide the lines in a WS97 syllable transcription

dict2mlf.pl - convert a dictionary file into an MLF-format"pseudo-transcription."

extract_examples.pl - extract training examples for a neural networkor support vector machine, from waveform files or MFCC files, basedon transcriptions given in an MLF.

extractmlflayer.pl - Extract one layer from a multi-layer MLF.

lat2praat.pl - Create a Praat transcription file that summarizes some of thealternative segmentations available in an HTK-format word lattice.

lattice_cheat.pl - Compute the distance between the dictionarypronunciation model for each edge in a lattice vs. a known phonemetranscription of the same data. Find the path through the latticethat minimizes this distance.

layermlfs.pl - Construct a multilayer MLF from several single-layerMLFs.

makescripts.pl - Search a directory tree to find examples of fileswith extensions matching a regular expression; list these files inan HTK-format "script" file.

mlf2dict.pl - Create a dictionary file based on the alignment betweenphoneme and word transcriptions in a multilayer MLF.

mlf2praat.pl - Convert an MLF to a Praat-format transcription.

mlf2tgt.pl - Create a binary-valued neural network target file basedon segmentation provided in an MLF.

normalize_phn.pl - Eliminate unwanted phoneme distinctions from anMLF.

normalize_swb_wrd.pl - Eliminate unwanted word distinctions from anMLF encoding Switchboard word transcripts.

phn2lm.pl - Convert phoneme transcription to a distinctive-featurebased landmark transcription.

pinch_lattice.pl - Pinch a word lattice: remove all nodes but thosein the maximum-likelihood path, and make each edge coterminous withone of the edges in the ML path.

praat2mlf.pl - Convert a Praat transcription to MLF.

prune_lattice.pl - Prune a word lattice to eliminate pathssignificantly less probable than the ML path.

rnc2mlf.pl - Convert transcriptions from the Boston Radio News corpusinto MLF.

score_lattice.pl - Compute the 1-best WER of a lattice, and thelattice error rate, by comparing to a known word-leveltranscription.

sort_txgd.pl - Sort tiers in a Praat transcription.

sphere2mlf.pl - Convert SPHERE transcriptions to MLF.

split_lattice.pl - Split a multi-utterance lattice file into manysingle-utterance lattice files.

timit2sphinx.pl - Convert TIMIT phoneme labels into SPHINX phonemelabels.

timit2swb.pl - Convert TIMIT phoneme labels into Switchboard phonemelabels.

xwaves2mlf.pl - Convert an XWaves label file (e.g., as used in theBoston Radio News corpus) into MLF.