Version 1 of the Carnegie Mellon University Statistical Language Modelingtoolkit was written by Roni Rosenfeld, and released in 1994. It is availableby ftp from here. Here is a excerpt from its README file:
Overview of the CMU SLM Toolkit, Rev 1.0
========================================
The Carnegie Mellon Statistical Language Modeling (CMU SLM) Toolkitis a set of unix software tools designed to facilitate languagemodeling work in the research community.
Some of the tools are used to process general textual data into:
- word frequency lists and vocabularies
- word bigram and trigram counts
- vocabulary-specific word bigram and trigram counts
- bigram- and trigram-related statistics
- various Backoff bigram and trigram language models
Others use the resulted language models to compute:
- perplexity
- Out-Of-Vocabulary (OOV) rate
- bigram- and trigram-hit ratios
- distribution of Backoff cases
- annotation of test data with language scores
Version 2 of the toolkit seeks to maintain the structure of version 1, toinclude all (or very nearly all) of the functionality of version 1, and toprovide useful improvements in terms of functionality and efficiency. Thekey differences between this version and version 1 are described in the nextsection.
