Skip to content

Dennis Klatt

Abstract

Dennis Klatt (1938–1988) spent twenty-three years at MIT working out how a human vocal tract shapes sound, and put the answer into a program. His formant synthesizer, published in full in 1980 so that anyone could build it, produced speech from rules rather than recordings; MITalk (1979) turned English text into those rules; and in 1984 Digital Equipment sold the two together as DECtalk, the first speech synthesizer a non-expert could plug in and use. Its default voice, Perfect Paul, was modelled on Klatt’s own. A close relative of that voice, on a Speech Plus board, spoke for Stephen Hawking from 1986 and became the most recognised synthetic voice in the world. Klatt died of cancer that December, having lost his own voice to the disease while his copy of it spoke for others.

DECtalk DTC01
A DECtalk DTC01, the 1984 unit that put Klatt’s synthesizer in a box with a serial port. Image: Emgaol, CC BY 3.0, via Wikimedia Commons.

Milwaukee to MIT

Dennis H. Klatt was born in Milwaukee on March 31, 1938. He took bachelor’s and master’s degrees in electrical engineering at Purdue in 1960 and 1961 and a doctorate in communication sciences at the University of Michigan in 1964, and in 1965 joined MIT’s Research Laboratory of Electronics as an assistant professor, in the speech group Kenneth Stevens had built around the acoustic theory of speech production. Stevens’s theory was that the vocal tract is a filter: the vocal cords produce a buzz, and the throat, mouth and nose shape it by resonance, the resonant frequencies being the formants that distinguish one vowel from another. If the theory was right, speech could be manufactured by computing the buzz and passing it through a set of resonators whose frequencies were set by rule. Klatt spent his career finding out how many rules it took.

The Synthesizer

The answer was published in March 1980 in the Journal of the Acoustical Society of America as “Software for a cascade/parallel formant synthesizer”: a complete description, with the Fortran source, of a synthesizer driven by about forty parameters updated every five milliseconds, and a set of tables saying how to set them. It was the standard reference design for the next twenty years, and it was free; laboratories around the world built Klatt synthesizers from the paper. MITalk, developed with Jonathan Allen and Sharon Hunnicutt and described in their 1987 book From Text to Speech, was the other half: the rules that took a stream of English text, worked out its pronunciation from a dictionary and morphological analysis, assigned stress and intonation, and produced the parameter stream the synthesizer needed. Klatt’s own version of the whole system, tuned over years of listening, was called Klattalk.

His 1987 survey, “Review of Text-to-Speech Conversion for English”, is the history of the field to that point, with recordings of every system he could find, from the Voder of 1939 onward, which he had collected and which remain the standard archive. The Acoustical Society gave him its Silver Medal in Speech Communication in 1987, and the Franklin Institute its Wetherill Medal the same year.

DECtalk

Digital Equipment Corporation licensed Klattalk in 1982 and announced DECtalk in December 1983; the DTC01, a 7-kilogram box with a serial port, shipped in early 1984 at about $4,000. Text went in over RS-232 and speech came out, in any of several voices defined by parameter sets: Perfect Paul, the default, built from measurements of Klatt’s own voice; Beautiful Betty, reportedly modelled on his wife’s (some accounts say a colleague’s wife’s); Huge Harry; Kit the Kid; and the rest. It was the first synthesizer that worked out of the box for anyone who could send it text, and it went into telephone information services, the National Weather Service’s radio broadcasts, airport weather announcements, reading machines for the blind, and the communication devices that gave a voice to people who had lost theirs. The larger story of those devices is in The Accessibility Revolution in Computing.

The best-known of them was not a DECtalk. Stephen Hawking, who had lost his speech to a tracheotomy in 1985, was given in 1986 a communication system whose voice came from a Speech Plus synthesizer, in the form he kept a CallText 5010 built in 1988, a board based on the MITalk rules and Klatt’s synthesis, and kept its voice for the thirty years that followed, refusing every upgrade because, he said, he had identified with it. When the hardware wore out, engineers reverse-engineered the board so that the voice could be preserved; the story is in The Voice Stephen Hawking Refused to Upgrade. The voice the world heard as Hawking’s was, at one remove, Klatt’s.

December 1988

Klatt had cancer through the mid-1980s; it took his voice, and in his last years the man whose voice spoke from thousands of machines could not speak himself. He continued working, and the 1987 review was written in that period. He died in Cambridge, Massachusetts, on December 30, 1988, aged 50. Formant synthesis was displaced in the 1990s by concatenative systems that spliced recorded speech, and those by neural networks, and the modern voices sound nothing like Perfect Paul; but the analysis in MITalk, how to get from text to pronunciation and prosody, is still the front end of every text-to-speech system, and the 1980 paper is still cited. DECtalk itself survived DEC, passing through Force Computers to Fonix and SpeechFX, and can still be bought.

📚 Sources