The History of Speech Synthesis
Abstract
Speech synthesis, making a machine produce human speech, has been attempted with bellows and reeds, with filters played like an organ, with rules derived from the physics of the vocal tract, with fragments cut from hours of recordings, and finally with neural networks that generate the sound wave one sample at a time. Wolfgang von Kempelen published a mechanical speaking machine in 1791; Bell Labs showed an electronic one, the Voder, at the New York World’s Fair in 1939 and made a computer sing “Daisy Bell” in 1961. Texas Instruments put a synthesizer on a single chip in the 1978 Speak & Spell. For decades the goal was intelligibility, and the users who needed it most were people who could not see a screen or could not speak. After DeepMind’s WaveNet in 2016 the goal became indistinguishability, and by 2023 a few seconds of someone’s recorded voice were enough to make it say anything, which turned a voice on the telephone from evidence of identity into a tool of fraud.
Talking Machines
In 1779 Christian Gottlieb Kratzenstein won a prize competition with a set of resonating tubes that produced five long vowels. Wolfgang von Kempelen, a Hungarian civil servant at the Habsburg court, better known for his chess-playing Turk (see the Mechanical Turk), had started a speaking machine in 1769 and worked on it for twenty years. In 1791 he published it in a 456-page book, Mechanismus der menschlichen Sprache. His third design used bellows as lungs, a single reed as the vocal cords, a rubber mouth shaped by the operator’s hand, and separate passages for nasal and hissing sounds, and it could speak short phrases in French, Italian and English in a monotone.
Charles Wheatstone built an improved copy in 1837. It inspired the young Alexander Graham Bell to build a speaking machine of his own, part of the interest in the physics of speech that led him to the telephone.
The Voder
At Bell Labs, Homer Dudley approached speech as a telephone engineer: as a signal that could be analysed into a few slowly changing parameters and rebuilt from them. His vocoder of the late 1930s did this automatically to compress speech for transmission, and during the Second World War it became part of SIGSALY, the encrypted telephone link between Washington and London. The Voder (Voice Operation Demonstrator) was the same idea played by hand. A wrist bar switched between a buzz for voiced sounds and a hiss for unvoiced ones, a foot pedal set the pitch, and ten keys controlled band-pass filters that shaped the sound into vowels and consonants.
It was shown at the New York World’s Fair and the Golden Gate International Exposition in San Francisco in 1939. It was hard to play. Operators trained for months before they could produce recognisable sentences; Helen Harper trained about twenty of them. The standard demonstration line was “Good afternoon, radio audience.” The Voder showed that speech could be built from a small number of controls. It did not show how a machine could decide on those controls by itself.
Daisy Bell
The computer answer came from the same laboratory. John L. Kelly Jr. and Carol Lochbaum at Bell Labs modelled the vocal tract as a series of tubes of different widths, and in 1961 they used an IBM 7090-series mainframe to sing the 1892 song “Daisy Bell (Bicycle Built for Two),” with an accompaniment programmed by Max Mathews. Arthur C. Clarke heard the recording on a visit to his friend John R. Pierce at Bell Labs in 1962, and in 2001: A Space Odyssey (1968) the computer HAL 9000 sings “Daisy” as it is being shut down.
The first general-purpose text-to-speech system for English was built in 1968 by Noriko Umeda and colleagues at the Electrotechnical Laboratory in Japan. Through the 1970s the main approach was formant synthesis: generate a buzz and pass it through resonators whose frequencies follow rules for each sound. Dennis Klatt at MIT wrote the most complete set of such rules, published his synthesizer in full in 1980, and saw it sold by Digital Equipment in 1984 as DECtalk. A version of that voice, on a Speech Plus board, spoke for Stephen Hawking from 1986 (see Hawking’s Voice). The same years produced the first Kurzweil Reading Machine (1976), which read printed books aloud.
Speak & Spell
The first talking machine most people owned was a toy. In 1976 a small team at Texas Instruments led by Paul Breedlove started a project with a budget of $25,000 to build a spelling game that spoke. It used linear predictive coding, a method from telephone research that describes each short slice of speech by a handful of filter coefficients, so that words recorded by a human speaker could be stored in a few kilobits and rebuilt on the fly. The chip that did the rebuilding, the TMC0280, was the first single-chip speech synthesizer. The Speak & Spell was shown at the summer Consumer Electronics Show in 1978. In 1982 it appeared in Steven Spielberg’s E.T. the Extra-Terrestrial as part of the alien’s device for phoning home, and in 2009 the IEEE named it a Milestone.
Recorded Voices
In the 1990s the field moved from modelling speech to reusing it. Concatenative synthesis stored real recordings of a speaker and joined short pieces to make new sentences. Andrew Hunt and Alan Black, at ATR in Japan, described unit selection in 1996: record many hours of one speaker, cut them into small units, and for each sentence search the whole database for the sequence of units that fits the target and joins most smoothly. The voices sounded human when the database had a good match and fell apart when it did not, and each new voice meant many hours in a recording studio.
Those weeks had consequences for the people who did them. In July 2005 the voice actor Susan Bennett spent four hours a day recording phrases and sentences for ScanSoft, later Nuance, without being told what they were for. When Apple launched Siri on the iPhone 4S in 2011, her voice was the American Siri; she made it public in 2013. The same technique could give a voice back. The film critic Roger Ebert lost his after cancer surgery, and the Scottish company CereProc built a synthetic Ebert from the commentary tracks he had recorded for DVDs of Casablanca and Citizen Kane; he used it on The Oprah Winfrey Show on 2 March 2010.
WaveNet
On 8 September 2016 DeepMind published WaveNet, a neural network that generated raw audio one sample at a time, 16,000 samples a second, each predicted from the ones before. In listening tests it cut the gap between the best existing synthesizers and real human speech by more than half, for both English and Mandarin. “The fact that directly generating timestep per timestep with deep neural networks works at all for 16kHz audio is really surprising,” its authors wrote.
The next step removed the recording studio. In January 2023 Microsoft researchers described VALL-E, a model that imitated a new speaker from a three-second sample. On 29 March 2024 OpenAI described its Voice Engine, which needed fifteen seconds, and said it would not release it widely because of the risk of misuse in an election year.
Dead End: The Voice as Proof
For a century a familiar voice on the telephone had been treated as identification: banks verified customers by voice, and employees acted on instructions from a boss they recognised. In 2019 the chief executive of a British energy company transferred €220,000 to a Hungarian supplier on the telephone instructions of a caller who sounded like his superior at the German parent company. The voice had been synthesised. The company’s insurer, Euler Hermes, described the case to the Wall Street Journal as the first it had seen of its kind.
In January 2024 voters in New Hampshire received automated calls in a synthetic voice resembling President Joe Biden’s, telling them to stay home rather than vote in the state’s primary. On 8 February 2024 the US Federal Communications Commission ruled unanimously that voices generated by AI count as “artificial” voices under the Telephone Consumer Protection Act of 1991, so that robocalls using them are illegal without the recipient’s prior consent. The law had been written against tape recorders. The technology Kempelen built to show how people speak had become good enough that hearing a voice no longer proved who was speaking.
📚 Sources
- Speech synthesis, Wikipedia (Kratzenstein 1779, Kempelen 1791, Wheatstone, Dudley, Pattern Playback, Umeda 1968, Speak & Spell, WaveNet)
- Wolfgang von Kempelen’s speaking machine, Wikipedia (1769 start, the 456-page 1791 book, the three designs, languages, Wheatstone’s 1837 replica, Bell)
- Voder, Wikipedia (Dudley, wrist bar, pedal, ten filters, the 1939 fairs, months of training, Helen Harper, “Good afternoon, radio audience”)
- “The IBM 7090 is The First Computer to Sing”, HistoryofInformation.com, and Daisy Bell, Wikipedia (1961, Kelly, Lochbaum and Mathews, Clarke’s 1962 visit, HAL 9000); “Daisy Bell (Bicycle Built for Two)”, National Recording Registry essay, Library of Congress
- Speak & Spell (toy), Wikipedia (Breedlove, 1976 and $25,000, TMC0280, LPC, summer CES 1978, E.T., IEEE Milestone 2009)
- Hunt, Andrew J. and Alan W. Black, “Unit selection in a concatenative speech synthesis system using a large speech database”, ICASSP 1996
- "‘I’m the original voice of Siri’", CNN, 4 October 2013 (July 2005, four hours a day, ScanSoft)
- “Amazingly, DVD Commentary Helped Give Roger Ebert His Voice Back”, TechCrunch, 2 March 2010
- van den Oord, Aäron and Sander Dieleman, “WaveNet: A generative model for raw audio”, DeepMind, 8 September 2016
- “Microsoft Debuts 3-Second Voice Cloning Tool VALL-E”, Voicebot.ai, 10 January 2023, and “Navigating the challenges and opportunities of synthetic voices”, OpenAI, 29 March 2024
- “A Voice Deepfake Was Used To Scam A CEO Out Of $243,000”, Forbes, 3 September 2019 (the Wall Street Journal report, €220,000, Euler Hermes)
- “FCC Makes AI-Generated Voices in Robocalls Illegal”, FCC, 8 February 2024, and the Declaratory Ruling FCC 24-17
- Image: TI Speak & Spell (1978) exhibited at “Game On 2.0” Technopolis in Athens, 2011.jpg by Tilemahos Efthimiadis (CC BY-SA 2.0), via Wikimedia Commons