Dead End: Verbmobil
Abstract
Verbmobil (1993 to 2000) was Germany’s attempt to build a machine that interprets spoken conversation: German, English and Japanese speakers arranging meetings and trips, each talking in their own language, with the computer translating in between. The research ministry paid 116 million Deutsche Mark, industry added another 52.6 million, and 31 partners on three continents built a system of 69 modules written in seven programming languages. It worked, within its three narrow domains, and it won the German Future Prize in 2001. It was never sold. In the project’s own final evaluation, the component that performed best was the statistical translator from RWTH Aachen, which learned from example sentence pairs and used no hand-written grammar; its error rate was less than half that of the hand-built linguistic engine at the centre of the design. One of its authors later led machine translation at Google.
The Brief
Before committing money, the German Federal Ministry for Research and Technology (BMFT) commissioned two feasibility studies. One came from a German consortium of industrial and academic groups. The other came from the Center for the Study of Language and Information (CSLI) at Stanford, written in August 1991 by Martin Kay, Jean Mark Gawron and Peter Norvig. Both recommended going ahead. A call for proposals went out in July 1992, a board of ten international experts reviewed the bids, and the first phase began in January 1993 with 60 million marks of ministry money for four years. The German Research Center for Artificial Intelligence (DFKI) in Saarbrücken got the coordination and the job of integrating the system, with Wolfgang Wahlster as scientific director.
Wahlster described the goal at the MT Summit in Kobe in July 1993: “a portable translation device that you can carry to a meeting with speakers of other foreign languages and it will translate what you say for them.” The first versions were meant to be modest about it. The planners assumed a German and a Japanese business partner would both speak some English and conduct most of the meeting in it, reaching for Verbmobil only when they got stuck. The device would listen along the whole time, so that when a German speaker said “Let’s meet again in June, außer am Pfingstmontag,” it could supply “except on Whit Monday” and finish the sentence in context. Wahlster called this “translation on demand.” He also stressed that Verbmobil, unlike earlier speech-translation projects, would not deal with telephone calls but with face-to-face meetings in a small room.
Japan was running its own programme. The ATR Interpreting Telecommunications Research Laboratories in Kyoto had started a seven-year speech translation project in March 1993 with 16 billion yen, and Verbmobil planned to collaborate with ATR on Japanese, and with Carnegie Mellon, CSLI and ICSI Berkeley on English.
Phase One
Phase one ran from January 1993 to December 1996 and was organised the German way: 16 subprojects, 135 work packages, 33 research groups from 28 institutions, about 125 people a year. The subprojects mirrored the textbook pipeline of the time, one for each stage from signal processing through syntax, semantic construction, transfer and generation to speech synthesis.
The first milestone, the Verbmobil Demonstrator, passed its review in February 1995. It recognised spontaneously spoken German about appointment scheduling, with a vocabulary of 1,292 word forms, and spoke the translation in English. It was shown to the public at CeBIT in Hannover on 13 March 1995. The research prototype followed in October 1996 with 2,500 German word forms, plus 400 Japanese ones, and was presented at CeBIT 1997.
Spontaneous speech was the hard part, and it was hard in ways written text never is. When the project later analysed its recordings it counted 35,000 hesitations, 85,000 breaths and 206,000 instances of background noise, and found that 21 percent of all turns contained at least one self-correction (“on Tuesday, no, Wednesday”). A grammar written by linguists expects sentences. People arranging a meeting produce fragments.
Phase Two
Phase two, from January 1997 to September 2000, narrowed the structure to eight subprojects and 23 partners and widened the task. Travel planning and remote PC maintenance joined appointment scheduling as domains. The vocabulary grew to 10,157 German, 6,871 English and 2,566 Japanese word forms. The system now also ran as a speech server that callers reached by telephone, the setting the 1993 plan had ruled out.
The final Verbmobil 1.0 consisted of 69 modules exchanging data through 224 message pools, with 2,380 point-to-point connections among them. The modules were written in C, C++, Lisp, Prolog, Fortran, Perl and Tcl/Tk, a list that records how many research groups with how many habits had to be made to cooperate. One distinctive module read prosody, using intonation and sentence melody to help work out what an utterance meant. Another wrote summaries of the dialogue. The whole system ran at an average of 3.4 to 5.3 times real time, so a ten-second utterance needed between 34 and 53 seconds of processing.
The heart of the design was deep translation: parse the utterance with a hand-written HPSG grammar (2,400 types for German, 2,015 for English, 1,184 for Japanese), build a semantic representation, apply transfer rules (22,782 of them), and generate the target sentence from the result with another 13,640 microplanning rules. This was the path the 1993 proposal had been built around.
Because deep analysis often failed on real speech, the project added shallower engines alongside it, and by the end the system ran five translation engines concurrently and selected among their results:
- semantic transfer, the deep linguistic path;
- a dialogue-act engine that classified each utterance into one of a small set of patterns (“suggest a date”) and filled in the slots;
- a case-based engine using 30,000 translation templates;
- a substring-based, example-driven engine;
- StatTrans, a statistical translator built by Hermann Ney’s group at RWTH Aachen, trained on 58,332 German-English utterance pairs and working with 691,583 learned “alignment templates,” short phrase pairs extracted automatically from the data.
The Bake-Off
In spring 2000 the University of Hamburg ran the final end-to-end evaluation. Two native speakers held a conversation with no contact except through Verbmobil. The speech recognisers made errors on about one word in four. Human evaluators then looked at each engine’s translation of each turn and answered one question: is this sentence approximately correct, yes or no? A missing translation counted as wrong. The German-to-English test covered 5,069 dialogue turns and the English-to-German test 4,136.
Ney, Franz Josef Och and Stephan Vogel published the result:
| Translation engine | Sentence error rate |
|---|---|
| Semantic transfer | 62% |
| Dialogue-act based | 60% |
| Example-based | 52% |
| Statistical | 29% |
The engine built on the grammars, the semantic database and the 22,782 transfer rules made the most errors. The engine that had been shown pairs of sentences made the fewest, by roughly half. Ney’s group put the reason plainly: the statistical approach “is able to avoid hard decisions at any level,” and it always produces some output, so a garbled recogniser result or an ungrammatical half-sentence still yields a translation. The transfer engine, when its parser failed, returned nothing, and nothing was scored as wrong.
With the engines combined, the project reported more than 80 percent approximately correct translations and a 90 percent success rate on the dialogue tasks, the figures it presented as having met its goals.
The Bill
The final accounting, published by DFKI’s Reinhard Karger and Wahlster in 2000:
| Item | Amount |
|---|---|
| BMBF funding, phase I (1993 to 1996) | 62.7 million DM |
| BMBF funding, phase II (1997 to September 2000) | 53.3 million DM |
| Industrial investment | 32.6 million DM |
| Related industrial R&D | about 20 million DM |
| Total | 168.6 million DM |
Universities and research institutes were funded at 100 percent; companies covered 60 percent of their own costs. The industrial partners included Siemens, Philips, DaimlerChrysler, Alcatel SEL and IBM Germany. Over seven years, 369 scientists held Verbmobil positions, and 919 more people passed through as students, research assistants and doctoral candidates, 164 of them writing PhD theses. The project produced 238 technical reports totalling 5,331 pages.
In November 2001 Federal President Johannes Rau presented Wahlster with the German Future Prize, the president’s award for technology and innovation, worth 500,000 marks, for Verbmobil.
The Dead End
Verbmobil remained a research prototype and was never sold as a product. The portable meeting device of 1993 did not appear in any form. Several things closed off the path.
The architecture could not leave the lab. A stationary system of 69 modules in seven programming languages, running at four or five times real time, was a research integration, not something that could be shrunk into a pocket device in 2000. Each new domain meant new grammar rules, new transfer rules and new dialogue-act patterns, so the three domains were a ceiling as much as a scope.
The field moved under it. The planning in 1991 and 1992 assumed that translation quality would come from better linguistic knowledge, and phase one was organised around that assumption. By the time phase two ended, the project’s own measurement showed that a data-driven component outperformed the knowledge-driven core. The bet on hand-built grammars was the same bet the expert-system era had made about knowledge in general, and it lost here for the same reason: hand-written rules were expensive to produce and broke on the fragments and repairs of real speech.
The winning ingredient was data, and data was about to become abundant elsewhere. StatTrans had 58,332 sentence pairs, recorded and transcribed at great cost in controlled sessions. Six years later, Google’s statistical translator was trained on billions of words of text. That scale was available to a search company with a crawler, not to a publicly funded consortium organised around linguistic subprojects.
What Verbmobil did leave is real but mostly indirect. Its speech recordings, 3,200 dialogues from 1,658 speakers and 181.6 hours of audio on 56 CDs, went to the Bavarian Archive for Speech Signals in Munich and remained available to researchers. The University of Tübingen’s treebanks of spoken German, English and Japanese came out of phase two. The trained people went into the German speech and language industry, into DFKI’s follow-on projects SmartKom and SmartWeb, and into the start-ups DFKI spun off in these years. The Deutsches Museum in Munich put the system into its hall of fame.
After Verbmobil
The statistical side kept going. Och finished his doctorate at Aachen in 2002 with a thesis on alignment templates, moved to the Information Sciences Institute of the University of Southern California, and by 2004 was at Google, where he led machine translation. His 2004 journal paper with Ney on the alignment template approach, written with the Google address, cites the Verbmobil evaluation as evidence for the method. On 28 April 2006 Och announced on Google’s research blog that Google’s statistical translation system was live, starting with Arabic and English, trained by feeding the computer “billions of words of text.” Norvig, the co-author of the 1991 feasibility study, had become Google’s director of research in 2005, in charge of the teams working on translation and speech recognition. In 2009 he, Alon Halevy and Fernando Pereira published “The Unreasonable Effectiveness of Data,” arguing that simple models trained on web-scale data beat elaborate models trained on little.
Speech translation itself arrived through the phone. Alex Waibel, Verbmobil’s deputy scientific director, co-founded Mobile Technologies, which released Jibbigo, a Spanish-English speech translator for the iPhone that ran entirely offline, in September 2009; Facebook bought the company in August 2013. Google added an experimental Conversation Mode to Google Translate for Android in January 2011, for English and Spanish, in which each speaker pressed a microphone button and heard the other’s words translated aloud: the scenario Verbmobil had been funded to build, on a device people already carried, with translation handled by a statistical system in a data centre. The later neural systems are covered in Machine Translation, and the people on the linguistic side of the story in Hans Uszkoreit and German Language Technology.
📚 Sources
- Wahlster, Wolfgang — “Verbmobil: Translation of Face-To-Face Dialogs”, Proceedings of MT Summit IV, Kobe, July 1993, pp. 127–136 (the portable-device vision, translation on demand, English as common language, face-to-face rather than telephone, the CSLI and German feasibility studies, the July 1992 call, 60 million DM for phase one, ATR’s 16-billion-yen project)
- Karger, Reinhard and Wolfgang Wahlster — “Facts and Figures about the Verbmobil Project”, in Wahlster (ed.), Verbmobil: Foundations of Speech-to-Speech Translation, Springer, 2000, pp. 22–30 (milestones and CeBIT dates, vocabularies, funding table, subprojects and partners, 69 modules and 224 pools, programming languages, real-time factor, rule and template counts, the five engines’ resources, corpus statistics, BAS, personnel)
- Wahlster, Wolfgang (ed.) — Verbmobil: Foundations of Speech-to-Speech Translation, Springer, 2000 (abstract via the MT Archive: five concurrent engines, three domains, over 80% approximately correct translations, 90% dialogue task success)
- Ney, Hermann, Franz Josef Och and Stephan Vogel — “The RWTH System for Statistical Translation of Spoken Dialogues”, Proceedings of HLT 2001 (the spring 2000 end-to-end evaluation at the University of Hamburg, 25% recogniser word error rate, 5,069 and 4,136 dialogue turns, sentence error rates 62/60/52/29%, 58,332 training sentence pairs)
- Och, Franz Josef and Hermann Ney — “The Alignment Template Approach to Statistical Machine Translation”, Computational Linguistics 30(4), 2004 (Och’s Google affiliation, the Verbmobil evaluation cited as support)
- Verbmobil — Wikipedia (German) (the Tübingen treebanks, the Deutsches Museum hall of fame)
- Verbmobil — Wikipedia (116 million DM public and 52 million DM industrial funding)
- “Am Anfang war das Verbmobil” — HNF Blog, 3 April 2018 (total 168.6 million DM, 369 scientists, Waibel as deputy director and his later Jibbigo)
- “Prof. Wahlster erhält den Deutschen Zukunftspreis 2001” — idw, 29 November 2001 (the prize presented by Johannes Rau, 500,000 DM)
- Verbmobil data collection overview — Institute of Phonetics, LMU Munich (the appointment-scheduling recordings of phase one)
- Franz Josef Och — Wikipedia (Aachen PhD 2002, ISI 2002–2004, head of machine translation at Google)
- Och, Franz — “Statistical machine translation live”, Google Research Blog, 28 April 2006 (Arabic-English launch, “billions of words of text”)
- Norvig, Peter — biography (Director of Research at Google from 2005, overseeing machine translation and speech recognition)
- Halevy, Alon, Peter Norvig and Fernando Pereira — “The Unreasonable Effectiveness of Data”, IEEE Intelligent Systems 24(2), 2009, pp. 8–12
- Jibbigo — Wikipedia (September 2009 Spanish-English offline release, Facebook acquisition August 2013)
- “Conversation Mode in Google Translate for Android” — Google Operating System blog, 12 January 2011 (experimental feature, English and Spanish only)
- Image: DFKI-SB.jpg by Renatoorsini (CC BY-SA 4.0), via Wikimedia Commons