Skip to content

Hans Peter Luhn

Abstract

Hans Peter Luhn (1896–1964) was a German printer’s son who sold thread-counting gauges to textile mills, joined IBM in his mid-forties, and in the twelve years before his retirement wrote the memo that described hashing (1953), filed the checksum that still validates every credit-card number (1954), had a computer write the first automatic abstract (1958), invented the KWIC index that reorganised scientific literature (1958), and coined the phrase “business intelligence” in the same year. He held about 80 patents, including one for a cocktail-recipe guide filed while Prohibition was still in force. Information retrieval as a machine discipline starts with him, and almost nobody outside the field knows his name.

Barmen

Luhn was born on July 1, 1896, in Barmen, now part of Wuppertal, where his father Johann was a master printer. He went to Switzerland to learn the printing trade with the intention of joining the family business, and the First World War took him instead into the German army as a communications officer. After the war he went into textiles rather than print, and in 1924 a textile firm sent him to the United States to look at mill sites. He stayed.

His first invention was the Lunometer, a pocket gauge for counting threads per inch in cloth, which he developed in 1927 and which H. P. Luhn & Associates, the engineering consultancy he founded, still sells. Through the 1930s he worked in textiles and as an independent engineering consultant and filed patents on whatever occurred to him: a foldable raincoat, a device for shaping women’s stockings, a game table, and a “Cocktail Oracle” recipe guide filed in 1933, before repeal.

IBM

IBM hired him in 1941 as a senior research engineer, in his mid-forties. In 1946 and 1947 he worked on machine-readable typewriting: a metallic ribbon in a typewriter put magnetic patterns on the paper as it typed, producing a document a machine could read back. Shortly afterward he was asked to help two MIT chemists, Malcolm Dyson and James Perry, search for chemical compounds recorded in coded form on punched cards. The result was the Luhn scanner, a photoelectric machine for finding cards by content rather than position. That problem, how to find a record when you know what it says but not where it is, occupied the rest of his career. He became manager of information retrieval research at IBM and produced about 70 patents for the company.

The 1953 Memo

In January 1953 Luhn wrote an internal IBM memorandum proposing that records be filed by a number computed from their contents. His example was a telephone number: split 314-159-2652 into pairs (31, 41, 59, 26, 52), add the two digits of each pair, keep the last digit of each sum, and the result, 45487, names the “bucket”; records whose keys land in the same bucket are chained together and searched in sequence. That is a hash table with chaining, described before the word “hashing” existed (it first appeared in print in a 1968 paper by Robert Morris), and it is the first known description of the technique. Around the same time Gene Amdahl, Elaine McGraw, Nathaniel Rochester and Arthur Samuel used hashing in the IBM 701 assembler, and Amdahl is credited with open addressing by linear probing; the full lineage is in Hashing and Data Structures.

The Checksum

On January 6, 1954, Luhn filed a patent for a “Computer for Verifying Numbers”, a hand-held mechanical device, and the arithmetic inside it outlived the machine. Starting from the right, double every second digit, subtract nine from any result above nine, add everything up; the number is valid if the total ends in zero. The Luhn algorithm catches every single-digit error and almost every transposition of adjacent digits, which is what a clerk typing account numbers actually gets wrong. The patent, US 2,950,048, was granted in 1960, and the scheme became the mod-10 check on credit-card numbers, Canadian Social Insurance Numbers, the IMEI of every mobile phone, and national identifiers in several countries. It is in the public domain and has been since the patent expired; it was never meant to be secret, only to be cheap.

Four Sentences in Three Minutes

The year 1958 was the one in which his ideas came out in public, three at once.

In April the IBM Journal of Research and Development published “The Automatic Creation of Literature Abstracts”. The method was statistical: count the words in an article, throw out the common ones, rank the rest by frequency, and score each sentence by how many high-frequency words it contains close together; the top-scoring sentences are the abstract. The New York Times later described the demonstration in his obituary: a 2,326-word Scientific American article on hormones of the nervous system went in on magnetic tape, and three minutes later the typewriter produced four sentences giving the gist. Extractive summarisation, as natural language processing later called it, was built on this paper for fifty years, and term-frequency weighting, which Karen Spärck Jones completed with the inverse-document-frequency half in 1972, starts here.

In October the same journal published “A Business Intelligence System”. The proposed system would abstract incoming documents automatically, match them against profiles of what each person in an organisation was working on, and route the relevant ones to them. Luhn’s name for the second half was Selective Dissemination of Information, and SDI services ran in corporate and academic libraries for decades on exactly that design. The phrase in the title is the first use of “business intelligence” in the sense it now carries; the analytics industry that adopted the name in the 1990s rarely cites the paper.

In November, at the International Conference on Scientific Information in Washington, Luhn showed the Keyword-in-Context index. A KWIC index takes every title in a bibliography, generates one line per significant word, with that word aligned in a column and the rest of the title wrapped around it, and sorts the lines alphabetically. A researcher scanning the column sees every title containing “polymerization” together, with enough context on either side to judge relevance. It was mechanical, produced on tabulating equipment, and needed no indexer. By the early 1960s it sat at the centre of hundreds of computerised indexing systems, including those of Chemical Abstracts Service, Biological Abstracts and the Institute for Scientific Information; one expert called it the greatest thing to happen in chemistry since the invention of the test tube. The write-up appeared in American Documentation in 1960.

Armonk

Luhn retired from IBM in 1961. The American Documentation Institute (now ASIS&T) gave him its Award of Merit in 1964, and he died of leukaemia in Armonk, New York, on August 19 of that year, aged 68. His working method had been the same in cloth and in text: find the measurable property, build a cheap device around it, and let the arithmetic do the judging. The hash table, the checksum, the frequency-ranked abstract and the permuted index are all that one move.

📚 Sources