Brewster Kahle and the Internet Archive
Abstract
The web forgets by default: the average page lives on the order of months before it moves or vanishes. That civilization’s fastest-growing medium had no memory struck Brewster Kahle (b. 1960) as an emergency, and he built the institution to fix it: the Internet Archive, founded in San Francisco in May 1996 with the motto “universal access to all knowledge.” Its Wayback Machine (named for the cartoon time machine in Rocky and Bullwinkle) reached one trillion archived web pages in October 2025, alongside millions of books, films, TV news broadcasts, sound recordings, and runnable vintage software. Kahle funded it the Valley way, selling two search startups to AOL and Amazon, then spent the fortune building a library. The library, in turn, spent the 2020s under siege: sued by book publishers and record labels, hacked, and leaned on, while becoming more indispensable each year as the rest of the web’s memory (GeoCities, Google’s cache) was switched off around it.
From the Connection Machine to WAIS
Kahle (born October 21, 1960, New York City) studied computer science at MIT (B.S. 1982) under Marvin Minsky and Danny Hillis, then followed Hillis to Thinking Machines (1983–1989) as a lead engineer on the Connection Machine, the massively parallel supercomputer of the AI-hardware era. Searching large text corpora on parallel hardware led him to his life’s theme: in 1989 he co-developed WAIS (Wide Area Information Servers), the first distributed search-and-retrieval system on the internet, a pre-web rival to Gopher and HTTP for how networked documents would be found (see Dead End: Gopher for how that contest ended). WAIS Inc. (1992) sold to AOL in 1995 for $15 million; Alexa Internet (1996, with Bruce Gilliat), which crawled the web to rank and recommend sites, sold to Amazon in 1999 for $250 million in stock. Kahle kept the crawls.
A Library for the Web
The Internet Archive, a nonprofit, was founded in May 1996, the same year as Alexa, deliberately: the for-profit crawled, the nonprofit preserved. For five years the archive accumulated in the dark; in October 2001 the Wayback Machine opened the collection to the public, type a URL, see the page as it was. It settled in 2009 into a former Christian Science church on Funston Avenue in San Francisco whose Greek-revival columns happen to look exactly like what it is: a library. A mirror lives at the Bibliotheca Alexandrina in Egypt, the new Library of Alexandria holding the backup of the new library of everything, a redundancy Kahle likes to point out the old Alexandria never had.
The collection grew into the web’s collective memory and then some, by late 2024: 866+ billion web pages, 42.5 million texts, 13 million videos, 3 million TV news broadcasts, 14 million audio recordings, 1.2 million software titles, over 100 petabytes. In October 2025 the Wayback Machine archived its trillionth page. The software collection runs in the browser via emulation (thousands of MS-DOS games and vintage programs, playable) making the Archive the field’s working museum of digital culture. Kahle also hedges digitally-only preservation with the physical: shipping containers in Richmond, California hold millions of scanned books, on the librarian’s principle that you keep the original after you copy it. The lineage is explicit: Vannevar Bush’s Memex, Ted Nelson’s Xanadu, and Wikipedia are the Archive’s intellectual siblings; Aaron Swartz, who worked on the Archive’s Open Library project, its martyr (see Aaron Swartz).
Institutional credentials followed the collection: Kahle entered the Internet Hall of Fame in 2012, and in July 2025 the Archive was designated a U.S. federal depository library. Its evidentiary role is now routine: Wayback captures are cited in courtrooms, journalism, and scholarship as the standard proof of “what the web said.” The Archive defended that independence early: served with an FBI National Security Letter demanding user records, Kahle fought it with the EFF and ACLU, and in 2008 the FBI withdrew the letter and lifted the gag, one of the very few NSLs ever beaten in the open (see The Privacy War).
The Siege Years
The 2020s tested whether a library for the digital age is legal in it.
- Books. The Archive lends scanned books under controlled digital lending, one physical copy owned, one digital copy loaned at a time. In March 2020, as COVID closed libraries, it dropped the waitlists and called it the National Emergency Library; four major publishers sued in June 2020 (Hachette v. Internet Archive). The district court ruled for the publishers in March 2023, the Second Circuit affirmed in September 2024, and the consent judgment removed hundreds of thousands of books from lending where commercial e-books exist. The core CDL theory (that a library may loan what it owns, digitally) lost in the Second Circuit (see Copyright and IP in the Digital Age).
- Music. Record labels led by UMG sued over the Great 78 Project (digitized shellac 78-rpm records), with statutory-damages exposure in the hundreds of millions, an existential number for a nonprofit. The case settled in September 2025.
- Attack. In October 2024 the Archive was breached and DDoSed: ~31 million user accounts exposed and services down for weeks, a reminder that the web’s memory is one under-funded nonprofit with a single headquarters.
Kahle’s consistent framing: every prior medium’s history survived because libraries bought, kept, and lent copies without asking permission; digital licensing replaces ownership with rental, and a library that can only rent cannot preserve. Publishers’ framing: scanning and lending entire in-print books is infringement at scale, whatever the mission. The courts, so far, agree with the publishers about books, while everyone, including the plaintiffs’ own lawyers, keeps citing the Wayback Machine.
⚠️ Dead End: The Web That Wasn’t Saved
The Archive’s importance is best measured by what happened where it wasn’t. GeoCities (38 million pages of the amateur 1990s web) was deleted by Yahoo! in 2009 with months of notice; only scramble efforts by the Archive and the volunteer Archive Team saved a fraction (see Dead End: Yahoo!). Google Cache, the de facto second copy of the live web, was quietly retired in 2024; Google now links to the Wayback Machine instead. Link rot is not an edge case but the norm: studies of U.S. Supreme Court opinions found roughly half their cited links dead, and a quarter of deep links in New York Times articles had rotted within a decade or two of publication. The counterfactual is stark: without one nonprofit’s crawlers, the primary sources of the internet era (the pages, the propaganda, the retracted claims, the deleted tweets of the powerful) would simply be gone. Every earlier chapter of this encyclopedia rests on institutions that kept the paper. For the web, there is essentially one such institution, a donation-funded nonprofit, permanently under-resourced for the size of its mandate, and already hacked once.
Fun Fact
The Wayback Machine is named after the WABAC Machine from the 1960s cartoon The Rocky and Bullwinkle Show, in which a boy and his time-traveling dog visit history. Inside the Archive’s church headquarters stand rows of three-foot ceramic statues, every employee who has worked there three years gets one, a terracotta army of librarians.
📚 Sources
- Wikipedia: Brewster Kahle · Internet Archive · Wayback Machine · Hachette v. Internet Archive
- Internet Archive: About / “universal access to all knowledge”
- Internet Archive blog: Wayback Machine to Hit One Trillion Web Pages Archived (announced July 2025; milestone reached October 22, 2025)
- EFF: Internet Archive and the FBI National Security Letter (2008)
- Second Circuit opinion: Hachette Book Group v. Internet Archive (September 2024, PDF)
- Zittrain, Albert, Lessig: “Perma: Scoping and Addressing the Problem of Link and Reference Rot in Legal Citations” (Harvard, 2014)
- Internet Hall of Fame: Brewster Kahle (2012)
- Image: Brewster Kahle 2009.jpg by Joi Ito (CC BY 2.0), via Wikimedia Commons