The Supercomputer Era: The Fastest Machines on Earth
Abstract
This article traces supercomputing from Seymour Cray’s machines in the 1960s through the vector era, the shift to massively parallel commodity clusters, and the race to exascale. It is also a story of national rivalry: Japanese vector machines and a US anti-dumping tariff of 454 percent in the 1990s, the Earth Simulator shock of 2002, China’s rise to the top of the Top500 and its withdrawal from the list under US export controls, Europe’s first exascale machine in Jülich in 2025, and a Chinese machine back at number one in 2026.
Seymour Cray and the First Supercomputers
Seymour Cray was an electrical engineer from Chippewa Falls, Wisconsin, who combined extraordinary technical intuition with an almost monastic focus. He worked in isolation, literally and figuratively: he built a lab under his house in Wisconsin, away from the corporate offices of the companies he worked for, and was known to say that the best way to design a computer was to hire a few good engineers and leave them alone.
At Control Data Corporation (CDC), Cray designed the CDC 6600 (1964), widely recognized as the first true supercomputer. It achieved 3 megaFLOPS (3 million floating-point operations per second), three times the speed of the previous record-holder, IBM’s 7030 Stretch. It did this through a novel architecture: ten peripheral processors handled input/output while a central processor focused exclusively on computation, and the central processor itself executed instructions in parallel using multiple functional units.
IBM’s chairman Thomas Watson Jr. wrote a memo in August 1963 asking how a laboratory of thirty-four people, fourteen of them engineers, had taken the lead from IBM’s own development organization. Cray’s response was one line: “It seems like Mr. Watson has answered his own question.”
Cray left CDC in 1972 to found Cray Research. The Cray-1 (1976) was his masterpiece: a $8 million machine that delivered 160 megaFLOPS, weighed 5.5 tons, and was shaped like a padded bench, the padding concealing the cooling system. Its defining innovation was the vector register: a hardware unit that could perform the same arithmetic operation on 64 numbers simultaneously, accelerating the loops that dominated scientific computation. Los Alamos National Laboratory bought the first unit for nuclear weapons simulation. The Cray-1 became the standard by which all scientific computation was measured.
Vector Processing vs. Modern Parallelism
A vector processor executes one instruction on many data elements simultaneously: “add these 64 numbers to those 64 numbers” in a single operation. This is powerful for the regular, predictable loops of scientific computation (weather simulation, fluid dynamics, matrix operations) and less useful for irregular, data-dependent code. Modern parallelism, by contrast, runs many independent threads of execution simultaneously on separate cores. Today’s supercomputers use both: many-core CPUs and GPUs for thread-level parallelism, combined with SIMD instructions for data-level parallelism within each core.
The Cold War and Scientific Computation
Supercomputers were not merely academic tools. The U.S. government’s primary customer for the most powerful machines was the nuclear weapons complex, specifically, the national laboratories at Los Alamos, Lawrence Livermore, and Sandia. After the Limited Test Ban Treaty (1963) prohibited atmospheric nuclear tests and the Comprehensive Test Ban Treaty (1996) halted underground tests, simulation became the only legal way to verify that nuclear weapons in the stockpile still worked as designed. The required simulation accuracy drove supercomputer procurement for decades.
The Cold War framing extended to export controls: the U.S. government restricted the export of supercomputers to Soviet-bloc countries and, later, China. The first commercial challenge to Cray came from Japan. Hitachi shipped the S-810 and Fujitsu the VP-200 in 1983, NEC the SX-2 in 1985, all vector machines in the Cray mould. American vendors complained that Japanese government agencies bought only Japanese machines and that the same firms undercut Cray abroad; in August 1987 Washington and Tokyo exchanged letters committing Japan to open public procurement, renewed in a broader agreement in March 1990. The dispute peaked in 1996, when the National Center for Atmospheric Research chose four NEC SX-4 systems in a $35 million lease. Cray Research petitioned the Commerce Department, claiming NEC was taking a $65 million loss on the deal; NEC declined to answer the investigation questionnaire, and in 1997 Commerce set an anti-dumping duty of 454 percent on NEC supercomputers, a figure calculated from “facts otherwise available”. The sale was shelved. See Japan’s Computing Industry.
The civilian scientific applications were equally consequential. Weather forecasting (global atmospheric simulation) required updating models faster than real time to be useful, a threshold that each generation of supercomputers moved further ahead. The European Centre for Medium-Range Weather Forecasts (ECMWF) and the U.S. National Weather Service became major supercomputer customers, and the improvement in forecast accuracy through the 1980s and 1990s is directly attributable to increases in computing power.
The Top500 List and the Benchmark Race
In 1993, Jack Dongarra at the University of Tennessee and Hans Meuer at the University of Mannheim began publishing the Top500 list, a twice-yearly ranking of the 500 most powerful supercomputers in the world, measured by performance on the Linpack benchmark: solving a dense system of linear equations.
Linpack performance is measured in FLOPS (floating-point operations per second). The trajectory of the Top500 list illustrates Moore’s Law applied to the largest machines:
| First list at #1 | Machine | Site | Country | Rmax |
|---|---|---|---|---|
| June 1993 | CM-5 (Thinking Machines) | Los Alamos | US | 59.7 GigaFLOPS |
| Nov 1993 | Numerical Wind Tunnel (Fujitsu) | National Aerospace Laboratory | Japan | 124 GigaFLOPS |
| June 1994 | Intel Paragon XP/S 140 | Sandia | US | 143 GigaFLOPS |
| June 1996 | SR2201 (Hitachi) | University of Tokyo | Japan | 220 GigaFLOPS |
| Nov 1996 | CP-PACS (Hitachi) | University of Tsukuba | Japan | 368 GigaFLOPS |
| June 1997 | ASCI Red (Intel) | Sandia | US | 1.07 TeraFLOPS |
| Nov 2000 | ASCI White (IBM) | Lawrence Livermore | US | 4.9 TeraFLOPS |
| June 2002 | Earth Simulator (NEC) | Yokohama | Japan | 35.9 TeraFLOPS |
| Nov 2004 | BlueGene/L (IBM) | Lawrence Livermore | US | 70.7 TeraFLOPS |
| June 2008 | Roadrunner (IBM) | Los Alamos | US | 1.03 PetaFLOPS |
| Nov 2010 | Tianhe-1A (NUDT) | Tianjin | China | 2.57 PetaFLOPS |
| June 2011 | K computer (Fujitsu) | RIKEN, Kobe | Japan | 8.2 PetaFLOPS |
| June 2013 | Tianhe-2 (NUDT) | Guangzhou | China | 33.9 PetaFLOPS |
| June 2016 | Sunway TaihuLight | Wuxi | China | 93 PetaFLOPS |
| June 2018 | Summit (IBM) | Oak Ridge | US | 122.3 PetaFLOPS |
| June 2020 | Fugaku (Fujitsu) | RIKEN, Kobe | Japan | 416 PetaFLOPS |
| June 2022 | Frontier (HPE) | Oak Ridge | US | 1.1 ExaFLOPS |
| Nov 2024 | El Capitan (HPE) | Lawrence Livermore | US | 1.74 ExaFLOPS |
| June 2026 | LineShine | Shenzhen | China | 2.2 ExaFLOPS |
Rmax is the performance on the list where the machine first took the top spot; several were upgraded later (BlueGene/L, K, Fugaku, Frontier and El Capitan all posted higher numbers on subsequent lists). Jaguar (Oak Ridge, November 2009), Sequoia (Lawrence Livermore, June 2012) and Titan (Oak Ridge, November 2012) held the spot briefly between the rows shown.
Frontier, installed at Oak Ridge National Laboratory in 2022, was the first machine to exceed one ExaFLOP ($10^{18}$ floating-point operations per second). It comprises 74 HPE Cray EX cabinets and 9,408 nodes, each node pairing one AMD EPYC CPU with four AMD Instinct MI250X GPU accelerators, consuming 21 megawatts of power, roughly the electricity consumption of a small city. El Capitan at Lawrence Livermore, built on AMD’s MI300A, took the top spot in November 2024 at 1.74 ExaFLOPS. By June 2026 five machines exceeded an ExaFLOP on Linpack: LineShine, El Capitan, Frontier, Aurora at Argonne, and JUPITER at Jülich.
The National Race
The Top500 was conceived as a statistics project. It turned into a scoreboard that governments read, and three regions treated the top spot as a matter of policy.
Japan
Japan’s vector builders held the top spot for most of the mid-1990s (Numerical Wind Tunnel, SR2201, CP-PACS), but the machine that changed American policy was the Earth Simulator. Planned from 1997 for climate and earthquake research and funded with roughly $350 to 400 million by the Japanese government, it went into operation at Yokohama on 11 March 2002: 640 NEC SX-6 nodes, 5,120 vector processors, 35.86 TeraFLOPS on Linpack, nearly five times the next machine on the list (ASCI White). The New York Times called it a “Computenik”, and Jack Dongarra said, “In some sense we have a Computenik on our hands.” Hans Meuer, co-founder of the Top500, noted that the Americans “were caught on the wrong foot” even though the completion date had been public for five years. In 2002 DARPA launched its High Productivity Computing Systems program, which funded Cray and IBM through three phases and produced the Chapel and X10 languages; at SC2002 in Baltimore, the Department of Energy announced a $216 to 267 million IBM contract for ASCI Purple and BlueGene/L. BlueGene/L took the top spot back in November 2004.
Japan returned to number one twice more, both times with Fujitsu machines at RIKEN in Kobe. The K computer (88,128 SPARC64 VIIIfx processors, 705,024 cores) led the June and November 2011 lists and was the first machine past 10 PetaFLOPS (10.51 in November 2011); it ran until 30 August 2019. Its successor Fugaku, built on Fujitsu’s ARM-based A64FX with 158,976 nodes and a programme cost of about ¥130 billion, led four consecutive lists from June 2020 to November 2021 (416, later 442 PetaFLOPS) and was the first system to top HPL, HPCG, HPL-AI and Graph500 at the same time. It did so without GPUs.
China
China’s first number one, Tianhe-1A at Tianjin in November 2010 (2.57 PetaFLOPS), paired 14,336 Intel Xeon X5670 CPUs with 7,168 Nvidia Tesla M2050 GPUs. Tianhe-2 at Guangzhou, built by the National University of Defense Technology (NUDT) with 32,000 Xeon E5 CPUs and 48,000 Xeon Phi coprocessors, held the top spot for six consecutive lists from June 2013 to November 2015 at 33.86 PetaFLOPS. Both ran on American silicon. In February 2015 the U.S. Commerce Department added NUDT and the national supercomputing centres in Changsha, Guangzhou and Tianjin to its Entity List, stating that the Tianhe machines were believed to be used for “nuclear explosive activities”, which blocked Intel from supplying chips for Tianhe-2’s planned upgrade. NUDT replaced the Xeon Phis with domestic Matrix-2000 accelerators.
The answer came in June 2016. Sunway TaihuLight at Wuxi ran on 40,960 Chinese-designed SW26010 processors (10,649,600 cores), reached 93 PetaFLOPS, and held first place for four lists until Summit in June 2018. Export controls widened: Sugon and its Hygon chip venture in June 2019, then in April 2021 seven more entities including Phytium, Sunway Microelectronics and the centres at Wuxi, Shenzhen, Jinan and Zhengzhou, cited by Commerce Secretary Gina Raimondo for supporting “military modernization efforts” and weapons programmes.
China then stopped submitting. Two systems provisionally entered for the 2019 list at roughly 260 and 315 PetaFLOPS were withdrawn before publication. At SC21 in November 2021, David Kahaner of the Asian Technology Information Program reported two exascale machines running in China: Sunway OceanLight at Qingdao (more than 103,600 SW26010Pro processors, about 1.05 ExaFLOPS on Linpack, completed March 2021) and Tianhe-3 at Tianjin (Phytium ARM CPUs with Matrix-2000+ accelerators, about 1.3 ExaFLOPS). Neither appeared on the Top500 or on China’s own Top100 list, but OceanLight won the 2021 Gordon Bell Prize for a random-quantum-circuit simulation that its authors described as closing the “quantum supremacy” gap. Dongarra, asked about the missing machines: “If everybody decides they don’t want to submit to the list, the list will stagnate and die at that point.”
The silence ended on 23 June 2026, when LineShine at the National Supercomputing Centre in Shenzhen debuted at number one with 2.198 ExaFLOPS, about 20 percent ahead of El Capitan: 13.79 million cores on 304-core LX2 processors at 1.55 GHz, a proprietary LingQi interconnect, Kylin OS, 42.2 megawatts, and no GPU or other accelerator at all. It was the first Chinese machine to lead the list since TaihuLight in 2017. See China’s Tech Industry.
Europe
The EuroHPC Joint Undertaking, established in 2018 with a budget of at least €8.2 billion for 2021 to 2027, changed Europe’s procurement model: the Joint Undertaking, funded by the EU and participating states, owns the machines it buys and co-funds them with the host country. JUPITER at Forschungszentrum Jülich, built by Eviden (Bull) and ParTec on BullSequana XH3000 racks with Nvidia GH200 Grace Hopper superchips, cost about €500 million, split between EuroHPC (€250 million), the German federal research ministry and the state of North Rhine-Westphalia (€125 million each). It debuted at number four in June 2025, was inaugurated on 5 September 2025 by Chancellor Friedrich Merz, and posted exactly 1.000 ExaFLOPS on the November 2025 list, the first European exascale machine and the first on the list outside the United States. Bull, separated from Atos in 2026, was the only European vendor in the June 2026 top ten.
Dead End: The Massively Parallel Processor and the Commodity Cluster
Through the late 1980s and early 1990s, a competing vision to Cray’s vector processors emerged: Massively Parallel Processing (MPP), connecting hundreds or thousands of ordinary processors with a fast network, each running its own program on its own memory.
Companies like Thinking Machines (CM-2, 65,536 processors, 1987) and nCUBE built MPP machines, and in Britain Inmos designed the transputer as a single-chip building block for them. The CM-2 in particular attracted enormous academic attention. But programming MPP machines was hard: the programmer had to explicitly manage which processor had which data and how processors communicated. And by the mid-1990s, the hardware advantage of specialized interconnects was evaporating as commodity Ethernet and later InfiniBand approached similar bandwidth.
The Beowulf Moment
In 1994, Thomas Sterling and Don Becker at NASA built Beowulf, a cluster of 16 commodity Intel 486 PCs connected by Ethernet, running Linux. It cost approximately $50,000 and delivered performance comparable to a $250,000 workstation cluster. Beowulf demonstrated that supercomputer-class performance could be assembled from commodity parts with open-source software. Thinking Machines filed for bankruptcy in 1994. The era of purpose-built supercomputer hardware (outside of the most extreme applications) effectively ended. Nearly every machine on the Top500 list is now a cluster of CPUs and GPUs connected by a fast network; the exceptions, Fugaku’s A64FX and LineShine’s LX2, are custom CPUs from countries that want to own the whole stack. Cray Research was acquired, eventually becoming part of HPE, and its brand survives primarily as a system integrator rather than a hardware innovator.
From Weather to Proteins: The Applications That Justify the Cost
The case for exascale computing is made by problems that are not metaphors for progress but direct, measurable challenges:
Climate modeling at sufficient resolution to capture regional effects requires simulating the ocean, atmosphere, land surface, and ice sheets simultaneously at scales of kilometers. Current climate models run at 25–100 km resolution; projecting regional rainfall, drought, and flooding with policy-relevant precision requires 1 km or finer, a computational demand that current machines can barely approximate.
Protein folding (predicting the three-dimensional structure of a protein from its amino acid sequence) was one of biology’s defining unsolved problems for fifty years. DeepMind’s AlphaFold2 (2020) solved it with deep learning; the training required the equivalent of thousands of GPU-years. Understanding why AlphaFold works, and extending it to drug design and protein engineering, requires simulation at scales beyond what AlphaFold alone provides.
Nuclear stockpile stewardship (maintaining confidence in aging weapons designs without physical testing) drives the U.S. Advanced Simulation and Computing program, which has been the primary funder of the most powerful U.S. supercomputers for thirty years.
For the parallel processors that now populate supercomputer nodes, see The GPU Revolution. For the transistors those processors are built from, see The Semiconductor Race.
📚 Sources
- Murray, Charles J.: The Supermen: The Story of Seymour Cray and the Technical Wizards Behind the Supercomputer (1997), Wiley
- Dongarra, Jack et al.: “The Linpack Benchmark: Past, Present and Future” — Concurrency and Computation, Vol. 15, No. 9 (2003)
- Sterling, Thomas; Salmon, John; Becker, Donald J. & Savarese, Daniel F.: How to Build a Beowulf: A Guide to the Implementation and Application of PC Clusters (1999), MIT Press
- Top500 Project: Top500 Supercomputer Sites — top500.org (1993–present)
- Oak Ridge National Laboratory: “Frontier” — system specifications (nodes, AMD EPYC/Instinct MI250X configuration, power draw)
- Top500: Top Systems (all number-one machines and their dates)
- Top500 lists: November 1996, November 2000, November 2004, June 2008, November 2010, November 2024, November 2025, June 2026 — Rmax figures in the table, JUPITER at 1.000 ExaFLOPS as the first European exascale system
- Top500: “LineShine Debuts at No. 1 as the TOP500 Enters a New Global Exascale Era” (23 June 2026) — LX2 processors, LingQi interconnect, 42.2 MW, CPU-only design, first Chinese number one since 2017, five exascale systems
- Meuer, Hans Werner: “The ‘Computenik’ from Yokohama” — HPCwire, 17 January 2003 — 35.86 TFLOPS, the nearly fivefold gap, the New York Times “Computenik”, $350–400 million, the 1997 start, ASCI Purple and BlueGene/L announced at SC2002 with the $216–267 million IBM contract
- Dongarra interview, HPCwire, 24 May 2002 — “In some sense we have a Computenik on our hands”
- Earth Simulator (Wikipedia) — 11 March 2002, 640 SX-6 nodes, 5,120 processors
- High Productivity Computing Systems (Wikipedia) — DARPA programme 2002–2010, Cray and IBM through all three phases, Chapel and X10
- IPSJ Computer Museum: Brief History of Japanese supercomputers — S-810 and VP-200 shipped 1983, SX-2 1985
- UPI, 7 August 1987: Japan agrees to open supercomputer procurement — the Yeutter–Matsunaga exchange of letters
- Christian Science Monitor, 18 August 1987: “US-Japan trade dispute moves to supercomputers”
- FCW, August 1996: “Cray alleges dumping following $35M award to NEC” — four SX-4 systems, the $65 million loss claim
- NEC Corp. v. United States, 151 F.3d 1361 (Fed. Cir. 1998) — the 454 percent dumping margin from “facts otherwise available”
- Numerical Wind Tunnel (Wikipedia) — 124 GFLOPS in November 1993
- K computer (Wikipedia) — 88,128 SPARC64 VIIIfx, 705,024 cores, 8.162 and 10.51 PFLOPS, shutdown 30 August 2019
- Fugaku (Wikipedia) — A64FX, 158,976 nodes, 416 then 442 PFLOPS, ¥130 billion, the four benchmarks
- Tianhe-1 (Wikipedia) — Tianhe-1A’s 14,336 Xeon X5670 and 7,168 Tesla M2050, 2,566 TFLOPS
- Tianhe-2 (Wikipedia) — 32,000 Xeon E5 and 48,000 Xeon Phi, 33.86 PFLOPS, the Matrix-2000 upgrade
- Federal Register, 18 February 2015: Addition of Certain Persons to the Entity List — NUDT and the Changsha, Guangzhou and Tianjin centres, “nuclear explosive activities”
- Sunway TaihuLight (Wikipedia) — SW26010, 10,649,600 cores, 93 PFLOPS
- Trader, Tiffany: “The US Places Seven Additional Chinese Supercomputing Entities on Blacklist” — HPCwire, 8 April 2021 — the 2019 Sugon listing and the 2021 seven, Raimondo quote
- Trader, Tiffany: “Three Chinese Exascale Systems Detailed at SC21” — HPCwire, 24 November 2021 — OceanLight and Tianhe-3 figures, the withdrawn 2019 submissions, the Gordon Bell Prize, Dongarra’s “stagnate and die”
- EuroHPC Joint Undertaking: Discover EuroHPC JU — established 2018, €8.2 billion for 2021–2027
- Forschungszentrum Jülich: JUPITER — GH200, BullSequana XH3000, Eviden and ParTec, the €500 million split
- EuroHPC JU: “JUPITER: Launching Europe’s Exascale Era” (5 September 2025) — inauguration with Chancellor Merz
- HPCwire, 17 September 2026: “CEO Le Roux Is ‘Bullish’ on French Supercomputer Maker’s Future” — Bull’s separation from Atos in 2026