Skip to content

The Supercomputer Era: The Fastest Machines on Earth

Abstract

This article traces supercomputing from Seymour Cray’s machines in the 1960s through the vector era, the shift to massively parallel commodity clusters, and the race to exascale. It is also a story of national rivalry: Japanese vector machines and a US anti-dumping tariff of 454 percent in the 1990s, the Earth Simulator shock of 2002, China’s rise to the top of the Top500 and its withdrawal from the list under US export controls, Europe’s first exascale machine in Jülich in 2025, and a Chinese machine back at number one in 2026.

Seymour Cray and the First Supercomputers

Seymour Cray was an electrical engineer from Chippewa Falls, Wisconsin, who combined extraordinary technical intuition with an almost monastic focus. He worked in isolation, literally and figuratively: he built a lab under his house in Wisconsin, away from the corporate offices of the companies he worked for, and was known to say that the best way to design a computer was to hire a few good engineers and leave them alone.

At Control Data Corporation (CDC), Cray designed the CDC 6600 (1964), widely recognized as the first true supercomputer. It achieved 3 megaFLOPS (3 million floating-point operations per second), three times the speed of the previous record-holder, IBM’s 7030 Stretch. It did this through a novel architecture: ten peripheral processors handled input/output while a central processor focused exclusively on computation, and the central processor itself executed instructions in parallel using multiple functional units.

IBM’s chairman Thomas Watson Jr. wrote a memo in August 1963 asking how a laboratory of thirty-four people, fourteen of them engineers, had taken the lead from IBM’s own development organization. Cray’s response was one line: “It seems like Mr. Watson has answered his own question.”

Cray left CDC in 1972 to found Cray Research. The Cray-1 (1976) was his masterpiece: a $8 million machine that delivered 160 megaFLOPS, weighed 5.5 tons, and was shaped like a padded bench, the padding concealing the cooling system. Its defining innovation was the vector register: a hardware unit that could perform the same arithmetic operation on 64 numbers simultaneously, accelerating the loops that dominated scientific computation. Los Alamos National Laboratory bought the first unit for nuclear weapons simulation. The Cray-1 became the standard by which all scientific computation was measured.

Vector Processing vs. Modern Parallelism

A vector processor executes one instruction on many data elements simultaneously: “add these 64 numbers to those 64 numbers” in a single operation. This is powerful for the regular, predictable loops of scientific computation (weather simulation, fluid dynamics, matrix operations) and less useful for irregular, data-dependent code. Modern parallelism, by contrast, runs many independent threads of execution simultaneously on separate cores. Today’s supercomputers use both: many-core CPUs and GPUs for thread-level parallelism, combined with SIMD instructions for data-level parallelism within each core.

The Cold War and Scientific Computation

Supercomputers were not merely academic tools. The U.S. government’s primary customer for the most powerful machines was the nuclear weapons complex, specifically, the national laboratories at Los Alamos, Lawrence Livermore, and Sandia. After the Limited Test Ban Treaty (1963) prohibited atmospheric nuclear tests and the Comprehensive Test Ban Treaty (1996) halted underground tests, simulation became the only legal way to verify that nuclear weapons in the stockpile still worked as designed. The required simulation accuracy drove supercomputer procurement for decades.

The Cold War framing extended to export controls: the U.S. government restricted the export of supercomputers to Soviet-bloc countries and, later, China. The first commercial challenge to Cray came from Japan. Hitachi shipped the S-810 and Fujitsu the VP-200 in 1983, NEC the SX-2 in 1985, all vector machines in the Cray mould. American vendors complained that Japanese government agencies bought only Japanese machines and that the same firms undercut Cray abroad; in August 1987 Washington and Tokyo exchanged letters committing Japan to open public procurement, renewed in a broader agreement in March 1990. The dispute peaked in 1996, when the National Center for Atmospheric Research chose four NEC SX-4 systems in a $35 million lease. Cray Research petitioned the Commerce Department, claiming NEC was taking a $65 million loss on the deal; NEC declined to answer the investigation questionnaire, and in 1997 Commerce set an anti-dumping duty of 454 percent on NEC supercomputers, a figure calculated from “facts otherwise available”. The sale was shelved. See Japan’s Computing Industry.

The civilian scientific applications were equally consequential. Weather forecasting (global atmospheric simulation) required updating models faster than real time to be useful, a threshold that each generation of supercomputers moved further ahead. The European Centre for Medium-Range Weather Forecasts (ECMWF) and the U.S. National Weather Service became major supercomputer customers, and the improvement in forecast accuracy through the 1980s and 1990s is directly attributable to increases in computing power.

The Top500 List and the Benchmark Race

In 1993, Jack Dongarra at the University of Tennessee and Hans Meuer at the University of Mannheim began publishing the Top500 list, a twice-yearly ranking of the 500 most powerful supercomputers in the world, measured by performance on the Linpack benchmark: solving a dense system of linear equations.

Linpack performance is measured in FLOPS (floating-point operations per second). The trajectory of the Top500 list illustrates Moore’s Law applied to the largest machines:

First list at #1 Machine Site Country Rmax
June 1993 CM-5 (Thinking Machines) Los Alamos US 59.7 GigaFLOPS
Nov 1993 Numerical Wind Tunnel (Fujitsu) National Aerospace Laboratory Japan 124 GigaFLOPS
June 1994 Intel Paragon XP/S 140 Sandia US 143 GigaFLOPS
June 1996 SR2201 (Hitachi) University of Tokyo Japan 220 GigaFLOPS
Nov 1996 CP-PACS (Hitachi) University of Tsukuba Japan 368 GigaFLOPS
June 1997 ASCI Red (Intel) Sandia US 1.07 TeraFLOPS
Nov 2000 ASCI White (IBM) Lawrence Livermore US 4.9 TeraFLOPS
June 2002 Earth Simulator (NEC) Yokohama Japan 35.9 TeraFLOPS
Nov 2004 BlueGene/L (IBM) Lawrence Livermore US 70.7 TeraFLOPS
June 2008 Roadrunner (IBM) Los Alamos US 1.03 PetaFLOPS
Nov 2010 Tianhe-1A (NUDT) Tianjin China 2.57 PetaFLOPS
June 2011 K computer (Fujitsu) RIKEN, Kobe Japan 8.2 PetaFLOPS
June 2013 Tianhe-2 (NUDT) Guangzhou China 33.9 PetaFLOPS
June 2016 Sunway TaihuLight Wuxi China 93 PetaFLOPS
June 2018 Summit (IBM) Oak Ridge US 122.3 PetaFLOPS
June 2020 Fugaku (Fujitsu) RIKEN, Kobe Japan 416 PetaFLOPS
June 2022 Frontier (HPE) Oak Ridge US 1.1 ExaFLOPS
Nov 2024 El Capitan (HPE) Lawrence Livermore US 1.74 ExaFLOPS
June 2026 LineShine Shenzhen China 2.2 ExaFLOPS

Rmax is the performance on the list where the machine first took the top spot; several were upgraded later (BlueGene/L, K, Fugaku, Frontier and El Capitan all posted higher numbers on subsequent lists). Jaguar (Oak Ridge, November 2009), Sequoia (Lawrence Livermore, June 2012) and Titan (Oak Ridge, November 2012) held the spot briefly between the rows shown.

Frontier, installed at Oak Ridge National Laboratory in 2022, was the first machine to exceed one ExaFLOP ($10^{18}$ floating-point operations per second). It comprises 74 HPE Cray EX cabinets and 9,408 nodes, each node pairing one AMD EPYC CPU with four AMD Instinct MI250X GPU accelerators, consuming 21 megawatts of power, roughly the electricity consumption of a small city. El Capitan at Lawrence Livermore, built on AMD’s MI300A, took the top spot in November 2024 at 1.74 ExaFLOPS. By June 2026 five machines exceeded an ExaFLOP on Linpack: LineShine, El Capitan, Frontier, Aurora at Argonne, and JUPITER at Jülich.

The National Race

The Top500 was conceived as a statistics project. It turned into a scoreboard that governments read, and three regions treated the top spot as a matter of policy.

Japan

Japan’s vector builders held the top spot for most of the mid-1990s (Numerical Wind Tunnel, SR2201, CP-PACS), but the machine that changed American policy was the Earth Simulator. Planned from 1997 for climate and earthquake research and funded with roughly $350 to 400 million by the Japanese government, it went into operation at Yokohama on 11 March 2002: 640 NEC SX-6 nodes, 5,120 vector processors, 35.86 TeraFLOPS on Linpack, nearly five times the next machine on the list (ASCI White). The New York Times called it a “Computenik”, and Jack Dongarra said, “In some sense we have a Computenik on our hands.” Hans Meuer, co-founder of the Top500, noted that the Americans “were caught on the wrong foot” even though the completion date had been public for five years. In 2002 DARPA launched its High Productivity Computing Systems program, which funded Cray and IBM through three phases and produced the Chapel and X10 languages; at SC2002 in Baltimore, the Department of Energy announced a $216 to 267 million IBM contract for ASCI Purple and BlueGene/L. BlueGene/L took the top spot back in November 2004.

Japan returned to number one twice more, both times with Fujitsu machines at RIKEN in Kobe. The K computer (88,128 SPARC64 VIIIfx processors, 705,024 cores) led the June and November 2011 lists and was the first machine past 10 PetaFLOPS (10.51 in November 2011); it ran until 30 August 2019. Its successor Fugaku, built on Fujitsu’s ARM-based A64FX with 158,976 nodes and a programme cost of about ¥130 billion, led four consecutive lists from June 2020 to November 2021 (416, later 442 PetaFLOPS) and was the first system to top HPL, HPCG, HPL-AI and Graph500 at the same time. It did so without GPUs.

China

China’s first number one, Tianhe-1A at Tianjin in November 2010 (2.57 PetaFLOPS), paired 14,336 Intel Xeon X5670 CPUs with 7,168 Nvidia Tesla M2050 GPUs. Tianhe-2 at Guangzhou, built by the National University of Defense Technology (NUDT) with 32,000 Xeon E5 CPUs and 48,000 Xeon Phi coprocessors, held the top spot for six consecutive lists from June 2013 to November 2015 at 33.86 PetaFLOPS. Both ran on American silicon. In February 2015 the U.S. Commerce Department added NUDT and the national supercomputing centres in Changsha, Guangzhou and Tianjin to its Entity List, stating that the Tianhe machines were believed to be used for “nuclear explosive activities”, which blocked Intel from supplying chips for Tianhe-2’s planned upgrade. NUDT replaced the Xeon Phis with domestic Matrix-2000 accelerators.

The answer came in June 2016. Sunway TaihuLight at Wuxi ran on 40,960 Chinese-designed SW26010 processors (10,649,600 cores), reached 93 PetaFLOPS, and held first place for four lists until Summit in June 2018. Export controls widened: Sugon and its Hygon chip venture in June 2019, then in April 2021 seven more entities including Phytium, Sunway Microelectronics and the centres at Wuxi, Shenzhen, Jinan and Zhengzhou, cited by Commerce Secretary Gina Raimondo for supporting “military modernization efforts” and weapons programmes.

China then stopped submitting. Two systems provisionally entered for the 2019 list at roughly 260 and 315 PetaFLOPS were withdrawn before publication. At SC21 in November 2021, David Kahaner of the Asian Technology Information Program reported two exascale machines running in China: Sunway OceanLight at Qingdao (more than 103,600 SW26010Pro processors, about 1.05 ExaFLOPS on Linpack, completed March 2021) and Tianhe-3 at Tianjin (Phytium ARM CPUs with Matrix-2000+ accelerators, about 1.3 ExaFLOPS). Neither appeared on the Top500 or on China’s own Top100 list, but OceanLight won the 2021 Gordon Bell Prize for a random-quantum-circuit simulation that its authors described as closing the “quantum supremacy” gap. Dongarra, asked about the missing machines: “If everybody decides they don’t want to submit to the list, the list will stagnate and die at that point.”

The silence ended on 23 June 2026, when LineShine at the National Supercomputing Centre in Shenzhen debuted at number one with 2.198 ExaFLOPS, about 20 percent ahead of El Capitan: 13.79 million cores on 304-core LX2 processors at 1.55 GHz, a proprietary LingQi interconnect, Kylin OS, 42.2 megawatts, and no GPU or other accelerator at all. It was the first Chinese machine to lead the list since TaihuLight in 2017. See China’s Tech Industry.

Europe

The EuroHPC Joint Undertaking, established in 2018 with a budget of at least €8.2 billion for 2021 to 2027, changed Europe’s procurement model: the Joint Undertaking, funded by the EU and participating states, owns the machines it buys and co-funds them with the host country. JUPITER at Forschungszentrum Jülich, built by Eviden (Bull) and ParTec on BullSequana XH3000 racks with Nvidia GH200 Grace Hopper superchips, cost about €500 million, split between EuroHPC (€250 million), the German federal research ministry and the state of North Rhine-Westphalia (€125 million each). It debuted at number four in June 2025, was inaugurated on 5 September 2025 by Chancellor Friedrich Merz, and posted exactly 1.000 ExaFLOPS on the November 2025 list, the first European exascale machine and the first on the list outside the United States. Bull, separated from Atos in 2026, was the only European vendor in the June 2026 top ten.

Dead End: The Massively Parallel Processor and the Commodity Cluster

Through the late 1980s and early 1990s, a competing vision to Cray’s vector processors emerged: Massively Parallel Processing (MPP), connecting hundreds or thousands of ordinary processors with a fast network, each running its own program on its own memory.

Companies like Thinking Machines (CM-2, 65,536 processors, 1987) and nCUBE built MPP machines, and in Britain Inmos designed the transputer as a single-chip building block for them. The CM-2 in particular attracted enormous academic attention. But programming MPP machines was hard: the programmer had to explicitly manage which processor had which data and how processors communicated. And by the mid-1990s, the hardware advantage of specialized interconnects was evaporating as commodity Ethernet and later InfiniBand approached similar bandwidth.

The Beowulf Moment

In 1994, Thomas Sterling and Don Becker at NASA built Beowulf, a cluster of 16 commodity Intel 486 PCs connected by Ethernet, running Linux. It cost approximately $50,000 and delivered performance comparable to a $250,000 workstation cluster. Beowulf demonstrated that supercomputer-class performance could be assembled from commodity parts with open-source software. Thinking Machines filed for bankruptcy in 1994. The era of purpose-built supercomputer hardware (outside of the most extreme applications) effectively ended. Nearly every machine on the Top500 list is now a cluster of CPUs and GPUs connected by a fast network; the exceptions, Fugaku’s A64FX and LineShine’s LX2, are custom CPUs from countries that want to own the whole stack. Cray Research was acquired, eventually becoming part of HPE, and its brand survives primarily as a system integrator rather than a hardware innovator.

From Weather to Proteins: The Applications That Justify the Cost

The case for exascale computing is made by problems that are not metaphors for progress but direct, measurable challenges:

Climate modeling at sufficient resolution to capture regional effects requires simulating the ocean, atmosphere, land surface, and ice sheets simultaneously at scales of kilometers. Current climate models run at 25–100 km resolution; projecting regional rainfall, drought, and flooding with policy-relevant precision requires 1 km or finer, a computational demand that current machines can barely approximate.

Protein folding (predicting the three-dimensional structure of a protein from its amino acid sequence) was one of biology’s defining unsolved problems for fifty years. DeepMind’s AlphaFold2 (2020) solved it with deep learning; the training required the equivalent of thousands of GPU-years. Understanding why AlphaFold works, and extending it to drug design and protein engineering, requires simulation at scales beyond what AlphaFold alone provides.

Nuclear stockpile stewardship (maintaining confidence in aging weapons designs without physical testing) drives the U.S. Advanced Simulation and Computing program, which has been the primary funder of the most powerful U.S. supercomputers for thirty years.

For the parallel processors that now populate supercomputer nodes, see The GPU Revolution. For the transistors those processors are built from, see The Semiconductor Race.


📚 Sources