Skip to content

Alex Krizhevsky

Abstract

Alex Krizhevsky trained the neural network that started the deep learning boom on two gaming graphics cards in his bedroom at his parents’ house in Toronto. AlexNet won the 2012 ImageNet competition by a margin large enough to end the argument about whether neural networks worked, and within a year the entire computer vision field had abandoned what it was doing. He had to be talked into the project by a fellow graduate student. Five years later, after Google acquired the three-person company built around the result, he left, saying he had lost interest in the work. He is the least visible person in the history of modern AI, and he wrote its founding artifact.

Tiny Images

Krizhevsky grew up in Canada and did his bachelor’s, master’s, and doctoral work at the University of Toronto, where Geoffrey Hinton supervised him.

His first widely used contribution was not a network but data. In 2009 he assembled and hand-labelled CIFAR-10 and CIFAR-100 out of the 80 Million Tiny Images collection: 60,000 colour images at 32 by 32 pixels, sorted into ten or a hundred classes. They were small enough to train on the hardware a graduate student could get and clean enough to compare methods on, and they became the standard benchmark for a decade of vision research. The accompanying technical report, Learning Multiple Layers of Features from Tiny Images, is one of the most cited documents in machine learning that was never published in a journal.

To train on them he wrote cuda-convnet, a convolutional network implementation in CUDA that ran on NVIDIA consumer graphics cards. This was unusual at the time. GPUs were for games and for a small amount of scientific computing; almost nobody in machine learning was writing their own CUDA kernels.

Being Talked Into It

Ilya Sutskever, a fellow Hinton student, had concluded that neural networks would work if given enough data, and that Fei-Fei Li’s ImageNet was the dataset that would prove it. Krizhevsky was not interested; he had a working system on CIFAR and no particular wish to scale it to 1.2 million images across a thousand categories. Sutskever convinced him.

The hardware was two NVIDIA GTX 580 cards, 3 gigabytes of memory each, in a machine in Krizhevsky’s bedroom at his parents’ house. That memory limit dictated the architecture: the network had to be split across the two cards, with the layers communicating only at certain points, because it did not fit on one. Krizhevsky extended cuda-convnet to handle the split and then spent roughly a year adjusting and retraining.

The finished network had eight learned layers, five convolutional and three fully connected, about 60 million parameters and 650,000 neurons. It used ReLU activations rather than the conventional sigmoid or tanh, which trained several times faster, and dropout in the fully connected layers to limit overfitting. A full training run took five to six days.

September 2012

The three entered the ImageNet Large Scale Visual Recognition Challenge as team SuperVision on September 30, 2012, and scored a top-5 error rate of 15.3 percent. The second-place entry, using the hand-engineered features that the field had spent fifteen years refining, scored 26.2 percent.

A margin of nearly eleven points in a competition where half a point was a result did not invite debate. Within a year essentially every serious computer vision group had switched to deep convolutional networks, and the industry began hiring the small number of people who knew how to train them. The consequences are traced in ImageNet and the Deep Learning Revolution and The GPU Revolution.

The paper, ImageNet Classification with Deep Convolutional Neural Networks by Krizhevsky, Sutskever, and Hinton, appeared at NeurIPS in December 2012 and is now among the most cited papers in computer science.

Two details of that result mattered as much as the architecture. It ran on consumer gaming hardware, which meant any lab could reproduce it, and NVIDIA’s business changed permanently as a result. And it was trained by one graduate student, which meant the barrier to the next such result was low.

The Auction

Hinton, Sutskever, and Krizhevsky incorporated DNNresearch in late 2012, a company with three employees, no product, and no revenue, and let the large technology firms bid for it. Google, Microsoft, Baidu, and DeepMind competed. Google won, at a price reported at $44 million, and the three joined the company in 2013.

Krizhevsky worked at Google Brain for four years. In September 2017 he left, and the reason he gave was that he had lost interest in the work. He went to Dessa, a small Toronto machine learning company later acquired by Square, and has kept out of public view since. He gives no interviews, holds no chair, and runs no lab, while Hinton collected a Nobel Prize and Sutskever co-founded OpenAI.

The 2012 AlexNet source code was released by Google and the Computer History Museum on March 20, 2025, after five years of negotiation.

What It Was and Was Not

AlexNet introduced no fundamentally new idea. Convolutional networks were Yann LeCun’s, from the 1980s; backpropagation was older; ReLU and dropout were recent but not Krizhevsky’s alone. What he supplied was the demonstration, at a scale nobody had attempted, that the old approach worked when given enough data and enough arithmetic.

That is the uncomfortable reading of 2012. The neural network winter had not been caused by a missing insight. It had been caused by machines too slow to show that the existing insight was right, and it ended when a student wired two gaming cards together in a bedroom. The people who had spent those decades arguing that symbolic methods were the only viable path had been wrong about the algorithms and right about nothing else; the argument was settled by hardware.

📚 Sources