Barto and Sutton: Learning from Reward
Abstract
In 1977 a fresh Michigan PhD, Andrew Barto (born 1948), was hired as a postdoc at the University of Massachusetts Amherst on an Air Force contract to evaluate a theory that neurons are hedonists. His graduate student Richard Sutton (born 1957 or 1958) had come from a Stanford psychology degree convinced that learning from reward was the thing to study. Over the next decade the two turned a fringe idea into reinforcement learning: the 1983 actor–critic paper that balanced a pole, Sutton’s 1988 temporal-difference method, and the 1998 textbook that defined the field. Their algorithms are inside AlphaGo and in the training of chatbots. They shared the 2024 Turing Award, and used the occasion to say the industry was building bridges and testing them by having people walk across.
The Hedonistic Neuron
Andrew Barto was born in 1948 and went to the University of Michigan to study naval architecture, switched to mathematics, took his BS with distinction in 1970 and a PhD in computer science in 1975 under Bernard Zeigler, with a thesis on cellular automata as models of natural systems. In 1977 he came to UMass Amherst as a postdoctoral researcher, “to join three faculty members who were working on neural networks.” The three, Michael Arbib, William Kilmer and Nico Spinelli, had a contract from the Air Force Office of Scientific Research to assess the ideas of A. Harry Klopf, a scientist at Wright-Patterson Air Force Base whose theory, published in 1982 as The Hedonistic Neuron, held that individual neurons seek to maximise their own reward and that intelligence is what emerges. Barto’s job was to find out whether there was anything in it.
Richard Sutton was born in Toledo, Ohio, grew up in Oak Brook, Illinois, and took a BA in psychology at Stanford in 1978. He arrived at Amherst that year as a graduate student, joined shortly by Charles Anderson, and the group that formed around Barto became the Adaptive Networks Laboratory, later the Autonomous Learning Laboratory. What they took from Klopf was the problem rather than the theory: an agent that acts, is rewarded or not, and must work out which of its past actions deserved the credit. Sutton’s 1984 dissertation was titled “Temporal Credit Assignment in Reinforcement Learning,” and the field had its name.
Two Papers
The first result was published in September 1983 in the IEEE Transactions on Systems, Man, and Cybernetics: “Neuronlike Adaptive Elements That Can Solve Difficult Learning Control Problems,” by Barto, Sutton and Anderson. Two adaptive units learned, by trial and error and with no model of the physics, to keep a pole balanced on a moving cart. One unit, the critic, learned to predict how well things were going; the other, the actor, learned what to do, using the critic’s prediction errors as its signal. The actor–critic architecture is still the shape of most reinforcement-learning systems.
The second was Sutton’s alone. “Learning to Predict by the Methods of Temporal Differences,” in Machine Learning in August 1988, set out TD learning: instead of waiting for the final outcome to correct a prediction, adjust each prediction toward the next one, so that learning proceeds from the difference between successive guesses. It was more efficient than waiting, it could be proved to converge, and the signal it computed, the difference between expected and received reward, turned out a decade later to match what dopamine neurons in the brain were doing. The wider history, from Bellman through TD-Gammon and Atari to AlphaGo, is in Reinforcement Learning. Sutton added the Dyna architecture (1991), which lets an agent learn a model of its world and plan with it, the options framework for actions that take time (1999, with Doina Precup and Satinder Singh), and the policy-gradient theorem (2000).
Sutton left academia for GTE Laboratories in 1985, came back to Amherst as a research scientist in 1995, moved to AT&T’s Shannon Laboratory in 1998, and in 2003 went to the University of Alberta, where he has been since. Barto stayed at Amherst for his whole career, professor from 1991, department chair from 2007 to 2011, emeritus in 2012.
The Book
Reinforcement Learning: An Introduction came out from MIT Press in 1998, with a second edition in 2018, and is the reason the field looks the way it does. It fixed the vocabulary (agent, environment, reward, value function, policy), organised the methods around the Bellman equation, and treated the subject as one thing rather than a scattering of results in control theory, psychology and AI. By 2025 it had been cited more than 75,000 times. Sutton keeps the full text free on his website.
Alberta, DeepMind, and the Bitter Lesson
At Alberta Sutton built a reinforcement-learning group and is a Fellow of the Alberta Machine Intelligence Institute; from 2017 to 2023 he also held a post at DeepMind, and in 2023 he joined John Carmack’s Keen Technologies. He became a Canadian citizen in 2015 and gave up his American citizenship in 2017.
His most read piece of writing is about 1,200 words long. “The Bitter Lesson,” posted on 13 March 2019, opens: “The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin.” Chess, Go, speech and vision had each gone the same way: researchers built in what they knew, this “always helps in the short term, and is personally satisfying to the researcher,” and then “in the long run it plateaus and even inhibits further progress,” until it is beaten by methods that just search and learn at scale. The essay became the standard citation for the case that scale beats design.
The Award
The ACM announced on 5 March 2025 that Barto and Sutton had received the 2024 Turing Award, “for developing the conceptual and algorithmic foundations of reinforcement learning.” ACM president Yannis Ioannidis said their work “is not a steppingstone that we have now moved on from.” Sutton had been an AAAI Fellow since 2001 and was elected to the Royal Society of Canada in 2016 and the Royal Society of London in 2021; Barto had received the IEEE Neural Networks Pioneer Award in 2004 and the IJCAI Award for Research Excellence in 2017.
They spent part of the press round criticising the people who now used their work. Barto told the Financial Times that “releasing software to millions of people without safeguards is not good engineering practice,” and the two compared the industry’s approach to “building a bridge and testing it by having people use it.”
Dead End: The Neuron That Wanted Things
Klopf’s hedonistic neuron did not become the basis of the field that grew out of the project set up to test it. What survived was the reframing. Instead of asking how a neuron learns, Barto and Sutton asked how an agent learns from delayed consequences, and that question turned out to have an answer in dynamic programming, a subject Bellman had finished with in the 1950s. Reinforcement learning spent the 1990s as the least fashionable branch of machine learning, too slow and too data-hungry for real problems, while supervised learning on labelled data took the money. The bitter lesson, when it came, cut in its favour: given enough computation, the slow method that learned from its own experience was the one that beat the world champion at Go.
📚 Sources
- Wikipedia: Richard S. Sutton — birth, degrees, dissertation, career dates, contributions, honours, citizenship
- Wikipedia: Andrew Barto — birth, Michigan degrees and advisor, UMass dates, laboratory, awards
- UMass Amherst CICS: Barto receives 2024 ACM Turing Award, 5 March 2025 — the announcement date, Barto’s “three faculty members” quotation, 75,000 citations, Ioannidis quotation, Sutton’s affiliations, Barto emeritus 2012
- Barto, Sutton & Anderson, “Neuronlike Adaptive Elements That Can Solve Difficult Learning Control Problems,” IEEE Transactions on Systems, Man, and Cybernetics SMC-13 (5), September 1983, pp. 834–846 (DOI 10.1109/TSMC.1983.6313077)
- Sutton, “Learning to Predict by the Methods of Temporal Differences,” Machine Learning 3 (1), August 1988, pp. 9–44 (DOI 10.1007/BF00115009)
- Sutton, Precup & Singh, “Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning,” Artificial Intelligence 112, August 1999 (DOI 10.1016/S0004-3702(99)00052-1)
- Sutton & Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press 2018 (free online)
- Rich Sutton, “The Bitter Lesson,” 13 March 2019 — the quotations
- Klopf, The Hedonistic Neuron: A Theory of Memory, Learning, and Intelligence, Hemisphere 1982 (Open Library)
- ACM: 2024 Turing Award — citation
- citiesabc, “Turing Award winners Barto and Sutton highlight dangers of AI development,” 6 March 2025 — the Financial Times “bridge” and “good engineering practice” quotations, as reported there
- Image: SD 2025 - Richard Sutton 01.jpg by Xuthoria (CC BY-SA 4.0), via Wikimedia Commons