Export citation

Export citation

Choose format for download:

Download Citation
  • Access by Xinjiang University

Nobel Lecture: Boltzmann machines*

Geoffrey Hinton

Geoffrey Hinton

Rev. Mod. Phys. 97, 030502 – Published 25 August, 2025

DOI: https://doi.org/10.1103/RevModPhys.97.030502

Abstract

To solve difficult computational tasks, artificial neural networks need to construct appropriate internal representations by adapting the weights on their connections in the direction that improves performance. In the 1980s, there were two promising techniques for computing the gradients required to adapt the weights. One technique was backpropagation, which is now used in almost all artificial intelligence systems. The other was the Boltzmann machine learning algorithm, which is no longer used. After describing two of the major successes of backpropagation, this Nobel Lecture explains the Boltzmann machine learning algorithm, which uses a property of the Boltzmann distribution to compute gradients in an elegant and unexpected way.

Physics Subject Headings (PhySH)

  • *The 2024 Nobel Prize for Physics was shared by John J. Hopfield and Geoffrey E. Hinton. This paper is the text of the address given in conjunction with the award.

Article Text

References (27)

  1. Ackley, D. H., G. E. Hinton, and T. J. Sejnowski, 1985, “A learning algorithm for Boltzmann machines,” Cognit. Sci. 9, 147–169.
  2. Amari, S.-i., 1993, “Backpropagation and stochastic gradient descent method,” Neurocomputing 5, 185–196.
  3. Anderson, J. R., and C. Peterson, 1987, “A mean field theory learning algorithm for neural networks,” Complex Syst. 1, 995–1019, https://content.wolfram.com/sites/13/2018/02/01-5-6.pdf.
  4. Crick, F., and G. Mitchison, 1983, “The function of dream sleep,” Nature (London) 304, 111–114.
  5. Deng, J., W. Dong, R. Socher, L.-J. Li, K. Li, and L. FeiFei, 2009, “Imagenet: A large-scale hierarchical image database,” in Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2009), Miami, 2009, pp. 248–255 (IEEE, New York).
  6. Freund, Y., and D. Haussler, 1991, “Unsupervised learning of distributions on binary vectors using two layer networks,” in Advances in Neural Information Processing Systems, Vol. 4, edited by J. Moody, S. Hanson, and R. P. Lippmann (Morgan Kaufmann, Burlington, MA).
  7. Geman, S., and D. Geman, 1984, “Stochastic relaxation, Gibbs distributions, and the Bayesian restoration of images,” IEEE Trans. Pattern Anal. Mach. Intell. PAMI-6, 721–741.
  8. Hinton, G. E., 2002, “Training products of experts by minimizing contrastive divergence,” Neural Comput. 14, 1771–1800.
  9. Hinton, G. E., J. L. McClelland, and D. E. Rumelhart, 1986, “Distributed representations,” in Parallel Distributed Processing, Vol. 1, edited by D. E. Rumelhart and J. L. McClelland (MIT Press, Cambridge, MA), pp. 77–109.
  10. Hinton, G. E., and R. R. Salakhutdinov, 2006, “Reducing the dimensionality of data with neural networks,” Science 313, 504–507.
  11. Hinton, G. E., et al., 2012, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Process. Mag. 29, 82–97.
  12. Hopfield, J. J., 1982, “Neural networks and physical systems with emergent collective computational abilities,” Proc. Natl. Acad. Sci. U.S.A. 79, 2554–2558.
  13. Kirkpatrick, S., C. D. Gelatt, and M. P. Vecchi, 1983, “Optimization by simulated annealing,” Science 220, 671–680.
  14. Krizhevsky, A., I. Sutskever, and G. E. Hinton, 2012, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems, Vol. 25, edited by F. Pereira, C. J. Burges, L. Bottou, and K. Q. Weinberger (Morgan Kaufmann, Burlington, MA).
  15. LeCun, Y., 1985, “Une procedure d’apprentissage pour reseau a seuil asymetrique [A learning scheme for asymmetric threshold networks],” in Proceedings of Cognitiva 85, Paris, 1985, pp. 599–604.
  16. LeCun, Y., 1989, “Generalization and network design strategies,” in Connectionism in Perspective, edited by R. Pfeifer, Z. Schreter, F. Fogelman-Soulié, and L. Steels (North-Holland, Amsterdam), pp. 143–155.
  17. Linnainmaa, S., 1970, “The representation of the cumulative rounding error of an algorithm as a Taylor expansion of the local rounding errors,” master’s thesis (University of Helsinki).
  18. Nair, V., and G. E. Hinton, 2010, “Rectified linear units improve restricted Boltzmann machines,” in Proceedings of the 27th International Conference on Machine Learning (ICML-10), Haifa, Israel, 2010, edited by J. Fürnkranz and T. Joachims (Omnipress, Madison, WI), pp. 807–814.
  19. Parker, D. B., 1985, “Learning-logic,” MIT Sloan School of Management Technical Report No. 47.
  20. Richards, I. A., 2017. Principles of Literary Criticism (Routledge, Abingdon, England).
  21. Rumelhart, D. E., G. E. Hinton, and R. J. Williams, 1986, “Learning representations by back-propagating errors,” Nature (London) 323, 533–536.
  22. Salakhutdinov, R., and G. Hinton, 2009, “Deep Boltzmann machines,” in Proceedings of Machine Learning Research: Artificial Intelligence and Statistics (AISTATS 2009), Clearwater Beach, FL, 2009 (PMLR), edited by D. van Dyk and M. Welling, pp. 448–455.
  23. Smolensky, P., 1986, “Information processing in dynamical systems: Foundations of harmony theory, in Parallel Distributed Processing, Vol. 1, edited by D. E. Rumelhart and J. L. McClelland (MIT Press, Cambridge, MA), pp. 194–281.
  24. Srivastava, N., G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, 2014, “Dropout: A simple way to prevent neural networks from overfitting,” J. Mach. Learn. Res. 15, 1929–1958, https://www.jmlr.org/papers/volume15/srivastava14a/srivastava14a.pdf.
  25. Tieleman, T., and G. Hinton, 2009, “Using fast weights to improve persistent contrastive divergence,” in Proceedings of the 26th Annual International Conference on Machine Learning, Montreal, 2009, edited by A. Danyluk, L. Bottou, and M. Littman (Association for Computing Machinery, New York), pp. 1033–1040.
  26. Vaswani, A., N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, 2017, “Attention is all you need,” in Advances in Neural Information Processing Systems, Vol. 30, edited by I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Curran Associates, Red Hook, NY), pp. 5998–6008.
  27. Werbos, P., 1974, “Beyond regression: New tools for prediction and analysis in the behavioral sciences,” Ph.D. thesis (Harvard University Committee on Applied Mathematics).

Outline

Information

Sign In to Your Journals Account

Filter

Filter

Article Lookup

Enter a citation