Machine Learning has been at the forefront of research in several disciplines for a number of years already. Given its roots in the statistical physics of learning, in this collection we highlight research at the intersection of machine learning and statistical physics, on the occasion of the two Statistical Physics Meets Machine Learning and the two Machine Learning Meets Statistical Physics sessions at the 2025 Global Physics Summit. The Physical Review E Special Collection Statistical Physics Meets Machine Learning - Machine Learning Meets Statistical Physics, guest-edited by David Schwab (CUNY, New York) and Yuhai Tu (IBM Watson Research Center, Yorktown Heights, NY) explores the latest insights gained looking at machine learning problems through the lens of statistical physics and at statistical physics through the lens of learning processes. We are confident that this synergy will bring to light novel perspectives and pave the way for future breakthroughs.

The Collection was guest-edited by David Schwab and Yuhai Tu. Every article published in this collection underwent a rigorous peer review process, adhering to the same high standards applied to all papers. The Physical Review E editorial team managed the peer review and made all editorial decisions.

See also the Physics Magazine Viewpoint by Hugo Cui covering papers in this Collection.

Motivated by developmental processes in biology, this study extends a classical optimal transport framework to processes that include growth. One key result presented here is the derivation of an analog of the Benamou-Brenier theorem, which identifies the conditions under which the transport map is described by stochastic dynamics on a potential landscape. The author exploits a mathematical connection between stochastic control theory, optimal transport and nonequilibrium thermodynamics. The framework can be used to find a potential landscape that describes the map from any (non-degenerate) initial distribution over cell states (of arbitrary dimensionality) to any target distribution in a finite time.

Placed at the intersection of theoretical neuroscience, machine learning, and statistical physics, this work extends earlier approaches to continuous decoding, presenting a geometric analysis of discriminability under neural variability. The results help understand how representational geometry governs task performance.

To understand the properties of the cross covariance between large-dimensional datasets, the authors extend the Marchenko–Pastur results for spectra of eigenvalues of self-covariance matrices to singular values of cross-covariance matrices. In particular, they analyze the spectrum of cross correlations when the dimensionality of the two variables is larger than the sample size, which is common in modern data science applications.

Dropout is a widely used regularization technique in training neural networks for which dropout rates are typically selected heuristically. In order to develop a principled framework for understanding their impact on learning dynamics, the authors present an analytic theory of dropout in two-layer neural networks trained via online stochastic gradient descent.

When describing the world, one seeks models with as few details as possible that are as accurate as possible; that is, one seeks good compressions. To solve this compression problem, the authors illustrate the minimax entropy principle (a direct generalization of the maximum entropy principle) and survey promising applications.

Transformers are a type of machine learning architecture. The authors investigate in this study how they can acquire an understanding of language structure when trained via next-token prediction. Notably, the authors find that when the data exhibit hierarchical structure, convolutional neural networks actually learn the task more efficiently than transformers. #MachineLearningSpotlight

This study reports that low-dimensional structures in training manifolds of deep neural networks arise not due to limited flexibility, but rather from the intrinsic low-dimensionality of the training task.

#MachineLearningSpotlight

A central paradox in machine learning concerns how lossy data transformations can improve generalization. The authors investigate this phenomenon using a solvable, high-dimensional regression model inspired by the renormalization group from statistical physics. They find that data coarse graining results in a new generalization phenomenon, distinct from the well-known double descent.

Sign In to Your Journals Account

Filter

Filter

Article Lookup

Enter a citation