Statistical physics and machine learning: A 30 years perspective
Naftali Tishby
Video
Slide
Machine learning and the renormalization group
Maciej Koch-Janusz
Physical systems differing in their microscopic details often display strikingly similar behaviour when probed at macroscopic scales. Those universal properties, largely determining their physical characteristics, are revealed by the renormalization group (RG) procedure, which systematically retains ‘slow’ degrees of freedom and integrates out the rest. We demonstrate a machine-learning algorithm based on a model-independent, information-theoretic characterization of a real-space RG capable of identifying the relevant degrees of freedom and executing RG steps iteratively without any prior knowledge about the system. We apply it to classical statistical physics problems in 1 and 2D: we demonstrate RG flow and extract critical exponents. We also prove results about optimality of the procedure.
Video
Slide
Opportunities for infusing physics into AI/ML algorithms
Animashree Anandkumar
There are rich opportunities for infusing knowledge from physics into AI/ML algorithms. Black-box deep learning has limitations in terms of generalizability, data requirements and robustness. By infusing domain knowledge, it is possible to overcome these limitations. We demonstrate it in stable drone-landing problem using learned dynamics. We present a novel deep-learning based robust nonlinear controller (Neural-Lander) that improves control performance of a quadrotor during landing. We employ a novel application of spectral normalization of weights of neural network to prove system stability with disturbance rejection. This is the first DNN-based nonlinear
feedback controller with stability guarantees that can utilize arbitrarily large neural nets. Experimental results demonstrate that the proposed controller significantly outperforms a baseline linear proportional-derivative (PD) controller in both 1D and 3D landing cases. I will also describe other opportunities for using physical knowledge in AI/ML algorithms.
Video
Slide
Bridging Many-Body Quantum Physics and Deep Learning via Tensor Networks
Yoav Levine
We establish a Tensor Network (TN) based common language between deep learning and many-body quantum physics, which allows us to offer bidirectional contributions. By showing that many-body wave-functions are structurally equivalent to mappings of ConvNets and RNNs, we construct their TN equivalents, and suggest quantum entanglement measures as natural quantifiers of dependencies in such networks. Accordingly, we propose a novel entanglement based deep learning design scheme. In the other direction, we identify that an inherent re-use of information in state-of-the-art deep learning architectures is a key trait that distinguishes them from standard TNs. Therefore, we employ a TN manifestation of information re-use and construct TNs corresponding to powerful architectures such as deep recurrent and overlapping convolutional networks. This allows us to demonstrate that the entanglement scaling supported by state-of-the-art deep learning architectures matches that of MERA TN in 1D, and that they support volume law entanglement in 2D, polynomially more efficiently than RBMs. We thus provide theoretical motivation to shift trending neural-network based wave-function representations closer to state-of-the-art deep learning architectures.
Video
Slide
On Learning Graph Inverse Problems with Neural Networks
Joan Bruna
Inverse Problems on graphs encompass many areas of physics, algorithms and statistics, and are a confluence of powerful methods, ranging from computational harmonic analysis and high-dimensional statistics to statistical physics. Similarly as with inverse problems in signal processing, learning has emerged as an intriguing alternative to regularization and other computationally tractable relaxations, opening up new questions in which high-dimensional optimization, neural networks and data play a prominent role. In this talk, I will argue that several tasks that are ‘geometrically stable’ can be well approximated with Graph Neural Networks, a natural extension of Convolutional Neural Networks on graphs. I will present recent work on supervised community detection, quadratic assignment, neutrino detection and beyond showing the flexibility of GNNs to extend classic algorithms such as Belief Propagation.
Video
Promise and Challenges of Machine Learning in Particle Physics, Astrophysics, and Cosmology
Kyle Cranmer
There is no doubt that there is a revolution going on in machine learning and artificial intelligence, but what does it mean for physics? Is it all hype, or will it transform the way we think about and do physics? After reviewing recent progress, I will discuss some of the challenges encountered including the integration of domain-specific knowledge, robustness, and interpretability. I will also identify some of the most promising strategies for addressing these challenges. In addition, I’ll comment on how these developments are disrupting the traditional norms of scholarly publishing in high-energy physics.
Video
Slide
Deep Learning for Science: Steps to opening the blackbox
Shirley Ho
Scientists have always attempted to identify and document analytic laws that underlie physical phenomena in nature. The process of finding natural laws has always been a challenge that requires not only experimental data, but also theoretical intuition. Often times, these fundamental physical laws are derived from many years of hard work over many generation of scientists. Automated techniques for generating, collecting, and storing data have become increasingly precise and powerful, but automated discovery of natural laws in the form of analytical laws or mathematical symmetries have so far been elusive.
Over the past few years, the application of deep learning to domain sciences – from biology to chemistry and physics is raising the exciting possibility of a data-driven approach to automated science, that makes laborious hand-coding of semantics and instructions that is still necessary in most disciplines seemingly irrelevant. The opaque nature of deep models, however, poses a major challenge. For instance, while several recent works have successfully designed deep models of physical phenomena, the models do not give any insight into the underlying physical laws. This requirement for interpretability across a variety of domains, has received diverse responses.
In this talk, I will present our analysis which suggests a surprising alignment between the representation in the scientific model and the one learned by the deep model.
Video
Understanding Neutrino Interactions Using Deep Learning
Adam Aurisano
The 2015 Nobel Prize in Physics was awarded for the discovery of neutrino oscillations, which indicates that neutrinos have mass. This phenomenon was unexpected and is one of the clearest signs of new physics beyond the Standard Model. The NOvA experiment aims to deepen our understanding of neutrino oscillations by measuring the properties of a muon neutrino beam produced at Fermi National Accelerator Laboratory at a Near Detector close to the beam source, and measuring the rate that muon neutrinos oscillate into electron neutrinos over an 810 km trip to a 14 kton Far Detector in Ash River, MN. Understanding this process may shed light on the origin of the matter-antimatter asymmetry of the universe. Performing this measurement requires high-precision methods for classifying neutrino interactions and estimating their energy. In this talk, I will discuss how Deep Learning methods used at NOvA for event classification and energy estimation, and I will describe progress transferring these techniques to the future long-baseline DUNE experiment, which will use liquid argon time-projection chamber technology to provide high resolution imaging of neutrino interactions.
Video
Slide
Statistical challenges in cosmological distance measurements
Markus Rau
In the era of precision cosmology, it becomes increasingly important to control sources of systematic error in distance measurements of nearby and faint galaxies. These errors are a dominant systematic for a variety of cosmological probes and can ultimately hinder our ability to accurately study dark energy and the cosmic expansion. Machine Learning and Bayesian modeling are a primary tool to improve the accuracy of these measurements while enabling the consistent parametrization of residual biases. As concrete examples, I will discuss how Machine Learning can be used to obtain accurate measurements of distance using Long Period variable stars to ultimately calibrate local supernovae samples. Connecting with complementary probes based on Weak Gravitational Lensing and Large-Scale-Structure, I will discuss how inaccurate distance, or redshift, measurements for samples of galaxies can bias ongoing and future large area photometric surveys like DES, KiDS, LSST and Euclid. I will present a Bayesian Hierarchical model that self-consistently parametrizes these errors and incorporates them into measurements of cosmological parameters.
Deep learning for generation of events and processes in particle physics, cosmology, and fluid dynamics
Karthik Kashinath
The fundamental sciences (including particle physics, cosmology and turbulence) generate exabytes of data from complex instruments and hi-fidelity simulations and analyze these to uncover the physics of the universe. Advancements in deep generative modeling is renewing interest in using high dimensional density estimators as computationally inexpensive emulators of full-fledged simulations. These generative models have the potential to make a dramatic shift in the field of scientific simulations. In this talk, we’ll demonstrate that deep generative models are capable of emulating simulation data, including generating full particle physics events, with the high statistical fidelity necessary for physics applications. This work explores state-of-the-art methods and exploits high-performance computing.
Video
Slide
Learning quantum states with generative models
Juan Carrasquilla
The technological success of machine learning techniques has motivated a research area in the condensed matter physics and quantum information communities, where new tools and conceptual connections between machine learning and many-body physics are rapidly developing. In this talk, I will discuss the use of generative models for learning quantum states. In particular, I will discuss a strategy for learning mixed states through a combination of informationally complete positive-operator valued measures and generative models. In this setting, generative models enable accurate learning of prototypical quantum states of large size directly from measurements mimicking experimental data.
Video
Slide
Neural-Network and String-Bond States: From Chiral Topological Order to Image Recognition
Ignacio Cirac
We show that there are strong connections between Neural-Network Quantum States in the form of Restricted Boltzmann Machines and some classes of Tensor-Network states in arbitrary dimensions. In particular we demonstrate that short-range Restricted Boltzmann Machines are Entangled Plaquette States, while fully connected Restricted Boltzmann Machines are String-Bond States with a nonlocal geometry and low bond dimension. These results shed light on the underlying architecture of Restricted Boltzmann Machines and their efficiency at representing many-body quantum states. String-Bond States also provide a generic way of enhancing the power of Neural-Network Quantum States and a natural generalization to systems with larger local Hilbert space. We compare the advantages and drawbacks of these different classes of states and present a method to combine them together. This allows us to benefit from both the entanglement structure of Tensor Networks and the efficiency of Neural-Network Quantum States into a single Ansatz capable of targeting the wave function of strongly correlated systems. While it remains a challenge to describe states with chiral topological order using traditional Tensor Networks, we show that Neural-Network Quantum States and their String-Bond States extension can describe a lattice Fractional Quantum Hall state exactly. Our results demonstrate the efficiency of neural networks to describe complex quantum wave functions and pave the way towards the use of String-Bond States as a tool in more traditional machine-learning applications, like image recognition.
Video
Slide
Predicting energies from electron densities: Machine learning for reactive molecular dynamics
Leslie Vogt
Density functional theory (DFT) is widely used to calculate the energy of molecular systems and is typically considered to be a computationally affordable electronic structure method. However, running Kohn-Sham DFT-based molecular dynamics (MD) with the energy recalculated at every step is still a costly approach, particularly for periodic materials. Using machine-learned density functionals that map potentials to energies via electron densities allow us to bypass the Kohn-Sham equations in 1D models and in molecular systems. The resulting models enable running MD simulations based on learned electronic structure energies rather than force fields and facilitate sampling conformational changes and reactive processes such as proton transfers.
Video
Slide
Neural Network Renormalization Group
Lei Wang
I will present a variational renormalization group (RG) approach using a deep generative model based on normalizing flows. The model performs a hierarchical of change-of-variables transformations from the physical space to a latent space with reduced mutual information. Conversely, it directly generates statistically independent physical configurations as a form of inverse RG flow. The generative model has an exact and tractable likelihood, which allows unbiased training and direct access to the renormalized energy function of the latent variables. To train the neural network, we employ the probability density distillation of the bare energy function, where the training loss provides a variational upper bound of the physical free energy. We demonstrate practical usage of the approach by identifying mutually independent collective variables of the Ising model and performing accelerated hybrid Monte Carlo sampling in the latent space. I will comment on the connection of the present approach to DeepMind’s WaveNet, OpenAI’s Glow, wavelet formulation of RG, and the modern pursuit of information preserving RG.
Video
Slide
From physics to machine-learning and back
Edgar A. Engel
The value of machine-learning (ML) approaches in the physical and biological sciences is undeniable — whether it is as surrogate models promising chemically-accurate properties predictions at the atomic scale (while sidestepping much of the computational cost of first-principles methods), or as a means of performing data-driven classifications and extracting important physical insight into the behaviour of complex systems and the structure-property relations of materials.
Against the backdrop of the proliferation of descriptors (and learning strategies), I will outline how exploiting fundamental physics and chemistry in the description of atomistic systems facilitates effective ML for chemistry and materials.
To this purpose I will introduce how a description of atomic structures based on atom densities leads, upon symmetrization, to the smooth overlap of atomic positions (SOAP) representation. This physically motivated descriptor enables regression of scalar, and can be extended to machine-learn also more complex properties such as tensors and the charge density. Furthermore, I will present an example of how machine learning can be beneficial for more complex tasks, namely the identification of stabilisable phases among databases of (computationally) locally-stable atomic structures, by means of a generalized convex hull construction.
Video
Slide
Optical random features for large-scale machine learning
Laurent Daudet
The propagation of coherent light through a thick layer of scattering material is an extremely complex physical process. However, it remains linear, and under certain conditions, if the incoming beam is spatially modulated to encode some data, the output as measured on a sensor can be modeled as a random projection of the input, i.e. its multiplication by an iid random matrix. One can leverage this principle for compressive imaging, and more generally for any data processing pipeline involving large-scale random projections. This talk will discuss recent technological developments of optical co-processors, and present a series of proof of concept experiments in machine learning, such as transfer learning, change point detection, or recommender systems.
Intelligent Controls for Particle Accelerators and Other Research and Industrial Infrastructures
Sandra Biedron
Despite the rush to incorporate artificial intelligence (AI) into every component of our lives, the truth is, not all systems (or sub-systems) require intelligent controls for reliable operation or for their interpretation. As the old saying goes – “Just because you can [build it/use it] doesn’t mean you should.” Some complex systems, however, can greatly benefit from the responsible and well-architected incorporation of control techniques with intelligent characteristics. In some cases that intelligent component encompasses the technique of machine learning.
Here we review several specific case examples of complex systems that can greatly benefit from the intelligent control that were otherwise unresponsive or less responsive than we hoped to simpler, less extravagant approaches. The case examples on which we will focus on particle accelerators that are used in many disciplines including fundamental discovery science and engineering.
We present our efforts and experiences in developing adaptive, artificial intelligence-based tools specifically to address control challenges found in particle accelerator systems over the past decade [See, for example 1-4]. We will discuss how we arrived at and down this AI path, and the opportunities and challenges we have and continue to face, especially for devices that will not necessarily be co-located with accelerator operators or experts. Finally, we will discuss the processes we use to determine when AI is needed (and not just wanted).
[1] E. Meier, S. G. Biedron, G. LeBlanc, M. J. Morgan. “Development of a Novel Optimization Tool for Electron Linacs Inspired by Artificial Intelligence Techniques in Video Games.” NIM A: 632. 1 (2011): 1-6.
[2] S.G. Biedron, A.L. Edelen, and S.V. Milton, “Advanced Controls for RF and Directed Energy Systems,” Eighteenth Annual Directed Energy Symposium, 7-11 March 2016.
[3] A.L. Edelen, S.G. Biedron et al., 2016, “Neural Networks for Modeling and Control of Particle Accelerators,” IEEE Transactions on Nuclear Science 63(2), 878-897.
[4] A. Edelen et al., “Recent Application of Neural Network-Based Approaches to Modeling and Control of Particle Accelerators,” to be published, 2018 International Particle Accelerator Conference, Vancouver B.C.
Machine learning in biology: Teaching an autonomous glider to soar like a bird
Gautam Reddy
Soaring birds often rely on ascending thermal plumes (thermals) in the atmosphere as they search for prey or migrate across large distances. How soaring birds find and navigate thermals within this complex landscape is unknown. We used reinforcement learning to train gliders in the field to navigate atmospheric thermals autonomously. Gliders of two-metre wingspan were equipped with a flight controller that precisely controlled the bank angle and pitch, modulating these at intervals with the aim of gaining as much lift as possible. A navigational strategy was determined solely from the gliders’ pooled experiences collected over several days in the field. The strategy relies on on-board methods to accurately estimate the local vertical wind accelerations and the roll-wise torques on the glider, which serve as navigational cues. We highlight the role of vertical wind accelerations and roll-wise torques as effective mechanosensory cues for soaring birds and provide a navigational strategy that is directly applicable to the development of autonomous soaring vehicles.
Video
Slide
Making quantum algorithms learn from data
Maria Schuld
An important question in the young discipline of quantum machine learning is what impact quantum computing technology will have on the field of machine learning. Can we accelerate known algorithms, or even contribute entirely new methods of data-driven decision making? How can in particular near-term devices - which run short and noisy quantum algorithms - be used to find patterns in data? One line of research, so called variational circuits, consider parameter-dependent quantum algorithms. These algorithms are interpreted as machine learning models that can be trained for a given task, for example to classify unseen data samples or to generate artificial data. In this talk I will give an overview of what we know and – more importantly – do not know about learning with variational circuits. I will cover ideas of how to train quantum algorithms, how to elegantly implement a neural network in a photonic quantum computer, and how we may potentially surpass purely classical models with this approach.
Video
Slide
From Reinforcement Learning to Spin Glasses: The Many Surprises on Quantum State Preparation
Pankaj Mehta
Video
Slide
Using Machine Learning for Analysis and Prediction of High-Dimensional Spatiotemporal Chaotic Dynamical Systems
Jaideep Pathak
We demonstrate the effectiveness of machine learning for analysis and prediction of spatiotemporal chaotic dynamical systems from data. Using a computationally efficient recurrent neural network called an echo state network or reservoir computer [1] we show that we can reconstruct the attractor of very high dimensional chaotic dynamical systems with unprecedented fidelity. This reconstruction allows us to determine the ergodic properties (e.g., the spectrum of Lyapunov exponents) of a dynamical system purely from data [2]. We also develop and introduce a computationally parallelized extension of a prediction scheme based on reservoir computing that allows us to obtain model-free predictions of spatiotemporal chaotic flows of arbitrarily large spatial extent and attractor dimension [3]. We obtain outstanding results using machine learning for these difficult tasks where traditional methods have had limited success. We demonstrate the scalability and computational efficiency of our approach using a popular toy model [4] often used for testing techniques in weather prediction (Lorenz 1996) and the spatiotemporally chaotic Kuramoto-Sivashinsky partial differential equation.
References
[1] Herbert Jaeger and Harald Haas, Harnessing nonlinearity: Predicting chaotic systems and saving energy in wireless communication. Science 304(5667):78–80, 2004.
[2] Jaideep Pathak, Zhixin Lu, Brian R Hunt, Michelle Girvan, and Edward Ott, Using machine learning to replicate chaotic attractors and calculate lyapunov exponents from data. Chaos: An Interdisciplinary Journal of Nonlinear Science 27(12):121102, 2017.
[3] Jaideep Pathak, Brian Hunt, Michelle Girvan, Zhixin Lu, and Edward Ott, Model-free prediction of large spatiotemporally chaotic systems from data: a reservoir computing approach. Phys. Rev. Lett. 120, 024102 (2018).
[4] Edward N Lorenz, Predictability: A problem partly solved. In Proc. Seminar on predictability, volume 1, 1996.
Video
Slide
Advances in machine learned potentials for molecular dynamics simulation
Kipton Barros
Recent machine learning techniques allow emulation of quantum physics with stunning fidelity. For example, deep neural networks can now predict molecular properties with accuracy comparable to density functional theory, and approaching that of coupled cluster theory, at a tiny fraction of the computational cost. We present methods for building machine learned potentials that will enable large-scale and highly accurate molecular dynamics simulations, e.g., for chemistry, materials science, and biophysics applications. Key ideas are:
1. Encoding known physical properties and symmetries in the neural network architecture.
2. Active learning to dynamically augment the training dataset with new quantum calculations in regions where the machine learned model is uncertain.
3. Transfer learning, such that we first train on a large quantity of relatively low-fidelity data, and then perform some final training iterations using a smaller quantity of high-fidelity data.
Video
Slide