- Open Access
- Access by Xinjiang University
Learning new physics from a machine
Phys. Rev. D 99, 015014 – Published 8 January, 2019
DOI: https://doi.org/10.1103/PhysRevD.99.015014
Abstract
We propose using neural networks to detect data departures from a given reference model, with no prior bias on the nature of the new physics responsible for the discrepancy. The virtues of neural networks as unbiased function approximants make them particularly suited for this task. An algorithm that implements this idea is constructed, as a straightforward application of the likelihood-ratio hypothesis test. The algorithm compares observations with an auxiliary set of reference-distributed events, possibly obtained with a Monte Carlo event generator. It returns a value, which measures the compatibility of the reference model with the data. It also identifies the most discrepant phase-space region of the data set, to be selected for further investigation. The most interesting potential applications are model-independent new physics searches, although our approach could also be used to compare the theoretical predictions of different Monte Carlo event generators, or for data validation algorithms. In this work we study the performance of our algorithm on a few simple examples. The results confirm the model independence of the approach, namely that it displays good sensitivity to a variety of putative signals. Furthermore, we show that the reach does not depend much on whether a favorable signal region is selected based on prior expectations. We identify directions for improvement towards applications to real experimental data sets.
Physics Subject Headings (PhySH)
Article Text
References (82)
- Particle Data Group, Review of particle physics, Chin. Phys. C 40, 100001 (2016).
- G. Choudalakis, On hypothesis testing, trials factor, hypertests and the BumpHunter, in Proceedings of the PHYSTAT 2011 Workshop on Statistical Issues Related to Discovery Claims in Search Experiments and Unfolding, 2011 (CERN, Geneva, Switzerland, 2011).
- B. Abbott et al. (D0 Collaboration), Search for new physics in data at D0 using Sherlock: A quasi model independent search strategy for new physics, Phys. Rev. D 62, 092004 (2000).
- V. M. Abazov et al. (D0 Collaboration), A quasi model independent search for new physics at large transverse momentum, Phys. Rev. D 64, 012004 (2001).
- A. Aktas et al. (H1 Collaboration), A general search for new phenomena in ep scattering at HERA, Phys. Lett. B 602, 14 (2004).
- F. D. Aaron et al. (H1 Collaboration), A general search for new phenomena at HERA, Phys. Lett. B 674, 257 (2009).
- P. Asadi, M. R. Buckley, A. DiFranzo, A. Monteux, and D. Shih, Digging deeper for new physics in the LHC data, J. High Energy Phys. 11 (2017) 194.
- T. Aaltonen et al. (CDF Collaboration), Model-independent and quasi-model-independent search for new physics at CDF, Phys. Rev. D 78, 012002 (2008).
- T. Aaltonen et al. (CDF Collaboration), Global search for new physics with at CDF, Phys. Rev. D 79, 011101 (2009).
- CMS Collaboration, MUSIC—An automated scan for deviations between data and Monte Carlo simulation, CERN, Report No. CMS-PAS-EXO-08-005.
- CMS Collaboration, Model unspecific search for new physics in pp collisions at , Report No. CMS-PAS-EXO-10-021.
- ATLAS Collaboration, A model independent general search for new phenomena with the ATLAS detector at , Report No. ATLAS-CONF-2017-001.
- ATLAS Collaboration, A general search for new phenomena with the ATLAS detector in pp collisions at , Report No. ATLAS-CONF-2012-107.
- ATLAS Collaboration, A general search for new phenomena with the ATLAS detector in pp collisions at , Report No. ATLAS-CONF-2014-006.
- L. de Oliveira, M. Kagan, L. Mackey, B. Nachman, and A. Schwartzman, Jet-images—Deep learning edition, J. High Energy Phys. 07 (2016) 069.
- A. Schwartzman, M. Kagan, L. Mackey, B. Nachman, and L. De Oliveira, Image processing, computer vision, and deep learning: New approaches to the analysis and physics interpretation of LHC events, J. Phys. Conf. Ser. 762, 012035 (2016).
- M. Kagan, L. d. Oliveira, L. Mackey, B. Nachman, and A. Schwartzman, Boosted jet tagging with jet-images and deep neural networks, EPJ Web Conf. 127, 00009 (2016).
- A. J. Larkoski, I. Moult, and B. Nachman, Jet substructure at the large hadron collider: A review of recent advances in theory and machine learning, arXiv:1709.04464.
- G. Louppe, K. Cho, C. Becot, and K. Cranmer, QCD-aware recursive neural networks for jet physics, arXiv:1702.00748.
- C. Shimmin, P. Sadowski, P. Baldi, E. Weik, D. Whiteson, E. Goul, and A. SÃÿgaard, Decorrelated jet substructure tagging using adversarial neural networks, Phys. Rev. D 96, 074034 (2017).
- P. Baldi, K. Bauer, C. Eng, P. Sadowski, and D. Whiteson, Jet substructure classification in high-energy physics with deep neural networks, Phys. Rev. D 93, 094034 (2016).
- D. Guest, J. Collado, P. Baldi, S.-C. Hsu, G. Urban, and D. Whiteson, Jet flavor classification in high-energy physics with deep neural networks, Phys. Rev. D 94, 112002 (2016).
- L. G. Almeida, M. BackoviÄĞ, M. Cliche, S. J. Lee, and M. Perelstein, Playing tag with ANN: Boosted top identification with pattern recognition, J. High Energy Phys. 07 (2015) 086.
- J. Barnard, E. N. Dawe, M. J. Dolan, and N. Rajcic, Parton shower uncertainties in jet substructure analyses with deep neural networks, Phys. Rev. D 95, 014018 (2017).
- G. Kasieczka, T. Plehn, M. Russell, and T. Schell, Deep-learning top taggers or the end of QCD?, J. High Energy Phys. 05 (2017) 006.
- A. Butter, G. Kasieczka, T. Plehn, and M. Russell, Deep-learned top tagging with a Lorentz layer, SciPost Phys. 5, 028 (2018).
- K. Datta and A. Larkoski, How much information is in a jet?, J. High Energy Phys. 06 (2017) 073.
- K. Datta and A. J. Larkoski, Novel jet observables from machine learning, J. High Energy Phys. 03 (2018) 086.
- K. Fraser and M. D. Schwartz, Jet charge and machine learning, J. High Energy Phys. 10 (2018) 093.
- A. Andreassen, I. Feige, C. Frye, and M. D. Schwartz, JUNIPR: A framework for unsupervised machine learning in particle physics, arXiv:1804.09720.
- S. Macaluso and D. Shih, Pulling out all the tops with computer vision and deep learning, J. High Energy Phys. 10 (2018) 121.
- ATLAS Collaboration, Performance of top quark and W boson tagging in run 2 with ATLAS, Report No. ATLAS-CONF-2017-064.
- CMS Collaboration, CMS phase 1 heavy flavour identification performance and developments, Report No. CMS-DP-2017-013, https://cds.cern.ch/record/2263802.
- ATLAS Collaboration, Optimisation and performance studies of the ATLAS -tagging algorithms for the 2017-18 LHC run, Technical Report No. ATL-PHYS-PUB-2017-013, CERN, Geneva, 2017.
- ATLAS Collaboration, Identification of hadronically-decaying W bosons and top quarks using high-level features as input to boosted decision trees and deep neural networks in ATLAS at , Technical Report No. ATL-PHYS-PUB-2017-004, CERN, Geneva, 2017.
- CMS Collaboration, Heavy flavor identification at CMS with deep neural networks, https://cds.cern.ch/record/2255736.
- ATLAS Collaboration, Identification of jets containing -hadrons with recurrent neural networks at the ATLAS experiment, Technical Report No. ATL-PHYS-PUB-2017-003, CERN, Geneva, 2017.
- ATLAS Collaboration, Quark versus gluon jet tagging using jet images with the ATLAS detector, Technical Report No. ATL-PHYS-PUB-2017-017, CERN, Geneva, 2017.
- CMS Collaboration, New developments for jet substructure reconstruction in CMS, https://cds.cern.ch/record/2275226.
- P. Baldi, K. Cranmer, T. Faucett, P. Sadowski, and D. Whiteson, Parameterized neural networks for high-energy physics, Eur. Phys. J. C 76, 235 (2016).
- S. Chang, T. Cohen, and B. Ostdiek, What is the machine learning?, Phys. Rev. D 97, 056009 (2018).
- T. Cohen, M. Freytsis, and B. Ostdiek, (Machine) learning to do more with less, J. High Energy Phys. 02 (2018) 034.
- J. Brehmer, K. Cranmer, G. Louppe, and J. Pavez, A guide to constraining effective field theories with machine learning, Phys. Rev. D 98, 052004 (2018).
- J. Brehmer, K. Cranmer, G. Louppe, and J. Pavez, Constraining Effective Field Theories with Machine Learning, Phys. Rev. Lett. 121, 111801 (2018).
- J. Brehmer, G. Louppe, J. Pavez, and K. Cranmer, Mining gold from implicit models to improve likelihood-free inference, arXiv:1805.12244.
- T. Roxlo and M. Reece, Opening the black box of neural nets: Case studies in stop/top discrimination, arXiv:1804.09278.
- J. H. Collins, K. Howe, and B. Nachman, CWoLa hunting: Extending the bump hunt with machine learning, arXiv:1805.02664.
- M. Paganini, L. de Oliveira, and B. Nachman, CaloGAN: Simulating 3D high energy particle showers in multilayer electromagnetic calorimeters with generative adversarial networks, Phys. Rev. D 97, 014021 (2018).
- L. de Oliveira, M. Paganini, and B. Nachman, Controlling physical attributes in GAN-accelerated simulation of electromagnetic calorimeters, J. Phys. Conf. Ser. 1085, 042017 (2018).
- M. Paganini, L. de Oliveira, and B. Nachman, Accelerating Science with Generative Adversarial Networks: An Application to 3D Particle Showers in Multilayer Calorimeters, Phys. Rev. Lett. 120, 042003 (2018).
- R. D. Ball et al. (NNPDF Collaboration), Parton distributions for the LHC run II, J. High Energy Phys. 04 (2015) 040.
- S. Forte, L. Garrido, J. I. Latorre, and A. Piccione, Neural network parametrization of deep inelastic structure functions, J. High Energy Phys. 05 (2002) 062.
- G. Cybenko, Approximation by superpositions of a sigmoidal function, Math. Control Signals Syst. 2, 303 (1989).
- V. Y. Kreinovich, Arbitrary nonlinearity is sufficient to represent all functions by neural networks: A theorem, Neural Netw. 4, 381 (1991).
- R. Hecht-Nielsen, Neural networks for perception, Vol. 2, http://dl.acm.org/citation.cfm?id=140639.140643.
- K. Hornik, M. Stinchcombe, and H. White, Multilayer feedforward networks are universal approximators, Neural Netw. 2, 359 (1989).
- S. Liang and R. Srikant, Why deep neural networks for function approximation?, arXiv:1610.04161.
- T. A. Poggio, H. Mhaskar, L. Rosasco, B. Miranda, and Q. Liao, Why and when can deep—but not shallow—networks avoid the curse of dimensionality: A review, arXiv:1611.00740.
- F. R. Bach, Breaking the curse of dimensionality with convex neural networks, arXiv:1412.8690.
- H. Montanelli and Q. Du, Deep ReLU networks lessen the curse of dimensionality, arXiv:1712.08688.
- C. M. Bishop, Pattern Recognition and Machine Learning (Springer, New York, 2006).
- I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning (MIT, Cambridge, MA, 2016).
- S. Haykin, Neural Networks: A Comprehensive Foundation (Prentice-Hall, Englewood Cliffs, NJ, 1998), Vol. 2.
- M. Kuusela, T. Vatanen, E. Malmi, T. Raiko, T. Aaltonen, and Y. Nagai, Semi-supervised anomaly detection—Towards model-independent searches of new physics, J. Phys. Conf. Ser. 368, 012032 (2012).
- J. Neyman and E. S. Pearson, On the problem of the most efficient tests of statistical hypotheses, Phil. Trans. R. Soc. A 231, 289 (1933).
- G. Cowan, Statistical Data Analysis (Clarendon, Oxford, 1998).
- G. Cowan, Lecture at the 2017 GGI school on the theory of fundamental interactions, https://www.youtube.com/watch?v=Y23Kxg61scc&list=PLDxsZU4NC6Z5DFFFx2bj03phoZpP1qrp2&index=2.
- http://neuralnetworksanddeeplearning.com/chap4.html.
- W. A. Rolke, A. M. Lopez, and J. Conrad, Limits and confidence intervals in the presence of nuisance parameters, Nucl. Instrum. Methods Phys. Res., Sect. A 551, 493 (2005).
- G. Cowan, K. Cranmer, E. Gross, and O. Vitells, Asymptotic formulae for likelihood-based tests of new physics, Eur. Phys. J. C 71, 1 (2011); Erratum, 73, 2501(E) (2013).
- http://www.cs.toronto.edu/~tijmen/csc321/slides/lecture_slides_lec6.pdf.
- S. S. Wilks, The large-sample distribution of the likelihood ratio for testing composite hypotheses, Ann. Math. Stat. 9, 60 (1938).
- A. Wald, Tests of statistical hypotheses concerning several parameters when the number of observations is large, Trans. Am. Math. Soc. 54, 426 (1943).
- S. Han, H. Mao, and W. J. Dally, Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding, arXiv:1510.00149.
- K. Cranmer, J. Pavez, and G. Louppe, Approximating likelihood ratios with calibrated discriminative classifiers, arXiv:1506.02169.
- https://www.wolfram.com/mathematica/.
- M. Z. Alom, T. M. Taha, C. Yakopcic, S. Westberg, M. Hasan, B. C. Van Esesn, A. A. S. Awwal, and V. K. Asari, The history began from AlexNet: A comprehensive survey on deep learning approaches, arXiv:1803.01164.
- T. Hastie, R. Tibshirani, and J. Friedman, The elements of statistical learning, arXiv:1011.1669v3.
- M. Augustine Cauchy, Méthode générale pour la résolution des systèmes d’équations simultanées, Comptes Rendus Hebd. Séances Acad. Sci. 25, 536 (1847).
- G. Cybenko, Approximation by superpositions of a sigmoidal function, Math. Control Signals Syst. 2, 303 (1989).
- K. Hornik, Approximation capabilities of multilayer feedforward networks, Neural Netw. 4, 251 (1991).
- V. Y. Kreinovich, Arbitrary nonlinearity is sufficient to represent all functions by neural networks: A theorem, Neural Netw. 4, 381 (1991).