- Open Access
Exploring generative AI assisted feedback writing for students’ written responses to a physics conceptual question with prompt engineering and few-shot learning
Phys. Rev. Phys. Educ. Res. 20, 010152 – Published 13 June, 2024
DOI: https://doi.org/10.1103/PhysRevPhysEducRes.20.010152
Abstract
Instructor’s feedback plays a critical role in students’ development of conceptual understanding and reasoning skills. However, grading student written responses and providing personalized feedback can take a substantial amount of time, especially in large enrollment courses. In this study, we explore using GPT-3.5 to write feedback on students’ written responses to conceptual questions with prompt engineering and few-shot learning techniques. In stage I, we used a small portion () of the student responses on one conceptual question to iteratively train GPT to generate feedback. Four of the responses paired with human-written feedback were included in the prompt as examples for GPT. We tasked GPT to generate feedback for another 16 responses and refined the prompt through several iterations. In stage II, we gave four student researchers (one graduate and three undergraduate researchers) the 16 responses as well as two versions of feedback, one written by the authors and the other by GPT. Students were asked to rate the correctness and usefulness of each feedback and to indicate which one was generated by GPT. The results showed that students tended to rate the feedback by human and GPT equally on correctness, but they all rated the feedback by GPT as more useful. Additionally, the success rates of identifying GPT’s feedback were low, ranging from 0.1 to 0.6. In stage III, we tasked GPT to generate feedback for the rest of the students’ responses (). The feedback messages were rated by four instructors based on the extent of modification needed if they were to give the feedback to students. All four instructors rated approximately 70% (ranging from 68% to 78%) of the feedback statements needing only minor or no modification. This study demonstrated the feasibility of using generative artificial intelligence (AI) as an assistant to generate feedback for student written responses with only a relatively small number of examples in the prompt. An AI assistant can be one of the solutions to substantially reduce time spent on grading student written responses.
Physics Subject Headings (PhySH)
Article Text
References (40)
- R. Beichner, An introduction to physics education research, in Getting Started in PER, edited by C. Henderson and K. Harper (2009), Vol. 2.
- L. S. Shulman, Those who understand: Knowledge growth in teaching, Educ. Res. 15, 4 (1986).
- K. A. Ericsson, R. T. Krampe, and C. Tesch-Romer, The role of deliberate practice in the acquisition of expert performance, Psychol. Rev. 100, 363 (1993).
- Louis Deslaurier, Ellen Schelew, and Carl Wieman, Improved learning in a large-enrollment class, Science 332, 862 (2011).
- P. Black and D. Wiliam, Assessment and classroom learning, Int. J. Phytorem. 5, 7 (1998).
- J. Larreamendy-Joerns, G. Leinhardt, and J. Corredor, Six online statistics courses: Examination and review, Am. Stat. 59, 240 (2005).
- D. Baidoo-Anu and L. Owusu Ansah, Education in the era of generative artificial intelligence (AI): Understanding the potential benefits of ChatGPT in promoting teaching and learning, J. AI 7, 52 (2023).
- E. Kasneci et al., ChatGPT for good? On opportunities and challenges of large language models for education, Learn. Individ. Differ. 103, 102274 (2023).
- Z. Li, C. Zhang, Y. Jin, X. Cang, S. Puntambekar, and R. J. Passonneau, Learning when to defer to humans for short answer grading, in Artificial Intelligence in Education, edited by N. Wang, G. Rebolledo-Mendez, N. Matsuda, O. C. Santos, and V. Dimitrova, Lecture Notes in Computer Science() Vol. 13916 (Springer, Cham, 2023), 10.1007/978-3-031-36272-9_34.
- C. Sung, T. Ma, T. I. Dhamecha, V. Reddy, S. Saha, and R. Arora, Pre-training BERT on domain resources for short answer grading, in Proceedings of the EMNLP-IJCNLP 2019–2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing, Hong Kong, China (Association for Computational Linguistics, Hong Kong, China, 2019), p. 6071.
- A. Ahmed, A. Joorabchi, and M. Hayes, On deep learning approaches to automated assessment: Strategies for short answer grading, in Proceedings of the 14th International Conference on Computer Supported Education (SCITEPRESS—Science and Technology Publications, 2022), pp. 85–94.
- A. Condor, Exploring automatic short answer grading as a tool to assist in human rating, in Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2020), Vol. 12164.
- P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing, ACM Comput. Surv. 55, 9 (2023).
- S. Steinert, K. E. Avila, S. Ruzika, J. Kuhn, and S. Küchemann, Harnessing large language models to enhance self-regulated learning via formative feedback, arXiv:2311.13984v2.
- G. Kortemeyer, Could an artificial-intelligence agent pass an introductory physics course, Phys. Rev. Phys. Educ. Res. 19, 010132 (2023).
- M. N. Dahlkemper, S. Z. Lahme, and P. Klein, How do physics students evaluate artificial intelligence responses on comprehension questions: A study on the perceived scientific accuracy and linguistic quality of ChatGPT, Phys. Rev. Phys. Educ. Res. 19, 010142 (2023).
- S. Küchemann, S. Steinert, N. Revenga, M. Schweinberger, Y. Dinc, K. E. Avila, and J. Kuhn, Can ChatGPT support prospective teachers in physics task development?, Phys. Rev. Phys. Educ. Res. 19, 020128 (2023).
- G. Polverini and B. Gregorcic, How understanding large language models can inform the use of ChatGPT in physics education, Eur. J. Phys. 45, 025701 (2023).
- F. Shi, X. Chen, K. Misra, N. Scales, D. Dohan, E. Chi, N. Schärli, and D. Zhou, Large language models can be easily distracted by irrelevant context, arXiv:2302.00093v3.
- M. Lee, A mathematical investigation of hallucination and creativity in GPT models, Mathematics 11, 2320 (2023).
- M. Zhang, O. Press, W. Merrill, A. Liu, N. A. Smith, and P. G. Allen, How language model hallucinations can snowball, arXiv:2305.13534.
- Y. Bang et al., A multitask, multilingual, multimodal evaluation of ChatGPT on reasoning, hallucination, and interactivity, arXiv:2302.04023v4.
- D. J. Woo, K. Guo, and H. Susanto, Cases of EFL secondary students’ prompt engineering pathways to complete a writing task with ChatGPT, arXiv:2307.05493.
- T. F. Heston and C. Khun, Prompt engineering in medical education, Int. Med. Educ. 2, 198 (2023).
- Y. Wang, Q. Yao, J. T. Kwok, and L. M. Ni, Generalizing from a few examples: A survey on few-shot learning, ACM Comput. Surv. 53, 1 (2020).
- M. Zong and B. Krishnamachari, Solving math word problems concerning systems of equations with GPT-3, in Proceedings of the 37th AAAI Conference on Artificial Intelligence (AAAI-23) (2023).
- J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou, Chain-of-thought prompting elicits reasoning in large language models, arXiv:2201.11903v6.
- E. Radiya-Dixit and X. Wang, How fine can fine-tuning be? Learning efficient language models, in Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, edited by S. Chiappa and R. Calandra (PMLR, 2020), pp. 2435–2443.
- E. Latif and X. Zhai, Fine-tuning ChatGPT for automatic scoring, Comput. Educ. 6, 100210 (2024).
- A. Elby, R. E. Scherr, T. McCaskey, R. Hodges, E. F. Redish, D. Hammer, and T. Bing, Open Source Tutorials in Physics Sensemaking: Suite I (2007), https://www.physport.org/curricula/MD_OST/.
- D. Hammer, Student resources for learning introductory physics, Am. J. Phys. 68, S52 (2000).
- S. Wheeler and R. E. Scherr, ChatGPT reflects student misconceptions in physics, presented at PER Conf. 2023, 10.1119/perc.2023.pr.Wheeler.
- R. E. Scherr and E. F. Redish, Newton’s zeroth law: Learning from listening to our students, Phys. Teach. 43, 41 (2005).
In one of the student responses with the correct conclusion, there was a minor mistake in the explanation, which we did not realize when we wrote the feedback. Both GPT-generated and human-written feedback stated that the response was correct on both the conclusion and explanation.
- C. Jones and B. Bergen, Does GPT-4 pass the turing test?, arXiv:2310.20216.
- M. Mitchell and D. C. Krakauer, The debate over understanding in AI’s large language models, Proc. Natl. Acad. Sci. U.S.A. 120, e2215907120 (2023).
- P. Tschisgale, P. Wulff, and M. Kubsch, Integrating artificial intelligence-based methods into qualitative research in physics education research: A case for computational grounded theory, Phys. Rev. Phys. Educ. Res. 19, 020123 (2023).
- D. Rozado, The political biases of ChatGPT, Soc. Sci. 12, 148 (2023).
- L. Lucy and D. Bamman, Gender and representation bias in GPT-3 generated stories, in Proceedings of the Third Workshop on Narrative Understanding, Virtual (Association for Computational Linguistics, 2021), p. 48.
- E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, On the dangers of stochastic parrots: Can language models be too big?, in Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT 2021, Virtual Event Canada (Association for Computing Machinery, New York, NY, 2021), p. 610.