Reuse & Permissions

It is not necessary to obtain permission to reuse this article or its components as it is available under the terms of the Creative Commons Attribution 4.0 International license. This license permits unrestricted use, distribution, and reproduction in any medium, provided attribution to the author(s) and the published article's title, journal citation, and DOI are maintained. Please note that some figures may have been included with permission from other third parties. It is your responsibility to obtain the proper permission from the rights holder directly for these figures.

Export citation

Export citation

Choose format for download:

Download Citation
  • Open Access

Evaluating IBM’s Watson natural language processing artificial intelligence as a short-answer categorization tool for physics education research

Jennifer Campbell, Katie Ansell, and Tim Stelzer

  • Department of Physics, University of Illinois Urbana-Champaign, 1110 West Green Street, Urbana, Illinois 61801, USA

Phys. Rev. Phys. Educ. Res. 20, 010116 – Published 22 March, 2024

DOI: https://doi.org/10.1103/PhysRevPhysEducRes.20.010116

Abstract

Recent advances in publicly available natural language processors (NLP) may enhance the efficiency of analyzing student short-answer responses in physics education research (PER). We train a state-of-the-art NLP, IBM’s Watson, and test its agreement with human coders using two different studies that gathered text responses in which students explain their reasoning on physics-related questions. The first study analyzes 479 student responses to a lab data analysis question and categorizes them by main idea. The second study analyzes 732 student answers to identify the presence or absence of each of the two conceptual themes. When training Watson with approximately one-third to half of the samples, we find that samples labeled with high confidence scores have similar accuracy to human agreement; yet for lower confidence scores, humans outperform the NLP’s labeling accuracy. In addition to studying Watson’s overall accuracy, we use this analysis to better understand factors that impact how Watson categorizes. Using the data from the categorization study, we find that Watson’s algorithm does not appear to be impacted by the disproportionate representation of categories in the training set, and we examine mislabeled statements to identify vocabulary and phrasing that may increase the rate of false positives. Based on this work, we find that, with careful consideration of the research study design and an awareness of the NLP’s limitations, Watson may present a useful tool for large-scale PER studies or classroom analysis tools.

View figure in article

Physics Subject Headings (PhySH)

Article Text

References (39)

  1. P. V. Engelhardt, An introduction to classical test theory as applied to conceptual multiple-choice tests, in Getting Started in PER, edited by C. Henderson and K. Harper (American Association of Physics Teachers, College Park, 2009), Vol. 2, https://www.compadre.org/Repository/document/ServeFile.cfm?ID=8807&DocID=1148.
  2. P. Ariwala, How is natural language processing applied in business?, https://marutitech.com/how-is-natural-language-processing-applied-in-business/ (2022).
  3. List of IBM Watson natural language understanding customers, Apps Run the World, Technical Report.
  4. Companies using IBM Watson natural language understanding, Technical Report.
  5. L. Cvetkovic, B. Milasinovic, and K. Fertalj, A tool for simplifying automatic categorization of scientific paper using Watson API, in Proceedings of the 40th International Convention on Information and Communication Technology, Electronics and Microelectronics (MIPRO) (IEEE, Opatija, Croatia, 2017), pp. 1501–1505.
  6. E. Yun, Review of trends in physics education research using topic modeling, J. Balt. Sci. Educ. 19, 388 (2020), https://eric.ed.gov/?id=EJ1264735.
  7. J. Campbell, K. Ansell, and T. Stelzer, Using IBM’s Watson to automatically evaluate student short answer responses, presented at PER Conf. 2022, Grand Rapids, MI, 10.1119/perc.2022.pr.Campbell.
  8. N. Chomsky, Three models for the description of language, IEEE Trans. Inform. Theory 2, 113 (1956).
  9. Conceptual Information Processing, edited by R. C. Schank, Fundamental Studies in Computer Science Vol. 3 (North-Holland, Amsterdam, 1975).
  10. Handbook of Mathematical Logic, edited by J. Barwise, Studies in Logic and the Foundations of Mathematics Vol. 90 (North-Holland, Amsterdam, 1977).
  11. J. Pearl, Bayesian networks: A model of self-activated memory for evidential reasoning, in Proceedings of the 7th conference of the Cognitive Science Society, University of California, Irvine, CA, USA (1985), http://ftp.cs.ucla.edu/pub/stat_ser/r43-1985.pdf.
  12. R. Kneser and H. Ney, Improved backing-off for M-gram language modeling, in Proceedings of the 1995 International Conference on Acoustics, Speech, and Signal Processing (IEEE, Detroit, MI, 1995), Vol. 1, pp. 181–184.
  13. J. L. Elman, Finding structure in time, Cogn. Sci. 14, 179 (1990).
  14. Y. Kim, Convolutional neural networks for sentence classification, in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) (Association for Computational Linguistics, Doha, Qatar, 2014), pp. 1746–1751.
  15. R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts, Recursive deep models for semantic compositionality over a sentiment treebank, in Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, edited by D. Yarowsky, T. Baldwin, A. Korhonen, K. Livescu, and S. Bethard (Association for Computational Linguistics, Seattle, WA, 2013), pp. 1631–1642.
  16. S. Burrows, I. Gurevych, and B. Stein, The eras and trends of automatic short answer grading, Int. J. Artif. Intell. Educ. 25, 60 (2015).
  17. G. Casalino, B. Cafarelli, L. Grilli, A. Guarino, P. Limone, D. Schicchi, and D. Taibi, Framing automatic grading techniques for open-ended questionnaires responses. A short survey, in Proceedings of the Second Workshop on Technology Enhanced Learning Environments for Blended Education (teleXbe2021) (Foggia, Italy, 2021), https://ceur-ws.org/.
  18. X. Zhai, L. Shi, and R. H. Nehm, A meta-analysis of machine learning-based science assessments: Factors impacting machine-human score agreements, J. Sci. Educ. Technol. 30, 361 (2021).
  19. X. Zhai, Y. Yin, J. W. Pellegrino, K. C. Haudek, and L. Shi, Applying machine learning in science assessment: A systematic review, Stud. Sci. Educ. 56, 111 (2020).
  20. A. Çinar, E. Ince, M. Gezer, and O. Yilmaz, Machine learning algorithm for grading open-ended physics questions in Turkish, Educ. Inf. Technol. 25, 3821 (2020).
  21. M. Mohler and R. Mihalcea, Text-to-text semantic similarity for automatic short answer grading, in Proceedings of the 12th Conference of the European Chapter of the Association for Computational Linguistics on EACL ’09 (Association for Computational Linguistics, Athens, Greece, 2009), pp. 567–575.
  22. M. A. Sultan, C. Salazar, and T. Sumner, Fast and easy short answer grading with high accuracy, in Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Association for Computational Linguistics, San Diego, CA, 2016), pp. 1070–1075.
  23. F. Zehner, C. Sälzer, and F. Goldhammer, Automatic coding of short text responses via clustering in educational assessment, Educ. Psychol. Meas. 76, 280 (2016).
  24. R. H. Nehm, M. Ha, and E. Mayfield, Transforming biology assessment with machine learning: Automated scoring of written evolutionary explanations, J. Sci. Educ. Technol. 21, 183 (2012).
  25. N. Madnani, A. Loukina, and A. Cahill, A large scale quantitative exploration of modeling strategies for content scoring, in Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications (Association for Computational Linguistics, Copenhagen, Denmark, 2017), pp. 457–467.
  26. J. A. Erickson, A. F. Botelho, S. McAteer, A. Varatharaj, and N. T. Heffernan, The automated grading of student open responses in mathematics, in Proceedings of the 10th International Conference on Learning Analytics and Knowledge (ACM, Frankfurt Germany, 2020), pp. 615–624.
  27. N. Süzen, A. N. Gorban, J. Levesley, and E. M. Mirkes, Automatic short answer grading and feedback using text mining methods, Proc. Comput. Sci. 169, 726 (2020).
  28. M. Ha, R. H. Nehm, M. Urban-Lurain, and J. E. Merrill, Applying computerized-scoring models of written biological explanations across courses and colleges: Prospects and limitations, CBE Life Sci. Educ. 10, 379 (2011).
  29. O. L. Liu, J. A. Rios, M. Heilman, L. Gerard, and M. C. Linn, Validation of automated scoring of science assessments, J. Res. Sci. Teach. 53, 215 (2016).
  30. D. M. Williamson, X. Xi, and F. J. Breyer, A framework for evaluation and use of automated scoring, Educ. Meas. 31, 2 (2012).
  31. P. G. Butcher and S. E. Jordan, A comparison of human and computer marking of short free-text student responses, Comput. Educ. 55, 489 (2010).
  32. Creating custom classification models (2022).
  33. J. Wilson, B. Pollard, J. M. Aiken, M. D. Caballero, and H. Lewandowski, Classification of open-ended responses to a research-based assessment using natural language processing, Phys. Rev. Phys. Educ. Res. 18, 010141 (2022).
  34. S. G. Pulman and J. Z. Sukkarieh, Automatic short answer marking, in Proceedings of the Second Workshop on Building Educational Applications Using NLP (Association for Computational Linguistics, Ann Arbor, Michigan, 2005), pp. 9–16.
  35. K. Crawford, Atlas of AI: Power, Politics, and the Planetary Costs of Artificial Intelligence, 1st ed. (Yale, CT; London, 2022).
  36. K. Ansell, Cultivating adaptive expertise in the introductory physics laboratory, Ph.D. thesis, University of Illinois Urbana-Champaign, 2020.
  37. K. Krippendorff, Content analysis: An introduction to its methodology, 3rd ed. (Sage, Los Angeles, CA; London, 2013).
  38. S. F. Itza-Ortiz, N. S. Rebello, D. A. Zollman, and M. Rodriguez-Achach, The vocabulary of introductory physics and Its implications for learning physics, Phys. Teach. 41, 330 (2003).
  39. https://github.com/mysteriousmartel/per_watson_coder.

Outline

Information

Sign In to Your Journals Account

Filter

Filter

Article Lookup

Enter a citation