Modeling Educational Test Results Based on MIRT Models
DOI:
https://doi.org/10.5281/zenodo.14576594Keywords:
MIRT, compensatory MIRT models, primary score matrices, testing, modeling, Hull criterion, EKC criterionAbstract
The article examines the challenges of modeling primary score matrices from test results using compensatory two-dimensional and three-dimensional 2-PL and 3-PL models of Multidimensional Item Response Theory (MIRT). These models are most commonly applied to the results of educational or psychological testing, as testing is typically aimed at identifying multiple competencies of examinees and can assess abilities across various dimensions. For instance, a single assessment in analytical geometry may evaluate both the ability to work with algebraic expressions and knowledge of the properties of vector, scalar, and mixed products. Consequently, a model that more accurately characterizes such test results is not unidimensional. Moreover, determining the exact dimensionality of an MIRT model is a fundamental step in analyzing the quality of test items.
The paper discusses potential challenges and limitations faced by educators when modeling test results using MIRT models, as well as ways to overcome these issues. The importance of creating a representative database of test result samples to address these challenges is emphasized.
The modeling process utilized the simdata function from the mirt package in the R programming language. A representative database of sample data was created for the selected models. For each primary score matrix, MIRT analysis was conducted: the dimensionality of the model was determined using the Hull criterion, parallel analysis, and the empirical Kaiser criterion; the best model was identified; parameter estimates for the selected model were obtained; errors were evaluated; and the model’s adequacy to the data was verified.
It was found that all the applied criteria for determining the dimensionality of the models yielded identical results, consistent with the input parameters of the simulation function. However, the mean parameter estimate vectors of the models differed from the input parameters when directly using the mirt package in R for evaluation. Solutions to this issue were proposed.
