Paper Summary
Share...

Direct link:

A Comparison of Several Item Response Theory Calibration Programs With Implications for Equating

Sun, April 6, 8:15 to 9:45am, Convention Center, Floor: 100 Level, 116

Abstract

The purpose of the current study was to compare item parameter estimation results from four different Item Response Theory (IRT) calibration programs (i.e., MULTILOG, PARSCALE, flexMIRT, and IRTPRO) and determine whether differences in results subsequently influence IRT true score and observed score equating relationships. Four IRT model combinations were fit to data from a national placement exam that consisted of both multiple-choice and free-response items. The two- and three-parameter logistic models were used in combination with Samejima’s (1969) graded response model and Muraki’s (1992) generalized partial credit model. Comparisons were made on the basis of (a) parameter estimation results, (b) MC, FR, and test characteristic curves, (c) MC, FR, and test information functions, and (d) estimated equating relationships.

Authors