Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Browse By Descriptor
Search Tips
Annual Meeting Theme
Exhibitors
About Philadelphia
About AERA
Personal Schedule
Sign In
X (Twitter)
The purpose of the current study was to compare item parameter estimation results from four different Item Response Theory (IRT) calibration programs (i.e., MULTILOG, PARSCALE, flexMIRT, and IRTPRO) and determine whether differences in results subsequently influence IRT true score and observed score equating relationships. Four IRT model combinations were fit to data from a national placement exam that consisted of both multiple-choice and free-response items. The two- and three-parameter logistic models were used in combination with Samejima’s (1969) graded response model and Muraki’s (1992) generalized partial credit model. Comparisons were made on the basis of (a) parameter estimation results, (b) MC, FR, and test characteristic curves, (c) MC, FR, and test information functions, and (d) estimated equating relationships.
Jaime Leigh Peterson, University of Iowa
Mengyao Zhang, The University of Iowa
Seohong Pak, The University of Iowa
Shichao Wang, University of Iowa
Wei Wang, Educational Testing Service
Michael J. Kolen, University of Iowa
Won-Chan Lee, University of Iowa