Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Help
About Vancouver
Personal Schedule
Sign In
As Common Core Standards assessments are developed, it is imperative that great care be placed on clarifying the way in which alternate tests are developed for students with the most significant disabilities and what psychometric models are appropriate for these assessments. Despite No Child Left Behind (NCLB) being passed in 2001, little research has been published utilizing item response theory (IRT) for assessments for students with significant cognitive disabilities, students assessed with Alternate Assessments based on Alternate Achievement Standards (AA-AAS). The only published piece of work was Dyehouse (2009), who begins the documentation of difficulties found while trying to implement IRT and NIRT techniques so that further research can persist on developing appropriate methods of dealing with this unique data. Dyehouse analyzed the Indiana Standards Tool for Alternate Reporting (ISTAR) by applying the generalized graded unfolding model (GGUM), and Mokken models and found that different models work well for different populations of lower ability students. At the CCSSO conference in 2010, Dean, Wyse, and Vansickle presented the problems encountered when they attempted to apply IRT methodologies to the Michigan alternate assessment.
States such as South Carolina and New Mexico use item response theory (IRT) to develop their tests while others use different techniques (e.g. raw score reporting).Several state assessments that did not employ IRT models when developing their tests were used to recover item parameters post-hoc. These assessments include selected response (SR) and constructed response (CR) items with and without level of independence (LOI) included within the scoring and reporting for the assessment. Sample sizes ranged between 275 and 700 students per grade per subject test. Assumptions were checked for dichotomized and polytomous data with the one parameter logistic model (1PL), graded response model (GRM), partial credit model (PCM), GGUM, and Mokken model and were found to not meet the criteria to produce accurate measurements. In addition, theoretical total score frequency distributions and information item function (IIF) curves did not align.
Several implications from this study can be applied to current reform such as recommending that states focus on developing test specifications aligned with the appropriate scoring and reporting model prior to implementation in order to assure accurate measurement. Lastly, future directions of research are suggested.