Individual Submission Summary
Share...

Direct link:

How Item Response Theory Can Improve Assessment Efficiency: An Illustration Using the Child Behavior Questionnaire

Thu, March 21, 4:00 to 5:30pm, Hilton Baltimore, Floor: Level 2, Key 4

Integrative Statement

The literature on child temperament is rapidly expanding as more researchers and practitioners come to appreciate the importance of early emerging individual differences in children’s personalities. The availability of assessment tools that are both robust and practical is therefore an increasingly pressing matter. One of the most popular temperament inventories, the Child Behavior Questionnaire (CBQ; Rothbart et al., 2001), is extremely comprehensive, but also quite lengthy. With 195 items making up 16 subscales the CBQ’s length can be prohibitive in many circumstances. Short forms of the CBQ exist (Putnam & Rothbart, 2006), but traditionally developed short forms may entail a non-trivial sacrifice of measurement precision and content coverage (e.g., Crede, Harms, Niehorster, & Gaye-Valentine, 2012). Modern psychometric techniques based on Item Response Theory (IRT), however, are well suited to reducing assessment length without compromising measurement quality. Accordingly, this study demonstrates these strengths of IRT and related techniques by using them to evaluate and streamline the CBQ. Shorter questionnaires that can match the psychometric and conceptual potency of longer forms benefit researchers who want to examine child functioning in fine detail, and informants and practitioners with a vested interest in parsimonious measurement.
The CBQ was first evaluated and edited in a calibration sample of 605 children (47.3% girls) aged 36-95 months (Mage = 52 months; SD age = 10.79 months) using graded response models, item factor analyses, and bi-factor models. The performance of the revised CBQ was then compared to the original and short forms across the calibration sample (maternal and paternal ratings) and 3 independent validation samples (Ns of 187, 468, and 590 children). When revising the scales weight was given to both general psychometric performance, and content coverage. If scales demonstrated multidimensionality, care was taken to ensure that each facet was represented in the revised scale. This helped maintain both the psychometric and substantive strengths of the original CBQ.
The initial evaluation of the CBQ suggested that its 16 scales generally function well, but contain a number of items that are weak and/or unnecessary. Overall, the CBQ was trimmed by 44% (Table 1). The revised form generally performed as well as the original across all samples. Measurement precision was similar (see Table 2 for marginal reliability across forms and samples), and the rank ordering of children was mostly preserved. The revised and original forms also demonstrated similar levels of inter-parent agreement, and the ability to predict problem behavior. Although exploratory factor analyses suggested some differences in latent structure across forms, the super-dimensional composites that subsume the CBQ’s 16 scales -- Effortful Control, Negative Affectivity, Surgency -- showed no functional differences. Altogether, this study highlights the potential of IRT and similar analytic techniques in satisfactorily reconciling competing concerns for psychometric quality, length, and content coverage.

Authors