Search
On-Site Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Housing and Travel
Sign In
1. Objectives or Purposes
In recent years, significant progress has been made in measuring college students’ critical thinking (CT) using authentic performance tasks (PT). However, challenges remain for the process of scoring student performance data (essays) reliably and efficiently. We addressed these challenges by applying an empirical strategy to reduce a scoring rubric for student essays.
2. Theoretical Framework
The feasibility of constructing PT derived from cognitive theory of CT (Evidence Centered Design – ECD, Mislevy & Haertel, 2006) has be demonstrated repeatedly (e.g., Braun et al., 2020; Davey et al., 2015). Zlatkin-Troitschanskaia et al. (2019) recently outlined the conceptual underpinnings of multifaceted scoring based on a latent factor model. Using sample data from the recently developed Wind Turbine PT (Shavelson et al. 2019), they used Confirmatory Latent Analysis (CLA) to cluster the underlying skill set of 23 indicators into four correlated but conceptually distinguishable latent constructs: (a) Identifying relevant and trustworthy information, (b) Weighing and evaluating (often contradictory) arguments, (c) Reflecting the consequences of the concluding decision/recommendation, and (d) Writing quality. While the four- factor model had good fit, the first two latent factors were highly correlated (r = .97). We thus opted for a three-factor solution.
3. Method
We optimized the three-factor model with regard to the most efficient indicator set based on the rationale of Composite Reliability (CR) in Structural Equation Modeling (SEM) suggested by Raykov (2004). A common misconception is that the reliability of a latent construct in SEM increases monotonically with additional indicators, which is only true for the special case of parallel measurement. For the commonly used “congeneric” measurement model (Joreskog, 1971), adding indicators with low loadings can reduce the latent reliability. Our goal was to reduce the set of 23 indicators variable to a subset of items with maximum reliability for the three factors with as few indicators as possible.
4. Data Source and Analysis
We ordered the indicators for each of the three factors using their reported standardized loadings (Zlatkin-Troitschanskaia et al., 2019, p. 10). We calculated reliabilities iteratively by adding the next highest loading indicator each time.
5. Results
Optimal reliability is reached with 11 out of 13 indicators for the first factor, 2 out of 4 for the second, and 4 out of 6 for the third factor (Table 1). While this would suggest a reduction to 17 indicators, the reliability for 6 indicators for the first factor yields a very strong reliability of .93. A latent factor in SEM is robustly identified with at least 3 indicators . We therefore argue for a 6/3/4 = 13 indicator solution.
6. Scholarly Significance
The CR approach can potentially reduce PT scoring costs with little to no reduction in the reliability of the multifaceted measurement of CT. This strategy can be used to optimize a scoring rubric to a specific task or to PTs in general, depending on the heterogeneity of the pool of PT under consideration.