Search
On-Site Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Housing and Travel
Sign In
Low test-taking engagement (TTE) is considered a validity threat, especially for large scale, low stakes assessments (Wise & Smith, 2016). Various methods for detecting (dis)engagement based on response time (RT) data have been proposed and evaluated for utility within different assessment contexts (Soland et al., 2021; Wise, 2017). We evaluated two different methods for detecting TTE within the 2019 NAEP Grade 12 Reading assessment (N=13,773), using data from four digitally-based assessment blocks (n=6,555). Specifically, we explored how TTE as detected by data-driven and theory-driven methods related to block performance, proficiency levels, test design, and student background.
In the Normative Method (NM; Wise & Ma, 2012), data-driven thresholds are established using a percentage of mean RT (e.g., below 30% of average item RT is considered disengaged). In contrast, Detecting Engagement Levels Using Cognitive Modeling (DELCOM; Arslan et al., 2021) uses theory-driven cognitive task analysis to estimate minimum item-level RTs, which can be computed a priori to data collection, drawing on ACT-R theory (Anderson & Lebiere, 1998; Anderson, 2007) and reading research (Carver, 1992) to inform timing estimates.
Altogether seven TTE thresholds were investigated (NM 10%, 15%, 20%, 25%, 30%; DELCOM Skimming, Scanning). NM 30% and DELCOM Skimming thresholds were selected for further analysis because these most strongly related to performance and self-reported effort. Because DELCOM can only be computed for selected response (SR) items, NM 30% was also computed excluding constructed-response (CR) items to facilitate comparison; the methods were highly comparable (Cohen’s kappa: .772-.847; exact agreement: .974-.983).
Test-takers were classified as “engaged” for approximately 90%-94% of items across TTE thresholds. These thresholds showed moderate-to-strong correlations with block scores and weak-to-moderate correlations with self-reported effort (Table 1). TTE also varied with proficiency level (National Center for Education Statistics, 2021), such that students performing below NAEP Basic showed significantly lower TTE across blocks (Table 2).
Regarding test design, we examined the impact of block position (1 vs. 2) and item position (1-10) on TTE. All four blocks received a higher proportion of TTE when in position 1 than in position 2 (Table 3). Across items, TTE decreased from approximately 94% (item 1) to 83% (item 10), consistent with prior research (Wise, 2017). CR items showed lower TTE than SR items.
Finally, we correlated TTE thresholds with student-level characteristics measured by the NAEP Student Questionnaire (National Center for Education Statistics, 2019). Weak-to-negligible correlations were observed between TTE and motivation variables (Table 4), with strongest relationships observed for reading confidence (r = .101-.191). TTE weakly correlated with prior experiences (Table 5), with strongest relations observed for self-reported frequency of reading text messages (r = .143-.206) and identifying main ideas of a passage (r = .127-.198).
These findings illustrate the value of comparing data-driven and theory-driven methods of detecting (dis)engagement; TTE thresholds are comparable across methods, but reveal nuances across blocks, proficiency levels, and student-level characteristics. Future work should capitalize on available sources of evidence (process data, surveys) to contextualize relationships among engagement and performance and to further our understanding of TTE within digital reading assessments.