Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Registraion, Housing and Travel
Personal Schedule
Sign In
Longitudinal assessment programs are currently being implemented or piloted by many member boards of the American Board of Medical Specialties. The purpose of these programs is to offer an alternative pathway to satisfy the Maintenance of Certification (MOC) framework’s Part III component, which has historically focused on assessing an individual’s medical knowledge. Longitudinal assessment programs depart from a traditional, singular testing event, in that they are more continuous, assess the individual at more frequent intervals, and are, generally, more flexible regarding administration (i.e., most programs being piloted to date offer remote or unproctored administrations). In moving away from a single testing event, the longitudinal programs attempt to leverage established findings in cognitive psychology, specifically the testing effect (see Roediger & Butler, 2011, for a review on literature dating back to 1917) and spacing effect (see Greene, 1989) to facilitate learning as part of the assessment process.
The shift in emphasis from solely summative decision making to more formative, educational outcomes has brought its share of psychometric challenges. For example, changes in the frequency of assessment and administration environment have brought on 1) a host of test security issues (Althouse, Diepenbrock, Dugosh, & Mulkey, 2018), and 2) significant challenges in maintaining the score scale over time via equating/linking (Furter, Qunbar, & Dwyer, 2018). The frequency of assessment and its continuity raises an important score reporting question as well, i.e., “What to report to the test-taker and when?” Because many of the longitudinal programs are continuous, with test-takers participating in the program year-after-year, waiting until the end of an assessment cycle to report how a test-taker performed is not acceptable from the test-taker’s standpoint. However, if preliminary scores or predictions are based on subsets or partial data, the imprecision in the estimate must be communicated in a way that accurately reflects the score and its accuracy.
The current paper focuses on the challenge of reporting scores within the context of longitudinal assessment. The paper begins by discussing the operational constraints and participation behaviors that can complicate psychometric operations, score reporting included, in a longitudinal setting. Next, a simulation study and a resampling study are presented to demonstrate the accuracy of reported scores at different points in the program timeline. Factors that would have an impact on the accuracy/precision of scores, such as program size, item drift, etc., were varied in the simulation study. The goal in running these studies was to best mimic the real-world contexts and rule systems set in place in a typical longitudinal assessment program. Finally, the findings of the studies are used to outline the implications of reporting scores at different time points to test-takers, and how studies such as these can inform policy decisions.