Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Help
About Vancouver
Personal Schedule
Sign In
Introduction
Objective structured clinical examinations (OSCEs) have been used for over a decade now in the assessment of medical students and residents, as well as in licensing practicing physicians. An OSCE consists of a series of medical cases in which examinees (i.e., medical students) interact with a standardized patient (SP), who portrays a scripted medical problem (i.e., appendicitis). The OSCE is an assessment tool designed to evaluate clinical competence and the interpersonal skills of examinees (Harden, Stevenson, Downie, & Wilson, 1975; Lawson, 2006).
Surprisingly, the reliability and validity of the OSCE has yet to be established. In fact, many articles using the OSCE will begin by stating the reliability and validity of the OSCE has been well established, often citing studies. However, many of these references have not provided adequate validity evidence for the OSCE, or cite other articles they believe provide OSCE validity evidence. This perpetuates the myth and inflates the problem, and it is hard to find a study that has validated the use and interpretation of OSCE scores thus far.
Objectives
The validity evidence regarding the OSCE needs to be investigated in greater detail. The aim of this article is to examine a) studies that are presented as validity or validation of the OSCE, and b) the extent to which the reported validity of the OSCE conforms to modern day validity theory. This will involve comparing what is being reported as validity to what the Standards (AERA, APA, & NCME, 1999) would classify and consider as validity evidence.
Methods
A systematic review using PubMed was conducted. Articles were included in the data analysis if the term ‘validity’ or ‘validation’ was in the title of the article, and either the term ‘OSCE’ or ‘objective structured clinical examination’ was in the title or abstract. Only 16 articles met this criterion.
Results
The most frequently reported sources of validity evidence were construct (50.0%) and content (50.0%) validity, followed by criterion-related concurrent validity (37.5%). The remaining sources of validity evidence were reported less frequently, which included face validity (25.0%), criterion-related predictive validity (18.8%), response processes (12.5%), and consequences (12.5%). A differences between what authors’ reported as validity evidence to how the same evidence would be classified using the Standards (AERA, APA, & NCME, 1999) was found.
Based on the Standards, a significant source of validity evidence regarding the OSCE were convergent (50%) and discriminant (50%) validity. The remaining sources of validity were reported less and included content (25.0%), test-criterion (18.8%), validity generalization (18.8%), internal structure (18.8%), response processes (12.5%) and consequences (6.3%).
Scholarly Significance
A difference exists in what authors report as OSCE validity evidence compared to the Standards (AERA, APA, & NCME, 1999). This discrepancy may arise from a lack of a clear definition of validity and its practices. This in turn limits the conclusions the authors can make regarding the validity and reliability of the OSCE. Using the current Standards will help overcome some of these issues and help guide future validation studies of the OSCE.
Tavinder K. Ark, The University of British Columbia
Neelam Ark, The University of British Columbia
Bruno D. Zumbo, The University of British Columbia