Paper Summary
Share...

Direct link:

Assessment Quality and Validity for the Training Assessment Framework

Fri, April 17, 12:00 to 1:30pm, Virtual Room

Abstract

This paper presents the technical quality of assessments, validity framework, and its empirical evidences. We also explore the relationships among various task formats, scoring systems, and estimates of performance. Furthermore, we have developed a new generalizable approach to performance assessment (PA), a long-sought supplement to interrater reliability, by deconstructing performance scores for procedurally oriented sequence and number of necessary steps involved in successfully solving problems. The future additional analysis plan includes exploring relationships among measures with Armed Services Vocational Aptitude Battery scores.

The elements of assessment validity framework are as follows: 1) criticality and accuracy of the assessment in its coverage of relevant topics; 2) credibility and utility in examination composition and its item responses being sensible; 3) independent subject matter experts’ (SMEs) judgement on item quality and relevance; 4) utility of the scoring rubrics; and 5) statistical quality of the items and statistical inference with small size sample. Technical quality tools include a feature analysis which involves reviewing and tagging items for specific parameters defined in the TAF and expert-novice analyses to set performance assessment standards.

Both qualitative and quantitative approaches were employed. Assessment quality were measured based on SMEs’ and course instructors’ item review questionnaires, consisting of questions on item accuracy, clarity, criticality, difficulty, and cognitive tasks, and content experts’ rating on various item features. Instructional sensitivity of the assessment was studied by comparing performance between pre- and post-instruction groups, and trainee’s performance was also compared to course instructors’ performance.

Data for supporting each of the validity elements was collected from three Navy training courses – electronics (ET), ship protection (SP), and personnel functions (PF). The participants in each course in 2018 study and 2019 study, were, respectively, 55 and 148 for ET, 22 and 46 for SP, and 40 in 2019 study for PF. The sample of course instructors and SMEs varied from 4 to 8 and 2 to 4, respectively. The performance data were collected from comprehensive test with 40 multiple choice items, 2 simulation-based PAs, and 2 knowledge maps (KMs).

Reliability of the comprehensive test ranged from .67 to .80, PAs from .72 to .95, and KMs from .72 to .77. SMEs review showed that very high level of criticality (i.e., 72.5% of items were 4 or 5 in 1 (least critical) to 5 (most critical) scale), clarity (90%), accuracy (89.8%), and a normal distribution of item difficulty and task complexity. The effect sizes (ES) calculated by comparing pre- vs. post-instruction groups were very large, especially in PAs for ET course showing the highest (ES=2.3), and ESs in the comprehensive test were 1.4 for PF, 1.2 for SP, and 1.4 for ET.

This study proposes a comprehensive and integrated approach for evaluating assessment quality and validity, utilizing both qualitative and quantitative methods, incorporating experts’ qualitative data into validity modeling to compensate for small sample issues. In addition, such an approach shows comparability of different assessments across different Navy courses.

Author