Paper Summary
Share...

Direct link:

Measuring Depth of Vocabulary Knowledge With Testlets: Simultaneously Assessing Six Aspects of Word Knowledge

Sun, April 15, 10:35am to 12:05pm, Vancouver Convention Centre, Floor: First Level, East Ballroom C

Abstract

Goals/Purposes. In this paper, we discuss the theoretical rationale and preliminary findings for our use of testlets to focus on multiple aspects of vocabulary knowledge. We are developing standardized assessments based on a rich conception of word knowledge that draws on the materials students are exposed to in schools in a large state. We describe the item types employed on the tests, discuss the psychometric model, and present preliminary empirical evidence about the utility of this approach.

Perspectives/Theoretical Framework. Although measures of vocabulary knowledge have been included in reading assessments since the 1920s, the traditional means of assessing vocabulary have come under fire for being "driven by tradition, convenience, psychometric standards and a quest for economy of effort” (Pearson, Hiebert & Kamil, 2007, p.282). This hampers both research and decision making, for most of the measures do not reflect our current understanding of the nature of word learning (Nagy & Scott, 2000; Scott, Hiebert & Lubliner, 2006). This test reflects semantic, morphological, contextual, syntactic, metacognitive, and definitional information about each word, all of which have been identified as important aspects of word knowledge in research (Baumann, Kame'enui & Ash, 2003; Carlisle, 2000; Kieffer & Lesaux, 2008; Farstrup & Samuels, 2008; Hiebert 2005; 2006; Landauer & Dumais, 1997; Scott, Nagy & Flinspach, 2008). Nagy & Hiebert (2009) cite frequency and dispersion as important factors to consider in selecting words for instruction; we also consider these to be important for assessment and address these factors through the word selection for the tests (see related paper on assessing depth of vocabulary with testlets, this session).

Methods/Techniques. Each word on the test has a series of sub-items related to it: a self-assessment of knowledge about the word, a question about the word in context, a semantic network question, a question about the words' definition, and a question concerning part of speech. We employ Thissen’s testlet method, in which the related sub-items for a word are scored as a single graded response. We scale using Samejima’s graded item response model.

Data Sources/Evidence. Data employed in the scaling are derived from a sample of 4,115 fourth and fifth grade students from four regions of a single large, highly populated state, with oversampling of underrepresented populations to allow for DIF analyses on a later form of the test. Over 300 words were tested. The words were carefully selected to ensure an appropriate and representative sample of the words from our four word banks.

Results. We present preliminary evidence about the utility of using testlets that capture these aspects of word knowledge. The method achieves high reliability (typically above .90) with a relatively short multiple-choice test. In this paper, we will discuss how evidence from different sub-item types are related and help provide information about student vocabulary ability. For example, the part-of-speech sub-item substantially increases information about high ability students.

Authors