Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Visiting Washington, D.C.
Personal Schedule
Sign In
X (Twitter)
This presentation will discuss validity for English learners (ELs) in large-scale assessment from the perspective of generalizability. The paper will discuss how a G theory-based perspective allows examination of the soundness of current practices concerning the testing of EL populations. The knowledge shared originates from empirical research the author has conducted applying G theory to develop testing models that are compatible with knowledge from the language sciences.
G theory can be viewed as a theory of sampling of observations. Among others, these observations involve the tasks or items included in a test, the raters who score students’ responses, the different languages or dialects in which students may be tested, and the format (e.g., with or without certain accessibility resources) in which items are administered. Accordingly, validity can be examined in terms of the extent to which students’ appropriate generalizations about the students’ skills or knowledge in a given domain can be made based on a given sample of observations.
Evidence on the use of G theory in the testing of ELs shows that a great deal of measurement error in EL testing is due to the fact that EL populations are tremendously heterogeneous—each EL has a unique set of strengths in English and in their first language. Test items vary tremendously in the sets of linguistic demands they pose to test takers. As a consequence, a considerable amount of measurement error in the testing of ELs is due to the interaction of these students’ linguistic skills and the linguistic features of test items.
Available evidence indicates that the technical qualities of tests should be examined disaggregating the target student populations by linguistic group. This disaggregation may reveal, for example, that the kinds of generalizations that can be made based on test scores may be appropriate for non-ELs but not for ELs. Appropriate adjustments can be made on the characteristics of tests (e.g., increasing the number of items included in a test with the purpose of increasing the generalizability of scores) to ensure high quality properties for diverse populations based on these analyses.
There is also evidence that speaks to the advantage of examining testing accommodations as a source of measurement error. While testing accommodations (modifications of the ways in which tests are administered) are used with the intent to support students in gaining access to the content of items, they also increase the cognitive demands during test taking and are based on assumptions about usability and effectiveness which may or may not be appropriate. Better informed decisions on the use of testing accommodations can be made based on examining the ratio between the amount of measurement error contributed by testing accommodations and the differences in scores obtained by students tested with and without testing accommodations.
The presentation will conclude with a set of recommendations for large-scale assessment systems.