Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Help
About Vancouver
Personal Schedule
Sign In
The Netherlands have a long tradition of cohort research on educational careers. Already from the late eighties longitudinal studies were conducted in either elementary education (PRIMA) or secondary education (VOCL). In both studies the achievements of different cohorts of children were systematically monitored by administration of a number of cognitive tests and questionnaires. In school year 2007/2008, the new cohort study COOL5-18 started. Cohorts involved in this study are monitored over a longer period of time. That is, from kindergarten until the end of secondary education. Currently, two waves of data collection occurred. In the present study, the achievements of 15-year-olds in reading comprehension and mathematics were analyzed in relation to the background variables of school track, gender, ethnicity and SES using the COOL5-18 and the VOCL cohort data.
For a meaningful comparison of learning results through time, it is conditional to have one single measurement scale. One possibility to satisfy this requirement is to administer the same test to all children at all measurement occasions. A potential disadvantage of such an approach, however, is that it is not possible to adapt to the ability levels of the participating children. This may not only lead to frustration in the children, but also to unreliable test results. In studies like COOL5-18 it is therefore common to use different test versions tailored to school track. The consequence of this approach is that comparisons between school tracks can no longer be made. The problem of the incomparability of test results can be solved by means of an equating procedure. A wide array of equating procedures has been proposed in the literature. The test versions of COOL5-18 and VOCL were equated within the framework of item response theory. If the chosen item response model holds, it becomes possible to translate the children’s raw scores into ability scores which are comparable across test versions.
For practical purposes, the item response model and the resulting equivalent scores must be considered true. In reality, however, a model is nothing more than a parsimonious description of the data. No single model will fit the data perfectly. Especially in studies like COOL5-18 many model violations may occur to a more or less extent due to the complexity of the research design. The validity of the underlying model is in question, for instance, if particular test items function differently for members of different groups (DIF) or if children’s behavior on test items changes in time (item drift). It is practically impossible to test for all model violations that might occur and to take them into account in the underlying model. In the present study, the robustness of the performed cross-cohort comparisons was therefore examined in a bootstrap procedure. In the bootstrap, the uncertainty about the item parameter estimates was explicitly included as possible source of variability in the results. In this manner, it is possible to draw more precise conclusions about children’s reading and mathematics achievements through time.