Search
Program Calendar
Browse By Day
Browse By Unit
Browse By Session Submission Type
Browse By Descriptors
Help
About Vancouver
Personal Schedule
Sign In
To stimulate higher order thinking in science education summative assessment must overcome two challenges. First, it has to move from the extensive use of multiple choice and short answer items to a wider range item formats (Hout & Elliott, 2011). Second, it has to improve its understanding of what higher order thinking means in the context of science education (Kind, Osborne, & Szu, in preparation). The present study is part of a project that aims to approach these two challenges jointly in what Bond (2003) calls a ‘theory-practice dialectic.’ It does so by advancing a rationale for scientific thinking using this to suggest variables for assessment scales. The variables were then operationalised and trialled, using a range of item formats.
The rationale is based on the work of Giere, Bickle and Mauldin (2006) and Klahr (2000) who separate scientific thinking into three main practices, hypothesising (explaining scientific phenomena), experimenting (gathering data and testing explanations) and evidence evaluation. This perspective sees these practices as reliant on domain-specific content, procedural and epistemic knowledge. Used together with a domain-general perspective (Bloom, 1956), it suggests that lower and higher-order thinking may be explained in terms of the use basic and advanced knowledge when solving problems posed by the need to engage in the three previously mentioned practices. For example, students may recall science concepts and experimental procedure when solving a simple experimental problem, but will have to synthesise these with epistemic knowledge to answer more advanced questions.
One way to address this conception of the distinction between lower and higher order cognitive demands is to define the three practices as main variables of any science assessment and use item formats as a way to control the level of thinking required. Doing so, however, raises a question of to what degree item format is independent of the construct being tested? Free response items, for example, using large text stems for students to read and interpret and thus may have more in common across scales than items using different formats within the same scale. To explore this issue and the overall validity of the scales, three different tests, one for each variable, were developed using a similar range of item formats and given to sample of 250 mixed ability 14 to 15 years old students.
Multiple methods are used in the data analysis. Rasch scaling is used as suggested by Smith (2004) to analyse construct validity, with particular focus on dimensionality and progression in the scales. Multitrait-Multitmethod (MTMM) analysis is used as suggested by Campbell and Fiske (1959), with a particular focus on ‘item-bias’ related to the different item formats. Both these methods are supported by confirmative factor analysis. In addition think-aloud interviews (N=20), as suggested by Reis and Judd (2000) explore how students interpreted and solved the items, supplementing the validity of our interpretation. Findings and their implications will be presented.