Paper Summary

Assessment Design: Network-Based Environments as and for Formative Assessment and Evaluation of Student Thinking

Sun, April 15, 8:15 to 10:15am, Sheraton Wall Centre, Floor: Grand Ballroom Level, North Grand Ballroom D

Abstract

Different approaches are needed to better assess student understanding of math and science ideas, especially with deeper and qualitatively different way of teaching and learning math and science. Traditional assessments don’t reflect the complexity and diversity of student understanding. For example, two students that are assigned a similar score might have very different ways of reasoning about the same problem. For several decades, researchers have supported the use of authentic assessments to help better elicit, document, and interpret more valid evidence of student understanding. Yet, the systemic implementation of authentic tasks to assess student understanding at a large scale is challenging. Difficulties include: authenticity of interpretation of students’ work (focus on student mathematical thinking and measuring what counts), cost, time consumption, and fairness (inter-rater agreement between several evaluators).
We ask evaluators to provide a non-judgmental description of student thinking as elicited in students’ response to an authentic task, prior to assigning a quantitative measure of their performance. The technology allowed a large group of evaluators to assess, in real time, middle school students’ work on a type of authentic modeling task in mathematics called model-eliciting activity (Lesh, Hoover, Hole, Kelly, & Post, 2000; Aliprantis & Carmona, 2003). Evaluators were asked to describe student mathematical thinking by marking a 1 or a 0, depending on whether they assessed that a given set of descriptors well-depicted (or not) students’ reasoning when they solved a particular model-eliciting activity. The descriptors were previously obtained by a systematic analysis of student mathematical thinking (Carmona, 2004). The aggregate of the group’s assessment was displayed in a public space, generating discussion of the descriptors and providing pictorial representations of student thinking through a spider graph. Evaluators were asked to score students’ work under two different conditions: (a) having previously depicted students’ work using the descriptors and network environment, and (b) directly providing a numerical score without the descriptive assessment. Evaluators’ quantitative scoring was not shared publicly. Inter-rater reliability among evaluators was calculated.
Results are as follows. First, students’ diversity and complexity in thinking was expressed through differences in the generated spider graphs, which allowed a characterization of student work by capturing features that are difficult to express by a numerical value in a scale. (Carmona, Stroup, & Davis, 2006; Carmona, 2007) An example of the displays is provided in Figure 2. Second, this process of continuous display of ideas, reflection, and analysis allows for a quick and reliable way to pull forward multiple evaluators’ ideas for discussion and negotiation of meaning. Third, changes in evaluators’ inter-rater reliability showed a significant improvement after evaluators used this tool to first describe student thinking, leading to a more fair evaluation system. These findings support the conclusion that assessing student thinking first, and evaluating it afterwards, could have a very positive impact in improving student evaluation when compared to only evaluating without assessing first. Further research needs to be conducted to extend this assessment design that focuses on the diversity and complexity of student thinking to a larger scale.

Author