Paper Summary

Automated Discovery of Commonsense Science Knowledge

Mon, April 16, 2:15 to 3:45pm, Sheraton Wall Centre, Floor: Third Level, South Beluga

Abstract

In research on science education, studies of students’ commonsense science knowledge (CSK) have played a central role. Although researchers studying CSK have employed a variety of techniques, many studies make use of one-on-one clinical interviews. Although these interview allow for a rich exploration of the CSK, the data that results – extended utterances, consisting of natural language, gestures, and drawings – pose a number of difficulties for analysis.
In this presentation, I will discuss how techniques from statistical natural language processing (SNLP) can be employed to analyze the data produced by clinical interviews about prior science conceptions. To illustrate these techniques, I will draw on a corpus of 54 interviews in which middle school students were asked to explain the seasons. My presentation will emphasize three innovative features of these new techniques:
1. Simple vector space models. In my work, I have explored the use of a family of techniques from SNLP known as vector space models. One type of vectors space model, Latent Semantic Analysis (LSA), has already been employed, with some success, in applications relevant to educational research (e.g., Foltz, Kintsch, & Landauer, 1998; Graesser, Lu, Jackson, & Mitchell, 2004; Shapiro & McNamara, 2000; Wade-Stein & Kintsch, 2004). However, LSA is a relatively sophisticated type of vector space models. I have found that much simpler methods give equivalent or better results, at least within my domain of application.
2. Verbal data. Most applications of SNLP in education are based on analysis of written text. (There are, however, exceptions, such as Magliano & Millis, 2003). In contrast, I am working from verbal protocol data produced by relatively open interviews between researchers and students
3. Use of cluster analysis to discover student conceptions. Ultimately what I am striving toward is the development of software algorithms that do the work of “seeing” student commonsense science conceptions. These algorithms would take, as input, a corpus consisting of interviews with multiple students about some science subject matter. Then they would output a set of “conceptions,” which it would discover, together with an interpretation of each interview in terms of these conceptions. In this manuscript, I describe just first steps toward this end. Small segments of transcripts are clustered in order to discover the “ideas” out of which student explanations are constructed. The results align well with human analyses of the same corpus

Author