Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Help
About Vancouver
Personal Schedule
Sign In
This paper discusses the application of tools and techniques from Educational Data Mining (EDM), an emerging field that spans the disciplines of computer science, statistics, data mining and educational measurement and focuses primarily on modeling and uncovering patterns in large datasets (Romero, Ventura, Pechenizkiy, & Baker, 2011; Hand, Blunt, Kelly. & Adams, 2000). EDM covers a wide range of methods and this paper focuses on some examples drawn from a selection of projects that are funded by the National Science Foundation and connected through the CADRE network.
As the number and complexity of e-learning environments increases to include such rich environments as simulations and games, so does the availability and depth of data on student actions and responses, driving a need for those interested in detecting patterns of learning from the datasets to seek out methodologies from a variety of sources and apply them in new ways (Quellmalz, Timms, & Schneider, 2009).
The paper will describe several examples of pattern detection in datasets derived from games and simulations. First, the research team at Worcester Polytechnic Institute has used data mining techniques to improve the use Bayesian Knowledge Tracing (KT) models (Pardos & Heffernan, 2010). The KT models are commonly used in cognitive tutoring in order to determine student knowledge based on four parameters: learn rate, prior, guess and slip. In the past the Expectation Maximization (EM) algorithm has been used to derive values for those parameters, but previous research showed that with four free parameters the standard KT model is prone to converging to erroneous degenerate states depending on the initial values of these four parameters. This work reports on how data were simulated from a model with known parameter values and then run a grid search over the parameter initialization space of KT to map out which initial values lead to erroneous learned parameters. From that analysis researchers developed an individualization model that has unique properties, which allow it to avoid the local maxima problem.
A second example derives from a multi-university research team at the University of Maryland, Temple University and Arizona State University. This team has developed a game environment, SAVE Science, to assess middle school children’s understanding of both science content and process. Children respond to a problem-based narrative by exploring the environment, collecting possible clues, and using their understanding of the topic and process to draw inferences about the problem (Ketelhut, Nelson, Schifter, & Kim, 2010). SAVE Science allows students to solve the problem in multiple ways, many of which are equally correct while others uncover misconceptions held by the student. This project initially is using cluster analysis to look for patterns in student actions and responses to directed questions.
This paper will report on both of these DRK12 projects, focusing on the rationale for using data mining techniques, the techniques themselves as well as findings.
Diane Jass Ketelhut, University of Maryland - College Park
Michael Timms, Australian Council for Educational Research