Paper Summary
Share...

Direct link:

Examining the Power of Multiple Data Sources in Predicting Academic Achievement in Undergraduate STEM Courses

Sun, April 7, 3:40 to 5:10pm, Sheraton Centre Toronto Hotel, Floor: Mezzanine, Pine East

Abstract

Purpose & Theoretical Framework
There is growing interest in the use of timely deployment of prediction algorithms that can identify students likely to perform poorly (Pistilli, Willis, & Campbell, 2014; Pritchard & Wilson, 2003). In order to predict students’ academic success, educational psychologists often collect self-reported motivations (e.g., Mega et al., 2014). Across the institution, researchers in higher education units develop prediction models using data available to their academic unit: 1) registrars use demographic and achievement information (Thiele, Singleton, Pope, & Stanistreet, 2016), learning technology administrators work with educational data miners to observe behaviors in learning management systems (LMS) and develop features for learning analytics and educational data mining analyses (Romero, Espejo, Zafra, Romero, & Ventura, 2013). Campus stakerholders without advanced expertise in learning theory often build prediction models by collecting expedient data and focusing on algorithm development, without sufficient consideration of educational theories and contexts that can make model features more meaningful and more powerful (Baker & Yacef, 2009). Few studies combine these data sources, fewer systematically compare the accuracy achieved by models with from multiple data sources, and fewer still conduct “feature engineering” steps to build theory-aligned traces from the raw data they have procured. This study compares the utility of individual, combined, and engineered data sets to predict achievement and inform future deployment of interventions to improve undergraduate achievement.
Unique to this study among prediction models are (1) the division of behavioral data from the LMS into “theory-aligned” and non-aligned variables and comparison of their ability to predict achievement (Dominguez, Uesbeck & Bernacki, 2016) and (2) investigation of patterns of learning events (Bannert, Reimann, & Sonnenberg, 2014; Winne, 2014).

Data & Methods
Anatomy (N=1326) and Calculus (N=448) students’ demographic and behavioral data were collected from University registrar and LMS systems (Table 1). Students’ goals and perceptions of value and cost were surveyed in the first two weeks of each face-to-face large lecture course. Behavioral data were further reclassified to distinguish non-theory-aligned events from those reflecting cognitive and metacognitive leanring processes (per Winne & Hadwin, 1998; Dominguez, et al. 2016, Table 2). Sequence pattern mining followed to identify frequent metacognitive event patterns. Regression analyses determined the variance explained by each data source’s variables, alone and in combination with other sources (see Table 3).


Results
Models including individual data sources, their combinations, and engineered behavioral features appear in Table 3. Each source predicted some variance; combinations demonstrate improved R2 and suggest combining data sources improves model accuracy. Later models (9, 10) further demonstrate that appraising digital resources to classify the behaviors they afford increases the usefulness of LMS data. In some course contexts (e.g. biology), mining patterns of events provides further explanatory power, but this seems dependent upon course design features (i.e., resources available [Table 2], and assessment practices).

Significance
Existing ambient data from university systems can provide value to institutions seeking to identify and support struggling learners. Partnership with learning experts to enrich data can improve accuracy of models and precision of interventions they afford.

Authors