Paper Summary
Share...

Direct link:

Predicting Academic Achievement Using Multiple Types of Campus Data in an Introductory Biology Course

Fri, April 17, 2:15 to 3:45pm, Virtual Room

Abstract

High attrition rates and slow progress towards degrees among undergraduate students in STEM disciplines threaten the health of the modern workforce (Dai & Cromley, 2014). Early prediction of student success is critical to inform interventions that can prevent students from earning poor grades that contribute to these trends. Most university courses include learning management system (LMS) sites where student engagement in learning activities can be captured and observed alongside institutional data (Macfadyen & Dawson, 2010) to inform models that predict student success. Campuses must invest resources to collect and model these data, and few studies have examined how institutional (admissions, transcript) and LMS data might contribute to accurate prediction models. This study examines how campus data sources can explain variance in students’ academic achievement through a systematic comparison of prediction models.

Method
Data from Introductory Biology courses (2014-2016; 5 semesters, Table 1) including demographic information, motivation via a questionnaire administered in week 1 (assessed self-efficacy; Bandura, 1977, achievement goals; Elliot & McGregor, 2001, and value and cost; Eccles et al., 1983), and LMS events were used to predict coure performance. LMS data were further classified into theory-aligned (i.e., Self-regulated learning, see Winne & Hadwin, 1998, 2008) and non-theory-aligned traces of learning events. Hierarchical regression analyses tested the predictive power of multiple types of data alone and in combination, with forward-selection used to determine significant variables amongst candidate predictors and prioritize model parsimony. We further examined the validity and sustainability of these prediction modeling strategies by applying the model to different semesters to gauge how robustly an algorithm predicts achievement over long periods of time, and semesters that vary in scheduling and course enrollment.

Results
Hierarchical regression analyses confirmed that each campus data source explained variance in achievement outcomes, and that LMS behaviors predict additional variance when interpreted and classified through learning theory. Additional interpretation via sequence mining and learning analytics approaches only improved variance explained in some cases (Table 2). Analyses assessing multiple semesters revealed that predictive accuracy was largely stable, with models explaining 28%-30% of final grades. A model dependent upon only data that accrue with no additional effort required of instructors to survey motivation (Model 6) demonstrated similar predictive power to those including motivational data, which suggests that campus decision-makers wishing to strategically leverage campus efforts can rely only on institutional and LMS data and still obtain robust prediction of student achievement as they develop student success initiatives to support those predicted to benefit from academic support services.

Significance
These findings demonstrate that – in accordance with the 2020 Annual Meeting theme – data partnerships between campus stakeholders and education researchers can benefit undergraduate students as they pursue challenging STEM careers. Prediction models that leverage existing campus data can explain academic achievement, and these ambient data from university systems can provide value to institutions seeking to identify and support struggling learners.

Authors