Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Registraion, Housing and Travel
Personal Schedule
Sign In
Objective and Theoretical Background. Understanding and supporting STEM persistence remain pressing issues. From a psychological perspective, researchers have considered how motivational processes, including students’ expectancies for success and value for the task, help to explain persistence in STEM majors (Wang & Degol, 2013). Expectancy-value theory (Eccles et al., 1983) proposes that these psychological variables combine in complex ways to predict persistence, but it is rarely modeled as such. This complex view is aligned with a dynamic systems theory perspective, which suggests that precursors and outcomes of human behavior are many and varied, and that the same life event may have a different impact depending on the context and timing of that event (Smith & Thelen, 2003). Traditional linear modeling may be inadequate to model such dynamic relations over time (Kaplan & Garner, 2017).
We explored machine learning as a potential alternative analytic approach for exploring the complex relations among achievement, course-taking, motivation, and persistence. Specifically, we asked: (1) How do demographic characteristics, initial motivation, credit enrollment, and academic achievement predict career intentions and major persistence? (2) Does motivation explain additional variance beyond institutional data? We compared findings obtained using machine learning versus traditional regression.
Method & Results. Undergraduate engineering students (N = 345) completed surveys assessing (a) their initial motivation for engineering (perceived competence, task-value, cost) immediately prior to their first-year in college (T1, August 2015), and (b) engineering career intentions 20 months later (T2, March 2017). Demographics, ACT, achievement and credit enrollment (Fall 2015, Spring 2016, Fall 2016), and Spring 2017 major were collected from university records.
We compared a machine learning technique (random forest) to traditional multiple linear regression to predict T2 major and career intentions. Random forest is an ensemble machine learning technique by which a series of randomly selected decision trees are “grown” simultaneously, and each decision tree’s accuracy is considered in aggregate (Breiman, 2001). The “random” elements are the random selection of data subsets for growing each tree and random features used to split each tree node. We trained and tested four random forest models, two for each dependent variable (one without and one with motivational variables) and then conducted parallel analyses using multiple regression.
The random forest models (out-of-sample RMSE for Career Intentions models = 2.404, 2.531; area under the curve (AUC) for Persistence models = .626, .626) performed about as well as multiple linear regression (out-of-sample RMSE for Career Intentions model = 2.471, 2.553; AUC for Persistence model = .571, .588), with neither model performing particularly well. Indeed, given a randomly selected “persister” and a randomly selected “leaver,” our model only assigns a higher probability to the persister in about 50% of the draws.
Significance. Although neither the machine learning approach nor the multiple regression approach explained a large proportion of variance in persistence, this study helps to demonstrate one potential approach to studying STEM persistence using machine learning. Findings will be used to illustrate machine learning as an approach for modeling complex psychological phenomena by estimating the effects of many variables simultaneously.
Emily Bovee, Michigan State University
Kristy A. Robinson, Michigan State University
S. Patrick Walton, Michigan State University
Daina Briedis, Michigan State University
Lisa Linnenbrink-Garcia, Michigan State University