Paper Summary
Share...

Direct link:

Latent Use Cases: A Topic Modeling Approach to Massive Open Online Course Student Behavior

Thu, April 16, 12:00 to 1:30pm, Marriott, Floor: Fourth Level, Armitage

Abstract

Objectives

From the fall of 2012 through the summer of 2013, HarvardX and MITx provided 17 courses drawing course enrollments from roughly 600,000 unique users from around the world. Such large numbers of users combined with their diverse cultural and educational backgrounds inherently implies variation in how users choose to interact with material. Tracking logs (also known as clickstream data) provide insight into this variation by storing records of user interactions with unique pieces of content, ranging from the unique number of resources accessed to the interactivity within a single lecture video.
Hidden within this collective data lies a great deal of insight into the students’ behavior, potentially explaining learning styles and optimal uses of course content. However, given participant diversity and the shear number of course components available, enumerating the possible patterns of student interactions is nearly impossible. In this paper, we make the problem of modeling student behavior more tractable by adapting the approach of Latent Dirichlet Allocation (LDA) (Blei, Ng, & Jordan, 2003) to user modeling.

Theoretical Framework

Latent Dirichlet Allocation (LDA) is a natural language processing (NLP) technique often utilized to extract topical themes from a collection of text-based documents. In this work, LDA is used to uncover behavioral patterns by drawing a close analogy between text documents (words) and the individual interactions of students (clicks) in a MOOC.

Methods and Data

This paper focuses on one MITx MOOC, Electricity and Magnetism (8.02x). Between January 2013 and September 2013, 43,758 people enrolled in 8.02x from a wide range of countries, ages, and levels of education. In this work, each student is represented as a “bag-of-interactions” using tracking log records, i.e., the number of times a student interacted with each of the core components of the course (chapters, sequences, verticals, problems, videos, and html pages). Students are then softly assigned to clusters and averaged based on their activity to derive “use cases” that define a probability distribution for interacting with the course components.

Findings

The application of LDA on 8.02x finds probability distributions that align with expert defined user subpopulations such as shoppers, completers, auditors, etc. as in DeBoer et al. paper (2014). This provides additional support for the impact of heterogeneity in MOOCs and gives higher granularity in characterizing students through the probability distributions themselves and the students’ soft assignments. These soft assignments also provide a way to predict learning outcomes such as certification and dropout.

Significance

Findings in this report demonstrate an unsupervised approach for uncovering behavioral patterns in MOOCs. These patterns provide intuition about which course components are being utilized and facilitate grouping students on the basis of their behavior.

References

Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent dirichlet allocation. the Journal of machine Learning research, 3, 993-1022.
DeBoer, J., Ho, A. D., Stump, G. S., & Breslow, L. (2014). Changing “Course” Reconceptualizing Educational Variables for Massive Open Online Courses. Educational Researcher, 0013189X14523038.

Author