Paper Summary
Share...

Direct link:

Embedding Computational Thinking Practices Into Data Science Education

Sat, April 29, 8:15 to 9:45am, Henry B. Gonzalez Convention Center, Floor: Meeting Room Level, Room 221 D

Abstract

Purpose. The nascent field of data science is described as the intersection of hacking skills, math/statistical knowledge, and domain knowledge. It would be beneficial to the fields of data science and computational thinking if the vague and intimidating phrase “hacking skills” is clarified and exemplified through computational thinking data practices as part of a joint curriculum. The purpose of this project is to establish and discuss elements of such curriculum with concrete hands-on activities.

Theoretical Framework. Computational thinking researchers are trying to establish a working definition for this concept that will lead toward its embedding and implementation in the curricula of fields different from computer science. A recent example is the taxonomy for computational thinking in mathematics and science proposed in Weintrop et al. (2016). This taxonomy is composed of four categories: data practices, modeling and simulation practices, computational problem solving practices, and systems thinking practices. Each category is further decomposed in a series of concrete practices. For example, the data category consists of collecting data, creating data, manipulating data, analyzing data, and visualizing data. Coincidentally, the steps for the cycle of data science process proposed by (Pfister, Blitzstein, Kaynig, 2015) are almost identical. Thus, mapping the concepts of both fields seems realistic.

Methods. 1) Annotation of data science projects (from undergraduate students) with elements of the taxonomies developed for computational thinking. 2) Restructuring of projects into instructional materials accompanied with questions that emphasize the role that computational thinking skills play in the data science cycle.

Data Sources. We have collected 70 student projects that were developed for a Data Science course at [Blinded] College. Projects revolved around datasets that were relevant to the students and for which they had domain expertise (important for the data science process). Concretely, in one project students downloaded all their emails and verified hypotheses such as: “I interact more with peers in my extra-curricular groups than the ones in my courses.” Since the email dataset was unstructured, students had to practice all steps of the data science cycle, while engaging in activities that are at the core of computational thinking, such as: sorting, filtering, cleaning, normalizing, structuring, etc.

Results. Students with more programming experience were able to create more sophisticated data structures from the unstructured dataset and ask more meaningful questions. We are annotating the use of CT practices used in these projects to make explicit how they contributed into a more meaningful data science project. They will become part of a shared repository of instructional materials.

Significance. By teaching data science through the lens of computational thinking, or considering it as one of the many domains for which computational thinking provides the toolkit to become successful, one can avoid equations such as data scientists = software engineer + statistician, which artificially create a duality where there shouldn’t be one. Every scientist is a data scientist too, and because computational thinking is now an unavoidable core scientific practice, the training of every scientist will by necessity include gaining mastery in many computational thinking practices.

Author