Individual Submission Summary
Share...

Direct link:

Redesigning Research for a Modern Machine Learning World: Merging Data Science and Developmental Science

Sat, March 23, 2:30 to 4:00pm, Hilton Baltimore, Floor: Level 2, Key 4

Integrative Statement

Recent methodological advances in data acquisition and analysis are expanding the possibility to explore and interpret ever-growing behavioral and social data streams. Using three empirical studies, we illustrate how “data mining” techniques are contributing to developmental researchers’ ability to both engage with new types of data and discover new knowledge in existing data. We introduce the conceptual bases and strategies employed in a new brand of developmental research that is emerging as knowledge from data science, including techniques and algorithms for data management, unsupervised and supervised learning, pattern recognition, and feature selection, is adapted for use in studies of human development.

First, we illustrate how modern data management pipelines incorporated into a laboratory-based study of children’s (N = 200, 50% female, age 24-60 months) self-regulation (Cole & Ram, 2018) have facilitated development and validation of new tasks, efficient handling of multimodal data structures, and timely analysis and dissemination of results. In doing so we highlight how such pipelines provide for increased data quality through quick response to anomalous events, open science through transparency/reproducibility/accessibility, cost/time efficiency in resource limited settings, and faster scaling-up to larger samples.

Second, we highlight the necessity of using machine learning algorithms in study of the increasing breadth of life experiences adolescents experience through the digital screens they carry with them at nearly all times. After introducing our new technology for capturing and viewing the entirety of an individuals’ screen-based life experience – the screenome (Reeves, Ram, Robinson et al., 2018), we illustrate the impossibility of coding and interpreting the millions of screenshots we obtain from individuals smartphones every 5 seconds that the devices are on – big data – and how machine learning algorithms provide for automatic sequencing and study of the digital lives of a small sample of low-income, Mexican-American teens (from Redwood City and East Palo Alto). We highlight how the rapid pace of screen life – which ranges from 20.5 minutes to 4.7 hours per day of smartphone use spread across 15 to 320 distinct sessions that average 10-15 seconds in length each – requires that developmental scientists interested in youths’ daily lives consider new research paradigms and data analysis methods.

Third, we highlight the affordances these same machine learning algorithms provide for discovery of new knowledge in existing longitudinal data. Using 7 waves of data from the Study of Early Child Care and Youth Development (N = 966, who completed the pubertal staging module, 48% female, 76% White) we demonstrate how decision tree ensembles and feature selection techniques are used to identify correlates of pubertal development. Specifically, we identify the 20 measures in the SECCYD that are maximally predictive of the timing and tempo of sexual maturity (Tanner stages). The findings illustrate how data mining methods are being used to uncover the complex and non-linear interplay of biological and psychological processes that shape and are shaped by puberty.

Together, these three studies demonstrate how the integration of data science and developmental science can reshape our descriptions and study of human development.

Authors