Search
Browse By Day
Browse By Time
Browse By Panel
Browse By Session Type
Browse By Topic Area
Browse Posters
Search Tips
Register for SRCD23
Personal Schedule
Welcome Letter
Program Guide
Change Preferences / Time Zone
Sign In
Empirical research with adolescents is difficult, time consuming, and costly. Some obstacles include accessing generalizable samples, maintaining respectable participation and completion rates, obtaining parental consent, and participant compensation. A profuse amount of text data exists on the internet that could be mined, providing researchers a unique opportunity to access publicly available data and avoid traditional obstacles to recruiting and data collection. The challenge, however, is to systematically collect and transform the copious text into usable data. Machine learning provides a way to accomplish this with language models. Language models are a complex conditional distribution over words that can be used to generate synthetic text that is indistinguishable from human data (Argyle et al., 2021). GPT-3 is a state-of-the-art 175 billion parameter language model, the largest publicly available language model built with all the text data that exists on the internet prior to 2020. Researchers are just beginning to see the potential of using language models to create synthetic social scientific data, but the validity of the data is still unknown. To address this issue, we used GPT-3 to generate N=200 “silicon subjects” to test GPT-3’s “algorithmic fidelity” by comparing the synthetic interviews to NSYR’s (National Survey of Youth and Religion) qualitative data investigating the reasons teenagers leave religion (Argyle et al., 2021). We used the same interview questions as the original study. GPT-3 was given specific conditioning contexts (demographics) to generate predictive text similar to the original interview data. Next, we used human coding for thematic analysis to determine if GPT-3 could replicate the NSYR reported themes. The NSYR data was collected in the early 2000s and GPT-3 was trained using text on the internet up to 2020. For this reason, we anticipated that some reasons would replicate, remaining stable across datasets, while new themes would also emerge to reflect the different cohort of data. Team-based coding was conducted using NVivo software (Marks et al., 2017). GPT-3 successfully replicated 3 out of 6 of the NSYR themes from the synthetic interviews. We will discuss the convergent and divergent themes that emerged (see Table 1). We will also demonstrate the utility of GPT-3 for data coding and visualization. Finally, limitations, future directions, and best practices for using language learning models to create silicon social sciences subjects will be discussed.