Search
Browse By Day
Browse By Time
Browse By Person
Browse By Committee or SIG
Browse By Session Type
Browse By Keywords
Browse By Geographic Descriptor
Search Tips
Personal Schedule
Change Preferences / Time Zone
Sign In
The purpose of this roundtable presentation is to describe what synthetic data is, why educational researchers might choose to create a synthetic dataset, how researchers go about doing so, and how other researchers can evaluate or critique research based on synthetic data. This presentation addresses the call to see technologies “as bridges rather than barriers” by discussing how synthetic data offers a potential collaborative tool for more inclusive research across contexts. The poster will address technical concerns about synthetic data, summarized below, while also aiming to educate new researchers about the issues of data bias and surveillance in both synthetic and real-world datasets, ultimately contributing to more equitable futures.
Why use synthetic data
Synthetic data is “generated using a purpose-built mathematical model or algorithm, with the aim of solving a (set of) data science task(s)” (Jordan et al, 2022, p. 5). Since its initial development in the 1990s, the use of synthetic data has been used largely by business and public policy areas (Berg et al., 2016), but there has been more recent interest in social, behavioral, and health sciences (Alloza et al., 2023; Grund et al., 2024; Ping et al., 2017). The current rise in interest for synthetic data is driven by two trends: the impact of privacy regulations and the growth of machine learning (Flanagan et al., 2022; Gu et al., 2025; Jordan et al, 2022).
The most probable reason for employing synthetic data in educational research is to address privacy concerns (Alloza et al., 2023; Jordan et al., 2022; Lui et al., 2025). Synthetic data offers an advantage to educational researchers as sharing real-world educational data often violates privacy regulations, and de-identification can be time-consuming and imperfect (Alloza et al., 2023; Howe et al., 2017). Synthetic data created to mimic a real-world dataset would allow for easier collaboration between data holders and educational researchers.
Creating and validating synthetic data
There are several techniques to create a synthetic dataset, with most employing some type of machine learning algorithm (Qualtrics, 2025). The most complex and comprehensive method is a fine-tuned foundational model in which a model is specifically trained using large-scale, real-world data. In one example of a data synthesizer tool (Ping et al., 2017), Bayesian networks are calculated to mimic correlations between variables, with cases created by the model to construct a new dataset with similar properties to the original.
Current synthetic datasets, however, are imperfect mimics of real-world data and complexity (Grund et al., 2024). Additionally, educational researchers need to be cautious, as the creation of a synthetic dataset does not guarantee privacy (Ivanova & Petrova, 2023; Kesteren, 2024; Lui et al., 2025). Data can be de-identified if the process does not completely anonymize rare events. Methods to support the correspondence between synthetic and real-world data include: holding back data in training; checks of differential prediction of subgroups; and calculation of measurements, such as means comparisons, distribution metrics, and data structure correlations (Alloza et al., 2023; Grund et al., 2024; Lui et al., 2025).