Search
Browse By Day
Browse By Time
Browse By Panel
Browse By Session Type
Browse By Topic Area
Search Tips
Register for SRCD21
Personal Schedule
Change Preferences / Time Zone
Sign In
X (Twitter)
Background: The ‘Big Data’ era has seen a rapid rise in the need for methods that can be used to analyse large corpuses of textual data from a plethora of sources (e.g., social media platforms, formal documentation, qualitative data sets, scientific literature). Traditional qualitative methods provide a well-established and robust approach for exploring meaning within textual data. However, a rare limiting factor for rigorous qualitative approaches has been the amount of data one has to analyse. Innovations such as Natural Language Processing techniques, including Latent Dirichlet Allocation (LDA) have become an increasingly popular solution for the synthesis of large amounts of textual data. LDA is a machine learning method that attempts to identify latent ‘topics’ based on patterns of co-occurring words within a set of documents in a corpus. Despite the appeal, LDA still requires researcher-directed interpretation of results, and few studies have compared the findings derived from an LDA analysis to those derived from the use of traditional qualitative methodologies. As such, it remains important to explore how results from LDA compares with different qualitative approaches, when applied to the same source of textual data.
Objective: This project aims to investigate how LDA complements and differs from qualitative methodologies when used on a large textual dataset derived from scientific journal articles on exploratory behaviour in children.
Methods: We used a pre-existing data set from a study investigating the different ways in which scientists understand exploratory behaviour in childhood (0 – 11.9 years). This data set comprised 409 text files containing paragraphs from the introduction sections of journal articles which contained the word “explor*”. Two independent analyses were conducted simultaneously; 1) textual data was pre-processed and analysed using LDA, and 2) data in its original form were analysed using template thematic analysis. Independent groups of researchers were used to conduct each part, with researchers blind to the findings from the other group. Findings derived from each approach were then compared and contrasted.
Results and Conclusions: As expected, the LDA analysis was quicker and required less researcher intensive time than the qualitative approach. Whilst there was some overlap in major topics and themes identified (e.g., exploration in specific areas of child functioning), the LDA approach was found to have some irrelevant topics and topics difficult to interpret. By contrast, the themes identified in the qualitative approach were more clearly defined and had clear illustration with supporting quotes. Despite their differences, the findings suggest a complementary approach whereby LDA could be used to derive a close approximation within a large corpus of text, with traditional qualitative approaches used to derive refined meaning from LDA results.
Lauryn Hagg, Deakin University
Presenting Author
Anna Klas, Deakin University
Non-Presenting Author
Subhadra Evans, Deakin University
Non-Presenting Author
Christopher Greenwood, Deakin University
Non-Presenting Author
Elizabeth Mary Westrupp, Deakin University
Non-Presenting Author
Matthew Fuller-Tyszkiewicz, Deakin University
Non-Presenting Author
Stephanie Merkouris, Deakin University
Non-Presenting Author
George Youssef, Deakin University
Non-Presenting Author