Search
On-Site Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Housing and Travel
Sign In
X (Twitter)
Objectives
Educational research has observed that (i) the progress of academic education and (ii) whether the object of information search and acquisition is a domain-specific topic or not both influence the information-seeking behavior of users of online media ([21]; [8]). Hence, different kinds of media and their domain-dependence are expected to effect students’ skill to critically use online information, hereafter called Critical Online Reasoning (COR; for details, see [3]), whereby a differentiation is made between (meta)cognitive requirements for COR in generic (GEN-COR) and domain-specific (DOM-COR) contexts — mainly based on the topic and degree of domain-specificity of task scenarios. Thereby, the question arises, whether COR skills developed during academic education for specific domains (medicine, law, teaching) are transferrable to generic tasks. We approach this question from a linguistic perspective: solving a COR task involves content evaluation of the queried source, which naturally draws on linguistic features. To date it is under-researched whether information sources regarded as potentially high-quality exhibit specific language-related commonalities, and if so, what these are ([7]). We approach this research desideratum using methods from computational linguistics, applied to experimental COR data.
Framework
Our approach is based on recent advances in computational linguistic assessments of educational media and materials. Authors ([1]) have previously established a correlation between linguistic features and the successful completion of educational assignments. Online resources were shown to exhibit a language-dependent “range” ([2]). Especially Wikipedia is ranked as ‘high-quality’ encyclopedic source regarding linguistic features and knowledge graphs ([5]). Such ‘computational educational linguistics’ is a (still emerging) corpus-based and algorithm-driven framework for content analysis, which is applied here for COR analysis.
Method
We analyze two collections of online texts from a study in which young professionals from three domains (medicine, law, and teaching) solved both GEN-COR (67 participants) and DOM-COR (47 participants) tasks (Table 1 in Appendix 2; for details, see [5]). Both sets of texts are pre-processed by means of state-of-the-art Natural Language Processing algorithms and investigated by means of sets of linguistic features including syntactic structure, readability, and content similarity. The computational classification is accompanied by a qualitative categorization of online texts using a modified scheme by Authors ([4]) (Figure 1, Appendix 2).
Results
Preliminary results suggest that online texts can be clustered for quality according to language-related features. However, clustering does not seem to coincide with manual categorization: the difference between GEN-COR and DOM-COR text types shown in Figure 1 cannot be reproduced on the level of texts, indicating that computational educational linguistic assessments should complement formal text type categorization — the ordering of text types according to quality using standard methods from computational methods departs from the point of view of DOM-COR and GEN-COR data. The paper will quantify the difference between those classifications.
Conclusion
We provide a computational linguistics foundation for the content analysis of online texts from GEN-COR and DOM-COR data, indicating that online texts can be clustered for quality in terms of language-related features. In future work, analyses are extended by incorporating procedural measurements such as dwell time ([3]).