Individual Submission Summary
Share...

Direct link:

Inter-observer Reliability of the OLOS Observation System with Researchers and Practitioners

Fri, April 9, 4:20 to 5:50pm EDT (4:20 to 5:50pm EDT), Virtual

Abstract

An important feature of early childhood observation systems is that they can be used reliably not only by trained researchers but also by practitioners to inform classroom practices. This presentation examines the circumstances under which observers are able to provide reliable classroom observations of individual students’ classroom experiences in domains that predict PreK-3rd grade children’s academic and self-regulation outcomes. The Optimizing Learning Opportunities for Students (OLOS) classroom observation system was designed for practitioners and researchers to record the classroom experiences of individual children. Observers use tablets to record the frequency and duration (in seconds) of specific types of instruction (e.g. code-focused literacy), context of instruction (whole class, small group, individual), who the student is with (with teacher, with peers, alone), and how students participate in that instruction (e.g. verbal response to question; reading text aloud).
We examine whether both research personnel and practitioners can reliably use OLOS and what factors (e.g. training type, researcher/practitioner status, time to complete the reliability process) were predictive of reliability scores.
During the 2018-19 and 2019-20 school years, OLOS was used to observe 1,588 students in 155 PreK-3rd grade classrooms. All observations were video-recorded as well as coded live in the field-based research team members. As part of this effort, 61 coders were trained using three methods: live training by the original study team, live training with a trained coder (train the trainer model), and independent self-paced, online training. Twenty-one of the coders were researchers internal to the OLOS research team, 23 were researchers external to that team, and 17 were teacher/practitioner partners.
Inter-observer reliability was assessed for duration codes that describe instruction and frequency codes that measure frequency of child participation. During the training and certification process, reliability was assessed using coding of ten-minute clips compared to a gold standard. A score was considered passing if the coder was at least 75% accurate in both duration and frequency coding.
During field data collection, four additional types of reliability were evaluated. Traditional interrater reliability was calculated when a gold standard coder randomly selected and recoded 20-40-minutes of video previously coded by other internal team members. Live coding was evaluated when two trained coders observed the same classroom live and their results compared. Live vs. video internal coding was evaluated when a coder who coded a classroom live-coded the same classroom from video, and results compared. Live vs. video external coding was evaluated when trained coders from the OLOS research team recoded the live coding of an external OLOS-trained coder.
Descriptive data are provided in Tables 1 and 2. For initial reliability, we will report the average number of practice and test videos coded prior to reaching reliability as well as average percent agreement on practice and test videos broken down by internal team members, external research partner coders, and practitioners, and type of training received. For the other four types of reliability, average duration and frequency reliability will be reported, with detailed information about areas of difficulty. Regression analyses will explore predictors of reliability scores.

Authors