Search
Program Calendar
Browse By Day
Browse By Time
Browse By Panel
Browse By Session Type
Browse By Topic Area
Search Tips
Virtual Exhibit Hall
Personal Schedule
Sign In
X (Twitter)
Collecting data from human subjects can be difficult, particularly when those subjects are preverbal infants. This problem has led to a field of small sample sizes (Oakes, 2017). Compounding the problem, many procedures used with infants often yield few data points per subject, resulting in sparse data. As a result, infant researchers need to recruit and test more infants to increase statistical power. Recent technological advances have helped this problem somewhat by making it possible to increase the number of observations collected from each infant. However, we have yet to fully understand the consequences of collecting more data from infant subjects, nor have we established best practices to determine how to make decisions about the trade-offs between increasing the number of infants tested and increasing the number of data points collected from each infant.
Thus, there are two solutions to address the problem of underpowered studies: 1) test more infants, and 2) collect more data from each infant. How do we know when to do each? In previous work, we aimed to characterize both intra-individual and inter-individual variability in datasets collected from looking time procedures. Our datasets varied in many ways (e.g., number of trials per infant, length of trials) yet we observed high variability within infants with little differences in performance between infants. For example, any two trials from one infant were no more similar to each other than two trials from another infant. This surprising finding has implications for how we address the problem of low power in infancy research, as well as the conclusions we draw from our observations about both group and individual differences in development.
The goal of the current research is to understand the impact of within- and between-infant variability on statistical power. Our previous work suggested that when the data do not capture individual differences, or individual differences are not present in the sample, we can increase power to detect a true effect by increasing the number of trials per infant. Although power can be improved by increasing the sample size, under some circumstances it may be more efficient to increase the number of observations per subject.
Using a collaborative approach, we extend this finding by characterizing datasets from other labs to identify contexts that yield reliable (lower within-infant variability) measures of infant looking behavior. We ask whether specific looking-time tasks are better suited to measure between-infant differences and if so, what is the effect of number of trials on power to detect these individual differences. Using simulations, we study the effects variability within and between infants; our variability estimates were obtained from diverse datasets collected in different labs from infants of different ages, in different tasks (i.e., visual paired comparison, looking-while-listening), and obtained using different procedures (i.e., eye-tracking, behavioral coding). With the goal of maximizing power, we use simulations to demonstrate the impact of these different sources of variability with the aim to better inform research design and build upon best practices for infancy research.