Individual Submission Summary
Share...

Direct link:

Using Machine Learning to Predict Individual Differences in Reading Achievement

Wed, April 7, 10:00 to 11:30am EDT (10:00 to 11:30am EDT), Virtual

Abstract

Introduction: After years of research showing that no single cause can explain all cases of reading disability, recent efforts to improve prediction of individual differences in reading ability have focused on creating multidimensional models that incorporate predictors from many levels of analysis. Taking a machine learning approach, the present study examines how child cognitive and behavioral characteristics and the many layers of home, neighborhood, and school environmental influences contribute to individual differences in children’s reading outcomes in order to choose a parsimonious multidimensional model to predict reading ability.

Sample: We used two samples of Florida public school students. Preliminary analyses for model building were run on a sample of 420 children who participated in Project KIDS (PKIDS), a two-component study that involved a literacy intervention and a follow-up questionnaire assessing children’s cognitive and behavioral abilities and home and neighborhood environments. Next, validation analyses will be run on a second sample from the Florida Twin Project on Reading (FTP), a cross-sequential study that combines questionnaire-based assessments of children’s behavior, environmental contexts, and achievement data. The reading outcomes used are a combination of standardized, parent-administered reading assessments collected by mail and standardized tests administered by school personnel as part of normal school attendance. For both samples we will also procure publicly-available school environmental data from the Florida Department of Education website, which will be matched to the study participants based on school names provided on their questionnaires. Due to a large age range in both samples, all reading scores are residualized on age.

Methods: For the preliminary analyses with the PKIDS sample, measures of the home environment and child characteristics were entered into a Classification and Regression Tree (CART) analysis. The data were split with 60% of the data used to train the model and 40% of the data used for prediction in order to determine the best set of predictors and prediction rules for children's reading performance. CART utilizes step-wise partitioning to create a series of if-then splits on each predictor kept in the model, which is based on a cut-point that maximizes the within-group similarities and between-group differences between two subgroups (see Figure 1). For the final analyses we will add school and neighborhood environmental predictors and use the FTP sample to cross-validate our PKIDS sample results and test the accuracy of the prediction model built with the FTP data.

Results: Results showed that the rank order of the variables kept in the model from most to least important was household income, ADHD symptoms, WM ability, Shifting ability, and parental print exposure. Household income emerged as the top-ranked predictor. ADHD symptoms interacted with various other variables and also emerged as one of the most discriminating predictors of reading. The preliminary model’s R2 = .04, showing that the combination of predictors in the model explained 4% of the variance in children’s reading scores, but more levels of analysis and environmental predictors will be added to our final analyses, and Mean Absolute Error will be used to assess the final prediction model’s accuracy.

Authors