Individual Submission Summary
Share...

Direct link:

Demographic Effects on the Composition of Children’s Early Vocabulary

Sat, March 23, 4:15 to 5:45pm, Hilton Baltimore, Floor: Level 2, Key 3

Integrative Statement

Children show substantial differences in their vocabularies based on demographic circumstances. Children from higher socioeconomic status (SES) families know more words than children from lower SES families (Fernald, Marchman, & Weisleder, 2013); girls know more words than boys (Eriksson et al., 2012); and first-born children know more words than later-born children (Goldfield & Reznick, 1990). Early differences in vocabulary are in turn predictive of educational attainment in later childhood (Marchman & Fernald, 2008). But vocabularies not only differ in size — they differ in composition. For example: girls say more words, but are they generally more likely to say all words, or are there specific words that they say more or less? Differences in the composition of children’s vocabularies provide a window into how variation in environmental input creates differences in uptake.
Here, we consider the possibility that individual words carry demographic signals. We use a large-scale, cross-linguistic sample of parent-report vocabulary data from the MacArthur-Bates Communicative Development Inventory (Fenson et al., 2007) and its cross-language adaptations, archived in Wordbank (Frank et al., 2016). We assess which words are learned differentially by girls vs. boys, by first-born vs. later-born children, and by children with varying levels of maternal education.
Our dataset consists of 7-25 languages (N children 86–11,129, median=1,065). Mixed-effects logistic regression, fit separately for each language, predicts whether children produce each word, with age and a given demographic variable as predictors. Fixed effect estimates reflect the overall contribution of a demographic variable to word learning, while random effect estimates reflect for each word, the additional contribution of a demographic variable.
Figure 1 shows the fixed effects for each language and demographic factor. The effects vary in magnitude across languages but are always in the same direction and are consistent with previous work.
Figure 2 shows the distribution of random effects across words in American English as an example. Across languages, there are striking consistencies in the linkages of specific words to assigned sex at birth (e.g., sports for males and garments for females). For birth order, differences are smaller but the words for “brother” and “sister” are, unsurprisingly, much more likely for later-born, across languages. Finally, words more likely to be understood by children of higher-educated mothers are often animal-related – these may be related to an emphasis on reading animal-related books. Negatively linked words include “candy” and “soda”, likely explained by a confound between birth order and SES. Later-born children might be from lower-SES families and have more environmental exposure to “junk foods”.
Demographics like sex, birth order, and maternal education are globally related to children’s vocabulary size, but are also specifically associated with particular vocabulary items. These findings have implications for measurement and test design in that demographically-linked items should be avoided to minimize test bias. Further, demographic differences in vocabulary size may emerge from the combination of specific topical and content differences and broader dynamics in the quantity and quality of language input that children receive.

Authors