Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Visiting Washington, D.C.
Personal Schedule
Sign In
X (Twitter)
The two most widely used measures of teacher effectiveness are Value Added (VA) scores based on student test outcomes and scores from observation protocols. Prior research with both VA scores and observation protocols includes discussion of the make-up of these scores. McCaffrey, Sass, Lockwood, & Milhay (2009) identify VA scores as encompassing a persistent teacher effect, a non-persistent teacher effect, and sampling error. Similarly, Hill, Charalambos, & Kraft (2012) discuss how “true” teacher quality along with variability related to the sampling of lessons, differences among raters, and evaluation systems all play a part in observation scores. This idea of persistent and true teaching quality is an implicit assumption that there is a stable component of teaching quality evident within each teacher that does not change, i.e. there are some teachers who are consistently and naturally better at teaching than others (McCaffrey et al., 2009).
Despite the fact that the idea of persistent employee quality is taken up not only in education but in workforce productivity literature at large, there is little empirical evidence to support the idea of stable employee quality. Goldhaber & Hansen (2013) suggest that there is evidence of persistent teaching quality in VA scores, but that the variability in scores due to that quality is much smaller than most assume. Although Goldhaber & Hansen’s findings are compelling, no such research exists with observational data. In fact, little work has been done that explores what exactly we measure with observational protocols and what, if any, evidence there is to show growth in observational scores. In this dissertation, I question the assumption of persistent teaching quality and investigate a method of capturing growth in observation scores. Specifically, I address the following three research questions:
1) What does the underlying factor structure of the Framework for Teaching (FFT) and Mathematical Quality of Instruction (MQI) protocols suggest we can measure about teaching quality with observation protocols?
2). In what ways and to what extent do we see evidence of growth in teaching quality through observation scores?
3). In what ways, if any, is growth in specific elements of teaching more prominent in early, mid, or late points in teachers’ careers?
Methods
In order to answer RQ 1, I conduct an exploratory factor analysis on MQI and FFT scores in the MET data both separately and conjointly. Preliminary findings not only suggest that the MQI and FFT both provide information about a single latent trait, but that those individual traits remain distinct from one another when the data are considered conjointly. That is, general teaching quality and mathematical teaching quality are two distinct factors and it is possible for a teacher to excel at teaching generally but struggle more as a math teacher specifically and vice versa. This information is important for understanding growth in teaching quality as teachers may grow differentially along the two factors.
After establishing the underlying factor structure of the MQI and FFT, I apply Item Response Theory (IRT) in an unconventional way to answer RQ 2. Here I conceptualize elements of observation protocols as “items” and estimate the difficulty of each. I then create observation scores for teachers weighted by the IRT item parameters. Teachers then receive more “credit” for high scores on “items” that are identified as more difficult. I create these weighted observation scores for each of the 8 occasions for which teachers appear in the MET dataset and then model growth trajectories to see if there is evidence of growth across all the occasions. Finally, in order to answer RQ3, I look to see if there are differences in teaching growth trajectories on these two factors (general teaching or mathematics teaching) at different points in teachers’ careers. For example, I hypothesize that it may be more common for early career teachers to show growth on general teaching while more seasoned teachers may have plateaued in general teaching but continue to grow in mathematics teaching.
Educational Significance
This project provides evidence for our ability to gain information about general teaching effectiveness apart from mathematical teaching effectiveness. In addition to gaining this information, this work provides a methodology for using IRT analysis and growth trajectories with observational scores in order to understand growth in teaching quality. This is particularly helpful as observational scores are used in many teacher evaluation systems and districts also use evaluation information to plan professional development opportunities for their teacher workforce.