Paper Summary
Share...

Direct link:

The Validity and Reliability of the Newark Public School's Framework for Effective Teaching

Sat, April 9, 8:15 to 9:45am, Convention Center, Floor: Level One, Room 146 A

Abstract

In order to make sound policy, curricular, or personnel decisions, it is imperative to base those decisions on data that are reliable and that support the uses for which the data were collected. This paper examines the validity and reliability of the teacher observation measure and evaluation scoring system which comprises the Newark Public Schools’ Framework for Effective Teaching (FET). The measures used by the FET include both a short and long classroom observation rubric, and mid-year and end-of-year evaluation rubric.

Validity can be defined as the degree to which evidence and theory support the use and interpretations of scores from instruments such as the FET. Reliability represents a necessary, but not sufficient component of validity, and can be defined as the extent to which scores (such as those based on the FET) are free from measurement error – that is, the extent to which ratings would be consistent if the FET were applied multiple times for a given teacher. To address the validity and reliability of the FET, the authors will examine the following:

• The reliability of the short and long observation scores, and the mid-year and end-of-year evaluation scores
• The decision consistency of the evaluation classifications (i.e., the percentage of teachers who, if re-evaluated would likely be in a different rating category – e.g., “ineffective instead of effective”)
• The extent to which a teacher’s ratings are influenced by the rater assigned to them
• The general psychometric characteristics of the observation and evaluation protocols (including dimensionality, item fit, differential item functioning, etc.)

These analyses in combination are policy relevant to both NPS and the broader education community in that they will provide evidence about the fairness of such an evaluation policy and whether differences in teachers’ effectiveness scores or classification are the result of true differences among teachers, or the result of measurement error or other sources of invalidity.

Data used for the study will be item-level observation and evaluation scores for classroom teachers in NPS during the 2013-14 school year. These data include approximately 2,895 end-of-year evaluations, 4,550 mid-year reviews, 5,023 short observations, and 5,161 long observations. The FET includes five competencies. Four of these competencies are scored for both the observations and the evaluations: lesson design and focus; rigor and inclusiveness; culture of achievement; and student progress towards mastery. The fifth competency (commitment towards personal and collective excellence) is only scored during mid-year and end-of-year evaluations.

All data will be analyzed using the Rasch partial credit model (Wright and Masters, 1982) and the many-facet Rasch model (Linacre, 1989). These models will be used to assess the general psychometric qualities of the observation and evaluation measures including reliability (along with standard Cronbach alpha indices of reliability). Bootstrap standard errors will also be created for scores calculated according to the FET system to support the decision consistency analyses.

Authors