Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Help
About Vancouver
Personal Schedule
Sign In
The availability of formative tools that assess the work of school leaders to advance student learning is growing. However, in order for evidence from these tools to be useful and trustworthy for practitioners, the validity and reliability of the evidence they produce must be established. This paper presents the results of statistical analyses that examine the validity and reliability of the Comprehensive Assessment of Leadership for Learning (CALL) formative assessment tool.
The CALL tool captures evidence of school leadership practices via a web survey that is completed by teachers and leaders within a school. The tool measures leadership practice in five domains which themselves are measured with multiple subdomain scales. This paper reports results from a pilot study of 74 schools located in 2 states. A total of 1,229 teachers and 98 leaders in these schools completed the CALL survey. This paper reports on results for teachers in the sample.
The validity of the CALL tool is assessed through four analyses. First, we examine the dimensionality of items intended to measure subdomain scales using confirmatory factor analyses. Second, we check the internal consistency in the way in which respondents report on leadership practice within a subdomain by conducting reliability analyses for all subdomain scales. A third validity check correlates the scores of subdomains within a domain with scores of subdomains in different domains. These analyses test whether our expectation that correlations of scales within a domain are stronger than correlations of scales across different domains. Since the CALL tool is designed to provide formative feedback to schools, an important piece of diagnostic information is the degree to which teachers in a school agree about leadership practice. To test this notion we examined variability of teachers’ subdomain scores within schools. Among 12 subdomain scales measured with 4 or more items we found that the reliabilities for these scales ranged from .526 to .844 with an average of .723. We further found that overall levels of agreement (as indicated by the school-level standard deviations for teacher subdomain scores) varied substantially across subdomains. In other words, some subdomains had much higher levels of within school agreement than others. Within subdomains, schools differed substantially in the level of agreement between teachers within the school. Implications of these results for instrument revision and for using the results for formative feedback are discussed.