
Teacher Observation and Reliability: Additional Insights Gathered from Inter-rater Reliability Analyses
Abstract
Using a newly created teacher evaluation instrument, Inter-rater Reliability (IRR) analyses were conducted on four teacher videos as a means to establish instrument reliability. Raters included 42 principals and assistant principals in a southern US school district. The videos used spanned the teacher quality spectrum and the IRR findings across these levels varied. Key findings suggest that while the overall IRR coefficient may be adequate to assess the validity of a classroom observation instrument, the overall coefficient may be unstable across the various teacher performance levels. Findings also strongly suggest that raters are much more likely to agree when they see high-quality teaching when compared to levels of agreement regarding low-quality teaching.
© 2019 Sally J. Zepeda, Albert M. Jimenez, published by UB ScholarWorks
This work is licensed under the Creative Commons Attribution 4.0 License.