Validating automated assessments of teaching effectiveness using multimodal data
Tim Fütterer, Ruikun Hou, Babette Bühler, Efe Bozkir, Courtney Bell, Enkelejda Kasneci, Peter Gerjets, Ulrich Trautwein
University of Tübingen Technical University of Munich University of Wisconsin–Madison Leibniz-Institut für Wissensmedien
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
For enhancing student learning in classrooms, high-quality teaching is essential. Research highlighted core dimensions of effective teaching, including classroom management, student support, and cognitive activation. However, traditional methods of assessing teaching effectiveness dimensions (e.g., student surveys) have limitations, including rating biases and resource intensiveness. To overcome these challenges, we explored machine learning (ML) algorithms for the automated assessment of teaching effectiveness. The study analyzed multimodal data—such as video, audio, and transcripts—from the Global Teaching Insights study, which included video recordings and transcripts from 46 teachers and 1,132 students in Germany. Scores for 18 teaching effectiveness subdimensions from three core dimensions were automatically generated by training attention-based ML models on multimodal features extracted from pretrained encoders. These ML-generated scores were compared with scores provided by human experts. A content validity study was conducted, where human experts evaluated the plausibility of ML-generated scores against human-generated scores. Structural equation models were used to assess the relationship between teaching effectiveness subdimensions and students’ tested achievement. ML-generated scores were more reliable for some subdimensions (e.g., nature of discourse), and they were also plausible and content valid. ML-generated scores achieved higher absolute accuracy than human scores in 11 of 18 subdimensions. Limitations include reliance on human ratings as ground truth and inconsistent predictive validity, underscoring the need for refined models to generate actionable insights, such as real-time feedback systems. The findings provide valuable insights for the development of automated feedback, enhancing the practical application of teaching effectiveness assessments. • First study to assess many teaching effectiveness subdimensions multimodally with machine learning (ML; text, audio, video). • ML-generated scores matched or exceeded human reliability in some subdimensions. • ML-generated scores were often rated as plausible by expert human raters. • The predictive validity of ML scores varied, highlighting modeling and study design challenges. • AI approaches show promise for scalable teaching assessment but need further refinement.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIEducational Technology and Pedagogy
Student Assessment and Feedback · Evaluation of Teaching Practices
参考文献 42
此处列出前 3 条
引用本文 3
按被引量排序,此处列出前 3 条