Beyond Pedagogical Principles: Multi-Horizon Preference Optimization for Efficient Socratic Tutoring
Xin Shi, Chao Zhang, Yifan Zhu, Xueqiao Zhang, Yawei Luo
Zhejiang University
内容与影响
The development of LLM-based tutor agents faces challenges in simultaneously ensuring adherence to pedagogical principles and achieving optimal pedagogical effectiveness, particularly in dynamic, multi-turn interactions.Existing methods are often constrained by static data or sparse reward signals in online settings.To address this gap, we propose Multi-Horizon Preference Optimization (MHPO), a novel framework that iteratively refines tutor agents using a multi-horizon reward function within a dynamic teacher-student simulation environment.Specifically, this reward function is designed to capture both turn-level pedagogical quality and trajectory-level pedagogical effectiveness, which is estimated via Monte Carlo rollouts.We further investigate two distinct strategies to aggregate these rewards for policy optimization.Our experiments demonstrate that MHPO significantly enhances base model performance, achieving a superior balance between principles and effectiveness compared to various baselines.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
回答优先基于摘要、文献信息与可获取全文;依据不足时会明确说明。
学术脉络
学科主题
计算机 / AIIntelligent Tutoring Systems and Adaptive Learning
Innovative Teaching and Learning Methods · Multi-Agent Systems and Negotiation