Physical Sciences › Computer Science › Artificial Intelligence
Intelligent Tutoring Systems and Adaptive Learning
597 papers indexed
Intelligent tutoring systems and adaptive learning explore how artificial intelligence models can personalize education based on learners' needs. This research focuses in particular on methods such as on-policy distillation, where a model refines its responses in real time, or on multi-agent architectures to generate pedagogical assessments from course materials. The goal is to enhance the quality of interactions, for instance by adjusting the behavior of large language models to encourage autonomous reasoning rather than mere answer delivery.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States46% · 175 papers
- China27% · 101 papers
- United Kingdom7.9% · 30 papers
- Japan5.3% · 20 papers
- Germany4.2% · 16 papers
- Hong Kong SAR China3.2% · 12 papers
- Australia3.2% · 12 papers
- South Korea2.9% · 11 papers
Across 378 papers on this subject with at least one lab located. 55 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- Latent-MOPD: Latent Multi-Teacher On-Policy Distillation
Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li, Koyoshi Shindo, Brandon Joseph Lwowski, Jing Li · 5 October 2026
On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrate…
- Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation
Rui Li, Liyang He, Zheng Zhang, Zhenya Huang, Linbo Zhu, Qi Liu · 5 October 2026
On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to …
- Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation
Chenglei Shen, Haoyang Yao, Weijie Yu, Song Jin, Xiao Zhang, Jun Xu · 5 October 2026
On-policy self-distillation (OPSD) uses reference solutions as privileged hindsight to supervise student-generated reasoning trajectories. However, reference-based guidance may explain a correct solution without addressing why the student's own reasoning fails. This reasoning mismatch between the gu…
- Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective
Han Cui, Jianhao Yan, Yun Luo, Hongbo Zhang, Zhizhang Fu, Yue Zhang · 5 October 2026
On-policy distillation (OPD) has become an important approach to language model post-training. However, despite its performance gains, OPD can also collapse into excessively long and repetitive generation, and the mechanism underlying these divergent outcomes remains poorly understood. We explain th…
- Learning to Revise Reasoning with Segment-wise On-Policy Distillation
Yuxiang Zhang, Ding Cao, Shuting Cui, Lei Wang, Weijieying Ren, Tianxiang Zhao · 5 October 2026
On-policy distillation (OPD) improves large language model reasoning by training students on their own rollouts with dense token-wise supervision from the teacher. However, token-wise OPD does not explicitly provide a coherent alternative reasoning step showing how the student's step could be revise…
- Spend Teacher Tokens Where They Matter: Success-Referenced On-Policy Distillation
Xiang Chen, Futao Su, Kong Wang, Jiayi Chen, TanLin Li · 5 October 2026
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher, but providing such supervision for every rollout requires substantial teacher computation. We introduce Success-Referenced On-Policy Distillation (SR-OPD), which reduces this cost by s…
- On-Premises Multi-Course RAG Tutoring for Business Education: Hardware-Software Trade-offs in a Campus AI Tutor
Sidney Shapiro, Joshua Lindemann · 5 October 2026
Campus AI tutors based on retrieval-augmented generation (RAG) must ground answers in assigned course materials while keeping textbooks and student dialogue on institutional infrastructure. We present CourseChat, an on-premises, multi-course RAG tutor for undergraduate business education, deployed b…
- PhysicsMate: A Curriculum-Grounded Bengali Benchmark for Secondary Physics QA with Small-Model Adaptation
Rashid Azraf Jahin, Saadman Sajid, Khan Raiyan Ibne Reza, Sumaiya Tabassum Nimi · 2 October 2026
Bengali secondary education lacks curriculum-grounded benchmarks for STEM question-solving, and general-purpose language models struggle with the precise terminology, unit conventions, and derivations that physics problems demand. We introduce PhysicsMate, a benchmark of 1834 question-answer pairs b…
- An Educator-Guided LLM Pedagogical Agent for Scaffolded Feedback in Conceptual Database Design
Sara Riazi, Pedram Rooshenas · 2 October 2026
We present an educator-guided LLM pedagogical agent for scaffolded feedback in conceptual database design. Integrated into an entity--relationship diagram (ERD) editor, the system grounds feedback in the student artifact, assignment requirements, educator-authored rubrics, and instructional resource…
- Is Better Teacher Supervision Enough? Unlocking Student-side Learning in Multimodal On-Policy Distillation
Siyuan Liu, Kanghui Tian, Yue Duan, Yutao He, Shangdong Yang, Jian Zhang, Yinghuan Shi · 1 October 2026
On-policy distillation (OPD) improves reasoning by providing token-level supervision from a teacher on a student's own trajectories. Existing methods primarily focus on enhancing this teacher-side guidance (e.g., by enriching teacher inputs and refining teacher feedback), yet we find that limited st…
- Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning
Md. Ismail Hossain, Humaira Kousar, Isidora Chara Tourni · 1 October 2026
On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate the…
- From Dissonance to Orchestration: Teacher Intervention in On-Policy Distillation
Yuhao Wang, Ruiyang Ren, Yinan Zhang, Ruiqing Zhang, Jing Liu, Chunyan Miao · 1 October 2026
On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student learns. Our controlled studies show that rollout quality alone is an inc…
- Learning from Think-Mode Advantage via On-Policy Distillation
Wanqi Ren, Jianxiang Wang, Danxuan Liu, Linyi Ding, Huaixiao Tou · 1 October 2026
Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode. We study learning from this think-mode advantage via on-policy distillation (OPD). OPD preserves student-generated trajectories and provides dense token-level teacher targets at student-visited prefix…
- Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models
Ayan Sengupta, Vaibhav Seth, Tanmoy Chakraborty · 1 October 2026
Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implicitly assuming the teacher to be a reliable oracle. In large language models (LLMs), this assumption often fails: teacher predictions can exhibit high…
- ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents
Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen · 1 October 2026
Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning si…
- Less Data, Better Timing: Student-Curriculum Coupling for VLM On-Policy Distillation in Temporal Video Grounding
Jiacheng Qiu, Yunsoo Kim, Ruichen Xu, Jian Luo, Petar M. Djuri\'c, Sima Mofakham · 1 October 2026
On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). However, existing pipelines typically construct the training curriculum from a fixed teacher…
- Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives
Qiwei Di, Xuheng Li, Kaixuan Ji, Chenggong Zhang, Heyang Zhao, Quanquan Gu · 1 October 2026
On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers,…
- On the Off-Policy Teacher in On-Policy Distillation
Langlin Huang, Hao Liu, Mononito Goswami, Xinyu Li, Prithwith Jana, Nikos Kanakaris, Patrick Bl\"obaum, Purak Jain · 1 October 2026
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the …
- Examining Variation in How Guided AI Tutors Resolve Student Impasses
Bakhtawar Ahtisham, Kirk Vanacore, Alessandra Napoli, Josh Arens, Ksenia Ionova, Clayton Cohn, Shima Salehi, Rene Kizilcec · 1 October 2026
When a student is stuck, a tutor faces the assistance dilemma: help given too early can hinder productive struggle, while help withheld too long leaves the student in a frustrating, persistent impasse (i.e., wheel spinning). Generative AI tutors increasingly use guardrails restricting answer-giving,…
- Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning
Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil, Md. Faiyaz Abdullah Sayeedi, Tasnim Mohiuddin, Md Mofijul Islam · 30 September 2026
On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding…
- How Well Can LLMs Simulate Real Learner Evaluations of Educational Feedback?
Momoka Furuhashi, Kouta Nakayama, Takashi Kodama, Saku Sugawara, Kyosuke Takami · 30 September 2026
While recent studies have explored human behavior and preference simulation using large language models (LLMs), it remains unclear how well LLMs can simulate subjective evaluations from real learners in educational settings. We investigate this question using real learner evaluation data on feedback…
- Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders
Zichao Yu, Qianshuo Ye, Xu Wang, Difan Zou · 30 September 2026
On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We study this question with sparse crosscoder…
- Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation
Yongliang Miao, Shuang Liu, Yanguang Liu, Yandong Bai, Mengnan Du · 30 September 2026
On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existi…
- TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment
Zhenyu Lei, Zihan Chen, Yaochen Zhu, Shangbin Feng, Zaiyi Zheng, Ruocheng Guo, Yushun Dong, Jundong Li · 30 September 2026
Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out cha…
- Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models
Zhenyu Wang, Tianze Wang, Linjun Zhang, Yifan Hu · 30 September 2026
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at differe…
Other topics in Artificial intelligence
The topics the OpenAlex classification attaches to the same theme, most active first.
- Large Language Models7,407 papers / 12 months+247%
- Adversarial Robustness in Machine Learning3,552 papers / 12 months+118%
- Reinforcement Learning in Robotics2,519 papers / 12 months+117%
- Explainable Artificial Intelligence (XAI)2,319 papers / 12 months+200%
- Domain Adaptation and Few-Shot Learning2,059 papers / 12 months+67%
- Advanced Graph Neural Networks1,926 papers / 12 months+38%
