Physical Sciences › Computer Science › Artificial Intelligence
Intelligent Tutoring Systems and Adaptive Learning
597 indexierte Paper
Intelligente Tutorensysteme und adaptives Lernen untersuchen, wie Modelle der künstlichen Intelligenz den Unterricht an die Bedürfnisse der Lernenden anpassen können. Diese Arbeiten befassen sich insbesondere mit Methoden wie der on-policy distillation, bei der ein Modell seine Antworten in Echtzeit verfeinert, oder mit Multi-Agent-Architekturen, um pädagogische Bewertungen aus Lehrmaterialien zu generieren. Ziel ist es, die Qualität der Interaktionen zu verbessern, beispielsweise durch die Anpassung des Verhaltens großer Sprachmodelle, um autonomes Denken statt bloßer Antwortvermittlung zu fördern.
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten46 % · 175 Artikel
- China27 % · 101 Artikel
- Vereinigtes Königreich7,9 % · 30 Artikel
- Japan5,3 % · 20 Artikel
- Deutschland4,2 % · 16 Artikel
- Australien3,2 % · 12 Artikel
- Sonderverwaltungsregion Hongkong3,2 % · 12 Artikel
- Südkorea2,9 % · 11 Artikel
Über 378 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 55 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- PhysicsMate: A Curriculum-Grounded Bengali Benchmark for Secondary Physics QA with Small-Model Adaptation
Rashid Azraf Jahin, Saadman Sajid, Khan Raiyan Ibne Reza, Sumaiya Tabassum Nimi · 2. Oktober 2026
Bengali secondary education lacks curriculum-grounded benchmarks for STEM question-solving, and general-purpose language models struggle with the precise terminology, unit conventions, and derivations that physics problems demand. We introduce PhysicsMate, a benchmark of 1834 question-answer pairs b…
- An Educator-Guided LLM Pedagogical Agent for Scaffolded Feedback in Conceptual Database Design
Sara Riazi, Pedram Rooshenas · 2. Oktober 2026
We present an educator-guided LLM pedagogical agent for scaffolded feedback in conceptual database design. Integrated into an entity--relationship diagram (ERD) editor, the system grounds feedback in the student artifact, assignment requirements, educator-authored rubrics, and instructional resource…
- Is Better Teacher Supervision Enough? Unlocking Student-side Learning in Multimodal On-Policy Distillation
Siyuan Liu, Kanghui Tian, Yue Duan, Yutao He, Shangdong Yang, Jian Zhang, Yinghuan Shi · 1. Oktober 2026
On-policy distillation (OPD) improves reasoning by providing token-level supervision from a teacher on a student's own trajectories. Existing methods primarily focus on enhancing this teacher-side guidance (e.g., by enriching teacher inputs and refining teacher feedback), yet we find that limited st…
- Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning
Md. Ismail Hossain, Humaira Kousar, Isidora Chara Tourni · 1. Oktober 2026
On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate the…
- From Dissonance to Orchestration: Teacher Intervention in On-Policy Distillation
Yuhao Wang, Ruiyang Ren, Yinan Zhang, Ruiqing Zhang, Jing Liu, Chunyan Miao · 1. Oktober 2026
On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student learns. Our controlled studies show that rollout quality alone is an inc…
- Learning from Think-Mode Advantage via On-Policy Distillation
Wanqi Ren, Jianxiang Wang, Danxuan Liu, Linyi Ding, Huaixiao Tou · 1. Oktober 2026
Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode. We study learning from this think-mode advantage via on-policy distillation (OPD). OPD preserves student-generated trajectories and provides dense token-level teacher targets at student-visited prefix…
- Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models
Ayan Sengupta, Vaibhav Seth, Tanmoy Chakraborty · 1. Oktober 2026
Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implicitly assuming the teacher to be a reliable oracle. In large language models (LLMs), this assumption often fails: teacher predictions can exhibit high…
- ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents
Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen · 1. Oktober 2026
Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning si…
- Less Data, Better Timing: Student-Curriculum Coupling for VLM On-Policy Distillation in Temporal Video Grounding
Jiacheng Qiu, Yunsoo Kim, Ruichen Xu, Jian Luo, Petar M. Djuri\'c, Sima Mofakham · 1. Oktober 2026
On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). However, existing pipelines typically construct the training curriculum from a fixed teacher…
- Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives
Qiwei Di, Xuheng Li, Kaixuan Ji, Chenggong Zhang, Heyang Zhao, Quanquan Gu · 1. Oktober 2026
On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers,…
- On the Off-Policy Teacher in On-Policy Distillation
Langlin Huang, Hao Liu, Mononito Goswami, Xinyu Li, Prithwith Jana, Nikos Kanakaris, Patrick Bl\"obaum, Purak Jain · 1. Oktober 2026
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the …
- Examining Variation in How Guided AI Tutors Resolve Student Impasses
Bakhtawar Ahtisham, Kirk Vanacore, Alessandra Napoli, Josh Arens, Ksenia Ionova, Clayton Cohn, Shima Salehi, Rene Kizilcec · 1. Oktober 2026
When a student is stuck, a tutor faces the assistance dilemma: help given too early can hinder productive struggle, while help withheld too long leaves the student in a frustrating, persistent impasse (i.e., wheel spinning). Generative AI tutors increasingly use guardrails restricting answer-giving,…
- Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning
Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil, Md. Faiyaz Abdullah Sayeedi, Tasnim Mohiuddin, Md Mofijul Islam · 30. September 2026
On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding…
- How Well Can LLMs Simulate Real Learner Evaluations of Educational Feedback?
Momoka Furuhashi, Kouta Nakayama, Takashi Kodama, Saku Sugawara, Kyosuke Takami · 30. September 2026
While recent studies have explored human behavior and preference simulation using large language models (LLMs), it remains unclear how well LLMs can simulate subjective evaluations from real learners in educational settings. We investigate this question using real learner evaluation data on feedback…
- Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders
Zichao Yu, Qianshuo Ye, Xu Wang, Difan Zou · 30. September 2026
On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We study this question with sparse crosscoder…
- Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation
Yongliang Miao, Shuang Liu, Yanguang Liu, Yandong Bai, Mengnan Du · 30. September 2026
On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existi…
- TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment
Zhenyu Lei, Zihan Chen, Yaochen Zhu, Shangbin Feng, Zaiyi Zheng, Ruocheng Guo, Yushun Dong, Jundong Li · 30. September 2026
Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out cha…
- Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models
Zhenyu Wang, Tianze Wang, Linjun Zhang, Yifan Hu · 30. September 2026
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at differe…
- Interactive-Policy Distillation with Bidirectional Propose-and-Verify
Shutong Wu, Xiwen Chen, Brendan Rappazzo, Daiheng Zhang, Anderson Schneider, Yuriy Nevmyvaka, Jiawei Zhang · 30. September 2026
On-policy distillation (OPD) trains a student model on its self-generated trajectories with dense token-level teacher feedback. However, naive OPD may suffer from teacher unanchoring, where the student's reasoning trajectory drifts far from the teacher, causing the teacher to be queried on states it…
- Learn Now, Use Next, Trust Later: Prequential Test-Time Learning for LLM Agents
Tong Zhao, Reed Li, Yuyang Hu, Yutao Zhu, Haijin Liang, Haibo Shi, Yu Lu, Zhicheng Dou · 30. September 2026
Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an ongoing interaction may therefore not be d…
- Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation
Jiaxuan Wang, Jiafei Lyu, Yuchen Cai, Siye Wu, Pengyuan Wang, Jiashun Liu, Xiang Cheng, Kai Yang, Yangkun Chen, Saiyong Yang, Lan-Zhe Guo · 30. September 2026
On-policy distillation (OPD) trains students using teacher feedback on their own sampled responses, yet how prompt choice shapes transfer across teacher-student pairs remains poorly understood. We systematically study prompt quantity, source, and selection across RL- and SFT-continuation pairs and c…
- Solving Without Stopping: On-Policy Distillation at Small Scale
Hongyang Li, Yiming Zhu, Xiao Li, Caesar Wu, Said Mammar, Pascal Bouvry · 30. September 2026
On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller models. We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, in thinking mode (reason at length, then …
- From Learner Behavior to Reusable Skills for Effective and Efficient Learner Simulation
Zijian Chen, Zheng Zhang, Miao Jia, Xingchen Hu, Weibo Gao, Linan Yue · 30. September 2026
Learner simulation aims to reproduce how a particular learner behaves on new tasks. Although Large Language Models (LLMs) can generate increasingly fine-grained learning behaviors, existing approaches often need to repeatedly process a growing interaction history to reconstruct the learner. This int…
- SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
Miteto Wei, Xiaohan Wang, Zehao Chen, Jiajun Chai, Sichao Liu, Li Wang, Haoyuan Xu, Zhaoyu Hu, Wei Lin, Guojun Yin · 30. September 2026
On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), whic…
- Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training
Emre Can Acikgoz, Yang Li, Zeyu Leo Liu, Srijan Bansal, Dilek Hakkani-T\"ur, Shafiq Joty, Semih Yavuz · 29. September 2026
Modern LLM post-training composes supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and on-policy distillation (OPD) into multi-stage pipelines, yet these stages are typically designed and evaluated in isolation. We show that this composition is consequential: a st…
Weitere Unterthemen aus Künstliche Intelligenz
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Large Language Models7.407 Papiere / 12 Monate+247 %
- Adversarial Robustness in Machine Learning3.552 Papiere / 12 Monate+118 %
- Reinforcement Learning in Robotics2.519 Papiere / 12 Monate+117 %
- Explainable Artificial Intelligence (XAI)2.319 Papiere / 12 Monate+200 %
- Domain Adaptation and Few-Shot Learning2.059 Papiere / 12 Monate+67 %
- Advanced Graph Neural Networks1.926 Papiere / 12 Monate+38 %
