Physical Sciences › Computer Science › Information Systems
Educational Technology and Assessment
15 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Últimos artículos
- In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners
Jaehoon Kim, Kwangwook Seo, Dongha Lee · 30 de septiembre de 2026
Stronger learning signals do not necessarily produce a stronger small reasoning model. Off-policy supervised fine-tuning (SFT) supplies stronger traces that may be incompatible with the student. On-policy post-training instead remains constrained by the quality of the student's own trajectories: rei…
- AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
Rishabh Agrawal, Hejie Cui, Shasha Li, Shanchan Wu, Sercan \"{O}. Ar{\i}k · 30 de septiembre de 2026
A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from …
- Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling
Bohao Wang, Xiaoyan Zhao, Yang Zhang, Jinghang Guo, Chun Chen, Can Wang, Jiawei Chen · 29 de septiembre de 2026
Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific f…
- PULSE: Identifying Demonstration-Utility Features with Sparse Autoencoders
Chenduo Hao, Chuanbao Gao, Pinjun Zeng, Jingze Zhu, Chonghan Liu, Zidong Liu, Xu Yang · 29 de septiembre de 2026
In-context learning is highly sensitive to demonstration choice, yet most methods select demonstrations using external query-demonstration similarity. Such criteria can miss model-specific signals: Similar demonstrations may activate different internal features and downstream behaviors. We introduce…
- Counterfactual Constraint-Conditioned On-Policy Distillation for Multi-Constraint Instruction Following
Yanzhao Zheng, Yuanqiang Yu, Tianze Xu, Chao Ma, Zhentao Zhang, Jihuai Zhu, Baohua Dong, Hangcheng Zhu, Ruohui Huang · 24 de septiembre de 2026
Multi-constraint instruction following requires a model to respond to a query under many simultaneously active constraints. Even strong instruction-tuned models still routinely violate some of them. Existing approaches either augment supervision with sequence- or token-level RL rewards from external…
- Score Centering Stabilizes Off-policy Reinforcement Learning
Martin Marek, Max Ryabinin · 18 de septiembre de 2026
Reinforcement learning (RL) of large language models is notoriously sensitive to small differences between training and inference engines, often referred to as the training-inference mismatch (TIM). However, completely eliminating TIM is impractical, as it would come at a major cost to rollout effic…
- Stress-testing Alignment Midtraining
Sid Baines, Jonathan Bostock, Maria Angelica Martinez, Andrew Draganov, David Africa, Daniel Tan · 18 de septiembre de 2026
When aligning frontier models through post-training techniques, it is not possible to directly demonstrate all of the behaviours we want a model to exhibit in all possible deployment environments; our model must generalise outside of the post-training distribution. One proposed solution is alignment…
- SkillLift: Learning Dense Rubrics from Sparse Oracles for Efficient Skill Evolution
Haoxiang Kang, Ming Wen · 15 de septiembre de 2026
LLM-based agents increasingly rely on persistent skills, i.e., reusable procedural prompts, to adapt without weight updates. Existing skill self-evolution methods directly revise skill text based on execution feedback, but each oracle evaluation requires a full agent rollout, creating a supervision …
- Behavior Quotient Learning for Low-Rank Adaptation of LLM Agents
Pengyang Zhou, Xiaobin Tu, Zhengxi Liu, Rongkun Xue, Haochen Li, Miancan Liu, Ziyuan Chen, Yinggui Wang, Jinkui Ren, Xiantao Zhang · 14 de septiembre de 2026
LLM-based agents rely on heterogeneous interaction capabilities to accomplish complex tasks. Existing approaches often distribute these capabilities across multiple LoRA adapters, which increases adapter storage requirements and introduces routing overhead during inference. A single LoRA avoids this…
- Inference-Time Nash Alignment
Hadi Hosseini, Debmalya Mandal, Duohan Zhang · 9 de septiembre de 2026
Preference-based fine-tuning methods such as RLHF and DPO require substantial compute and large preference datasets. They also need direct access to the model parameters which are not provided by many state-of-the art models. Inference-time alignment offers a cost-effective alternative without updat…
- When and What to Teach: Budget-Aware Online Adaptation for Web Agents
Jianwei Zhang, Sihan Cao, Pengcheng Zheng, Ya Wen, Pei Ke, Kuien Liu, Shen Gao, Wei Dong, Yang Yang, Chaoning Zhang · 9 de septiembre de 2026
Web agents have achieved significant success in automating complex internet tasks but deploying them in real-world environments requires continuous online adaptation. Given that deploying powerful proprietary models remains commercially cost-prohibitive, practitioners must rely on lightweight local …
- Online Self-Weighted Fine-Tuning
Haiquan Wen, Yiwei He, Bei Peng, Guangliang Cheng · 2 de septiembre de 2026
Standard supervised fine-tuning (SFT) assigns the same explicit loss weight to every expert demonstration, regardless of the model's changing competence over training queries. Reinforcement learning (RL) based methods adapt update strength using model-generated rollouts, but often require substantia…
- Instruction Quality Matters: Refining Instructions for Effective Preference Learning
Seohyeong Lee, Hwaran Lee, Buru Chang · 28 de agosto de 2026
Preference learning optimizes models using response pairs, yet the informativeness of these pairs is fundamentally shaped by the instructions from which they are generated. We identify instruction quality as a hidden bottleneck in preference learning: low-quality or ambiguous instructions restrict t…
- Disentangled Skill Representations for Predictive Human Modeling
Mariah Schrum, Deepak Gopinath, Srijan Srivatsa, Guy Rosman, Tiffany Chen · 26 de agosto de 2026
Understanding human skill is important for AI systems that collaborate with, coach, or assist people. Unlike typical latent variable estimation problems which rely on single observations, skill is a persistent, compositional, and behaviorally grounded construct that must be inferred from patterns ov…
- Agentic Reinforcement Learning with Self-Distilled Reward Shaping
Ranxu Zhang, Guinan Chen, Chenshaodong, Jinghao Lin, Xiaozhou Xu, Sunzhe, Yanyong Zhang, Chao Wang · 5 de agosto de 2026
Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying which intermediate decisions deserve credit. Training-only privileged skills can provide denser supervision by allowing the same frozen policy snapsh…
- An Effective Strategy for Modeling Score Ordinality and Non-uniform Intervals in Automated Speaking Assessment
Tien-Hong Lo, Szu-Yu Chen, Yao-Ting Sung, Berlin Chen · 23 de junio de 2026
A recent line of research on automated speaking assessment (ASA) has benefited from self-supervised learning (SSL) representations, which capture rich acoustic and linguistic patterns in non-native speech without underlying assumptions of feature curation. However, speech-based SSL models capture ac…
- MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation
ChangSu Choi, Hoyun Song, Dongyeon Kim, WooHyeon Jung, Minkyung Cho, Sunjin Park, NohHyeob Bae, Seona Yu, KyungTae Lim · 19 de junio de 2026
Distilling the tool-use capabilities of large language models (LLMs) into small language models (SLMs) is essential for their practical application. The predominant approach, supervised fine-tuning (SFT), suffers from poor out-of-domain (OOD) generalization due to its rigid alignment with static tea…
- JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory
Louie Hong Yao, Nicholas Jarvis, Tiffany Zhan, Saptarshi Ghosh, Linfeng Liu, Tianyu Jiang · 16 de junio de 2026
Standard LLM evaluation practices compress diverse abilities into single scores, obscuring their inherently multidimensional nature. We present JE-IRT, a geometric item-response framework that embeds both LLMs and questions in a shared space. For question embeddings, the direction encodes semantics …
- Fine-Tuning a 7B Advisor on Free-Tier GPUs: An Adapter-Handoff Recipe and a Synthetic-Data Reliability Caution
Md Millat Hosen · 16 de junio de 2026
Fine-tuning a 7B language model for specialized advising is attractive in resource-constrained settings, but multi-epoch runs routinely exceed the wall-clock limits of the free-tier GPUs (Kaggle, Colab) such users rely on. We report two things. First, a practical recipe: a three-epoch QLoRA fine-tun…
- An Efficient Learning Method to Connect Observables
Hang Yu, Takayuki Miyagi · 26 de mayo de 2026
Constructing fast and accurate surrogate models is a key ingredient for making robust predictions in many topics. We introduce a new model, the Multiparameter Eigenvalue Problem (MEP) emulator. The new method connects emulators and can make predictions directly from observables to observables. We pr…
- Test-Time Alignment via Hypothesis Reweighting
Yoonho Lee, Jonathan Williams, Henrik Marklund, Archit Sharma, Eric Mitchell, Anikait Singh, Chelsea Finn · 21 de abril de 2026
Reward models trained on aggregate preferences often fail to capture individual users' values, but existing adaptation methods such as fine-tuning or long-context conditioning are too costly for real-time personalization. We propose Hypothesis Reweighting (HyRe), which enables real-time personalizat…
- Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
Zhepeng Cen, Haolin Chen, Shiyu Wang, Zuxin Liu, Zhiwei Liu, Jielin Qiu, Ding Zhao, Silvio Savarese, Caiming Xiong, Huan Wang, Weiran Yao · 13 de abril de 2026
Large Language Models (LLMs) have achieved remarkable success through imitation learning on vast text corpora, but this paradigm creates a training-generation gap and limits robust reasoning. Reinforcement learning (RL) offers a more data-efficient solution capable of bridging this gap, yet its appl…
- GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNA
Zhichao Wang · 9 de abril de 2026
This paper proposes \textit{Group-relative Implicit Fine-Tuning (GIFT)}, a reinforcement learning framework for aligning large language models (LLMs) that unifies on-policy optimization with implicit preference learning. GIFT combines three key elements: (1) group-based sampling and normalization fr…
- Adaptive $Q$-Aid for Conditional Supervised Learning in Offline Reinforcement Learning
Jeonghye Kim, Suyoung Lee, Woojun Kim, Youngchul Sung · 16 de marzo de 2026
Offline reinforcement learning (RL) has progressed with return-conditioned supervised learning (RCSL), but its lack of stitching ability remains a limitation. We introduce $Q$-Aided Conditional Supervised Learning (QCS), which effectively combines the stability of RCSL with the stitching capability …
- Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
Linghao Zhu, Yiran Guan, Dingkang Liang, Jianzhong Ju, Zhenbo Luo, Bin Qin, Jian Luan, Yuliang Liu, Xiang Bai · 12 de febrero de 2026
Reinforcement learning (RL) has emerged as an effective post-training paradigm for enhancing the reasoning capabilities of multimodal large language model (MLLM). However, current RL pipelines often suffer from training inefficiencies caused by two underexplored issues: Advantage Collapsing, where m…
Otros asuntos del tema Sistemas de información
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Software Engineering Research853 artículos / 12 meses+218 %
- Information Retrieval and Search Behavior727 artículos / 12 meses+506 %
- Recommender Systems and Techniques600 artículos / 12 meses+154 %
- Expert finding and Q&A systems205 artículos / 12 meses+1175 %
- Information and Cyber Security169 artículos / 12 meses+1500 %
- Blockchain Technology Applications and Security146 artículos / 12 meses+650 %
