Physical Sciences › Computer Science › Artificial Intelligence
Domain Adaptation and Few-Shot Learning
2,059 papers indexed
Adapting artificial intelligence models to new contexts or scarce data remains a core challenge in the field. Research explores how to adjust algorithms trained on one dataset so they maintain performance on others - often very different - without requiring a large volume of additional examples. Between domain adaptation, which aims to reduce the gap between distinct data distributions, and few-shot learning, which seeks to learn from very few samples, these studies address questions such as preserving acquired knowledge, balancing model plasticity and stability, or optimizing internal representations for diverse tasks.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- China38% · 530 papers
- United States33% · 461 papers
- United Kingdom6.5% · 92 papers
- Canada5.9% · 83 papers
- South Korea5.7% · 80 papers
- Germany4.8% · 68 papers
- India4.6% · 65 papers
- Japan3.9% · 55 papers
Across 1,413 papers on this subject with at least one lab located. 70 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- AIBL: Augmented Instance-Based Learning with Structured Memory and Neural Embeddings
Radha Poovendran, Andrea Stocco, Linda Bushnell · 5 October 2026
Sequential learning systems often make decisions from accumulated experience while receiving high-dimensional inputs whose distribution may change over time. Instance-Based Learning Theory (IBLT) provides a principled case-based framework for such settings through stored situation-decision-utility i…
- Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models
Zhenghao Zhao, Gaowen Liu, Zhiling Lan, Yan Yan · 5 October 2026
Data selection is already a central bottleneck in large-language-model training, where web-scale corpora are noisy and token budgets are finite. In continual pre-training (CPT), it becomes a forgetting-control problem: a poorly chosen target-domain corpus can overwrite capabilities encoded in the pr…
- BaCP: Backbone Contrastive Pruning for Preserving Representations in Extremely Sparse Neural Networks
Mohammad Haroon Khawaja, Muhammad Haseeb, Mohammad Fatim Shoaib, Muhammad Tahir · 5 October 2026
Unstructured pruning at extreme sparsity often suffers from representational collapse, causing sharp drops in accuracy. To address this, we study Backbone Contrastive Pruning (BaCP), which regularizes the sparse network's embedding space by aligning it with pretrained, fine-tuned, and historical sna…
- MuLoRA: Spectrally Balanced Low-Rank Adaptation for Continual Learning
Junkang Liu · 5 October 2026
Low-rank adaptation (LoRA) provides a parameter-efficient approach to continual learning, but its nominal rank can conceal a loss of effective adaptation capacity. We identify \emph{spectral plasticity collapse}: during sequential adaptation, update energy becomes concentrated in a small subset of s…
- CLEAN: Psychometrically Consistent Incremental Cognitive Diagnosis under Concept-Space Expansion via Architectural Isolation
Tao He, Jinxing Xiang, Fan Jiang · 5 October 2026
Cognitive diagnosis (CD) is a fundamental task in intelligent education that profiles learner proficiency over knowledge concepts. In real-world learning platforms, newly added items continually introduce previously unseen concepts, necessitating dynamic expansion of the underlying concept space. Ye…
- Platonic Task Arithmetic
Junghwan Park, Woojin Cho · 5 October 2026
Distinct pre-trained models specialized for the same task converge to closely similar behavior, yet the parameter updates that produce it share no common coordinate system. Weight-space task arithmetic is therefore confined to a single model, and transporting an update between models requires a stru…
- GRAFT: Growing Agglomerative Foundation Models via Continual Teacher Distillation
Zhenghao Zhao, Chi Zhang, Qingshuang Chen, Yelin Kim · 5 October 2026
Vision foundation models such as DINOv2, SigLIP2, and MASt3R develop complementary capabilities from different pretraining objectives, yet their knowledge remains distributed across separate, specialized models. Multi-teacher knowledge distillation offers a path toward consolidating these capabiliti…
- What Should World Models Forget? Stratified Retention for Continual Adaptation
Nishit Anand, Ramani Duraiswami, Dinesh Manocha · 5 October 2026
Continual learning treats degradation on previously seen data as evidence of failure, a convention inherited from settings with a stationary prediction target, where a correct label remains correct indefinitely. World models do not satisfy this condition. Their prediction target is the environment, …
- Measure Less, Know More: Self-Supervised Test-Time Feature Acquisition
Eeshaan Jain, Linus Bleistein, Bart Deplancke, Charlotte Bunne · 5 October 2026
Recent progress in multimodal, high-dimensional learning has enabled foundation models to process heterogeneous, large-scale data. However, at test time, acquiring all features or modalities can be prohibitively costly and often redundant. Sequentially selecting informative modalities is therefore c…
- How to Find and Reuse Policies for Continuous Adaptation in Lifelong Reinforcement Learning
Saptarshi Nath, Inish M. D'Souza, Antonio Carta, Soheil Kolouri, Andrea Soltoggio · 5 October 2026
In lifelong reinforcement learning, retaining previously learned policies is not sufficient for effective transfer to a new task. Useful knowledge may be distributed across several prior policies, and its relevance may change as the learner acquires experience. One hypothesis is that task similarity…
- Learning Style, Forgetting Semantics: A Case Study of SFT and RFT on Classification Tasks
Haodong Liang, Yanhao Jin, Krishnakumar Balasubramanian, Lifeng Lai · 5 October 2026
Why does supervised fine-tuning (SFT) lead to more forgetting than reinforcement fine-tuning (RFT), even when all teacher demonstrations are semantically correct? We study this question on classification tasks where tokens within each semantic class express the same semantic answer in different styl…
- Slow-Fast Multi-Teacher On-Policy Distillation for Capability Preservation
Xiaofei Yin, Tong Chu, Jiyuan Fu, Jun Lan, Shuheng Zhou, Huijia Zhu · 5 October 2026
Foundation multimodal large language models are designed to support a broad spectrum of capabilities across diverse domains. Multi-teacher on-policy distillation (MOPD) provides an effective framework for consolidating domain-specific expertise into a single student model. However, MOPD training gra…
- Safeguarding Mutual Correction in Source-Free Domain Adaptation via Cut Statistics
Seongjun Lee, Changhee Lee · 5 October 2026
Source-Free Domain Adaptation (SFDA) aims to adapt a source-pretrained model to an unlabeled target domain without access to the original source domain. While early single-model approaches rely on self-refinement, they are inherently susceptible to confirmation bias and struggle to correct their own…
- Task-Oriented Rank Adaptation for Continual Learning in Text Classification
Rey Sanchez Lopez, Eduardo Morales Manzanares, Hugo Jair Escalante · 2 October 2026
Continual learning (CL) in text classification faces two critical challenges: catastrophic forgetting and negative transfer across sequential tasks. Parameter-Efficient Fine-Tuning (PEFT) methods such as LoRA enable efficient adaptation by learning low-rank updates of the model parameters. However, …
- Local Support Learning
Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes · 2 October 2026
We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this …
- Capturing In-Context Learning Dynamics with Task Operators
Guangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha, Wenqian Ye, Aidong Zhang · 2 October 2026
In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requires processing the full set of examples, resulting in inefficient deployments, and how ICL works mechanistically is not fully understood. Prior work co…
- Rethinking Data Augmentation under Covariate Shift: Invariant-Guided Diffusion and Prototype Reweighting
Hongyu Cao, Xinyuan Wang, Arun Vignesh Malarkkan, Kunpeng Liu, Haifeng Chen, Yanjie Fu · 2 October 2026
In many industrial applications, 1) tabular data is scarce and imbalanced and thus requires synthetic expansion; 2) input distributions drift between training and deployment (covariate shift); 3) validation sets often diverge from unseen test environments; or 4) standard generative models simply mim…
- On-the-fly Weight Generation: A Hypernetwork Proof of Concept on ARC-1D
Fabio J. Fehr, Philip Torr · 2 October 2026
General-purpose models can adapt to many tasks from context, while specialised models can execute individual functions with less capacity. Yet obtaining such specialists requires task-specific training or adaptation. We ask whether they can instead be generated directly from a few demonstrations. Us…
- Restoring without Forgetting: Filter-Level Continual Image Restoration via Parameter-Space Integrated Gradients
Xin Feng, Jin Zhao, Yizhen Zhang, Wenjie Pei, Fanglin Chen, Guangming Lu · 1 October 2026
Adapting image restoration models to a stream of new tasks without revisiting past data remains challenging due to catastrophic forgetting. In this work, we propose Restoring without Forgetting (RwF), a filter-level continual adaptation framework for image restoration built upon a critical observati…
- Orthogonal Yet Coupled: Decoupling Geometric Components for Model Merging
Zijing Wang, Yongkang Liu, Mingyang Wang, Ercong Nie, Mengjie Zhao, Yunpu Ma, Kang Liu, Zihan Wang, Shi Feng, Daling Wang, Hinrich Sch\"utze · 1 October 2026
Merging pretrained models has emerged as an effective approach for consolidating diverse capabilities into a single unified model. However, prevailing merging methods typically treat each task vector as an indivisible merging unit, overlooking the heterogeneous geometric changes encoded within it. T…
- Unmerge: Efficient Machine Unlearning via Task Arithmetic
Haoran Tang, Andrew Tan, Rajiv Khanna · 1 October 2026
Approximate machine unlearning seeks to remove the influence of a forget set from a trained model without full retraining. Existing gradient-based methods require data-dependent hyperparameter search, struggle when forget and retain knowledge are entangled, and offer little insight into where unlear…
- Disentangling Self-Distillation: Measuring and Modeling Acquisition and Retention
Luis Zuin, Alexis Huet, Dario Rossi, Zied Ben Houidi · 1 October 2026
Self-distillation with privileged context adapts a language model from demonstrations by letting the model, once conditioned on a reference response, teach its context-free copy token by token. Our taxonomy reveals existing methods differ along three entangled axes: (i) the rollout source (student o…
- Learning What to Forget: Distributional Unlearning for LLM Representation Spaces
Pinaki Mohanty, Haoran Tang, Maggie Makar, Rajiv Khanna · 1 October 2026
Machine learning systems increasingly face the need to remove the influence of entire data domains, such as toxic language, harmful behavior, or topical content, rather than isolated records. Recent work formalizes this problem as \emph{distributional unlearning}: selecting a subset of a forget doma…
- Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features
Cunchun Li, Haonan He, Yifan Gao, Minglei Li, Jingqi Ye, Qingyu Yang, Peng Ye · 30 September 2026
Supervised fine-tuning (SFT) learns most aggressively from tokens that the model deems least likely. This helps acquire new behaviors, but also amplifies noisy or conflicting supervision and can overwrite useful pretrained knowledge. Through a unified policy-loss view, we revisit existing token-rewe…
- Width Expansion as a Method for Class Incremental Learning
A. L. S. Conde, Y. Elkhatib, C. M. Ranieri · 30 September 2026
Class Incremental Learning (Class-IL) requires models to learn new classes over time while preserving previously acquired knowledge without access to past data or task identity. This setting intensifies the stability-plasticity dilemma and makes catastrophic forgetting a central challenge. Existing …
Other topics in Artificial intelligence
The topics the OpenAlex classification attaches to the same theme, most active first.
- Large Language Models7,407 papers / 12 months+247%
- Adversarial Robustness in Machine Learning3,552 papers / 12 months+118%
- Reinforcement Learning in Robotics2,519 papers / 12 months+117%
- Explainable Artificial Intelligence (XAI)2,319 papers / 12 months+200%
- Advanced Graph Neural Networks1,926 papers / 12 months+38%
- Natural Language Processing Techniques1,595 papers / 12 months+69%
