Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
2,542 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume — last 12 months
Latest papers
- Prediction horizon shapes representations in predictive learning
Aviv Ratzon, Omri Barak · 6 May 2026
Predictive learning has emerged as a central paradigm for training models across diverse data domains and is increasingly viewed as a foundation for modern artificial intelligence. A common intuition for this success is that accurate prediction requires models to capture the underlying dynamics of t…
- Flow Sampling: Learning to Sample from Unnormalized Densities via Denoising Conditional Processes
Aaron Havens, Brian Karrer, Neta Shaul · 6 May 2026
Sampling from unnormalized densities is analogous to the generative modeling problem, but the target distribution is defined by a known energy function instead of data samples. Because evaluating the energy function is often costly, a primary challenge is to learn an efficient sampler. We introduce …
- A Few-Step Generative Model on Cumulative Flow Maps
Zhiqi Li, Duowen Chen, Yuchen Sun, Bo Zhu · 6 May 2026
We propose a unified, few-step generative modeling framework based on \emph{cumulative flow maps} for long-range transport in probability space, inspired by flow-map techniques for physical transport and dynamics. At its core is a cumulative-flow abstraction that connects local, instantaneous update…
- Tempered Guided Diffusion
Andreas Makris, Paul Fearnhead, Chris Nemeth · 6 May 2026
Training-free conditional diffusion provides a flexible alternative to task-specific conditional model training, but existing samplers often allocate computation inefficiently: independent guided trajectories can vary widely in quality, and additional function evaluations along a single trajectory m…
- Component-Aware Self-Speculative Decoding in Hybrid Language Models
Hector Borobia, Elies Segu\'i-Mas, Guillermina Tormo-Carb\'o · 6 May 2026
Speculative decoding accelerates autoregressive inference by drafting candidate tokens with a fast model and verifying them in parallel with the target. Self-speculative methods avoid the need for an external drafter but have been studied exclusively in homogeneous Transformer architectures. We intr…
- Disentangled Anatomy-Disease Diffusion (DADD) for Controllable Ulcerative Colitis Progression Synthesis
Umut Dundar, Alptekin Temizel · 6 May 2026
Synthesizing longitudinal medical images at controllable disease stages while preserving patient-specific anatomy is hindered by the entanglement of pathological textures and structural features. We address this challenge for ulcerative colitis (UC) endoscopy, where severity follows a continuous ord…
- Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges
Eitan Kosman, Gabriele Serussi, Chaim Basking · 6 May 2026
Modality translation is inherently under-constrained, as multiple cross-modal mappings may yield the same marginals. Recent work has shown that diffusion bridges are effective for this task. However, most existing approaches rely on fully paired datasets, thereby imposing a single data-driven constr…
- PerFlow: Physics-Embedded Rectified Flow for Efficient Reconstruction and Uncertainty Quantification of Spatiotemporal Dynamics
Hao Zhou, Rui Zhang, Han Wan, Hao Sun · 6 May 2026
Reconstructing PDE-governed fields from sparse and irregular measurements is challenging due to their ill-posed nature. Deterministic surrogates are trained on dense fields that struggle with limited measurements and uncertainty quantification. Generative models, by learning distributions over spati…
- Distribution-Free Pretraining of Classification Losses via Evolutionary Dynamics
Meng Xiang, Yan Pei · 6 May 2026
We propose Evolutionary Dynamic Loss (EDL), a framework that learns a transferable classification loss in the probability space using unlimited synthetic prediction-label pairs, without accessing real samples during the main loss pretraining stage. EDL parameterizes the loss as a lightweight network…
- Memorization In Stable Diffusion Is Unexpectedly Driven by CLIP Embeddings
Bumjun Kim, Albert No · 6 May 2026
Understanding how textual embeddings contribute to memorization in text-to-image diffusion models is crucial for both interpretability and safety. This paper investigates an unexpected behavior of CLIP embeddings in Stable Diffusion, revealing that the model disproportionately relies on specific emb…
- IConFace: Identity-Structure Asymmetric Conditioning for Unified Reference-Aware Face Restoration
Axi Niu, Jinyang Zhang, Senyan Qing · 6 May 2026
Blind face restoration is highly ill-posed under severe degradation, where identity-critical details may be missing from the degraded input. Same-identity references reduce this ambiguity, but mismatched pose, expression, illumination, age, makeup, or local facial states can lead to overuse of refer…
- Decision Boundary-aware Generation for Long-tailed Learning
Jiacheng Yang, Ruichi Zhang, Chikai Shang, Mengke Li, Xinyi Shang, Junlong Gao, Yonggang Zhang, Yang Lu · 6 May 2026
Long-tailed data bias decision boundaries toward head classes and degrade tail class accuracy. Diffusion-based generative augmentation address this problem by generating additional data, while head-to-tail transfer further mitigate the generator bias inherit from long-tailed dataset. However, we sho…
- FEAT: Fashion Editing and Try-On from Any Design
Soye Kwon, Keonyoung Lee, Dahuin Jung, Jaekoo Lee · 6 May 2026
Fashion design aims to express a designer's creative intent and to depict how garments interact with the human body. Recent methods condition on multimodal inputs to support garment editing and virtual try-on. However, existing methods still (i) confine design to garment-related images, excluding cr…
- Motion-Aware Caching for Efficient Autoregressive Video Generation
Jing Xu, Yuexiao Ma, Songwei Liu, Xuzhe Zheng, Shiwei Liu, Chenqian Yan, Xiawu Zheng, Rongrong Ji, Fei Chao, Xing Wang · 6 May 2026
Autoregressive video generation paradigms offer theoretical promise for long video synthesis, yet their practical deployment is hindered by the computational burden of sequential iterative denoising. While cache reuse strategies can accelerate generation by skipping redundant denoising steps, existi…
- Rethink MAE with Linear Time-Invariant Dynamics
Zice Wang · 6 May 2026
Standard representation probing for visual models relies on mathematically permutation-invariant operations like Global Average Pooling (GAP) or CLS tokens, treating patch representations as an unstructured bag-of-words. We challenge this paradigm by demonstrating that token order is a critical, exp…
- AsymK-Talker: Real-Time and Long-Horizon Talking Head Generation via Asymmetric Kernel Distillation
Yuxin Lu, Qian Qiao, Jiayang Sun, Min Cao, Guibo Zhu · 6 May 2026
Recent advances in diffusion models have markedly enhanced the visual fidelity of audio-driven talking head generation. Nevertheless, existing methods are constrained by three critical limitations: causal inefficiency that impedes real-time inference, incompatibility with temporally coherent conditi…
- Ortho-Hydra: Orthogonalized Experts for DiT LoRA
Seunghyun Ji · 6 May 2026
LoRA fine-tuning of diffusion transformers (DiT) on multi-style data suffers from \emph{style bleed}: a single low-rank residual cannot represent several distinct artist fingerprints, and the optimizer converges to their average. Mixture-of-experts LoRA in the HydraLoRA style replaces the up-project…
- Transformers with Selective Access to Early Representations
Skye Gunasekaran, T\'ea Wright, Rui-Jie Zhu, Jason Eshraghian · 6 May 2026
Several recent Transformer architectures expose later layers to representations computed in the earliest layers, motivated by the observation that low-level features can become harder to recover as the residual stream is repeatedly transformed through depth. The cheapest among these methods add stat…
- TimeTok: Granularity-Controllable Time-Series Generation via Hierarchical Tokenization
Seokhyun Lee, Jaeho Kim, Changjun Oh, Mihaela van der Schaar, Changhee Lee · 6 May 2026
Time-series generative models often lack control over temporal granularity, forcing users to accept whatever granularity the model produces. To enable truly user-driven generation, we introduce TimeTok, a unified framework for Granularity-Controllable Time-Series Generation (GC-TSG), which generates…
- DataEvolver: Let Your Data Build and Improve Itself via Goal-Driven Loop Agents
Qisong Zhang (School of Artificial Intelligence, Beijing University of Posts and Telecommunications), Wenzhuo Wu (School of Artificial Intelligence, Beijing University of Posts and Telecommunications), Zhuangzhuang Jia (School of Artificial Intelligence, Beijing University of Posts and Telecommunications), Yunhao Yang (School of Artificial Intelligence, Beijing University of Posts and Telecommunications), Huayu Zhang (Institute of Artificial Intelligence), Xianghao Zang (Institute of Artificial Intelligence), Zhixiang He (Institute of Artificial Intelligence), Zhongjiang He (Institute of Artificial Intelligence), Kongming Liang (School of Artificial Intelligence, Beijing University of Posts and Telecommunications), Zhanyu Ma (School of Artificial Intelligence, Beijing University of Posts and Telecommunications) · 6 May 2026
Constructing controllable visual data is a major bottleneck for image editing and multimodal understanding. Useful supervision is rarely produced by a single rendering pass; instead it emerges through iterative generation, inspection, correction, filtering, and export. We present DataEvolver, a clos…
- Perturb and Correct: Post-Hoc Ensembles using Affine Redundancy
Eleanor Quint · 5 May 2026
Models that are indistinguishable on in-distribution data can behave very differently under distribution shift. We introduce Perturb-and-Correct (P&C), a post-hoc method for constructing epistemically diverse predictors from a single pretrained network. P&C applies random hidden layer perturbations …
- LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images
James Flora, Kowshik Thopalli, Akshay R. Kulkarni, Weng-Keen Wong, Shusen Liu · 5 May 2026
We present LatentDiff, a scalable framework for semantic dataset comparison that operates directly in the latent space of pretrained vision encoders. By combining sparse autoencoder-based divergence testing with density ratio estimation, LatentDiff identifies interpretable semantic differences betwe…
- Skipping the Zeros in Diffusion Models for Sparse Data Generation
Phil Sidney Ostheimer, Mayank Nagda, Andriy Balinskyy, Gabriel Vicente Rodrigues, Jean Radig, Carl Herrmann, Stephan Mandt, Marius Kloft, Sophie Fellenz · 5 May 2026
Diffusion models (DMs) excel on dense continuous data, but are not designed for sparse continuous data. They do not model exact zeros that represent the deliberate absence of a signal. As a result, they erase sparsity patterns and perform unnecessary computation on mostly zero entries. With Sparsity…
- Statistically-Lossless Quantization of Large Language Models
Michael Helcig, Eldar Kurtic, Dan Alistarh · 5 May 2026
Model quantization has become essential for efficient large language model deployment, yet existing approaches involve clear trade-offs: methods such as GPTQ and AWQ achieve practical compression but are lossy, while lossless techniques preserve fidelity but typically do not accelerate inference. Th…
- Focus and Dilution: The Multi-stage Learning Process of Attention
Zheng-An Chen, Pengxiao Lin, Zhi-Qin John Xu, Tao Luo · 5 May 2026
Transformer-based models have achieved remarkable success across a wide range of domains, yet our understanding of their training dynamics remains limited. In this work, we identify a recurrent focus-dilution cycle in attention learning and provide a rigorous explanation in a one-layer Transformer s…
