Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
2.542 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
Xuanyu Zhu, Yan Bai, Yang Shi, Yihang Lou, Yuanxing Zhang, Jing Jin, Yuan Zhou · 13. Mai 2026
Representation autoencoders that reuse frozen pretrained vision encoders as visual tokenizers have achieved strong reconstruction and generation quality. However, existing methods universally extract features from only the last encoder layer, discarding the rich hierarchical information distributed …
- Autoregressive Learning in Joint KL: Sharp Oracle Bounds and Lower Bounds
Yunbei Xu, Yuzhe Yuan, Ruohan Zhan · 13. Mai 2026
We study the fundamental and timely problem of learning long sequences in autoregressive modeling and next-token prediction under model misspecification, measured by the joint Kullback--Leibler (KL) divergence. Our goal is to characterize how the sequence horizon \(H\) affects both approximation and…
- Is Monotonic Sampling Necessary in Diffusion Models?
Muhammad Haris Khan · 13. Mai 2026
Diffusion models generate samples by iteratively denoising a Gaussian prior, traversing a sequence of noise levels that, in every published sampler, decreases monotonically. Six years of intensive work has refined nearly every aspect of this recipe, including the corruption operator, the training ob…
- VASR: Variance-Aware Systematic Resampling for Reward-Guided Diffusion
Shivanshu Shekhar, Sagnik Mukherjee, Jia Yi Zhang, Tong Zhang · 13. Mai 2026
Sequential Monte Carlo (SMC) samplers for reward-guided diffusion models often suffer from rapid lineage collapse: a few high-reward particles dominate the population within a handful of resampling steps, destroying diversity and degrading sample quality. We propose a variance-decomposition framewor…
- LOFT: Low-Rank Orthogonal Fine-Tuning via Task-Aware Support Selection
Lanxin Zhao, Bamdev Mishra, Pratik Jawanpuria, Lequan Lin, Dai Shi, Junbin Gao, Andi Han · 13. Mai 2026
Orthogonal parameter-efficient fine-tuning (PEFT) adapts pretrained weights through structure-preserving multiplicative transformations, but existing methods often conflate two distinct design choices: the subspace in which adaptation occurs and the transformation applied within that subspace. This …
- QDSB: Quantized Diffusion Schr\"odinger Bridges
Tobias Fuchs, Florian Kalinke, Nadja Klein · 13. Mai 2026
Learning generative models in settings where the source and target distributions are only specified through unpaired samples is gaining in importance. Here, one frequently-used model are Schr\"odinger bridges (SB), which represent the most likely evolution between both endpoint distributions. To acc…
- Sequential Behavioral Watermarking for LLM Agents
Hyeseon An, Shinwoo Park, Dongsu Kim, Yo-Sub Han · 13. Mai 2026
LLM-based agents act through sequences of executable decisions, but their trajectories provide little evidence of which agent or policy produced them, making provenance, ownership, and unauthorized reuse difficult to establish from observed behavior alone. This motivates watermarking signals embedde…
- Sharpen Your Flow: Sharpness-Aware Sampling for Flow Matching
Aditi Gupta, Soon Hoe Lim, Annan Yu, N. Benjamin Erichson · 13. Mai 2026
Flow matching models generate samples by numerically integrating a learned velocity field, with each integration step requiring a neural network evaluation. Fast generation therefore requires using a small fixed evaluation budget effectively: the key question is not only how to integrate the flow, b…
- L2P: Unlocking Latent Potential for Pixel Generation
Zhennan Chen, Junwei Zhu, Xu Chen, Jiangning Zhang, Jiawei Chen, Zhuoqi Zeng, Wei Zhang, Chengjie Wang, Jian Yang, Ying Tai · 13. Mai 2026
Pixel diffusion models have recently regained attention for visual generation. However, training advanced pixel-space models from scratch demands prohibitive computational and data resources. To address this, we propose the Latent-to-Pixel (L2P) transfer paradigm, an efficient framework that directl…
- DriftXpress: Faster Drifting Models via Projected RKHS Fields
Ali Falahati, Elliot Creager, Gautam Kamath, Shubhankar Mohapatra · 13. Mai 2026
Drifting Models have emerged as a new paradigm for one-step generative modeling, achieving strong image quality without iterative inference. The premise is to replace the iterative denoising process in diffusion models with a single evaluation of a generator. However, this creates a different trade-…
- Enforcing Constraints in Generative Sampling via Adaptive Correction Scheduling
Noah Trupin, Yexiang Xue · 13. Mai 2026
Hard constraints in generative sampling are typically enforced by projection, applied either once at the end of sampling or after every update. This binary framing overlooks a fundamental issue: projection changes the distribution of states which future updates depend on. As a result, delayed projec…
- A Composite Activation Function for Learning Stable Binary Representations
Seokhun Park, Choeun Kim, Kwanho Lee, Sehyun Park, Insung Kong, Yongdai Kim · 13. Mai 2026
Activation functions play a central role in neural networks by shaping internal representations. Recently, learning binary activation representations has attracted significant attention due to their advantages in computational and memory efficiency, as well as interpretability. However, training neu…
- One-Step Generative Modeling via Wasserstein Gradient Flows
Jiaqi Han, Puheng Li, Qiushan Guo, Renyuan Xu, Stefano Ermon, Emmanuel J. Cand\`es · 13. Mai 2026
Diffusion models and flow-based methods have shown impressive generative capability, especially for images, but their sampling is expensive because it requires many iterative updates. We introduce W-Flow, a framework for training a generator that transforms samples from a simple reference distributi…
- StyleVAR: Controllable Image Style Transfer via Visual Autoregressive Modeling
Liqi Jing, Dingming Zhang, Peinian Li, Lichen Zhu, Yang Xu, Hanyu Xing · 13. Mai 2026
We build on the Visual Autoregressive Modeling (VAR) framework and formulate style transfer as conditional discrete sequence modeling in a learned latent space. Images are decomposed into multi-scale representations and tokenized into discrete codes by a VQ-VAE; a transformer then autoregressively m…
- Sobolev Regularized MMD Gradient Flow
Chenyang Tian, Bharath K. Sriperumbudur, Arthur Gretton, Zonghao Chen · 13. Mai 2026
We propose Sobolev-regularized Maximum Mean Discrepancy (SrMMD) gradient flow, a regularized variant of maximum mean discrepancy (MMD) gradient flow based on a gradient penalty on the witness function. The proposed regularization mitigates the non-convexity of the MMD objective and yields provable \…
- TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment
Jiaming Li, Chenyu Zhu, Zhiyuan Ma, Nanxi Yi, Youjun Bao, Li Sun, Quanying Lv, Xiang Fang, Daizong Liu, Jianjun Li, Kun He, Bowen Zhou · 13. Mai 2026
Reinforcement learning (RL) has shown extraordinary potential in aligning diffusion models to downstream tasks, yet most of them still suffer from significant reward hacking, which degrades generative diversity and quality by inducing visual mode collapse and amplifying unreliable rewards. We identi…
- MaskTab: Scalable Masked Tabular Pretraining with Scaling Laws and Distillation for Industrial Classification
Bo Zheng, Yudong Chen, Zihua Xiong, Shuai Fang, Peidong He, Yang Yang, Sheng Guo · 13. Mai 2026
Tabular data forms the backbone of high-stakes decision systems in finance, healthcare, and beyond. Yet industrial tabular datasets are inherently difficult: high-dimensional, riddled with missing entries, and rarely labeled at scale. While foundation models have revolutionized vision and language, …
- Generative Diffusion Prior Distillation for Long-Context Knowledge Transfer
Nilushika Udayangani, Kishor Nandakishor, Marimuthu Palaniswami · 13. Mai 2026
While traditional time-series classifiers assume full sequences at inference, practical constraints (latency and cost) often limit inputs to partial prefixes. The absence of class-discriminative patterns in partial data can significantly hinder a classifier's ability to generalize. This work uses kn…
- Search Your Block Floating Point Scales!
Tanmaey Gupta, Hayden Prairie, Xiaoxia Wu, Reyna Abhyankar, Qingyang Wu, Austin Silveria, Pragaash Ponnusamy, Jue Wang, Ben Athiwaratkun, Leon Song, Tri Dao, Daniel Y. Fu, Chris De Sa · 13. Mai 2026
Quantization has emerged as a standard technique for accelerating inference for generative models by enabling faster low-precision computations and reduced memory transfers. Recently, GPU accelerators have added first-class support for microscaling Block Floating Point (BFP) formats. Standard BFP al…
- Follow the Mean: Reference-Guided Flow Matching
Pedro M. P. Curvo, Maksim Zhdanov, Floor Eijkelboom, Jan-Willem van de Meent · 13. Mai 2026
Existing approaches to controllable generation typically rely on fine-tuning, auxiliary networks, or test-time search. We show that flow matching admits a different control interface: adaptation through examples. For deterministic interpolants, the velocity field is solely governed by a conditional …
- LatentHDR: Decoupling Exposure from Diffusion via Conditional Latent-to-Latent Mapping for Text/Image-to-Panoramic HDR
Pedram Fekri, WenChen Li, William Chen, Peter Altamirano · 13. Mai 2026
High Dynamic Range (HDR) generation remains challenging for generative models, which are largely limited to low dynamic range outputs. Recent diffusionbased approaches approximate HDR by generating multiple exposure-conditioned samples, incurring high computational cost and structural inconsistencie…
- Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning
Hyeonjin Kim, Hangyeol Jung, Heechan Yun, Sungjun Yun, Dong-Jun Han · 13. Mai 2026
Unlearning specific concepts in text-to-image diffusion models has become increasingly important for preventing undesirable content generation. Among prior approaches, sparse autoencoder (SAE)-based methods have attracted attention due to their ability to suppress target concepts through lightweight…
- Efficient Adjoint Matching for Fine-tuning Diffusion Models
Jeongwoo Shin, Dongsoo Shin, Joonseok Lee, Jaewoong Choi, Jaemoo Choi · 13. Mai 2026
Reward fine-tuning has become a common approach for aligning pretrained diffusion and flow models with human preferences in text-to-image generation. Among reward-gradient-based methods, Adjoint Matching (AM) provides a principled formulation by casting reward fine-tuning as a stochastic optimal con…
- FeatMap: Understanding image manipulation in the feature space and its implications for feature space geometry
Elias B. Krey, Nils Neukirch, Nils Strodthoff · 13. Mai 2026
Intermediate feature representations represent the backbone for the expressivity and adaptability of deep neural networks. However, their geometric structure remains poorly understood. In this submission, we provide indirect insights into this matter by applying a broad selection of manipulations in…
- EPIC: Efficient Predicate-Guided Inference-Time Control for Compositional Text-to-Image Generation
Sunung Mun, Sunghyun Cho, Jungseul Ok · 13. Mai 2026
Recent text-to-image (T2I) generators can synthesize realistic images, but still struggle with compositional prompts involving multiple objects, counts, attributes, and relations. We introduce EPIC (Efficient Predicate-Guided Inference-Time Control), a training-free inference-time refinement framewo…
