Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
2,557 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume — last 12 months
Latest papers
- DRAGON: Distributional Rewards Optimize Diffusion Generative Models
Yatong Bai, Jonah Casebeer, Somayeh Sojoudi, Nicholas J. Bryan · 18 November 2025
We present Distributional RewArds for Generative OptimizatioN (DRAGON), a versatile framework for fine-tuning media generation models towards a desired outcome. Compared with traditional reinforcement learning with human feedback (RLHF) or pairwise preference approaches such as direct preference opt…
- X-VMamba: Explainable Vision Mamba
Mohamed A. Mabrok, Yalda Zafari · 18 November 2025
State Space Models (SSMs), particularly the Mamba architecture, have recently emerged as powerful alternatives to Transformers for sequence modeling, offering linear computational complexity while achieving competitive performance. Yet, despite their effectiveness, understanding how these Vision SSM…
- MeanFlow Transformers with Representation Autoencoders
Zheyuan Hu, Chieh-Hsin Lai, Ge Wu, Yuki Mitsufuji, Stefano Ermon · 18 November 2025
MeanFlow (MF) is a diffusion-motivated generative model that enables efficient few-step generation by learning long jumps directly from noise to data. In practice, it is often used as a latent MF by leveraging the pre-trained Stable Diffusion variational autoencoder (SD-VAE) for high-dimensional dat…
- PID-controlled Langevin Dynamics for Faster Sampling of Generative Models
Hongyi Chen, Jianhai Shu, Jingtao Ding, Yong Li, Xiao-Ping Zhang · 18 November 2025
Langevin dynamics sampling suffers from extremely low generation speed, fundamentally limited by numerous fine-grained iterations to converge to the target distribution. We introduce PID-controlled Langevin Dynamics (PIDLD), a novel sampling acceleration algorithm that reinterprets the sampling proc…
- Stabilizing Self-Consuming Diffusion Models with Latent Space Filtering
Zhongteng Cai, Yaxuan Wang, Yang Liu, Xueru Zhang · 18 November 2025
As synthetic data proliferates across the Internet, it is often reused to train successive generations of generative models. This creates a ``self-consuming loop" that can lead to training instability or \textit{model collapse}. Common strategies to address the issue -- such as accumulating historic…
- MixAR: Mixture Autoregressive Image Generation
Jinyuan Hu, Jiayou Zhang, Shaobo Cui, Kun Zhang, Guangyi Chen · 18 November 2025
Autoregressive (AR) approaches, which represent images as sequences of discrete tokens from a finite codebook, have achieved remarkable success in image generation. However, the quantization process and the limited codebook size inevitably discard fine-grained information, placing bottlenecks on fid…
- Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation
Xin Zhao, Xiaojun Chen, Bingshan Liu, Zeyao Liu, Zhendong Zhao, Xiaoyan Gu · 18 November 2025
Generative vision-language models like Stable Diffusion demonstrate remarkable capabilities in creative media synthesis, but they also pose substantial risks of producing unsafe, offensive, or culturally inappropriate content when prompted adversarially. Current defenses struggle to align outputs wi…
- Statistically Accurate and Robust Generative Prediction of Rock Discontinuities with A Tabular Foundation Model
Han Meng, Gang Mei, Hong Tian, Nengxiong Xu, Jianbing Peng · 18 November 2025
Rock discontinuities critically govern the mechanical behavior and stability of rock masses. Their internal distributions remain largely unobservable and are typically inferred from surface-exposed discontinuities using generative prediction approaches. However, surface-exposed observations are inhe…
- On Powerful Ways to Generate: Autoregression, Diffusion, and Beyond
Chenxiao Yang, Cai Zhou, David Wipf, Zhiyuan Li · 18 November 2025
Diffusion language models have recently emerged as a competitive alternative to autoregressive language models. Beyond next-token generation, they are more efficient and flexible by enabling parallel and any-order token generation. However, despite empirical successes, their computational power and …
- D$^{3}$ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMs
Shuochen Chang, Xiaofeng Zhang, Qingyang Liu, Li Niu · 18 November 2025
Diffusion-based multimodal large language models (Diffusion MLLMs) have recently demonstrated impressive non-autoregressive generative capabilities across vision-and-language tasks. However, Diffusion MLLMs exhibit substantially slower inference than autoregressive models: Each denoising step employ…
- SineLoRA$\Delta$: Sine-Activated Delta Compression
Cameron Gordon, Yiping Ji, Hemanth Saratchandran, Paul Albert, Simon Lucey · 18 November 2025
Resource-constrained weight deployment is a task of immense practical importance. Recently, there has been interest in the specific task of \textit{Delta Compression}, where parties each hold a common base model and only communicate compressed weight updates. However, popular parameter efficient upd…
- To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance
Wanlong Fang, Tianle Zhang, Alvin Chan · 18 November 2025
Multimodal learning often relies on aligning representations across modalities to enable effective information integration, an approach traditionally assumed to be universally beneficial. However, prior research has primarily taken an observational approach, examining naturally occurring alignment i…
- Selecting Fine-Tuning Examples by Quizzing VLMs
Tenghao Ji, Eytan Adar · 18 November 2025
A challenge in fine-tuning text-to-image diffusion models for specific topics is to select good examples. Fine-tuning from image sets of varying quality, such as Wikipedia Commons, will often produce poor output. However, training images that \textit{do} exemplify the target concept (e.g., a \textit…
- Optimizing Input of Denoising Score Matching is Biased Towards Higher Score Norm
Tongda Xu · 18 November 2025
Many recent works utilize denoising score matching to optimize the conditional input of diffusion models. In this workshop paper, we demonstrate that such optimization breaks the equivalence between denoising score matching and exact score matching. Furthermore, we show that this bias leads to highe…
- AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation
Jiayin Zhu, Linlin Yang, Yicong Li, Angela Yao · 18 November 2025
Optimization-based text-to-3D methods distill guidance from 2D generative models via Score Distillation Sampling (SDS), but implicitly treat this guidance as static. This work shows that ignoring source dynamics yields inconsistent trajectories that suppress or merge semantic cues, leading to "seman…
- Hierarchical Schedule Optimization for Fast and Robust Diffusion Model Sampling
Aihua Zhu, Rui Su, Qinglin Zhao, Li Feng, Meng Shen, Shibo He · 18 November 2025
Diffusion probabilistic models have set a new standard for generative fidelity but are hindered by a slow iterative sampling process. A powerful training-free strategy to accelerate this process is Schedule Optimization, which aims to find an optimal distribution of timesteps for a fixed and small N…
- Adaptive Stepsizing for Stochastic Gradient Langevin Dynamics in Bayesian Neural Networks
Rajit Rajpal, Benedict Leimkuhler, Yuanhao Jiang · 18 November 2025
Bayesian neural networks (BNNs) require scalable sampling algorithms to approximate posterior distributions over parameters. Existing stochastic gradient Markov Chain Monte Carlo (SGMCMC) methods are highly sensitive to the choice of stepsize and adaptive variants such as pSGLD typically fail to sam…
- Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
Kabir Khan, Manju Sarkar, Anita Kar, Suresh Ghosh · 18 November 2025
Large generative models (for example, language and diffusion models) enable high-quality text and image synthesis but are hard to train or adapt in cross-device federated settings due to heavy computation and communication and statistical/system heterogeneity. We propose FedGen-Edge, a framework tha…
- Functional Mean Flow in Hilbert Space
Zhiqi Li, Yuchen Sun, Greg Turk, Bo Zhu · 18 November 2025
We present Functional Mean Flow (FMF) as a one-step generative model defined in infinite-dimensional Hilbert space. FMF extends the one-step Mean Flow framework to functional domains by providing a theoretical formulation for Functional Flow Matching and a practical implementation for efficient trai…
- Adaptive Symmetrization of the KL Divergence
Omri Ben-Dov, Luiz F. O. Chamon · 17 November 2025
Many tasks in machine learning can be described as or reduced to learning a probability distribution given a finite set of samples. A common approach is to minimize a statistical divergence between the (empirical) data distribution and a parameterized distribution, e.g., a normalizing flow (NF) or a…
- Low-Bit, High-Fidelity: Optimal Transport Quantization for Flow Matching
Dara Varam, Diaa A. Abuhani, Imran Zualkernan, Raghad AlDamani, Lujain Khalil · 17 November 2025
Flow Matching (FM) generative models offer efficient simulation-free training and deterministic sampling, but their practical deployment is challenged by high-precision parameter requirements. We adapt optimal transport (OT)-based post-training quantization to FM models, minimizing the 2-Wasserstein…
- DiffPro: Joint Timestep and Layer-Wise Precision Optimization for Efficient Diffusion Inference
Farhana Amin, Sabiha Afroz, Kanchon Gharami, Mona Moghadampanah, Dimitrios S. Nikolopoulos · 17 November 2025
Diffusion models produce high quality images but inference is costly due to many denoising steps and heavy matrix operations. We present DiffPro, a post-training, hardware-faithful framework that works with the exact integer kernels used in deployment and jointly tunes timesteps and per-layer precis…
- Neural Local Wasserstein Regression
Inga Girshfeld, Xiaohui Chen · 17 November 2025
We study the estimation problem of distribution-on-distribution regression, where both predictors and responses are probability measures. Existing approaches typically rely on a global optimal transport map or tangent-space linearization, which can be restrictive in approximation capacity and distor…
- Evolutionary Retrofitting
Mathurin Videau (TAU), Mariia Zameshina (LIGM), Alessandro Leite (TAU), Laurent Najman (LIGM, KUSTAR), Marc Schoenauer (TAU), Olivier Teytaud (TAU) · 17 November 2025
AfterLearnER (After Learning Evolutionary Retrofitting) consists in applying evolutionary optimization to refine fully trained machine learning models by optimizing a set of carefully chosen parameters or hyperparameters of the model, with respect to some actual, exact, and hence possibly non-differ…
- Stochastic Variational Inference with Tuneable Stochastic Annealing
John Paisley, Ghazal Fazelnia, Brian Barr · 17 November 2025
We exploit the observation that stochastic variational inference (SVI) is a form of annealing and present a modified SVI approach -- applicable to both large and small datasets -- that allows the amount of annealing done by SVI to be tuned. We are motivated by the fact that, in SVI, the larger the b…
