Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
2542 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- A Flow Matching Algorithm for Many-Shot Adaptation to Unseen Distributions
Tyler Ingebrand, Ruihan Zhao, Kushagra Gupta, David Fridovich-Keil, Sandeep P. Chinchali, Ufuk Topcu · 8 de mayo de 2026
While generative modeling has achieved remarkable success on tasks like natural language-conditioned image generation, enabling model adaptation from example data points remains a relatively underexplored and challenging problem. To this end, we propose Function Projection for Flow Matching (FP-FM),…
- DBMSolver: A Training-free Diffusion Bridge Sampler for High-Quality Image-to-Image Translation
Sankarshana Venugopal (Seoul National University), Mohammad Mostafavi (Seoul National University), Jonghyun Choi (Seoul National University) · 8 de mayo de 2026
Diffusion-based image-to-image (I2I) translation excels in high-fidelity generation but suffers from slow sampling in state-of-the-art Diffusion Bridge Models (DBMs), often requiring dozens of function evaluations (NFEs). We introduce DBMSolver, a training-free sampler that exploits the semi-linear …
- Taming the Entropy Cliff: Variable Codebook Size Quantization for Autoregressive Visual Generation
Bowen Zheng, Weijian Luo, Guang Yang, Colin Zhang, Tianyang Hu · 8 de mayo de 2026
Most discrete visual tokenizers rely on a default design: every position in the sequence shares the same codebook. Researchers try to scale the codebook size $K$ to get better reconstruction performance. Such a constant-codebook design hits a fundamental information-theoretic limit. We observe that …
- FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation
Mingfeng Lin, Jiakun Chen, Liang Han, Liqiang Nie · 8 de mayo de 2026
Pixel-space diffusion has re-emerged as a promising alternative to latent-space generation because it avoids the representation bottleneck introduced by VAEs. Yet most existing methods still treat image generation as a frequency-homogeneous process, overlooking the distinct roles and learning dynami…
- Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models
Amir Rezaei Balef, Mykhailo Koshil, Katharina Eggensperger · 8 de mayo de 2026
Transformer-based tabular foundation models (TFMs) dominate small to medium tabular predictive benchmark tasks, yet their inference mechanisms remain largely unexplored. We present the first large-scale mechanistic study of layerwise dynamics in 6 state-of-the-art tabular in-context learning models.…
- Spherical Flows for Sampling Categorical Data
Jannis Chemseddine, Gregor Kornhardt, Gabriele Steidl · 8 de mayo de 2026
We study the problem of learning generative models for discrete sequences in a continuous embedding space. Whereas prior approaches typically operate in Euclidean space or on the probability simplex, we instead work on the sphere $\mathbb S^{d-1}$. There the von Mises-Fisher (vMF) distribution induc…
- Time-Inhomogeneous Preconditioned Langevin Dynamics
Alexander Falk, Laurenz Nagler, Andreas Habring, Thomas Pock · 8 de mayo de 2026
Langevin sampling from distributions of the form $p(x) \propto \exp(-\Psi(x))$ faces two major challenges: (global) mode coverage and (local) mode exploration. The first challenge is particularly relevant for multi-modal distributions with disjoint modes, whereas the second arises when the potential…
- Learning Discrete Autoregressive Priors with Wasserstein Gradient Flow
Bowen Zheng, Yihong Luo, Tianyang Hu · 8 de mayo de 2026
Discrete image tokenizers are commonly trained in two stages: first for reconstruction, and then with a prior model fitted to the frozen token sequences. This decoupling leaves the tokenizer unaware of the model that will later generate its tokens. As a result, the learned tokens may preserve image …
- The Interplay of Data Structure and Imbalance in the Learning Dynamics of Diffusion Models
Flavio Nicoletti, Chenxiao Ma, Enrico Ventura, Luca Saglietti, Stefano Sarao Mannelli · 8 de mayo de 2026
Real-world datasets are inherently heterogeneous, yet how per-class structural differences and sampling imbalance shape the training dynamics of diffusion models-and potentially exacerbate disparities-remains poorly understood. While models typically transition from an initial phase of generalizatio…
- Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
Nilaksh, Saurav Jha, Artem Zholus, Sarath Chandar · 8 de mayo de 2026
World model-based policy evaluation is a practical proxy for testing real-world robot control by rolling out candidate actions in action-conditioned video diffusion models. As these models increasingly adopt latent diffusion modeling (LDM), choosing the right latent space becomes critical. While the…
- Order-Agnostic Autoregressive Modelling with Missing Data
Ignacio Peis, Pablo M. Olmos, Jes Frellsen · 8 de mayo de 2026
Order-Agnostic autoregressive models have demonstrated strong performance in deep generative modeling, yet their use in settings with incomplete data remains largely unexplored. In this work, we reinterpret them through the lens of missing data. First, we show that their standard training procedure …
- Flow Matching with Arbitrary Auxiliary Paths
Xin Peng, Ang Gao · 8 de mayo de 2026
We introduce a new generative modeling framework, \textbf{Flow Matching with Arbitrary Auxiliary Paths (AuxPath-FM)}, which generalizes conditional flow matching by incorporating an auxiliary variable drawn from an arbitrary distribution into the probability path. Unlike prior methods that restrict …
- Diverse Sampling in Diffusion Models with Marginal Preserving Particle Guidance
Gal Vinograd, Idan Achituve, Ethan Fetaya · 8 de mayo de 2026
We present EDDY (Exact-marginal Diversification via Divergence-free dYnamics), a guidance mechanism for diffusion and flow matching models that promotes diversity among samples generated while maintaining quality. EDDY exploits symmetries of the Fokker-Planck equation, using drift perturbations that…
- ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters
Philippe Hansen-Estruch, Jiahui Chen, Vivek Ramanujan, Orr Zohar, Yan Ping, Animesh Sinha, Markos Georgopoulos, Edgar Schoenfeld, Ji Hou, Felix Juefei-Xu, Sriram Vishwanath, Ali Thabet · 8 de mayo de 2026
Vision Transformer (ViT) autoencoders have emerged as compelling tokenizers for images, offering improved reconstruction over convolutional tokenizers. However, existing ViT tokenizers cannot explore this landscape as performance degrades outside training resolutions, and reliance on adversarial los…
- DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
Akash Haridas, Utkarsh Saxena, Parsa Ashrafi Fashi, Mehdi Rezagholizadeh, Vikram Appia, Emad Barsoum · 8 de mayo de 2026
Diffusion Transformers rely on static patchify tokenization, assigning the same token budget to smooth backgrounds, detailed object regions, noisy early timesteps, and late-stage refinements. We introduce the Dynamic Chunking Diffusion Transformer (DC-DiT), which replaces fixed patchification with a…
- Information-Preserving Domain Transfer with Unlabeled Data in Misspecified Simulation-Based Inference
Joon Jang, Eunho Jeong, Kyu Sung Choi, Hyeonjin Kim · 8 de mayo de 2026
Simulation-based inference (SBI) provides amortized Bayesian parameter inference from simulator-generated data without requiring explicit likelihood evaluation. Its reliability can degrade under model misspecification, where real-world observations are not well represented by the simulator used for …
- Autoregressive Visual Generation Needs a Prologue
Bowen Zheng, Weijian Luo, Guang Yang, Colin Zhang, Tianyang Hu · 8 de mayo de 2026
In this work, we propose Prologue, an approach to bridging the reconstruction-generation gap in autoregressive (AR) image generation. Instead of modifying visual tokens to satisfy both reconstruction and generation, Prologue generates a small set of prologue tokens prepended to the visual token sequ…
- Physical Fidelity Reconstruction via Improved Consistency-Distilled Flow Matching for Dynamical Systems
Sicheng Ma, Tianyue Yang, Xiuzhe Wu, Xiao Xue · 8 de mayo de 2026
Reconstructing high-fidelity flow fields from low-fidelity observations is a central problem in scientific machine learning, yet recent diffusion and flow-matching models typically rely on iterative sampling, making them costly for latency-sensitive workflows such as ensemble forecasting, real-time …
- The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks
Taehun Cha, Daniel Beaglehole, Adityanarayanan Radhakrishnan, Donghun Lee · 8 de mayo de 2026
Understanding how deep neural networks learn representations remains a central challenge in machine learning theory. In this work, we propose a feature-centric framework for analyzing neural network training by relating weight updates to feature evolution. We introduce a simple identity, the Feature…
- Energy Generative Modeling: A Lyapunov-based Energy Matching Perspective
Yixuan Wang, Wenqian Xue, Warren E. Dixon · 8 de mayo de 2026
Generative models based on static scalar energy functions represent an emerging paradigm in which a single time independent potential drives sample generation through its gradient field, eliminating the need for time conditioning entirely. We unify the training and sampling phases of this paradigm, …
- End-to-End Identifiable and Consistent Recurrent Switching Dynamical Systems
Carles Balsells-Rodas, Zhengrui Xiang, Xavier Sumba, Yingzhen Li · 8 de mayo de 2026
Learning identifiable representations in deep generative models remains a fundamental challenge, particularly for sequential data with regime-switching dynamics. Existing approaches establish identifiability under restrictive assumptions, such as stationarity or limited emission models, and typicall…
- Feature Starvation as Geometric Instability in Sparse Autoencoders
Faris Chaudhry, Keisuke Yano, Anthea Monod · 8 de mayo de 2026
Sparse autoencoders (SAEs) are used to disentangle the dense, polysemantic internal representations of large language models (LLMs) into interpretable, monosemantic concepts. However, standard $\ell_1$-regularized SAEs suffer from feature starvation (dead neurons) and shrinkage bias, often requiring…
- Inference-Time Refinement Closes the Synthetic-Real Gap in Tabular Diffusion
Eugenio Lomurno, Filippo Balzarini, Francesco Benelle, Francesca Pia Panaccione, Matteo Matteucci · 8 de mayo de 2026
Diffusion-based generators set the current state of the art for synthetic tabular data. These methods approach but rarely exceed real-data utility, and closing this synthetic-real gap has so far been pursued exclusively at training time, via architectural advances, scaling, and retraining of monolit…
- Threshold-Guided Optimization for Visual Generative Models
Jinbin Bai, Yu Lei, Qingyu Shi, Aosong Feng, Yi Xin, Zhuoran Zhao, Fei Shen, Kaidong Yu, Jason Li · 7 de mayo de 2026
Aligning large visual generative models with human feedback is often performed through pairwise preference optimization. While such approaches are conceptually simple, they fundamentally rely on annotated pairs, limiting scalability in settings where feedback is collected as independent scalar ratin…
- Taming Outlier Tokens in Diffusion Transformers
Xiaoyu Wu, Yifei Wang, Tsu-Jui Fu, Liang-Chieh Chen, Zhe Gan, Chen Wei · 7 de mayo de 2026
We study outlier tokens in Diffusion Transformers (DiTs) for image generation. Prior work has shown that Vision Transformers (ViTs) can produce a small number of high-norm tokens that attract disproportionate attention while carrying limited local information, but their role in generative models rem…
