Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
2.542 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Training-Free Generative Sampling via Moment-Matched Score Smoothing
Zhenyu Yao, Daniel Paulin · 15. Mai 2026
Diffusion models generate samples by denoising along the score of a perturbed target distribution. In practice, one trains a neural diffusion model, which is computationally expensive. Recent work suggests that score matching implicitly smooths the empirical score, and that this smoothing bias promo…
- Covariance-aware sampling for Diffusion Models
Andrea Schioppa, Tim Salimans · 15. Mai 2026
We present a covariance-aware sampler that improves the quality of pixel-space Diffusion Model (DM) sampling in the few-step regime. We hypothesize that in the few-step regime samplers fail because they rely solely on the predicted mean of the reverse distribution, while our solution explicitly mode…
- Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning
Ido Sobol, Kihyuk Sohn, Yoav Blum, Egor Zakharov, Max Bluvstein, Andrea Vedaldi, Or Litany · 15. Mai 2026
We often aim to generate images that are both photorealistic and 3D-consistent, adhering to precise geometry, material, and viewpoint controls. Typically, this is achieved by fine-tuning an image generator, pre-trained on billions of real images, using renders of synthetic 3D assets, where annotatio…
- RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
Xiang Fan, Yuheng Wang, Bohan Fang, Zhongzheng Ren, Ranjay Krishna · 15. Mai 2026
Video generation powers a vast array of downstream applications. However, while the de facto standard, i.e., latent diffusion models, typically employ heavily conditioned denoising networks, their decoders often remain unconditional. We observe that this architectural asymmetry leads to significant …
- The Racial Character of Computer Graphics Research
Theodore Kim, Alexa Schor, Julian Posada, Alka V. Menon · 15. Mai 2026
Computer graphics algorithms for generating photorealistic imagery are widely perceived to be universal, and capable of conjuring anything that a filmmaker or game designer can imagine. However, recent works have suggested that 3D algorithms for depicting synthetic humans are far from generic, and i…
- Generalizing Score-based generative models for Heavy-tailed Distributions
Tiziano Fassina, Gabriel Cardoso, Sylvan Le Corff, Thomas Romary · 15. Mai 2026
Score-based generative models (SGMs) have achieved remarkable empirical success, motivating their application to a broad range of data distributions. However, extending them to heavy-tailed targets remains a largely open problem. Although dedicated models for heavy-tailed distributions have been pro…
- HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts
Tao Zhong, Dongzhe Zheng, Christine Allen-Blanchette · 15. Mai 2026
Sparse Mixture-of-Experts (MoE) layers route tokens through a handful of experts, and learning-free compression of these layers reduces inference cost without retraining. A subtle obstruction blocks every existing compressor in this family: three experts can each be pairwise compatible yet form an i…
- Separating Intrinsic Ambiguity from Estimation Uncertainty in Deep Generative Models for Linear Inverse Problems
Yuxin Guo, Dongrui Deng, Pulkit Grover · 15. Mai 2026
Recently, deep generative models have been used for posterior inference in inverse problems, including high-stakes applications in medical imaging and scientific discovery, where the uncertainty of a prediction can matter as much as the prediction itself. However, posterior uncertainty is difficult …
- Supersampling Stable Diffusion and Beyond: A Seamless, Training-Free Approach for Scaling Neural Networks Using Common Interpolation Methods
Md Abu Obaida Zishan, Jannatun Noor, Annajiat Alim Rasel · 15. Mai 2026
Stable Diffusion (SD) has evolved DDPM (Denoising Diffusion Probabilistic Model) based image generation significantly by denoising in latent space instead of feature space. This popularized DDPM-based image generation as the cost and compute barrier was significantly lowered. However, these models c…
- Pro-DG: Procedural Diffusion Guidance for Architectural Facade Generation
Aleksander Plocharski, Jan Swidzinski, Przemyslaw Musialski · 15. Mai 2026
We use hierarchical procedural rules for the generation of control maps within the stable diffusion framework to produce photo-realistic architectural facade images. Starting from a single input image and its segmentation, we apply an inverse procedural module to identify the facade's hierarchical l…
- Action-Inspired Generative Models
Eshwar R. A., Debnath Pal · 15. Mai 2026
We introduce Action-Inspired Generative Models (AGMs), a dual-network generative framework motivated by the observation that existing bridge-matching methods assign uniform regression weight to every stochastic transition in the transport landscape, regardless of whether a given bridge sample lies a…
- MCLR: Improving Conditional Modeling via Inter-Class Likelihood-Ratio Maximization and Unifying Classifier-Free Guidance with Alignment Objectives
Xiang Li, Yixuan Jia, Xiao Li, Jeffrey A. Fessler, Rongrong Wang, Qing Qu · 14. Mai 2026
Diffusion models achieve strong performance in generative modeling, but their success often relies heavily on classifier-free guidance (CFG), an inference-time heuristic that modifies the sampling trajectory. In theory, diffusion models trained with standard denoising score matching (DSM) should rec…
- Path-independent Flow Matching for Multi-parameter Generative Dynamics
Francisco T\'ellez, AmirHossein Zamani, Philippe Martin, Shuang Ni, Guy Wolf, Eugene Belilovsky, Sina Sanjari, Yanlei Zhang · 14. Mai 2026
Flow Matching is a powerful framework for learning transport maps between probability distributions. Yet its standard single-parameter formulation is not designed to capture multi-parameter variations where the resulting transport should be path-independent. Path independence is crucial because it e…
- Twincher: Bijective Representation Learning for Robust Inversion of Continuous Systems
Arkady Gonoskov · 14. Mai 2026
Recent advances in AI have been primarily driven by large-scale neural architectures that excel at function approximation, rather than by tailored inductive biases and inference or learning strategies that could be important for resource-efficient real-world perception and planning through the solut…
- The Expressivity Boundary of Probabilistic Circuits: A Comparison with Large Language Models
Zhiyu Zhao, Xuejie Liu, Muhan Zhang, Anji Liu · 14. Mai 2026
Probabilistic Circuits (PCs) are deep generative models that support exact and efficient probabilistic inference. Yet in autoregressive language modeling, PCs still lag behind Transformer-based large language models (LLMs), suggesting an important expressivity gap. In this work, we compare PCs and L…
- Constraint-Aware Flow Matching: Decision Aligned End-to-End Training for Constrained Sampling
Jacob K. Christopher, James E. Warner, Ferdinando Fioretto · 14. Mai 2026
Deep generative models provide state-of-the-art performance across a wide array of applications, with recent studies showing increasing applicability for science and engineering. Despite a growing corpus of literature focused on the integration of physics-based constraints into the generation proces…
- Support-Conditioned Flow Matching Is Kernel Smoothing
Daniel Matsui Smola · 14. Mai 2026
Generative models are often conditioned on a small set of examples via cross-attention. Under the Gaussian optimal-transport path, we show that the exact velocity field induced by a finite support set is a Nadaraya--Watson kernel smoother whose bandwidth decreases with flow time, from broad averagin…
- Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
Yatin Dandi, Matteo Vilucchio, Luca Arnaboldi, Hugo Tabanelli, Florent Krzakala · 14. Mai 2026
Understanding how deep neural networks learn useful internal representations from data remains a central open problem in the theory of deep learning. We introduce Neural Low-Degree Filtering (Neural LoFi), a stylized limit of gradient-based training in which hierarchical feature learning becomes an …
- SPOT: Selective Prompt Projection via Total Variation for Inference-Only Safe Text-to-Image Generation
Minhyuk Lee, Hyekyung Yoon, Myungjoo Kang · 14. Mai 2026
Text-to-Image (T2I) diffusion models enable high quality open ended synthesis, but practical use requires suppressing unsafe generations while preserving behavior on benign prompts. We study this tension relative to the frozen generator, using its prompt conditioned distribution as the preservation …
- ArcVQ-VAE: A Spherical Vector Quantization Framework with ArcCosine Additive Margin
Jaeyung Kim, YoungJoon Yoo · 14. Mai 2026
Vector Quantized Variational Autoencoder (VQ-VAE) has become a fundamental framework for learning discrete representations in image modeling. However, VQ-VAE models must tokenize entire images using a finite set of codebook vectors, and this capacity limitation restricts their ability to capture ric…
- Understanding and Accelerating the Training of Masked Diffusion Language Models
Chunsan Hong, Sanghyun Lee, Chieh-Hsin Lai, Satoshi Hayakawa, Yuhta Takida, Yuki Mitsufuji, Seungryong Kim, Jong Chul Ye · 14. Mai 2026
Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models (ARMs) for language modeling. However, MDMs are known to learn substantially more slowly than ARMs, which may become problematic when scaling MDMs to larger models. Therefore, we ask the following questio…
- Accelerating Particle-based Energetic Variational Inference
Xuelian Bao, Lulu Kang, Chun Liu, Yiwei Wang · 14. Mai 2026
In this work, we propose a new particle-based variational inference (ParVI) method for accelerating the Energetic Variational Inference with Implicit scheme (EVI-Im) introduced in Ref. \cite{wang2021particle}. Inspired by energy quadratization (EQ) and operator splitting techniques for gradient flow…
- Do Heavy Tails Help Diffusion? On the Subtle Trade-off Between Initialization and Training
Hamza Cherkaoui, H\'el\`ene Halconruy, Antonio Ocello · 14. Mai 2026
Recent works have proposed incorporating heavy-tailed (HT) noise into diffusion- and flow-based generative models, with the goals of better recovering the tails of target distributions and improving generative diversity. This motivation is intuitive: if the data are heavy-tailed, HT noise may appear…
- When Diffusion Breaks Constraints: Sequential Autoregressive Generation with RL and MCTS
Zirui Zhao, Boye Niu, Harold Soh, David Hsu, Wee Sun Lee · 14. Mai 2026
Data-driven generative models excel in language and vision, but diffusion models often fail in constrained planning and design tasks, exhibiting severe constraint violations in engineering inverse design, molecular generation, multi-robot planning, and floorplan/scene synthesis even with projection …
- Cross Modality Image Translation In Medical Imaging Using Generative Frameworks
Giulia Romoli, Alessia Capoccia, Filippo Ruffini, Francesco Di Feola, Luca Boldrini, Arturo Chiti, Renato Cuocolo, Tugba Akinci D'Antonoli, Fatemeh Darvizeh, Marcello Di Pumpo, Bradley J. Erickson, Liu Fang, Deborah Fazzini, Paola Feraco, Fabrizia Gelardi, Francesco Gossetti, Ana Isabel Hern\'aiz Ferrer, Michail E. Klontzas, Seyedmehdi Payabvash, Katrine Riklund, Sara N. Strandberg, Valerio Guarrasi, Paolo Soda · 14. Mai 2026
Medical image-to-image (I2I) translation enables virtual scanning, i.e. the synthesis of a target imaging modality from a source one without additional acquisitions. Despite growing interest, most proposed methods operate on 2D slices, are evaluated on isolated tasks with different experimental set-…
