Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
2 529 papiers indexés
Volume mensuel — 12 derniers mois
Derniers papiers
- FILLER: Feature Imputation via Latent Location Exploration and Retrieval
Santu Mondal, Chayan Maitra, Rajat K. De · 28 juillet 2026
In real-world machine learning applications, incomplete observations create a fundamental challenge. Researchers have come up with several ideas to address this crucial problem. However, current models still face challenges in balancing scalability and structural consistency. This study proposes a f…
- Soft-Constrained Optimization of Latent Space in Variational Autoencoders
Ye Shi · 28 juillet 2026
The usefulness of a variational autoencoder (VAE) depends on two properties of its latent space that are hard to obtain together: high encoding capacity in the individual latent variables, and a low-dimensional, disentangled organization of those variables. Weakening the Kullback-Leibler regularizat…
- All in One: Generative Modeling as Mean-Field Game Design
Kun Zhao, Xu Chen · 28 juillet 2026
Mean-field games (MFGs) offer a unifying lens on continuous-time generative modeling: a cost tuple recovering twelve prominent models---Continuous Normalizing Flows, OT-Flow, Score-based Models, Schr\"{o}dinger Bridges, and more---as special cases of one variational problem. Yet two dimensions of th…
- From Score Learning to Discretized Sampling: An End-to-End Generalization Analysis of Diffusion Models
Jinshu Huang, Yiming Jiang, Chunlin Wu · 28 juillet 2026
Despite the empirical success of score-based diffusion models, a complete theoretical understanding of how finite-sample learning, network parameterization, and numerical discretization jointly dictate generative quality remains underdeveloped. Existing sampling analyses often evaluate the generativ…
- Logit-Coordinate Generative Models for Mixed Continuous-Categorical Tabular Data
Yuefei Shen, Xiaotong Shen · 28 juillet 2026
Mixed continuous--categorical data pose a representation problem for continuous generative models. Flow Matching and Gaussian diffusion operate in Euclidean spaces, whereas categorical laws lie on probability simplices and may be highly imbalanced. We study a logit-coordinate framework that encodes …
- Learning Sampling Parameters for Diffusion Models
Arisrei Lim, Yossi Gandelsman · 28 juillet 2026
Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, and noise schedules. These parameters are typically manually chosen once and then held fixed across prompts and denoising timesteps, even though differ…
- Joint Flow Matching for Generator-Consistent Classification
Hayden McAlister, Lech Szymanski · 28 juillet 2026
We introduce Joint Flow Matching (JFM), a training framework for continuous normalising flows over multiple variables. Standard flow matching transports variables from noise to data simultaneously, offering no natural mechanism for forward and reverse conditional inference from a shared joint model.…
- UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective
Xiaoyi Jiang, Jingyuan Li, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu · 28 juillet 2026
Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling. However, adapting AR checkpoints across corruption kernels remains challenging because exis…
- A Coulomb Particle Model for Learning Kernel Attention in Transformers
Masoud Badiei Khuzani, Sharath Honnaiah, Atiq Islam, Alex Cozzi, Abraham Bagherjeiran · 28 juillet 2026
Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature distribution. We propose a particle-based method that learns this distribution by optimizing kernel-target alignment while regularizing particles with a Riesz/Coul…
- Beyond ICA: Identifiability by Symmetry Breaking
Pengzhou Wu · 28 juillet 2026
We prove the identifiability of deep generative models (DGMs) with piecewise-affine (PWA) decoders and Gaussian mixture model (GMM) priors, in a purely unsupervised setting. We introduce three algebraic contrast principles for symmetry breaking: domain contrast, which trivializes the mixture symmetr…
- Physics-Guided Generative AI for Property-Targeted 3D Porous Media Design
Peng Wang · 28 juillet 2026
Inverse design of three-dimensional porous media is central to applications in filtration, catalysis, energy storage, fuel cells, thermal management, and biomedical scaffolds, but remains challenging because many distinct pore geometries can share similar porosity or permeability while small structu…
- Exploration of the generative capabilities of Boltzmann machines applied to social systems under the majority rule
Mauricio A. Valle, Gonzalo A. Ruz · 28 juillet 2026
We study the generative capabilities of Boltzmann machines to recover systems governed by the majority rule under critical conditions. To this end, we train deep belief networks (DBNs) with different configurations, where the first layer can use Gaussian visible units with more than two states (i.e.…
- dRAE: Representation Autoencoder with Hyper-Spherical Codes
Tianren Ma, Lin Long, Chuyan Chen, Mu Zhang, Junbo Zhao, Tong Zhang, Qixiang Ye · 27 juillet 2026
In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric …
- Time-Reversed Imaging: A Multimodal Benchmark and Framework for Reconstructing Past Human-Environment Interactions
Jorge Bacca, Kebin Contreras, Luis Toscano-Palomino, Mauro Dalla Mura · 27 juillet 2026
We introduce time-reversed imaging, a new paradigm that infers what just happened in a scene from fading multimodal traces. Instead of extrapolating or interpolating video frames, our goal is to infer past human-environment interactions from residual physical imprints observable in thermal, ultravio…
- Persistent Computational State: A Session-Centric Runtime for Generative World Models
Zhen Lin · 27 juillet 2026
Generative world models are increasingly driven as simulators: a planner forks a state, rolls out futures, backtracks, and returns to a visited viewpoint. Recent benchmarks establish that current video world models fail this usage, and attribute it to the model, prescribing new architectures and tra…
- TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward
Debottam Dutta, Jaehoon Hahm, Jianchong Chen, Romit Roy Choudhury · 27 juillet 2026
Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present TILT, a training-free framework for compositional text-to-imag…
- Wavelet Phase Diffusion for Structurally and Semantically Consistent Sim-to-Real Translation
Kaiwen Wang, Frank Bieder, Yinzhe Shen, Carlos Fernandez, Jan-Hendrik Pauls, Omer Sahin Tas · 27 juillet 2026
Simulation-to-reality translation must bridge the appearance gap between synthetic and real domains while preserving structural and semantic consistency. Conditioning-based methods achieve spatial alignment but introduce computationally expensive control modules. Paired-data methods achieve realism …
- An Explainable FFT-Based Spatial-Frequency Fusion Framework for Deepfake Detection
Pamela Kirui, Cho Hyuk, Qingzhong Liu, Haodi Jiang · 27 juillet 2026
Deepfake generation has raised growing concerns regarding digital media authenticity, misinformation, identity fraud, and public trust. Recent studies show that combining spatial and frequency features leads to stronger detection results than using independently. This paper presents MSCA-FFT, a Fast…
- From Scalars to Time Series: Rethinking Implicit Neural Representations for Time-Varying Volumetric Data
Weihan Zhang, Xuan Zhao, Yenwen Peng, Yuqi Chen, Jun Tao · 24 juillet 2026
Implicit neural representations (INRs) for time-varying volumetric data are typically trained using dense sampling over spatiotemporal coordinates, where each observation corresponds to a single point in space and time. This coordinate-wise formulation requires extensive sampling during optimization…
- PhantomFill: When the Form Demands an Answer, Language Models Invent One
Rana Muhammad Usman · 24 juillet 2026
Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the form itself causes hallucination. We ask thirteen models the same question about the same input and change only the answer format. The inputs are built so the …
- Expanding Flow Maps
Sophia Tang, Pranam Chatterjee · 24 juillet 2026
Flow-based generative models have enabled remarkable progress in fast and controllable generation across continuous and discrete state spaces, yet existing parameterizations are constrained to fixed dimensions or fixed sequence lengths. Here, we introduce Expanding Generative Flows (EFlows), which d…
- DART: A Degradation-Aware Recurrent Transformer for Archival Film Restoration
Miko{\l}aj Jastrz\k{e}bski, Wojciech Koz{\l}owski, Kamil Adamczewski · 24 juillet 2026
Archival film restoration is a challenging problem because historical footage contains compound degradations such as scratches, dust, blur, noise, flicker, and photometric aging, while clean reference videos are unavailable. Existing video restoration methods largely treat these degradations implici…
- M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data
Francesca Pia Panaccione, Carlo Sgaravatti, Marco Venere · 24 juillet 2026
Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular profiles, is crucial for comprehensive disease understanding. However, gene expression data acquisition remains constrained by high costs and privacy concerns, limiting its use in multimodal …
- Context-weighted Discrete Flow Matching
Daniil Cherniavskii, Daniel Severo, Karen Ullrich · 24 juillet 2026
Discrete flow matching provides a flexible framework for generative modeling on discrete structures. However, the standard factorized training objective exposes the model to targets of varying difficulty, mixing well-conditioned, predictable tokens with ambiguous, high-entropy ones. We empirically d…
- Source-Prior-Driven Selective Adaptation for Efficient Diffusion Model Finetuning
Yi Xiong, Yuan-Yuan Cheng, Xiao-Ming Fu · 24 juillet 2026
Fine-tuning large diffusion models for new domains or styles involves a trade-off: improving target-specific generation often degrades the pretrained model's broad generative capability. Existing full and parameter-efficient fine-tuning methods typically handle this trade-off only implicitly. In thi…