Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
2542 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- Learning Normalized Energy Models for Linear Inverse Problems
Nicolas Zilberstein, Santiago Segarra, Eero Simoncelli, Florentin Guth · 18 de mayo de 2026
Generative diffusion models can provide powerful prior probability models for inverse problems in imaging, but existing implementations suffer from two key limitations: $(i)$ the prior density is represented implicitly, and $(ii)$ they rely on likelihood approximations that introduce sampling biases…
- Entropic Auto-Encoding via Implicit Free-Energy Minimization
Hazhir Aliahmadi, Irina Babayan, Greg van Anders · 18 de mayo de 2026
Despite their ubiquity, variational autoencoders (VAEs) inherently suffer from posterior collapse, a failure mode in which latent variables are effectively ignored. This failure arises because explicit prior imposition drives optimization toward loss landscape regions corresponding to uninformative …
- COCO-Inpaint: A Benchmark for Detecting and Localizing Inpainting-Based Image Manipulations
Haozhen Yan, Yan Hong, Jiahui Zhan, Suning Lang, Yikun Ji, Huijia Zhu, Jun Lan, Jianfu Zhang · 18 de mayo de 2026
Recent advances in image manipulation have enabled highly photorealistic content generation, but also lowered the barrier to arbitrary editing, raising concerns about multimedia authenticity and security. Existing Image Manipulation Detection and Localization (IMDL) methods mainly target splicing or…
- Don't Stop Me Yet: Sampling Loss Minima via Dissipative Riemannian Mechanics
Albert Kj{\o}ller Jacobsen, Leo Uhre Jakobsen, Johanna Marie Gegenfurtner, Georgios Arvanitidis · 18 de mayo de 2026
The minima of modern neural network loss functions are typically not isolated, rather they form connected components of reparameterization invariant solutions on the training data. Analytically characterizing these solutions is a hard problem, but sampling approaches are feasible. By construction, e…
- Mind Dreamer: Untethering Imagination via Active Latent Intervention on Latent Manifolds
Shaojun Xu, Xiaoling Zhou, Yihan Lin, Yapeng Meng, Xinglong Ji, Luping Shi, Rong Zhao · 18 de mayo de 2026
Model-Based Reinforcement Learning (MBRL) leverages latent imagination for sample efficiency, yet remains constrained by Historical Tethering: imagination is typically initialized from observed states. This creates a learning asymmetry, where the world model's manifold discovery outpaces the policy'…
- FSCM: Frequency-Enhanced Spatial-Spectral Coupled Mamba for Infrared Hyperspectral Image Colorization
Tingting Liu, Yuan Liu, Guiping Chen, Xiubao Sui, Qian Chen · 18 de mayo de 2026
Thermal infrared imaging is robust to illumination variations and smoke interference, making it important for all-weather perception. However, the lack of natural color and fine texture limits target recognition, human visual interpretation, and the transfer of visible-light models. Existing infrare…
- Ada-Diffuser: Latent-Aware Adaptive Diffusion for Decision-Making
Fan Feng, Selena Ge, Minghao Fu, Zijian Li, Yujia Zheng, Zeyu Tang, Yingyao Hu, Biwei Huang, Kun Zhang · 18 de mayo de 2026
Recent work has framed decision-making as a sequence modeling problem using generative models such as diffusion models. Although promising, these approaches often overlook latent factors that exhibit evolving dynamics, elements that are fundamental to environment transitions, reward structures, and …
- Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models
Yishun Lu, Wes Armour · 18 de mayo de 2026
Autoregressive next-token training offers a unified formulation for image generation and text understanding, but it also creates strong modality competition that destabilizes optimization and limits large-batch scaling. We show that first-order optimizers such as AdamW are vulnerable to cross-modali…
- Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
Vincent-Daniel Yun, Junhyuk Jo, Sai Praneeth Karimireddy, Sunwoo Lee · 18 de mayo de 2026
Layer pruning removes entire Transformer decoder blocks from large language models, but introduces a mismatch between the hidden state received by the next surviving layer and the distribution it was trained to process, leading to significant performance degradation. We propose Ghosted Layers, a tra…
- Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
Xiaolong Fu, Lichen Ma, Zipeng Guo, ShiPing Dong, Lan Yang, Tan Lit Sin, Gaojing Zhou, Yu He, Jingling Fu, Shizhe Zhou, Junshi Huang, Jason Li · 18 de mayo de 2026
The integration of Reinforcement Learning (RL) into flow matching models for text-to-image (T2I) generation has driven substantial advances in generation quality. However, these gains often come at the cost of exhaustive exploration and inefficient sampling strategies due to slight variation in the …
- Martingale Neural Operators: Learning Stochastic Marginals via Doob-Meyer Factorization
Kai Hidajat · 18 de mayo de 2026
Neural operators excel as deterministic surrogates, but inevitably collapse to the conditional mean when applied to stochastic PDEs, discarding the variance and tail structure upon which uncertainty quantification depends. Recovering this structure typically requires Monte Carlo rollouts or grafted …
- Entropy-Based Characterisation of the Polarised Regime in Latent Variable Models
Peter Clapham, Lisa Bonheme, Marek Grzes · 18 de mayo de 2026
Variational Autoencoders (VAEs) often exhibit a polarised regime in which latent variables separate into active, passive, and mixed subsets. Existing criteria for identifying active dimensions depend on a Gaussian prior, limiting their applicability to variational models and specific priors. We prop…
- PanoWorld: Geometry-Consistent Panoramic Video World Modeling
Le Jiang, Xiangyu Bai, Bishoy Galoaa, Shayda Moezzi, Caleb James Lee, Tooba Imtiaz, Edmund Yeh, Jennifer Dy, Yanzhi Wang, Sarah Ostadabbas · 18 de mayo de 2026
We present PanoWorld, a panoramic video world model that generates geometry-consistent 360$\degree$ video from a single image and a caption. Existing panoramic video methods optimize primarily for visual realism and do not explicitly constrain the underlying 3D scene state, producing outputs that ap…
- Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance
Song Wu, Xinyu Chen, Qian Wang, Liang Li, Zili Yi, Junlan Feng · 18 de mayo de 2026
Video editing poses a significant challenge. While a series of tuning-free methods circumvent the need for extensive data collection and model training, they often underutilize the rich information embedded within noisy latent, leading to unsatisfactory results. To address this, we propose a \textit…
- Constrained latent state modeling: A unifying perspective on representation learning under competing constraints
Gwenol\'e Quellec · 18 de mayo de 2026
Learning latent representations from complex data is central to modern machine learning, spanning temporal, multimodal, and partially observed systems. In such settings, representations are better understood as latent states capturing underlying system dynamics, rather than as mere compressed summar…
- How Data Augmentation Shapes Neural Representations
Tianxiao He, Alex H. Williams, Sarah E. Harvey · 18 de mayo de 2026
Data augmentation is widely recognized for improving generalization in deep networks, yet its impact on the geometry of learned representations remains poorly understood. In this work, we characterize how different data augmentation strategies reshape internal representations in neural networks. Usi…
- Entropy Across the Bridge: Conditional-Marginal Discretization for Flow and Schr\"odinger Samplers
Bruno Trentini, Dejan Stancevic, Michael M. Bronstein, Alexander Tong, Luca Ambrogioni · 18 de mayo de 2026
For a fixed flow-based generative model under a small inference budget, sample quality can depend strongly on where the sampler spends its few function evaluations. Flow matching and Schr\"odinger bridges define probability paths, yet their inference grids are usually heuristic or inherited from one…
- Navigating Potholes with Geometry-Aware Sharpness Minimization
Simon Dufort-Labb\'e, Mehrab Hamidi, Razvan Pascanu, Ioannis Mitliagkas, Damien Scieur, Aristide Baratin · 18 de mayo de 2026
Sharpness-aware minimization (SAM) encourages flat minima by perturbing parameters along directions of high loss curvature, but treats all parameter directions uniformly, ignoring the underlying loss geometry. We introduce LLQR+SAM, which combines SAM with a learned preconditioner obtained from the …
- LoCO: Low-rank Compositional Rotation Fine-tuning
An Nguyen, Jaesik Choi, Anh Tong · 18 de mayo de 2026
Parameter-efficient fine-tuning (PEFT) has emerged as an critical technique for adapting large-scale foundation models across natural language processing and computer vision. While existing methods such as low-rank adaptations achieve parameter efficiency via low-rank weight updates, they are limite…
- Autoguided Online Data Curation for Diffusion Model Training
Valeria Pais, Luis Oala, Daniele Faccio, Marco Aversa · 18 de mayo de 2026
The costs of generative model compute rekindled promises and hopes for efficient data curation. In this work, we investigate whether recently developed autoguidance and online data selection methods can improve the time and sample efficiency of training generative diffusion models. We integrate join…
- DiLA: Disentangled Latent Action World Models
Tianqiu Zhang, Muyang Lyu, Yufan Zhang, Fang Fang, Si Wu · 18 de mayo de 2026
Latent Action Models (LAMs) enable the learning of world models from unlabeled video by inferring abstract actions between consecutive frames. However, LAMs face a fundamental trade-off between action abstraction and generation fidelity. Existing methods typically circumvent this issue by using two-…
- Pretraining Objective Matters in Extreme Low-Data FGVC: A Backbone-Controlled Study
Alexander Hackett, Srikanth Thudumu, Ginny Fisher, Mahule Roy, Aisha Sartaj, Jason Fisher · 18 de mayo de 2026
Extreme low-data fine-grained classification is common in expert domains where labeling is expensive, yet practitioners still need principled guidance for selecting pretrained encoders. We study emerald inclusion grading with a custom dataset of labeled images across three classes and ask: under mat…
- Time-Varying Deep State Space Models for Sequences with Switching Dynamics
Sanja Karilanova, Subhrakanti Dey, Ay\c{c}a \"Oz\c{c}elikkale · 18 de mayo de 2026
The identification and modeling of time-varying systems is a fundamental challenge in signal processing and system identification. To address this challenge, we propose a class of time-varying state-space model (SSM) based neural networks in which the neurons' states are governed by time-varying dyn…
- RaPD: Resolution-Agnostic Pixel Diffusion via Semantics-Enriched Implicit Representations
Yanhao Ge, Shanyan Guan, Weihao Wang, Ying Tai, Mingyu You · 18 de mayo de 2026
Natural images are continuous, yet most generative models synthesize them on discrete grids, limiting resolution-flexible generation. Continuous neural fields enable resolution-free rendering, but prior methods introduce continuity only at the decoding stage as an interpolation module, leaving the g…
- FFAvatar: Few-Shot, Feed-Forward, and Generalizable Avatar Reconstruction
Thuan Hoang Nguyen, Jiahao Luo, Yinyu Nie, Hao Li, Gordon Guocheng Qian, Jian Wang · 18 de mayo de 2026
Avatar reconstruction has traditionally relied on per-subject optimization that requires hours of computation or on expensive preprocessing that limits scalability. We introduce FFAvatar, a generalizable feed-forward framework that reconstructs high-quality, animatable 3D Gaussian head avatars from …
