Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
2.542 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)
Ankit Hemant Lade, Sai Krishna Jasti, Indar Kumar, Aman Chadha · 12. Mai 2026
A Mamba state-space model trained only for next-step prediction appears to recover Granger-causal structure through a simple readout $S = |W_{out} W_{in}|$, with early experiments suggesting the phenomenon generalized across architectures and benefited from interventional data at $p < 10^{-5}$. We p…
- Towards Robust Sequential Decomposition for Complex Image Editing
Zilai Zeng, Mingdeng Cao, Zijie Li, Xiaochen Lian, Yichun Shi, Peihao Zhu, Chen Sun, Peng Wang · 12. Mai 2026
Recent advances in visual generative models have enabled high-fidelity image editing guided by human instructions. However, these models often struggle with complex instructions involving combinatorial editing operations or inter-step dependencies. This difficulty stems from the limitations of two c…
- NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training
Fang Wu, Haokai Zhao, Da Xing, Hanqun Cao, Tinson Xu, Yanchao Li, Xiangru Tang, Zehong Wang, Aaron Tu, Kuan Pang, Hanchen Wang, Hongbin Lin, Zeqi Zhou, Yinxi Li, Peng Xia, Li Erran Li, Molei Tao, Jure Leskovec, Aditya Joshi, Yejin Choi · 12. Mai 2026
Diffusion models have achieved remarkable success across a wide range of generative tasks, yet their training paradigm largely treats injected noise as uniformly informative. In this work, we challenge this assumption and introduce NoiseRater, a meta-learning framework for instance-level noise valua…
- Geometry Guided Self-Consistency for Physical AI
Yinwei Dai, Zhuofu Chen, Lijie Yang, Ravi Netravali · 12. Mai 2026
State-of-the-art physical AI models generate a chunk of actions per inference through diffusion or flow matching, iteratively refining an initial noise sample into an action trajectory. Because this inference process is inherently stochastic, committing to a single trajectory per round is brittle, a…
- SWIFT: Prompt-Adaptive Memory for Efficient Interactive Long Video Generation
Shanwen Tan, Hao Li, Jingtao Zhang, Xiaosong Jia, Xue Yang, Shaofeng Zhang, Yanyong Zhang · 12. Mai 2026
Streaming long-video generation faces a central challenge in continuous semantic switching, requiring adaptive memory to preserve coherent visual evolution. Current approaches rely on cache rebuilding at prompt boundaries or fixed memory budgets, but they introduce redundant computation and limit fl…
- Entropy-informed Decoding: Adaptive Information-Driven Branching
Benjamin Patrick Evans, Sumitra Ganesh, Leo Ardon · 12. Mai 2026
Large language models (LLMs) achieve remarkable generative performance, yet their output quality is dependent on the decoding strategy. While sampling-based methods (e.g., top-k, nucleus) and search-and-select based methods (e.g., beam search, best-of-n, majority voting) can improve upon greedy deco…
- What Will Happen Next: Large Models-Driven Deduction for Emergency Instances
Zhengqing Hu, Dong Chen, Junkun Yuan, Liang Liu, Hua Wang, Zhao Jin, Yingchaojie Feng, Wei Chen, Mingliang Xu · 12. Mai 2026
Traditional simulation methods reproduce occurred emergency instances through presetting to assist people in risk assessment and emergency decision-making. However, due to the lack of randomness and diversity, existing simulation systems struggle to fully explore the potential risk as emergency inst…
- Primal-Dual Guided Decoding for Constrained Discrete Diffusion
Federico Tomasi, Dmitrii Moor, Alice Wang, Mounia Lalmas · 12. Mai 2026
Discrete diffusion models generate structured sequences by progressively unmasking tokens, but enforcing global property constraints during generation remains an open challenge. We propose primal-dual guided decoding, an inference-time method that formulates constrained generation as a KL-regularise…
- Why Do DiT Editors Drift? Plug-and-Play Low Frequency Alignment in VAE Latent Space
Xiaoce Wang, Sifan Zhou, Kaifei Wang, Leli Xu, Xuerui Qiu, Tao He, Ming Li · 12. Mai 2026
Recent advances in diffusion transformers (DiTs) have enabled promising single-turn image editing capabilities. However, multi-turn editing often leads to progressive semantic drift and quality degradation.In this work, we study this problem from a latent-space frequency perspective by decomposing t…
- Probing the Impact of Scale on Data-Efficient, Generalist Transformer World Models for Atari
Jooyeon Kim · 12. Mai 2026
Developing generalist systems that retain human-like data efficiency is a central challenge. While world models (WMs) offer a promising path, existing research often conflates architectural mechanisms with the independent impact of model \emph{scale}. In this work, we use a minimalist transformer wo…
- On Variance Reduction in Learning Mean Flows
Juanwu Lu, Ziran Wang · 12. Mai 2026
One-step generative modeling has emerged as a leading approach to amortize the inference cost of diffusion and flow-matching models. Among distillation-free methods, MeanFlow training is notoriously unstable, with non-decreasing loss and unbounded gradient variance. In this work, we establish a theo…
- Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Models
Kai Zhao, Dongliang Nie, Yuchen Lin, Zhehan Luo, Yixiao Gu, Deng-Ping Fan, Dan Zeng · 12. Mai 2026
Joint-Embedding Predictive Architectures (JEPAs) provide a simpleframework for learning world models by predicting future latent representations.However, JEPA training is subject to a bias-variance tradeoff.Without sufficient structural constraints, excessive representationalvariance causes the mode…
- When Few Steps Are Enough: Training-Free Acceleration of Identity-Preserved Generation
Dongqi Zheng · 12. Mai 2026
Identity-preserved image generation is typically built on many-step diffusion backbones, making personalized generation expensive at deployment time. We show that this cost is often unnecessary for identity-conditioned FLUX generation. A frozen InfuseNet identity adapter trained with dev transfers d…
- Spectral Transformer Neural Processes
Xianhe Chen, Hao Chen, Yingzhen Li · 12. Mai 2026
Time series, spatial data, and images are natural applications of Neural Processes. However, when such data exhibit strong periodicity and quasi-periodicity, existing methods often suffer from underfitting and generalise poorly beyond the training distribution. In this work, we propose Spectral Tran…
- Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria
Juanxi Tian, Fengyuan Liu, Jiaming Han, Yilei Jiang, Yongliang Wu, Yesheng Liu, Haodong Li, Furong Xu, Wanhua Li · 12. Mai 2026
Aligning multimodal generative models with human preferences demands reward signals that respect the compositional, multi-dimensional structure of human judgment. Prevailing RLHF approaches reduce this structure to scalar or pairwise labels, collapsing nuanced preferences into opaque parametric prox…
- Lattice Deduction Transformers
Liam Davis, Leopold Haller, Alberto Alfarano, Mark Santolucito · 12. Mai 2026
We introduce the Lattice Deduction Transformer (LDT), a recurrent transformer that approximates logically sound deduction by projecting its latent state through a lattice between forward passes. We train on-policy in a process that mirrors deduction in a search-based constraint solver and supervise …
- Deterministic Decomposition of Stochastic Generative Dynamics
Xingyu Song, Yuan Mei, Naoya Takeishi · 12. Mai 2026
Modern generative models can be understood as probability transport from a simple base distribution to a target data distribution. Deterministic transport models offer tractable velocity-field parameterizations, whereas stochastic generative models capture richer density evolution through drift and …
- Constant-Target Energy Matching: A Unified Framework for Continuous and Discrete Density Estimation
Zhijun Zeng, Yixuan Jiang, Pipi Hu, Zuoqiang Shi · 12. Mai 2026
Density estimation is a central primitive in probabilistic modeling, yet continuous, discrete, and mixed-variable domains are often treated by separate objectives, limiting the ability to exploit a common statistical structure across data types. Continuous score-based methods rely on log-density gra…
- MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing
Liwei Cheng, Zirui Song, Shibo Feng, Lunjie Zhou, Yixuan Guan, Dayan Guan · 12. Mai 2026
Text-in-image editing has become a key capability for visual content creation, yet existing benchmarks remain overwhelmingly English-centric and often conflate visual plausibility with semantic correctness. We introduce MULTITEXTEDIT, a controlled benchmark of 3,600 instances spanning 12 typological…
- The Safety-Aware Denoiser for Text Diffusion Models
Amman Yusuf, Zhejun Jiang, Mijung Park · 12. Mai 2026
Recent work on text diffusion models offers a promising alternative to autoregressive generation, but controlling their safety remains underexplored. Existing safety approaches are geared toward autoregressive models and typically rely on post-hoc filtering or inference-time interventions. These are…
- LaWM: Least Action World Models for Long-Horizon Physical Consistency from Visual Observations
Qixin Xiao, Maani Ghaffari · 12. Mai 2026
Learning predictive world models from visual observations is a core problem in embodied AI, with applications to model-based reinforcement learning and robotic planning. Existing latent world models typically generate future states with unconstrained neural transition functions, while modern video g…
- What Cohort INRs Encode and Where to Freeze Them
Vasiliki Sideri-Lampretsa, Sophie Starck, Robbie Holland, Julian McGinnis, Daniel Rueckert · 12. Mai 2026
Reusing the early layers of cohort-trained INRs as initialization for new signals has been shown to accelerate and improve signal fitting, yet it remains unclear which layers of the shared encoder learn transferable representations and what those representations encode. We address both questions for…
- VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
Kuanwei Lin, Wenhao Zhang, Ge Li · 11. Mai 2026
Video large multimodal models increasingly face a scalability bottleneck: long videos produce excessively long visual-token sequences, which sharply increase memory and latency during inference. While existing compression methods are effective in specific settings, most are either weakly query-aware…
- Inference-Time Attribute Distribution Alignment for Unconditional Diffusion
Hao Luan, See-Kiong Ng, Chun Kai Ling · 11. Mai 2026
Inference-time controllable generation is essential for real-world applications of unconditional diffusion models. However, most existing techniques focus on individual samples, struggling in applications that require the sample population to follow specific attribute distributions (e.g., demographi…
- Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement
Roussel Desmond Nzoyem, Mauro Comi · 11. Mai 2026
Training world models on vast quantities of unlabelled videos is a critical step toward fully autonomous intelligence. However, the prevailing paradigm of encoding raw pixels into opaque latent spaces and relying on heavy decoders for reconstruction leaves these models computationally expensive and …
