Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
2 529 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- Causal Evidence of Stack Representations in Modeling Counter Languages Using Transformers
Nishit Singh · 3 juin 2026
Formal languages have proven to be effective conduits to understand the inner mechanisms of transformers. Past work has shown that transformers trained on next token prediction over counter languages learn representations consistent with an underlying stack structure. Beyond representational analysi…
- AugMask: Training Diffusion Models on Incomplete Tabular Data via Stochastic Augmentation and Masking
Jungkyu Kim, Taeyoung Park, Kibok Lee · 3 juin 2026
Score-based diffusion models have emerged as prominent deep generative models; however, their application to tabular data remains challenging because their backbones assume fully specified inputs, whereas real-world tabular data often contain missing values. We propose AugMask, a plug-and-play train…
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
Zhiyuan Liu, Yicun Yang, Yaojie Zhang, Junjie Chen, Chang Zou, Qingyuan Wei, Shaobo Wang, Yichen Zhu, Linfeng Zhang · 3 juin 2026
Autoregressive Models (ARMs) have long dominated the landscape of Large Language Models. Recently, a new paradigm has emerged in the form of diffusion-based Large Language Models (dLLMs), which generate text by iteratively denoising masked segments. This approach has shown significant advantages and…
- A Factorized Low-Rank RNN Framework for Uncovering Independent Neural Latent Dynamics and Connectivity
Chengrui Li, Yunmiao Wang, Yule Wang, Weihan Li, Dieter Jaeger, Anqi Wu · 3 juin 2026
Low-rank recurrent neural networks (lrRNNs) are a class of models that uncover low-dimensional latent dynamics underlying neural population activity. Although their functional connectivity is low-rank, it lacks independence interpretations, making it difficult to assign distinct computational roles …
- High-Dimensional Latents Should Be Diagnosed Through Phase Structure
Alejandro Ascarate, Leo Lebrat, Rodrigo Santa Cruz, Clinton Fookes, Olivier Salvado · 3 juin 2026
We study autoencoder and variational-autoencoder latent spaces through the lens of spin-glass theory. The paper has two components. First, we formalize a latent-space spin-glass dictionary: for a fixed decoder, the reconstruction term together with a hyperspherical coordinates prior induces a Hamilt…
- Conditional Latent Diffusion Model with Fourier-based Motion Modelling for Virtual Population Synthesis
Shaokun Lan, Haoran Dou, Jinghan Huang, Arezoo Zakeri, Fengming Lin, Zherui Zhou, Jinming Duan, Alejandro F. Frangi · 3 juin 2026
In-silico trials of medical devices require the generation of virtual populations of anatomies. In cardiovascular applications, virtual anatomy is typically represented as a 3D+t mesh sampled from a generative model. However, most existing mesh generators focus on static anatomy, while sequence mode…
- GuidedBridge: Training-freely Improving Bridge Models with Prior Guidance
Zehua Chen, Yucheng Yang, Binjie Yuan, Kaiwen Zheng, Jun S. Liu, Jun Zhu · 3 juin 2026
Guidance methods, such as classifier-free guidance (CFG) and auto-guidance (AG), have advanced noise-to-data generation in diffusion models. Recently, bridge models have introduced a data-to-data generative process that can exploit an instructive clean prior. In this work, inspired by previous metho…
- Formalizing the Binding Problem
Lianghuan Huang, Yihao Li, Saeed Salehi, Yingshan Chang, Ansh Soni, Konrad P. Kording · 3 juin 2026
Representations of the world, arguably, contain information about features (e.g. something is blue, something is a circle) but also information about which features are part of the same object (e.g. the circle is blue), which we call binding information. Any system with the ability to understand sce…
- TreeFlash: Parallel AR-Approximation for Faster Speculative Decoding
Peer Rheinboldt, Fr\'ed\'eric Berdoz, Roger Wattenhofer · 3 juin 2026
One-shot block drafters for speculative decoding generate the full draft in a single forward pass, achieving strong throughput by eliminating sequential token generation. However, they predict each draft token conditioned only on the prefix context, with no dependence on previously drafted tokens. T…
- Training-Free Multi-Concept LoRA Composition with Prompt-Aware Weighting
Georgios Tsoumplekas, Stella Bounareli, Vasileios Argyriou · 3 juin 2026
Low-Rank Adaptation (LoRA) successfully enables personalization in text-to-image generation by adapting pre-trained diffusion models to specific visual concepts and styles. However, extending such models to multi-concept customization remains challenging. Naively combining multiple LoRA weights or t…
- Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs
Wei Ding, Yudong Zhang, Ruobing Xie, Xingwu Sun, Jiansheng Chen, Yu Wang · 3 juin 2026
As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a prerequisite for principled design. Yet large vision-language models (LVLMs) currently lack the tools to do so, and parameter-efficient encoder confi…
- Low-Frequency Shortcuts in Texture-Driven Visual Learning
Utku \c{S}irin, Cathy Hou, David Alvarez-Melis, Stratos Idreos · 3 juin 2026
Neural networks suffer from shortcut learning, where learned features generalize well to the training set but not to in-distribution (ID) or out-of-distribution (OOD) test sets. Existing studies are all based on a few standard benchmarks, which are shape-driven. Numerous application domains, however…
- A Geometric Lens on Physics-Aligned Data Compression
Aleix Segui, Wesley Armour · 3 juin 2026
In AI for Science, physics-informed losses are increasingly used to train learned compressors for scientific data, but their rate-distortion implications remain poorly understood. At fixed bitrate, these objectives often improve preservation of a target physical observable while degrading standard r…
- What Do Students Learn? A Feature-Level Analysis of Dark Knowledge
Seungu Kang, Songkuk Kim · 3 juin 2026
Knowledge Distillation (KD) is a powerful tool for model compression, yet the precise mechanisms by which student models acquire feature representations remain underexplored. In this work, we analyze student feature learning using the Interaction Tensor framework. Our analysis reveals that effective…
- MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data
Teng Hu, Mingchun Lu, Yating Wang, Jiangning Zhang, Jinkun Hao, Ye Pan, Ran Yi, Lizhuang Ma, Dacheng Tao · 3 juin 2026
Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a single perspective. Extending these models to multi-agent settings introduces two critical challenges: data scarcity (coord…
- Are we really tilting? The mechanics of reward guidance in flow and diffusion models
Sanjit Dandapanthula, Nicholas M. Boffi · 3 juin 2026
Reward guidance algorithms steer a learned generative process toward the reward-tilted measure at inference time. While empirically powerful, these methods are prone to reward hacking: the guided model over-optimizes the reward at the cost of fidelity to the learned distribution. Prior work has attr…
- You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
Kairan Zhao, Eleni Triantafillou, Peter Triantafillou · 2 juin 2026
Generative models have been shown to "memorize" certain training data, leading to verbatim or near-verbatim generating images, which may cause privacy concerns or copyright infringement. We introduce Guidance Using Attractive-Repulsive Dynamics (GUARD), a novel framework for memorization mitigation …
- Geometric Erasure by Contrastive Velocity Matching in Rectified Flows
Jonas Henry Grebe, Tobias Braun, Anna Rohrbach, Marcus Rohrbach · 2 juin 2026
While the rapid adoption of multimodal generative models offers immense potential, it has also increased the risks of harmful content synthesis, deepfakes, and copyright infringements. To address these challenges, concept erasure has emerged as a prospective safeguard. However, as the field graduall…
- Hallucination-Aware Diffusion Sampling for Inverse Problems via Robust Prior Updates
Pengfei Jin, Yiqi Tian, Kailong Fan, Bingjie Qi, Quanzheng Li · 2 juin 2026
Diffusion-based inverse problem solvers can produce realistic reconstructions, but realism alone does not ensure that the recovered details are supported by the measurement. We study this failure as measurement-conditioned hallucination: visually meaningful content that is either implausible or inco…
- Initialization is Half the Battle: Generating Diverse Images from a Guidance Potential Posterior
Xiang Li, Dianbo Liu, Kenji Kawaguchi · 2 juin 2026
Despite the remarkable fidelity of generative models, they frequently suffer from mode collapse. Existing strategies for enhancing diversity predominantly focus on intervening during the generation trajectory. We identify a critical oversight that the standard Gaussian initialization often causes tr…
- Guidance for Low-Level Perceptual Editing in Unconditional Diffusion Models
Shreyansh Modi, Akshat Tomar, Aarush Aggarwal · 2 juin 2026
Unconditional diffusion models offer powerful generative priors, yet steering them toward aesthetically enhanced outputs remains largely unexplored. We show that h-space patching, the dominant paradigm for training-free diffusion editing, systematically fails for global, low-level transformations re…
- OctoT2I: A Self-Evolving Agentic Text-to-Image Router
Xu Jiang, Bin Chen, Gehui Li, Yule Duan, Ronggang Wang, Jian Zhang · 2 juin 2026
The explosive growth of Text-to-Image (T2I) models, from large-scale versions to lightweight, real-time ones, now faces diminishing marginal returns from single-model scaling. Agentic T2I methods emerged to alleviate this bottleneck by using multiple models. However, existing agentic T2I methods suf…
- Theoretical Analysis of Engression and Reverse Markov Engression
Jiaqi Huang, Gongjun Xu, Ji Zhu · 2 juin 2026
Engression is a recently proposed and effective framework for conditional distribution learning. Its multi-step Reverse Markov extension further improves generative flexibility by decomposing complex conditional sampling into sequential reverse transitions. Despite their strong empirical performance…
- Efficient Synthetic Network Generation via Latent Embedding Reconstruction
Feifan Jiang, Yinan Bu, Shihao Wu, Gongjun Xu, Ji Zhu · 2 juin 2026
Network data are ubiquitous across the social sciences, biology, and information systems. Generating realistic synthetic network data has broad applications from network simulation to scientific discovery. However, many existing black-box approaches for network generation tend to overfit observed da…
- How Optimality Structures Sparse Dictionaries: A Theory for Understanding SAE Representations
William Dorrell · 2 juin 2026
Sparse Autoencoders (SAEs) have found success parsing neural representations into interpretable concepts, providing a basis for understanding and control. However, what exactly SAEs extract, and, correspondingly, the scientific conclusions we can draw from them, are not obvious. Empirically, the pro…
