Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
4992 artículos indexados
Los métodos de generación de imágenes mediante inteligencia artificial exploran arquitecturas donde dos redes compiten o colaboran para producir contenidos visuales. Entre estos enfoques, modelos como los Generative Adversarial Networks y los diffusion models buscan controlar de manera precisa la síntesis de imágenes, ya sea mediante mecanismos de guía, operadores de composición o técnicas de atención dispersa. Los trabajos recientes se centran en la optimización de las etapas de aprendizaje, la manipulación de los espacios latentes o incluso la integración de restricciones como la causalidad en la generación de videos o conjuntos coherentes como atuendos.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- China46 % · 1644 artículos
- Estados Unidos35 % · 1261 artículos
- Reino Unido6,8 % · 244 artículos
- Corea del Sur6,2 % · 223 artículos
- Alemania5,3 % · 191 artículos
- RAE de Hong Kong (China)4,6 % · 165 artículos
- Canadá4 % · 144 artículos
- Singapur4 % · 142 artículos
Sobre 3586 artículos de este tema con al menos un laboratorio localizado. 88 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Know When to Hold 'em: Correct-Token Retention in Uniform-State Diffusion Language Models
Mojtaba Nafez, James Henderson · 2 de octubre de 2026
Uniform-state diffusion models (USDMs) can revise any token at any denoising step, which lets them correct their own mistakes, a key advantage over masked diffusion. Self-correction, however, requires both revising incorrect tokens and retaining correct ones, and we show that current USDMs lack the …
- Diffusion Editing with Soft Mask: Pixel Level Redo of Image and Video with Adjustable Strength
Candi Zheng, Yuan Lan · 2 de octubre de 2026
Diffusion models with prompt and reference image-guided editing have seen rapid progress, yet they remain too coarse for pixel-level control. One promising direction is to incorporate a soft mask that specifies spatially varying edit strengths but training such fine-grained control demands expensive…
- BAM! Bayesian Anything Model: a foundation model for generative computational imaging
Alessio Spagnoletti, Charlesquin Kemajou Mbakam, Jonathan Spence, Andr\'es Almansa, Marcelo Pereyra · 1 de octubre de 2026
Generative models are transforming Bayesian computational imaging, yet the field still lacks physics-aware foundation models. Current practice falls into two camps. Large foundation image models are deployed as plug-and-play priors with zero-shot approximate likelihood guidance, which introduces sig…
- Diffusable Latents from Structure-Agnostic Distillation
Adrien Ramanana Rahary, Nicolas Dufour, Patrick P\'erez, David Picard · 1 de octubre de 2026
Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and reach higher sample quality. Standard distillation aligns the latent at each position to a co-located teacher feature, tying the latent layout to the…
- TexTailor: Texture-Preserving Video Virtual Try-On via Adaptive Garment Conditioning
Zijing Qin, Jun Zhou, Ruicheng Zhang, Jiaqi Hou, Zunnan Xu, Ronghui Li, Zhenyu Xie, Xiu Li · 1 de octubre de 2026
Video virtual try-on has attracted increasing attention due to its broad potential in digital fashion and intelligent e-commerce. However, existing methods primarily focus on low-resolution settings and still face substantial challenges when extended to high-resolution scenarios. These limitations c…
- DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence
Xu Huang, Ye Huang, Zijun Liao, Yuwei Niu, Xiaojie Li, Menghan Zhou, De Wen Soh, Xiaotong Li, Daquan Zhou · 1 de octubre de 2026
High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff between reconstruction fidelity and generation efficiency: high compression image encoder always increases the learning difficulty of diffusion trainin…
- TripleFlow: Training-Free Video Object Removal by Bridging Residual Editing and Native Generation
Songhe Wang, Lifu Wei, Shuolin Xu, Charles A. Kamhoua, David Miller · 1 de octubre de 2026
Video object removal presents a uniquely difficult editing challenge. Because a removal prompt specifies only what to erase rather than what to generate, the model must infer and reconstruct a highly specific occluded background entirely from the surrounding context. Existing training-free methods s…
- Uncertainty-Aware Consistency Distillation for Few-Step Video Generation
Lingyu Liu, Yaxiong Wang, Li Zhu, Zhedong Zheng · 1 de octubre de 2026
We study few-step video generation, i.e., distilling a multi-step video generator, which typically requires tens of sampling steps, incurring substantial latency and compute, into a few-step student. Consistency distillation is a common recipe, in which a multi-step teacher provides the consistency …
- EPIC: Epipolar-Consistent 360{\deg} Immersive Stereo Video Generation
Debabrata Mandal, Dongdong Fu, Jonathon Miller, William Villareal, Xi Peng, Praneeth Chakravarthula · 1 de octubre de 2026
Immersive displays can enable rich and diverse virtual experiences. Manually authoring every possible experience to realize this potential, however, is prohibitively expensive, difficult to scale, and impractical. Generative AI models could remove this bottleneck, but today's models are built for co…
- Curating Synthetic Data for Task-Specific Visual Perception
Saptarshi Neil Sinha, Paul Julius K\"uhn, Michael Weinmann · 1 de octubre de 2026
Synthetic data are most valuable where general-purpose datasets cannot provide the domain-specific priors a task requires, and where manual annotation is expensive, imprecise, or infeasible. In this article we argue that the central question for specialized vision systems is not how to generate more…
- ExploreNet: Learning Where to Explore in Diffusion GRPO
Shuyue Stella Li, Xiaochuang Han, Yulia Tsvetkov, Luke Zettlemoyer · 1 de octubre de 2026
Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This noise decides which rollouts the model learns from, yet it perturbs every channel and spatial position of the latent equally. In this paper, we inste…
- Masked Swingers: Harnessing Data Augmentation to Advance Autoencoders for Self-Supervised Learning
Anthony Fuller, Scott C. Lowe, Daniel G. Kyrollos, Graham W. Taylor, Evan Shelhamer, James R. Green · 1 de octubre de 2026
Self-supervised learning (SSL) removes the need for annotations and makes models that are capable across more domains than supervised learning. The autoencoder SSL framework learns by reconstructing its own input after information loss through a bottleneck or noise injection. Masked autoencoders (MA…
- E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models
Arseny Ivanov, Alexander Kolesov, Alexander Korotin, Ivan Oseledets, Mikhail Goncharov · 1 de octubre de 2026
Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most…
- FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models
Darshan Thaker, Lachlan Ewen MacDonald, Ren\'e Vidal · 1 de octubre de 2026
Gradient-based reward guidance provides a flexible way to use downstream reward models to control masked diffusion language models at inference time. However, its computational cost remains high as each decoding iteration incurs expensive diffusion model forward passes and reward model backpropagati…
- D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders
Xinyue Xu, Jiahao Zhang, Lijie Hu, Peter Hase, Hao Wang · 1 de octubre de 2026
Sparse autoencoders (SAEs) reveal visual structure in diffusion transformers (DiTs), but interpreting a feature does not establish whether it can be used to control generation. We introduce D-Scope (Diffusion Scope), a framework that connects feature interpretation to generation control through shar…
- Steering Fields: Adaptive Vector Fields for Safe Image Generation and Beyond
Simone Facchiano, Jan Eric Lenssen, Bernt Schiele, Wolfgang Stammer, Fabio Galasso, Jonas Fischer · 1 de octubre de 2026
As state-of-the-art text-to-image flow models achieve near-photorealistic quality, controlling their outputs, e.g., suppressing harmful content while promoting benign alternatives, has become a central challenge. The current steering paradigm consists of adding a global steering vector to selected a…
- Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models
Junjian Li, Xiaolong Liu, Peng Sun, Liantao Wu, Linghan Chen, Yudong Gao, Honglong Chen · 1 de octubre de 2026
Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack …
- PartiCam: Camera Controlled Video Generation with Reward Guidance
Amine Ouasfi, Runjia Li, Junlin Han, Eric Marchand, Philip H. S. Torr, Adnane Boukhayma · 1 de octubre de 2026
We present PartiCam, a training-free Particle filtering rooted method for improved Camera controlled video generation. Generating videos that follow a precisely specified camera trajectory remains challenging for large video diffusion models. Training-free approaches are backbone-agnostic and avoid …
- CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models
Shu Yu, Chaochao Lu · 1 de octubre de 2026
Online reinforcement learning has been extended to flow matching for diffusion model (DM) image generation. However, this paradigm faces three limitations: (1) Window selection. Existing methods manually set the stochastic differential equation (SDE) sampling window, i.e., the denoising steps where …
- Distilling Diffusion Score Discrepancy for Efficient Training Data Attribution
Shixuan Liu, Joan Serr\`a, Kin Wai Cheuk, Jinju Kim, Woosung Choi, Yukara Ikemiya, Wei-Hsiang Liao, Jiaqi W. Ma, Yuki Mitsufuji · 1 de octubre de 2026
Training data attribution for diffusion models aims to identify the training samples that influence a generated instance, but existing methods either require costly per-sample gradient computation or query-specific model optimization. Moreover, most methods attribute changes in a proxy loss rather t…
- ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images
Shubhang Bhatnagar, Ishan Bhatnagar, Viraj Shah, Narendra Ahuja · 1 de octubre de 2026
Text-to-image diffusion models are personalized to a subject by DreamBooth fine-tuning on a handful of its images. Increasingly, these images come from a diffusion model rather than a camera. We show that fine-tuning on such synthetic images degrades subject fidelity, producing oversaturated color a…
- Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion
Chong Wang, Zixuan Fu, Shiqi Huang, Siyuan Yang, Hao Cheng, Bihan Wen · 30 de septiembre de 2026
Pixel-space diffusion Transformers (DiTs) directly operate on high-dimensional visual data, yet their hidden representations typically undergo uniform refinement across depth. Natural images, however, are inherently organized at different levels of granularity. Global structure can often be represen…
- Tokens Change, Structure Endures: Spectral Watermarking for Generated Speech
Kanghwi Lee, Kyeongseok Jeong, Jeongmin Liu · 30 de septiembre de 2026
Watermarking is a promising tool for establishing the provenance of AI-generated speech. While many neural audio watermarking methods rely on a separately trained watermark generator, token-level watermarking is a training-free alternative that operates directly during generation. Its main weakness …
- From Scores to Samples: Elastic Forcing for Autoregressive Video Generation
Chi Zhang, Yueyi Liu, Haoyang Shi, Ruichuan An, Haoyu Li, Yuhang Wu, Sen Cui, Miao Liu · 30 de septiembre de 2026
Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-traini…
- Flow-JEPA: Robust Latent Dynamics for JEPA World Models via Flow Matching
Yanchen Huo, Ziying Song, Yadan Luo · 30 de septiembre de 2026
Joint-Embedding Predictive Architectures (JEPAs) provide a powerful framework for latent world modeling and planning in a reconstruction-free manner. Although numerous JEPA-based approaches have been proposed to mitigate representation collapse, our experiments on localized, out-of-distribution visu…
Otros asuntos del tema Visión por computador y reconocimiento de formas
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Multimodal Machine Learning Applications8069 artículos / 12 meses+191 %
- Advanced Neural Network Applications2354 artículos / 12 meses+48 %
- Advanced Vision and Imaging841 artículos / 12 meses+78 %
- Human Pose and Action Recognition836 artículos / 12 meses+457 %
- Face recognition and analysis482 artículos / 12 meses+88 %
- Image Enhancement Techniques463 artículos / 12 meses+20 %
