Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
4 992 papiers indexés
Les méthodes de génération d’images par intelligence artificielle explorent des architectures où deux réseaux s’affrontent ou collaborent pour produire des contenus visuels. Parmi ces approches, les modèles comme les Generative Adversarial Networks et les diffusion models cherchent à contrôler finement la synthèse d’images, que ce soit par des mécanismes de guidage, des opérateurs de composition ou des techniques d’attention éparse. Les travaux récents portent sur l’optimisation des étapes d’apprentissage, la manipulation des espaces latents, ou encore l’intégration de contraintes comme la causalité dans la génération de vidéos ou d’ensembles cohérents comme des tenues vestimentaires.
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Pays des laboratoires
- Chine46 % · 1 644 articles
- États-Unis35 % · 1 261 articles
- Royaume-Uni6,8 % · 244 articles
- Corée du Sud6,2 % · 223 articles
- Allemagne5,3 % · 191 articles
- R.A.S. chinoise de Hong Kong4,6 % · 165 articles
- Canada4 % · 144 articles
- Singapour4 % · 142 articles
Sur 3 586 articles de ce sujet dont au moins un laboratoire est situé. 88 pays représentés.
Il s'agit du pays du laboratoire, jamais de la nationalité des personnes. Un article signé depuis plusieurs pays compte pour chacun d'eux, les parts dépassent donc 100 % au total. La couverture est partielle et le manque n'est pas aléatoire : un chercheur dont l'institution est inconnue publie en général peu, ce qui sur-représente les laboratoires établis.
Derniers papiers
- IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models
Vladislav Gromadskii, David Li, Samson Gourevitch, Yazid Janati, Eric Moulines, Maxim Panov, Alexander Korotin · 5 octobre 2026
Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generator…
- Soundwich: Video Generation with Layered and Controllable Audio
Zhuo Ning, AmirHossein Naghi Razlighi, Sagi Polaczek, Daniel Cohen-Or, Ali Mahdavi-Amiri · 5 octobre 2026
Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workflows, where speech, music, sound effects, and ambient sounds are repr…
- Branch-Centric Tokenization and Test-Time Augmentation for Skeleton Generation
Zhengyuan Li, Chuanyu Pan, Yuanming Hu, Raymond A. Yeh · 5 octobre 2026
Automatic skeleton generation involves predicting both joint positions and skeletal connectivity. However, existing approaches struggle to encode branch structures into token sequences and do not use test-time computation effectively. We study these choices within a unified autoregressive framework.…
- DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation
Jiahao Zhan, Yan Wang, Yongrui Ma, Qunliang Xing, Ruchang Yao, Runtao Liu, Shijie Zhao, Tianfan Xue · 5 octobre 2026
Streaming video generation has benefited from distribution matching distillation (DMD), which matches the joint distribution of video frames to a video teacher's approximation of the real video distribution. Although this joint matching mitigates drift during autoregressive rollouts, limitations rem…
- ProgressNet: Sketching and Prompting with a Frozen Text-to-Image Model
Arkaprabha Basu, Chaitat Utintu, Yi-Zhe Song · 5 octobre 2026
Humans draw progressively: a few strokes, a look at the result, a stroke erased, a prompt revised. Image generators do not work this way. They typically take a finished sketch and produce the image in a single pass, so every edit starts the picture again, and the models that do keep state across tur…
- UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation
Daikun Liu, Xin Zhan, Teng Wang, Xiaoping Wang, Changyin Sun · 5 octobre 2026
We propose UniDynamics, a diffusion-based framework for future 4D dynamic scenes (RGB, depth, and optical flow) generation from a single event-RGB pair, without requiring long histories or control priors as in existing methods, while explicitly modeling future motion fields. The core idea is to leve…
- VDOT++: Unified Few-Step Video Generation via Unbalanced Optimal Transport Distillation
Yutong Wang, Xingtong Ge, Enhuai Liu, Yunke Wang, Tianfan Xue, Yu Qiao, Yaohui Wang, Xinyuan Chen, Chang Xu · 5 octobre 2026
Video creation spans text-to-video (T2V), image-to-video (I2V), and condition-based generation, yet video diffusion models remain costly because they repeatedly evaluate large backbones during sampling. Distribution matching distillation (DMD) reduces this cost, but its reverse Kullback--Leibler (KL…
- Does Physics Live in the Activations? Localizing Physical Quantities in Video Diffusion Models
Jonas Kneifl, Jakub Skalski, Bart{\l}omiej Twardowski, Kamil Deja · 5 octobre 2026
Video generation models produce strikingly realistic sequences and are increasingly proposed as world models, yet recent benchmarks reveal pronounced deficits in their physical reasoning. This raises the question of whether these models internalize physical principles or merely reproduce familiar mo…
- In-Distribution Forcing for Long Video Generation at Test Time
Jeongwoo Shin, Youngyoon Choi, Sangwoo Jo, Hyunmog Kim, Sungjoon Choi, Joonseok Lee, Jaewoong Choi, Jaemoo Choi · 5 octobre 2026
Modern autoregressive (AR) video diffusion models excel at short-horizon video generation, yet generating long videos remains challenging due to drifting, where colors and textures shift, and motion dynamics decay. Existing works primarily rely on KV conditioning, which selects or modifies cached ke…
- Custom Forcing: Training-Free Subject Customization for Autoregressive Video Generation
Yunseung Ok (Kyung Hee University), Hyunsoo Kim (The University of Texas at Austin), Minseo Kim (Kyung Hee University), Suhyun Kim (Kyung Hee University) · 5 octobre 2026
Autoregressive video models can generate minute-long videos in real time, but they produce generic subjects from text rather than specific subjects from user-provided images. Existing customization methods either require costly per-subject optimization or use pretrained conditioning networks that jo…
- TRAC: Trajectory-aware Reuse and Adaptive Correction for Efficient Autoregressive Video Generation
Jiaxing Song, Weiqi Yan, You Huang, Mingte Qiu, Huazhong Liu, Xiaofeng Zhu, Yunshan Zhong · 5 octobre 2026
In this paper, we present trajectory-aware reuse and adaptive correction (TRAC), a training-free framework for efficient autoregressive (AR) video generation. Existing acceleration methods mainly target single-trajectory generation with bidirectional attention. AR video generation, by contrast, sequ…
- VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
Huawei Lin, Tony Geng, Zhaozhuo Xu, Weijie Zhao · 5 octobre 2026
Autoregressive (AR) models have recently shown strong performance in image generation, where a critical component is the visual tokenizer (VT) that maps continuous pixel inputs to discrete token sequences. The quality of the VT largely defines the upper bound of AR model performance. However, curren…
- VIDiff: Translating Videos via Multi-Modal Instructions with Diffusion Models
Zhen Xing, Shuyuan Tu, Qi Dai, Zihao Zhang, Hui Zhang, Han Hu, Zuxuan Wu, Yu-Gang Jiang · 5 octobre 2026
Diffusion models have achieved significant success in image and video generation. This motivates a growing interest in video editing tasks, where videos are edited according to provided text descriptions. However, most existing approaches only focus on video editing for short clips and rely on time-…
- LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation
Ziqi Ma, Shreya Sharma, Mohamed El Banani, Katja Schwarz, Chongjie Ye, Chao-Yuan Wu, Li Fei-Fei, Ben Mildenhall, Georgia Gkioxari, Justin Johnson, Gowthami Somepalli · 5 octobre 2026
Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Existing post-training techniques, which assign a single scalar reward to …
- Contextual Flow Matching: Adaptive Step Selection in Flow Models for Efficient Visual Generation
Divya Jyoti Bajpai, Arun Verma, Manjesh Kumar Hanawal · 5 octobre 2026
Flow Matching enables high-quality visual generation via continuous-time dynamics, but inference remains costly due to multiple sequential function evaluations. Existing acceleration methods reduce the number of function evaluations but often introduce additional training overhead, degrade quality, …
- Correcting Guided Diffusion Trajectories with Spectral Alignment
Gihoon Kim, Taesup Kim · 5 octobre 2026
The practical success of conditional image generation hinges on fine-grained differences in condition alignment and visual fidelity. Classifier-free guidance (CFG) is central to this success, but its lack of an explicit criterion makes it difficult to assess whether the guided trajectory is progress…
- DAGS: Disentangled Appearance-and-Geometry Steering of a Frozen Image DiT for Temporally Stabilized Generative Rendering
Karthik Mohan Kumar, Damian Andrysiak, Pedro Antonio Pena, Kunal Tyagi, Rama Harihara · 5 octobre 2026
Diffusion transformers (DiTs) generate high-fidelity images from text and image conditions, but their outputs carry large variance and their faithfulness to a desired target depends heavily on how the condition is supplied. We present DAGS, a lightweight, attention-free, disentangled appearance and …
- Depth as Time in One-Step Generative Models
Arnold Caleb Asiimwe, William Yang, Sanghyuk Chun, Esin Tureci, Olga Russakovsky · 5 octobre 2026
The recent wave of one-step generative models, which compress the multi-step trajectory of diffusion via either distillation or learned flow maps, has reached an inflection point where they can generate high-quality images. Here, we ask a natural question that follows from these advances: what happe…
- Know When to Hold 'em: Correct-Token Retention in Uniform-State Diffusion Language Models
Mojtaba Nafez, James Henderson · 2 octobre 2026
Uniform-state diffusion models (USDMs) can revise any token at any denoising step, which lets them correct their own mistakes, a key advantage over masked diffusion. Self-correction, however, requires both revising incorrect tokens and retaining correct ones, and we show that current USDMs lack the …
- Diffusion Editing with Soft Mask: Pixel Level Redo of Image and Video with Adjustable Strength
Candi Zheng, Yuan Lan · 2 octobre 2026
Diffusion models with prompt and reference image-guided editing have seen rapid progress, yet they remain too coarse for pixel-level control. One promising direction is to incorporate a soft mask that specifies spatially varying edit strengths but training such fine-grained control demands expensive…
- BAM! Bayesian Anything Model: a foundation model for generative computational imaging
Alessio Spagnoletti, Charlesquin Kemajou Mbakam, Jonathan Spence, Andr\'es Almansa, Marcelo Pereyra · 1 octobre 2026
Generative models are transforming Bayesian computational imaging, yet the field still lacks physics-aware foundation models. Current practice falls into two camps. Large foundation image models are deployed as plug-and-play priors with zero-shot approximate likelihood guidance, which introduces sig…
- Diffusable Latents from Structure-Agnostic Distillation
Adrien Ramanana Rahary, Nicolas Dufour, Patrick P\'erez, David Picard · 1 octobre 2026
Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and reach higher sample quality. Standard distillation aligns the latent at each position to a co-located teacher feature, tying the latent layout to the…
- TexTailor: Texture-Preserving Video Virtual Try-On via Adaptive Garment Conditioning
Zijing Qin, Jun Zhou, Ruicheng Zhang, Jiaqi Hou, Zunnan Xu, Ronghui Li, Zhenyu Xie, Xiu Li · 1 octobre 2026
Video virtual try-on has attracted increasing attention due to its broad potential in digital fashion and intelligent e-commerce. However, existing methods primarily focus on low-resolution settings and still face substantial challenges when extended to high-resolution scenarios. These limitations c…
- DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence
Xu Huang, Ye Huang, Zijun Liao, Yuwei Niu, Xiaojie Li, Menghan Zhou, De Wen Soh, Xiaotong Li, Daquan Zhou · 1 octobre 2026
High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff between reconstruction fidelity and generation efficiency: high compression image encoder always increases the learning difficulty of diffusion trainin…
- TripleFlow: Training-Free Video Object Removal by Bridging Residual Editing and Native Generation
Songhe Wang, Lifu Wei, Shuolin Xu, Charles A. Kamhoua, David Miller · 1 octobre 2026
Video object removal presents a uniquely difficult editing challenge. Because a removal prompt specifies only what to erase rather than what to generate, the model must infer and reconstruct a highly specific occluded background entirely from the surrounding context. Existing training-free methods s…
Autres sujets du thème Vision par ordinateur et reconnaissance de formes
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Multimodal Machine Learning Applications8 069 papiers / 12 mois+191 %
- Advanced Neural Network Applications2 354 papiers / 12 mois+48 %
- Advanced Vision and Imaging841 papiers / 12 mois+78 %
- Human Pose and Action Recognition836 papiers / 12 mois+457 %
- Face recognition and analysis482 papiers / 12 mois+88 %
- Image Enhancement Techniques463 papiers / 12 mois+20 %
