Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
2542 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- Measuring the Robustness of Audio Deepfake Detection under Real-World Corruption
Xiang Li, Pin-Yu Chen, Wenqi Wei · 7 de julio de 2026
Deepfakes have emerged as a widespread and rapidly escalating concern in generative AI, spanning images, audio, and videos. Among these, audio deepfakes are particularly alarming due to the growing accessibility of high-quality voice synthesis tools and the ease with which synthetic speech can be di…
- FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation
Weichen Qin, Yufan Xie, Peihao Wang, Chia-Jui Chou, Minghui Du, Peng Xu, Ziren Luo, Yi Yang, Jingyi Yu, Bo Liang, Jiakai Zhang · 7 de julio de 2026
Simulation-Based Inference (SBI) is critical for scientific discovery, with generative models offering a promising path toward efficient inference. However, existing methods struggle with effective multimodal modeling. They often rely on brute-force fusion strategies that ignore the structural dispa…
- Constrained Flow Matching via Lagrangian Dual Flows
Vince Kurtz, Alexander Davydov · 7 de julio de 2026
Flow matching is a powerful tool for generative modeling, but emerging applications in robotics, planning, and physics require inference-time constraints on generated outputs. Such constraints are often complex and highly nonlinear. As a result, methods designed for linear constraints like image inp…
- LGQ: Learnable Geometric Quantization for Image Tokenization
Idil Bilge Altun, Mert Onur Cakiroglu, Elham Buxton, Mehmet Dalkilic, Hasan Kurban · 7 de julio de 2026
Recent collapse-free quantizers such as FSQ achieve stable training by replacing the learnable codebook with an engineered geometry: a fixed scalar grid whose structure is dictated by the codebook size K. We show this trade-off is unnecessary. We introduce Learnable Geometric Quantization (LGQ), whi…
- Optimal Mixture-of-Experts Model Averaging for Conditional Generative Models
Shijin Gong, Baihua He, Xinyu Zhang · 7 de julio de 2026
Conditional generative models have emerged as powerful tools for sampling from target conditional distributions, driving substantial advances across a wide range of scientific and applied domains. As these models proliferate, practitioners often face multiple plausible generators whose performance c…
- ELBO-T2IAlign: A Generic ELBO-Based Method for Calibrating Pixel-level Text-Image Alignment in Diffusion Models
Qin Zhou, Zhiyang Zhang, Jinglong Wang, Xiaobin Li, Jing Zhang, Qian Yu, Lu Sheng, Dong Xu · 7 de julio de 2026
Diffusion models excel at image generation. Recent studies have shown that these models not only generate high-quality images but also encode text-image alignment information through attention maps or loss functions. This information is valuable for various downstream tasks, including segmentation, …
- Reflected Schr\"odinger Bridge Matching
Marcus H\"aggbom, Viktor Nilsson, Pierre Nyquist, Joakim and\'en · 7 de julio de 2026
Recent advances in generative modeling have enabled the efficient computation of Schr\"odinger bridges (SB) in high-dimensional settings by leveraging partially simulation-free training methods inspired by flow matching. However, these have not covered SBs with reflecting dynamics, a useful model ch…
- Tightening the Score Matching Gap for Diffusion Models
Benjamin Dupuis, Tyler Farghly, Maxime Haddouche, Alain Durmus, Umut Simsekli · 7 de julio de 2026
Diffusion models (DMs) are a state-of-the-art generative method to approximately sample from an unknown distribution. Their training and evaluation primarily rely on an Evidence Lower Bound (ELBO), which relates the Kullback-Leibler (KL) divergence of model samples to the score matching loss along t…
- Non-asymptotic Convergence of Stochastic Gradient Descent in Score-based Generative Models
Stanislas Strasman (SU, LPSM), Sobihan Surendran (SU, LPSM), Sylvain Le Corff (SU, LPSM) · 7 de julio de 2026
Score-based Generative Models (SGMs) have achieved impressive performance in data generation across a wide range of applications. While the statistical properties of their sampling procedures are increasingly well understood, the optimization dynamics underlying their training remain less explored. …
- Is Generation Required for Data-Efficient Perception?
Jack Brady, Bernhard Sch\"olkopf, Thomas Kipf, Simon Buchholz, Wieland Brendel · 7 de julio de 2026
It has been hypothesized that achieving the data efficiency of human visual perception requires a generative approach in which internal representations result from inverting a decoder. Yet today's most successful vision models are non-generative, relying on an encoder that maps images to representat…
- CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion
Adam Fisch, Daniel Deutsch, Joshua Maynez, Alekh Agarwal, Jonathan Berant, William Cohen, Amir Globerson, Jacob Eisenstein · 7 de julio de 2026
Evaluating generative AI models is a routine, but resource-intensive, process that is conducted over and over again during the course of model development. In this work, we propose Collaborative Evaluation (CollabEval), a simple, effective, and principled method for exploiting dependencies between h…
- SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
Linxi Li, Yuncong Yu, Qianwei Guo, Liwei Jin, Yechen Wang, Carsten Maple · 7 de julio de 2026
While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects. Existing environmental audio datasets such as EnvSDD provide important initial resources, but remain limited in scale and generation provenance for studying iso…
- Vidu S1: A Real-Time Interactive Video Generation Model
Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, Yang Luo, Yuji Wang, Dechuang Chen, Jungang Li, Chengyang Ye, Marco Chen, Hongzhou Zhu, Min Zhao, Yuxuan Jiang, Zhengkun Huang, Chendong Xiang, Kaiwen Zheng, Haoxu Wang, Xiaohang Wang, Qi Jia, Xin Chen, Yimin Chen, Youhe Jiang, Fangcheng Fu, Zhijie Deng, Fan Bao, Jianfei Chen, Jun Zhu · 7 de julio de 2026
We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation content at any moment through voice instructions. Vidu S1 supports infinite-length real-time video generation without blurring, drift, or visual dis…
- DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics
Silin Gao, Hao Zhao, Zeming Chen, Sepideh Mamooler, Antara Raaghavi Bhattacharya, Qiyu Wu, Hiromi Wakaki, Yuki Mitsufuji, Li Mi, Syrielle Montariol, Antoine Bosselut · 7 de julio de 2026
Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require models to predict or simulate multiple levels of dynamic constituents, such as actions taken in the visual sequence, and the associated changes to the visu…
- Efficient bias mitigation in T2I diffusion models using Concept Graphs
Mansi, Avinash Kori, Francesco Leofante · 7 de julio de 2026
Text-to-Image diffusion models often propagate harmful bias inherited from the training data. Existing bias mitigation techniques typically intervene only at the text encoder or provide inference-time guidance, often leading to generations that collapse into semantically incoherent outputs. To addre…
- Observable- and Positional-Encoding-Dependent Symmetry Readout from Neural Network Weights
Naoya Chiba, Satoshi Sugiyama, Yuki Uranishi · 7 de julio de 2026
Post-hoc analysis of trained neural network weights often seeks to recover geometric structure directly from the parameters. We show that, for positional-encoding-equipped neural fields, the symmetry visible from weights is not the true symmetry group itself, but an observable symmetry set determine…
- Steering Optimisation Trajectories in Diffusion Representation Learning
Rajat Rasal, Avinash Kori, Tian Xia, Ben Glocker · 7 de julio de 2026
We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures. We trace this behaviour to optimisation dynamics; we analyse curves of image reconstruction against latent representation quality, revealing trajectories that organise arou…
- Towards Diverse and Comprehensive Benchmarks for Mutual Information Estimation
Alberto Foresti, Ivan Butakov, Alexander Tolmachev, Giulio Franzese, Alexey Frolov, Pietro Michiardi · 7 de julio de 2026
Mutual information (MI) estimation is a central problem in machine learning and statistics; however, existing benchmarks typically evaluate estimators on simplified, low-dimensional distributions, leaving their performance on complex, realistic data largely unexplored. We address this gap with a com…
- Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting
Disheng Liu, Tuo Liang, Chaoda Song, Yu Yin · 7 de julio de 2026
Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models. Existing approaches to exploiting this potential typically involve 1) training or fine-tuning generators, or 2) using lightweight post-hoc adaptation like prompt engin…
- HyperVAttention: Efficient Sparse Attention with Spatio-Temporal Clustering for Video Diffusion
Dongyeun Lee, Amir Zandieh, Vahab Mirrokni, Junmo Kim, Insu Han · 7 de julio de 2026
Video Diffusion Transformers (VDiTs) have demonstrated significant capabilities in high-fidelity video generation. However, their ability to produce long-duration videos is fundamentally constrained by the quadratic complexity of the self-attention mechanism. Recent clustering-based sparse attention…
- SPLIT: Training-Free AI-Generated and Partially Edited Video Detection via Spatial Patch-Level Incoherence and Temporal Roughness
Jongyeop Hyun, Hyounghun Kim · 7 de julio de 2026
Deploying AI-generated video detectors in real-world services demands an ultra-low false positive rate (FPR) on real videos to avoid falsely rejecting authentic content, a regime where standard metrics such as AUROC fail to reflect actual operating behavior. We introduce Spatial Patch-Level Incohere…
- Inclusive KL Gradient Flows: Otto-Wasserstein, Fisher-Rao-Gaussian, and Local-Estimator Dynamics
Jia-Jie Zhu · 7 de julio de 2026
Otto's Wasserstein gradient flow of the inclusive (forward) Kullback--Leibler (KL) divergence offers a principled framework for analyzing statistical inference algorithms, yet algorithms targeting the exclusive (reverse) KL divergence are rarely studied with such tools. We establish a unified gradie…
- CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation
Xuyao Huang, Zelai Deng, Xu Wang, Xizhong Xiao, Zhijie Deng · 7 de julio de 2026
The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dolly zoom, slow motion, etc. While Diffusion Transformers (DiTs) exhibit strong performance in video generation, their large parameter sizes and multi-step iterati…
- Separating Representation from Reconstruction Enables Scalable Text Encoders
Megi Dervishi, Mathurin Videau, Yann LeCun · 7 de julio de 2026
While decoders have rapidly scaled, encoders have remained largely unchanged since BERT. We revisit this disparity by frozen backbone evaluation via probing. Under this lens, the representations of BERT encoders become increasingly $\textit{unexploitable}$ by frozen probes, despite improved perplexi…
- CORA: Per-Slice Coherent Orthogonal Rotation for SVD-based Low-Rank Adaptation
Pengcheng Wang, Ziran Liu, Wei Wang, Wei Jiang · 7 de julio de 2026
Parameter-Efficient Fine-Tuning (PEFT) commonly adapts pretrained weights through low-rank updates, and recent methods further exploit the singular value decomposition (SVD) of the base weight for initialization or subspace selection. However, these methods do not explicitly preserve the coupled geo…
