Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
2542 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- dRAE: Representation Autoencoder with Hyper-Spherical Codes
Tianren Ma, Lin Long, Chuyan Chen, Mu Zhang, Junbo Zhao, Tong Zhang, Qixiang Ye · 27 de julio de 2026
In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric …
- Persistent Computational State: A Session-Centric Runtime for Generative World Models
Zhen Lin · 27 de julio de 2026
Generative world models are increasingly driven as simulators: a planner forks a state, rolls out futures, backtracks, and returns to a visited viewpoint. Recent benchmarks establish that current video world models fail this usage, and attribute it to the model, prescribing new architectures and tra…
- An Explainable FFT-Based Spatial-Frequency Fusion Framework for Deepfake Detection
Pamela Kirui, Cho Hyuk, Qingzhong Liu, Haodi Jiang · 27 de julio de 2026
Deepfake generation has raised growing concerns regarding digital media authenticity, misinformation, identity fraud, and public trust. Recent studies show that combining spatial and frequency features leads to stronger detection results than using independently. This paper presents MSCA-FFT, a Fast…
- Time-Reversed Imaging: A Multimodal Benchmark and Framework for Reconstructing Past Human-Environment Interactions
Jorge Bacca, Kebin Contreras, Luis Toscano-Palomino, Mauro Dalla Mura · 27 de julio de 2026
We introduce time-reversed imaging, a new paradigm that infers what just happened in a scene from fading multimodal traces. Instead of extrapolating or interpolating video frames, our goal is to infer past human-environment interactions from residual physical imprints observable in thermal, ultravio…
- Wavelet Phase Diffusion for Structurally and Semantically Consistent Sim-to-Real Translation
Kaiwen Wang, Frank Bieder, Yinzhe Shen, Carlos Fernandez, Jan-Hendrik Pauls, Omer Sahin Tas · 27 de julio de 2026
Simulation-to-reality translation must bridge the appearance gap between synthetic and real domains while preserving structural and semantic consistency. Conditioning-based methods achieve spatial alignment but introduce computationally expensive control modules. Paired-data methods achieve realism …
- TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward
Debottam Dutta, Jaehoon Hahm, Jianchong Chen, Romit Roy Choudhury · 27 de julio de 2026
Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present TILT, a training-free framework for compositional text-to-imag…
- DART: A Degradation-Aware Recurrent Transformer for Archival Film Restoration
Miko{\l}aj Jastrz\k{e}bski, Wojciech Koz{\l}owski, Kamil Adamczewski · 24 de julio de 2026
Archival film restoration is a challenging problem because historical footage contains compound degradations such as scratches, dust, blur, noise, flicker, and photometric aging, while clean reference videos are unavailable. Existing video restoration methods largely treat these degradations implici…
- Expanding Flow Maps
Sophia Tang, Pranam Chatterjee · 24 de julio de 2026
Flow-based generative models have enabled remarkable progress in fast and controllable generation across continuous and discrete state spaces, yet existing parameterizations are constrained to fixed dimensions or fixed sequence lengths. Here, we introduce Expanding Generative Flows (EFlows), which d…
- M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data
Francesca Pia Panaccione, Carlo Sgaravatti, Marco Venere · 24 de julio de 2026
Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular profiles, is crucial for comprehensive disease understanding. However, gene expression data acquisition remains constrained by high costs and privacy concerns, limiting its use in multimodal …
- Zero-Flow Two-Sample Tests
Yakun Wang, Leyang Wang, Song Liu, Taiji Suzuki · 24 de julio de 2026
We propose a new approach to two-sample testing for deciding whether two sets of samples are drawn from the same distribution. The test is built on a statistical discrepancy based on the zero-flow criterion, termed zero-flow discrepancy (ZFD). We prove the validity of ZFD and propose a practical tes…
- PhantomFill: When the Form Demands an Answer, Language Models Invent One
Rana Muhammad Usman · 24 de julio de 2026
Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the form itself causes hallucination. We ask thirteen models the same question about the same input and change only the answer format. The inputs are built so the …
- Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy
Jingyuan Li, Xiaoyi Jiang, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu · 24 de julio de 2026
Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios. While positivity guarantees nonnegative reverse jump rates, it does not ensure Bayes realizability: ratios at a noisy state need not be jointly induced by any clean-token poster…
- Context-weighted Discrete Flow Matching
Daniil Cherniavskii, Daniel Severo, Karen Ullrich · 24 de julio de 2026
Discrete flow matching provides a flexible framework for generative modeling on discrete structures. However, the standard factorized training objective exposes the model to targets of varying difficulty, mixing well-conditioned, predictable tokens with ambiguous, high-entropy ones. We empirically d…
- From Scalars to Time Series: Rethinking Implicit Neural Representations for Time-Varying Volumetric Data
Weihan Zhang, Xuan Zhao, Yenwen Peng, Yuqi Chen, Jun Tao · 24 de julio de 2026
Implicit neural representations (INRs) for time-varying volumetric data are typically trained using dense sampling over spatiotemporal coordinates, where each observation corresponds to a single point in space and time. This coordinate-wise formulation requires extensive sampling during optimization…
- Source-Prior-Driven Selective Adaptation for Efficient Diffusion Model Finetuning
Yi Xiong, Yuan-Yuan Cheng, Xiao-Ming Fu · 24 de julio de 2026
Fine-tuning large diffusion models for new domains or styles involves a trade-off: improving target-specific generation often degrades the pretrained model's broad generative capability. Existing full and parameter-efficient fine-tuning methods typically handle this trade-off only implicitly. In thi…
- Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervision in Video Face Swapping
Logan Robbins · 24 de julio de 2026
Video face swapping has no natural paired supervision: no real footage exists of one person's face performing another person's video. The strongest current answer, DreamID-V's SyncID-Pipe, mints pairs by replacing the identity in exactly two frames of a real clip -- the first and the last -- and reg…
- PC-Edit: Prompt-Contrastive Region Discovery and Region-Guided Editing
Jian Zhang, Zhijun Zhang · 24 de julio de 2026
Replacing an object with one that differs in category or shape requires complete source removal, natural target formation unconstrained by the source silhouette, and preservation of unrelated content. Existing training-free editors either localize edits from terminal predictions under source and tar…
- SESaMo: Symmetry-Enforcing Stochastic Modulation for Normalizing Flows
Janik Kreit, Dominic Schuh, Kim A. Nicoli, Lena Funcke · 24 de julio de 2026
Deep generative models have recently garnered significant attention across various fields, from physics to chemistry, where sampling from unnormalized Boltzmann-like distributions represents a fundamental challenge. In particular, autoregressive models and normalizing flows have become prominent due…
- ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing
Yueyi Liu, Chi Zhang, Sen Cui, Miao Liu · 24 de julio de 2026
Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-mapping nature of generative models and the single-point optimization of standard TTT. In this paper, we demonstrate that…
- Bayesian Wind Tunnels for Model Selection
Siddhartha R Dalal, Vishal Misra, Abhay Parekh · 23 de julio de 2026
Prior work has shown that transformers can perform exact Bayesian filtering within a fixed hypothesis class. Can they also perform Bayesian model selection -- identifying the correct hypothesis class from data? We introduce model-selection Bayesian wind tunnels: controlled environments where g…
- Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions
David R. Wessels, Farhad Ramezanghorbani, David W. Romero, Alireza Moradzadeh, Olivia Viessmann, Maksim Zhdanov, John St. John, Ken Janik, David M Knigge, Yucheng Tang, Erik J Bekkers, Saee Gopal Paliwal · 23 de julio de 2026
Subquadratic alternatives to attention require compromises when applied to multi-dimensional data: standard convolutions lack global receptive fields and input dependency, while recurrent models require rasterizing data such as images, volumes, and partial differential equation (PDE) into an ad-hoc …
- Enhancing next token prediction based pre-training for jet foundation models
Joschka Birk, Anna Hallin, Gregor Kasieczka, Nikol Madzharova, Ian Pang, David Shih · 23 de julio de 2026
Next token prediction is an attractive pre-training task for jet foundation models, in that it is simulation free and enables excellent generative capabilities that can transfer across datasets. Here we study multiple improvements to next token prediction, building on the initial work of OmniJet-$\a…
- Analytic Distribution of Classifier-Free Guidance for Schedule Design
Enze Jiang, Zheng Ma · 23 de julio de 2026
Classifier-free guidance (CFG) is the default mechanism for conditional generation in diffusion models, but the distribution sampled by its deterministic guided dynamics is not captured by the usual product-distribution heuristic $p_0^\omega q_0^{1-\omega}$. We analyze CFG through the probability fl…
- Directional Kernel Mean Difference: A Fast Signed Statistic for Univariate Distribution Comparison
Shijie Zhong, Jiangfeng Fu · 23 de julio de 2026
We introduce the Directional Kernel Mean Difference (DKMD), a signed statistic for univariate distribution comparison that preserves the direction of distributional shifts. Unlike the squared Maximum Mean Discrepancy (MMD), which discards directional information by squaring the RKHS distance, DKMD i…
- HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation
Jinliang Shen, Lianghao Su, Zheming Li, Kang He, ZiLiang Lai, Yanbing Jiang, Chengru Song · 23 de julio de 2026
Autoregressive (AR) video diffusion models have become a promising paradigm for long and streaming video synthesis, but the continuously growing Key-Value (KV) cache makes attention the dominant inference cost, especially at high resolution where each frame contributes many tokens. Existing remedies…
