Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Generative Adversarial Networks and Image Synthesis
2 529 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- Structured Nonparametric Variational Inference for Dependent Latent Modeling
Yuda Shao, Zhiling Gu, Shan Yu · 16 juin 2026
Variational inference (VI) is a core engine of modern AI, enabling scalable approximate Bayesian learning and uncertainty-aware training of large probabilistic and generative models. In this paper, we propose Structured Nonparametric Variational Inference (SN-VI), a novel framework for modeling comp…
- SPRI: SVD-Partitioned Residual Initialization for Data-Constrained MoE Upcycling
Weiqiao Shan, Ruixiang Mao, Yuang Li, Yuhao Zhang, Yingfeng Luo, Tong Zheng, Chen Xu, Yucheng Qiao, Chunxiang Jin, Yi Yuan, Jingdong Chen, Tong Xiao, Jingbo Zhu · 16 juin 2026
Mixture-of-Experts (MoE) models enable efficient scaling, but training them from scratch remains prohibitively expensive. MoE upcycling mitigates this cost by converting pretrained dense models into sparse MoE models. However, existing upcycling methods typically rely on large-scale continued traini…
- Probabilistic Signature Inversion: Learning Conditional Distributions from Truncated Signatures
Junoh Kang, Kiseop Lee, Bohyung Han · 16 juin 2026
The signature transform is a principled feature map for continuous-time paths, valued for its uniqueness and universality. Recovering a path from its truncated signature is, however, structurally ill-posed because the truncated signature map is not injective. We therefore reframe truncated signature…
- Exact Posterior Score Estimation for Solving Linear Inverse Problems
Abbas Mammadov, Ozgur Kara, Kaan Oktay, Iskander Azangulov, Adil Kaan Akan, Hyungjin Chung, James Matthew Rehg, Yee Whye Teh · 16 juin 2026
Diffusion and flow-based models learn powerful data priors by training a denoiser to reverse Gaussian corruption. To use this prior to solve a linear inverse problem, one needs to sample from the posterior, but the score that the prior provides is the unconditional score, not the posterior score. Ex…
- Distilling Drifting Transformers with Representation Autoencoders
Jiawei Zhang, Mengfei Xia, Gen Li, Yuantao Gu · 16 juin 2026
Representation Autoencoders (RAEs) have improved diffusion and flow models by semantically richer latent space owing to the strongly label-wise clustered DINO features in the pretrained encoders. Yet in the distillation stage, the severe anisotropy and large curvatures caused by the rich semantic re…
- RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos
Zhenhua Wu, Yun Pang, Mingkun Chang, Yuwei Ning, Liangzhi Wang, Yi Xiao, Guanbin Li · 16 juin 2026
Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect and reproduce at scale. Editable 3D Gaussian Splatting (3DGS) simulation offers a promising alternative by reconstructing real driving scenes and supporting controllable scene editin…
- Towards a Unified Generative Model for Scarce Time Series with Domain Experts
Zihao Yao, Qi Zheng, Jiankai Zuo, Yaying Zhang · 16 juin 2026
Synthesizing realistic time series with generative models has wide-ranging applications in real-world scenarios. Despite recent progress, most existing methods are trained under the assumption of abundant training data, which substantially limits their effectiveness in data-scarce settings. In this …
- Temporal Difference Learning for Diffusion Models
Qizhen Ying, Yangchen Pan, Victor Adrian Prisacariu, Junfeng Wen · 16 juin 2026
Diffusion models are typically trained with objectives that focus on local denoising targets at individual time steps (or adjacent pairs), which do not enforce consistency between predictions along the denoising trajectory. This lack of cross-time consistency can degrade performance, especially for …
- Stop the Sampler! Classifier-Based Adaptive Stopping for Sampling Kernels
Kirill Korolev, Nikita Morozov, Stepan Pavlenko, Esmeralda S. Whitammer, Sergey Samsonov · 16 juin 2026
Sampling from complex, unnormalized probability densities is a fundamental challenge in Bayesian inference and probabilistic modeling. While Markov chain Monte Carlo (MCMC) methods provide asymptotic guarantees, they often suffer from slow mixing and high computational costs due to fixed or manually…
- PhysGuard: Fisher-Guided Gradient Projection for Sim-to-Real Neural PDE Surrogates
Changjian Zhou, Junfeng Fang, Negin Yousefpour, Peng Wu, Bin Yan, Guillermo A Narsilio · 16 juin 2026
Neural operator models trained on simulation data often lose accuracy when applied to experimental measurements due to the sim-to-real gap. Standard fine-tuning with limited real data can reduce this gap, but it may also damage the core physics-relevant representations learned during pretraining. Al…
- Temporally Consistent and Controllable Video Generation of 2D Cine CMR via Latent Space Motion Modeling
Yiheng Cao (SyCoIA - IMT Mines Al\`es), Gustavo Andrade-Miranda (SyCoIA - IMT Mines Al\`es), Jiatian Zhang, Guillaume Sall\'e, Xin Gao · 16 juin 2026
Cine cardiac magnetic resonance is the gold standard for assessing cardiac function, but the scarcity of public datasets limits the development of advanced data-driven models. To address this limitation, we propose a generative method for synthesizing temporally coherent and anatomically consistent …
- Selective Synergistic Learning for Video Object-Centric Learning
WonJun Moon, Jae-Pil Heo · 16 juin 2026
Typical video object-centric learning (VOCL) approaches employ slot-based frameworks that rely on reconstruction-driven encoder-decoder architectures, where learning is mediated by two spatial maps: attention maps from the encoder and object maps from the decoder. As these two distinct maps exhibit …
- DifFRACT: Diffusion Feature Reconstruction and Attribution for Circuit Tracing
Artyom Mazur, Nina Konovalova, Aibek Alanov · 16 juin 2026
Mechanistic interpretability seeks to explain neural network behavior by decomposing model computations into interpretable features and circuits. While transcoder-based circuit tracing has recently enabled detailed causal analyses of large language models, multimodal diffusion transformers for image…
- Quantifying the Impact of Lossy Compression on Neural Generative Surrogate Modeling
Zhimin Li, Harshitha Menon, Charles Jekel, Valerio Pascucci, Peter Lindstrom · 16 juin 2026
Neural networks are used as generative surrogate models for scientific discovery, which are trainable approximations of scientific simulations. These models enable users to replace time-consuming numerical simulations with learned alternatives, providing quick solutions. However, high-fidelity gener…
- Teacher-Student Structure for Domain Adaptation in Ensemble Audio-Visual Video Deepfake Detection
Elham Abolhasani, Maryam Ramezani, Hamid R. Rabiee · 16 juin 2026
The rapid advancement of generative AI models is leading to more realistic deepfake media, encompassing the manipulation of audio, video, or both. This raises severe privacy and societal concerns. Numerous studies in this area have yielded promising intra-domain results; however, these models freque…
- Sub-Semantic Image Segmentation
Aviad Cohen Zada, Nadav Orenstein, Shai Avidan, Gal Oren · 16 juin 2026
Images can be segmented based on visual cues (i.e., texture segmentation) or into objects (i.e., semantic segmentation). We propose a new category of sub-semantic image segmentation that blurs the line between the two. In sub-semantic image segmentation, language is not used to name whole objects. I…
- Efficient Flow Matching using Latent Variables
Anirban Samaddar, Yixuan Sun, Viktor Nilsson, Sandeep Madireddy · 16 juin 2026
Flow matching models have shown great potential in image generation tasks among probabilistic generative models. However, most flow matching models in the literature do not explicitly utilize the underlying clustering structure in the target data when learning the flow from a simple source distribut…
- Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)
Ankit Hemant Lade, Sai Krishna Jasti, Indar Kumar, Aman Chadha · 16 juin 2026
A Mamba state-space model trained only for next-step prediction appears to recover Granger-causal structure through a simple readout $S = |W_{out} W_{in}|$, with early experiments suggesting the phenomenon generalized across architectures and benefited from interventional data at $p < 10^{-5}$. We p…
- Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech
Adarsh Arigala, Arjun Gangwar, S Umesh, Yova Kementchedjhieva · 16 juin 2026
Recent advances in pixel-based text modeling show that representing text as images enables models to exploit visual cues for language understanding. Grounding text in its visual form allows structurally similar characters with different Unicode encodings to produce similar embeddings, benefiting cro…
- The Information-Theoretic Benefit of Shared Representations under Orthogonality Constraints
Thomas Dittrich, Oliver Potocki, Philipp Grohs · 16 juin 2026
Modern deep learning architectures are increasingly multi-task and multi-modal, using a pretrained foundation model combined with task-specific, fine-tuned models. Empirically, exploiting similarity across different problems, instead of solving them individually, can significantly improve overall pe…
- MA-SBI: Misspecification-Aware Simulation-Based Inference via Side-Channel Guidance
Arunkumar V, Manoranjan Gandhudi, Gangadharan G. R., Arun Prakash, S. Senthilkumar · 16 juin 2026
Simulation-based inference (SBI) of latent parameters is often hindered by simulator misspecification, the mismatch between simulated and real-world observations caused by inherent modeling simplifications. RoPE, the recent state-of-the-art for robust SBI, addresses this through optimal transport be…
- BRICKS-WM: Building Reusability via Interface Composition Kinetics for Structured World Models
Shaowei Zhang, Jiahan Cao, Xunlan Zhou, Shenghua Wan, De-Chuan Zhan · 16 juin 2026
Model-based Reinforcement Learning (MBRL) has achieved remarkable success in continuous control by leveraging latent world models. However, prevailing approaches typically rely on monolithic latent dynamics, entangling environment dynamics into a coupled process. This coupling severely limits reusab…
- Where Does Texture Evidence Live in SAM? Features, Proposal Masks, and Texture Segmentation
Nadav Orenstein, Aviad Cohen Zada, Shai Avidan, Gal Oren · 16 juin 2026
Texture segmentation stresses foundation segmentation because meaningful regions are defined by material or repeated appearance rather than object identity. Segment Anything Models (SAMs) often fail by default on such texture-defined partitions, but this failure is ambiguous: the texture evidence ma…
- Divide-and-Denoise: A Game-Theoretic Method for Fairly Composing Diffusion Models
Abhi Gupta, Polina Barabanshchikova, Vikas Garg, Samuel Kaski, Tommi Jaakkola · 16 juin 2026
The abundance of pre-trained diffusion models provides an opportunity for composition. Combining several models, however, runs the risk of one model dominating or models disagreeing with each other. Here, we propose Divide-and-Denoise, a method for coordinating multiple pre-trained diffusion models …
- Shift-and-Sum Quantization for Visual Autoregressive Models
Jaehyeon Moon, Bumsub Ham · 16 juin 2026
Post-training quantization (PTQ) enables efficient deployment of deep networks using a small set of data. Its application to visual autoregressive models (VAR), however, remains relatively unexplored. We identify two key challenges for applying PTQ to VAR: (i) large reconstruction errors in attentio…
