Search
Search: diffusion models
Words are combined with AND. Use quotes for an exact phrase, a leading dash to exclude a word.
Papers
Page 18 of 40
More than 1,000 papers match: here are the 1,000 most recent, ranked by relevance.
- Q-ARVD: Quantizing Autoregressive Video Diffusion Models
Siao Tang, Xinyin Ma, Gongfan Fang, Xingyi Yang, Xinchao Wang · 21 May 2026 · Image and Video Quality Assessment
Autoregressive video diffusion models (ARVDs) have emerged as a promising architecture for streaming video generation, paving the way for real-time interactive video generation and world modeling. Despite their potential, the substantial inference cost of ARVDs remains a major obstacle to practical …
- AttriStory: Fine-grained Attribute Realization for Visual Storytelling with Diffusion Models
Manogna Sreenivas, Rohit Kumar, Soma Biswas · 21 May 2026 · Multimodal Machine Learning Applications
Visual storytelling with diffusion models has made impressive strides in maintaining character consistency across narrative scenes. However, a critical gap remains: while these methods ensure a character remains consistent across scenes, they provide no systematic method to ensure if fine-grained at…
- Diffusion Models Memorize in Training -- and Generalize in Inference
Tim Kaiser, Markus Kollmann · 21 May 2026 · Generative Adversarial Networks and Image Synthesis
Diffusion models generalize well in practice. However, an optimal diffusion model fully memorizes the training data and therefore fails to generalize, raising the question of what induces generalization in a real diffusion model. We show that, despite generalizing at the sample level, diffusion mode…
- Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models
Taesung Kwon, Jonghyun Park, Hyungjin Chung, Jong Chul Ye · 21 May 2026 · Generative Adversarial Networks and Image Synthesis
Diffusion models provide powerful priors for zero-shot video inverse problems, but their real-time deployment is hindered by two inefficiencies: high initial latency caused by holistic video restoration, and low throughput resulting from multiple VAE passes to enforce measurement consistency in pixe…
- Tippett-minimum Fusion of Representation-space Diffusion Models for Multi-Encoder Out-of-Distribution Detection
Neelkamal Bhuyan · 21 May 2026 · Domain Adaptation and Few-Shot Learning
We address out-of-distribution (OOD) detection across the full spectrum of distribution shifts -- global domain changes, semantic divergence, texture differences, and covariate corruptions -- through a multi-encoder fusion of per-encoder representation-space diffusion models (RDMs). We statistically…
- AirfoilGen: A valid-by-construction and performance-aware latent diffusion model for airfoil generation
Zhijie Yang, Min Tang, Qiang Zou · 21 May 2026 · Model Reduction and Neural Networks
Airfoil shape design is a fundamental task in aerospace engineering, with a direct impact on flight stability and fuel consumption. Deep learning has recently emerged as a promising tool for this task, but existing deep generative approaches remain limited in both geometric validity and physical con…
- Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
Wei Huang, Andi Han, Mingyuan Bai, Huanjian Zhou, Qixin Zhang, Taiji Suzuki, Kenji Fukumizu · 21 May 2026 · Generative Adversarial Networks and Image Synthesis
Diffusion models generate high-dimensional data with remarkable quality, yet how their training efficiently learns the score function, bypassing the curse of dimensionality when data is supported on low-dimensional manifolds, remains theoretically unexplained. We identify a collapse-and-refine mecha…
- Landscape-Awareness for Geometric View Diffusion Model
Yan-Ting Chen, Hao-Wei Chen, Tsu-Ching Hsiao, Chun-Yi Lee · 20 May 2026 · Advanced Vision and Imaging
Accurate camera viewpoint estimation under sparse-view conditions remains challenging, particularly in two-view scenarios. Recent approaches leverage diffusion models such as Zero123 to synthesize novel views conditioned on relative viewpoint, showing promising results when repurposed for viewpoint …
- When Preference Labels Fall Short: Aligning Diffusion Models from Real Data
Weiyan Chen, Weijian Deng, Yao Xiao, Weijie Tu, ZiYi Dong, Ibrahim Radwan, Liang Lin, Pengxu Wei · 20 May 2026 · Generative Adversarial Networks and Image Synthesis
Preference alignment aims to guide generative models by learning from comparisons between preferred and non-preferred samples. In practice, most existing approaches rely on preference pairs constructed from model-generated images. Such supervision is inherently relative and can be ambiguous when bot…
- Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention
Wenhu Zhang, Yiming Wu, Huanyu Wang, Yaoyang Liu, Huanzhang Dou, Senqiao Yang, Sitong Wu, Hanbin Zhao, Jiaya Jia · 20 May 2026 · Multimodal Machine Learning Applications
Diffusion Language Models (DLMs) enable globally coherent, bidirectional, and controllable text generation, offering advantages over traditional autoregressive LLMs, while scaling to ultra-long sequences remains costly. Many existing block-sparse attention methods select blocks by fixed sampling pat…
- Awakening the Hydra: Stabilizing Multi-Concept Backdoor Injection in Text-to-Image Diffusion Models
Kai Wang, Jiale Zhang, Chengcheng Zhu, Chuang Ma, Songze Li · 20 May 2026 · Adversarial Robustness in Machine Learning
Text-to-image diffusion models are increasingly developed through open-source reuse and repeated downstream fine-tuning, where reused checkpoints are difficult to verify and thus more susceptible to hidden backdoor behaviors. In such ecosystems, a single pretrained model may be sequentially adapted …
- Boosting Text-to-Image Diffusion Models via Core Token Attention-Based Seed Selection
Yunzhe Zhang, Hongfu Liu, Pengyu Hong · 20 May 2026 · Generative Adversarial Networks and Image Synthesis
Text-to-image diffusion models can synthesize high-quality images, yet the outcome is notoriously sensitive to the random seed: different initial seeds often yield large variations in image quality and prompt-image alignment. We revisit this "seed effect" and show that attention dynamics over prompt…
- Reducing Diffusion Model Memorization with Higher Order Langevin Dynamics
Benjamin Sterling, M\'onica F. Bugallo, Tom Tirer · 20 May 2026 · Generative Adversarial Networks and Image Synthesis
Diffusion/score-based models have emerged as powerful generative models, capable of generating high-quality samples that mimic the training data distribution. However, it has been observed that they are prone to reproducing training samples-known as "memorization"-potentially violating copyright and…
- Noise scheduling and linear dynamics in diffusion models on Lie groups
Javad Komijani · 20 May 2026 · Markov Chains and Monte Carlo Methods
We investigate the role of the noise schedule in diffusion processes on Lie groups, with particular emphasis on applications to lattice gauge theory. We show that a specific noise schedule leads to a linear decay of the expectation value of the Wilson action as a function of diffusion time. We compa…
- LIFT and PLACE: A Simple, Stable, and Effective Knowledge Distillation Framework for Lightweight Diffusion Models
Hyunsoo Han, Sangyeop Yeo, Jaejun Yoo · 20 May 2026 · Generative Adversarial Networks and Image Synthesis
We demonstrate that in knowledge distillation for diffusion models, the teacher network's highly complex denoising process - stemming from its substantially larger capacity - poses a significant challenge for the student model to faithfully mimic. To address this problem, we propose a coarse-to-fine…
- Inference-Time Scaling in Diffusion Models through Iterative Partial Refinement
Taegu Kang, Jaesik Yoon, Sungjin Ahn · 20 May 2026 · Generative Adversarial Networks and Image Synthesis
Inference-time scaling has emerged as a major approach for improving reasoning capabilities, and has been increasingly applied to diffusion models. However, existing inference-time scaling methods for diffusion models typically rely on external verifiers or reward models to rank and select samples, …
- Designing streetscapes from street-view imagery using diffusion models
Yuzhou Chen, Yuebing Liang, Lingqian Hu, Kailai Sun, Qingqi Song, Chang Zhao, Shenhao Wang · 19 May 2026 · Automated Road and Building Extraction
Street-view imagery (SVI) is widely used to quantify key indicators of urban environment, such as green- ery, sky, or road view indices. However, existing studies largely focus on measuring current streetscapes and rarely support the generation of alternative and non-existing urban scenarios, which …
- DreamEdit3D: Personalization of Multi-View Diffusion Models for 3D Editing
Jinxin Ai, Matthias Nießner, Ziya Erkoç · 19 May 2026 · 3D Shape Modeling and Analysis
While 2D diffusion models have achieved remarkable success in identity-preserving personalization, extending this capability to 3D assets remains a significant challenge due to the complexities of multi-view consistency and spatial control. Inspired by these 2D advancements, we present a novel perso…
- HighSync: High-Quality Lip Synchronization via Latent Diffusion Models
Saeed Firouzi Daghigh, Majid Iranpour Mobarekeh, Mostafa Alavi, Mehdi Bagheri · 19 May 2026 · Speech and Audio Processing
We present HighSync, an end-to-end diffusion-based framework for high-fidelity lip synchronization that generates photorealistic talking-face videos aligned with arbitrary input audio. Existing approaches consistently struggle to reconcile image quality with synchronization accuracy, producing eithe…
- Face inpainting with Identity Preserving Latent Diffusion Models
João Santos, Carlos Santiago, Manuel Marques · 19 May 2026 · Generative Adversarial Networks and Image Synthesis
Face inpainting techniques recover missing or occluded facial regions in a visually realistic manner, but preserving the identity in the final output remains a fundamental challenge. Identity consistency is crucial for downstream applications such as face recognition, digital forensics, and human-co…
- Geometry-Aware Attention Guidance for Diffusion Models via Modern Hopfield Dynamics
Kwanyoung Kim · 19 May 2026 · Model Reduction and Neural Networks
Classifier-Free Guidance (CFG) improves sample quality in diffusion models, but its dual-pass inference and reliance on null-condition training limit its use in few-step regimes. Attention-space guidance has emerged as a complementary paradigm that addresses this gap, yet why prior sparse-vs-dense a…
- SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models
Xiaomeng Yang, Mengping Yang, Junyan Wang, Zhijian Zhou, Zhiyu Tan, Hao Li · 19 May 2026 · Multimedia Communication and Technology
Preference learning has garnered extensive attention as an effective technique for aligning diffusion models with human preferences in visual generation. However, existing alignment approaches such as Diffusion-DPO suffer from two fundamental challenges: training instability caused by high gradient …
- Simple Approximation and Derivative Free Inference-Time Scaling for Diffusion Models via Sequential Monte Carlo on Path Measures
Chenyang Wang, Weizhong Wang, Yinuo Ren, Jose Blanchet, Yiping Lu · 19 May 2026 · Generative Adversarial Networks and Image Synthesis
iffusion-based generative models increasingly rely on inference-time guidance, adding a drift term or reweighting mixture of experts, to improve sample quality on task-specific objectives. However, most existing techniques require repeated score or gradient evaluations, introducing bias, high comput…
- Diffusion Models, Denoiser Architecture and Creativity
Itamar Levine, Yair Weiss · 19 May 2026 · Creativity in Education and Neuroscience
The creativity of diffusion models refers to their ability to generate highly realistic images that are different from their training data. Creativity is somewhat surprising since it is known that if the denoiser used in the diffusion model is the Bayes optimal denoiser for a given training set, the…
- Dual-Rate Diffusion: Accelerating diffusion models with an interleaved heavy-light network
Grigory Bartosh, David Ruhe, Emiel Hoogeboom, Jonathan Heek, Thomas Mensink, Tim Salimans · 19 May 2026 · Generative Adversarial Networks and Image Synthesis
Diffusion models achieve state-of-the-art generative performance but suffer from high computational costs during inference due to the repeated evaluation of a heavy neural network. In this work, we propose Dual-Rate Diffusion, a method to accelerate sampling by interleaving the execution of a heavy …
The search covers titles only, not the text of the abstracts. To query the content of the papers, the research assistant searches the indexed abstracts.
