Suche
Suche: diffusion models
Die Wörter werden mit UND verknüpft. Anführungszeichen für eine exakte Wortfolge, ein Bindestrich davor schließt ein Wort aus.
Paper
Seite 6 von 40
Mehr als 1.000 Paper passen: hier die 1.000 neuesten, nach Relevanz sortiert.
- Controllable Clothing: Precise Labels and Generation for Virtual Try-On with Latent Diffusion Models
Max Rehman Linder · 7. August 2026 · Generative Adversarial Networks and Image Synthesis
In this technical report, I present a new method for guiding image generation in the context of Virtual- Try-On (VITON). The proposed method leverages new open source Ai models to augment the image data with labels, such as lengths and styles. By training adapters with these labels paired with image…
- Enhancing Low Back Pain Assessment with Diffusion Models for Lumbar Spine MRI Segmentation
Maria Monzon, Thomas Iff, Ender Konukoglu, Catherine R. Jutzeler · 6. August 2026 · Medical Imaging and Analysis
This study introduces a diffusion-based framework for robust and accurate semantic segmentation of lumbar spine MRI scans from patients with low back pain (LBP), regardless of whether the scans are T1- or T2-weighted. We compared with advanced models for segmenting vertebrae, intervertebral discs (I…
- STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models
Qingyan Wei, Guangzhao Li, Xiaobing Tu, Yinggui Wang, Xiantao Zhang, Jinkui Ren, Xiaohong Liu, Linfeng Zhang · 6. August 2026 · Generative Adversarial Networks and Image Synthesis
On-policy distillation (OPD) has become an effective approach for consolidating multiple task-specialized image generation models into a single student. However, existing OPD methods optimize the student mainly to match the teacher's output velocity, making the teacher the upper limit of the optimiz…
- When Diffusion Models Forget Who You Are: Identity Preservation in Face Inpainting under Large Occlusions
Feng Ding, Shuhuai Xie, Yue Zhou, Yulan Zhang, Guopu Zhu, Mengyao Xiao · 6. August 2026 · Generative Adversarial Networks and Image Synthesis
Face inpainting with diffusion models has recently achieved impressive visual quality, yet preserving identity fidelity under significant occlusion and conflicting text guidance remains a major challenge. To address this issue, we present Reference Semantic Inpainting for Face (ReSem-Face), a cascad…
- UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
Haiyang Zhou, Wangbo Yu, Chaoran Feng, Xunyu Zhou, Yonghong Tian, Li Yuan · 6. August 2026 · Advanced Vision and Imaging
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent vi…
- Regularization can make diffusion models more efficient
Mahsa Taheri, Johannes Lederer · 6. August 2026 · Advanced Mathematical Modeling in Engineering
Diffusion models are one of the key architectures of generative AI. Their main drawback, however, is the computational costs. This study indicates that the concept of sparsity, well known especially in statistics, can provide a pathway to more efficient diffusion pipelines. Our mathematical guarante…
- Intrinsic-Hybrid Latent Diffusion Models for Generative Modeling on Unknown Manifolds
Yizhu Wang, Mu Niu, Xiaochen Yang · 6. August 2026 · Generative Adversarial Networks and Image Synthesis
We introduce the Intrinsic Hybrid Latent Diffusion Model (ILDM), a generative framework that integrates probabilistic dimensionality reduction with geometry-aware diffusion on unknown manifolds. While diffusion models (DMs) have achieved state-of-the-art results in high-dimensional data synthesis, t…
- Curiosity-Diffuser: Curiosity Guide Diffusion Models for Reliability
Zihao Liu, Xing Liu, Yuhang Dong, Haitao Chang, Zhengxiong Liu, Panfeng Huang · 6. August 2026 · Reinforcement Learning in Robotics
One of the bottlenecks in robotic intelligence is the instability of neural network models. This leads to risks when applying intelligence in the physical world. Specifically, imitation policy based on neural network may generate hallucinations, leading to inaccurate behaviors that impact the safety…
- S$^3$-Diff: Structural Semantic Synergy Diffusion Model for High Fidelity Super Resolution of Pathological Images
Jiaming Liang, QiHui Han, Guangye Ou, Jiawen Liu, Haolin Chen, Xi Zhong, Jiazhou Chen, Xiaoqi Sheng, Hongmin Cai · 5. August 2026 · AI in cancer detection
Digital pathology relies on high-resolution whole slide images for accurate diagnosis, yet limitations in imaging devices, storage, and transmission often make lower-resolution pathology images more common in clinical workflows. Current super-resolution techniques often tend to smooth diagnostically…
- SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
Shanghao Liu, Renze Chen, Size Zheng, Yuanqiang Liu, Yun, Liang, Hailong Yang · 5. August 2026 · Generative Adversarial Networks and Image Synthesis
Video diffusion transformers (vDiTs) generate high quality but pay quadratic self-attention cost, making inference prohibitive at video-token scales. The challenge is input-adaptive sparsity: selecting critical Q/K/V tokens with negligible overhead and executing them for end-to-end gains. We present…
- TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models
Seokho Han, Dongwei Wang, Jinhee Kim, Yiran Chen, Kang Eun Jeon, Huanrui Yang, Jong Hwan Ko · 5. August 2026 · Parallel Computing and Optimization Techniques
Static quantization assigns one weight precision to every denoising step. To preserve quality, that precision must accommodate the most quantization-sensitive step, even though many other steps can tolerate fewer bits. The resulting model may satisfy its memory budget, but it repeatedly pays worst-c…
- Divide-and-Conquer: Towards Generalizable Amortized Bayesian Inference for the Drift Diffusion Model
Yufei Wu, Shanqing Gao, Andreas Voss, Francis Tuerlinckx · 5. August 2026 · Fault Detection and Control Systems
The drift diffusion model (DDM) is a cornerstone of cognitive decision-making research. Although numerous estimation methods exist, researchers continue to seek inference approaches that are both fast and flexible across diverse study designs. Amortized Bayesian inference (ABI) can provide nearly in…
- Assessment of Conditional Diffusion Model for Synthetic Histopathology Image Generation
Seyed Kahaki, Shijie Li, Weijie Chen, Nicholas Petrick · 5. August 2026 · AI in cancer detection
Synthetic histopathology image generation has emerged as an approach that may address data scarcity in computational pathology, yet current evaluation methodologies may not fully assess synthetic data quality for medical applications. This work investigates and addresses limitations in existing eval…
- Simulation-free and finite-time diffusion model
Kentaro Kaba, Masayuki Ohzeki, Yuki Sughiyama · 5. August 2026 · Advanced Mathematical Modeling in Engineering
The performance of generative diffusion models is determined by the choice of the reference diffusion process connecting the empirical and prior distributions. Conventional approaches typically trade off simulation-free training against finite-time generation. We propose a framework for designing th…
- Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors
Alexander Scheinker · 4. August 2026 · Generative Adversarial Networks and Image Synthesis
Autoregressive models accumulate error over long rollouts, yet at deployment there is no ground truth to measure it against. We train a single conditional latent diffusion model that steps a dynamical system forward or backward in time via a direction flag, and show that this bidirectionality suppli…
- Pseudorandom Streams within Diffusion Models Act as Learnable Inputs That Affect Generation Quality
Shengzhi Deng, Chenqi Ye, Yanze Guo · 4. August 2026 · Advanced Neuroimaging Techniques and Applications
Diffusion models rely on stochastic inputs, yet on finite-precision hardware, the "randomness" they consume is realized as deterministic numerical orbits generated by pseudorandom rules. Accessible orbit structure can become a learnable input and affect both training and generation because the reali…
- Music Restoration via Latent Operator Optimization and Diffusion Model Priors
Michal \v{S}vento, Eloi Moliner, Valtteri Kallinen, Lauri Juvela, Vesa V\"alim\"aki, Pavel Rajmic · 4. August 2026 · Speech and Audio Processing
Music restoration seeks to recover a clean signal from an observed recording degraded by an unknown effect, distortion, or corruption. Existing systems often rely on paired training data and distortion-specific supervision, which limits their use when the forward process is not known in advance. We …
- Exploring More to Solve More: Boosting Diversity in Text Diffusion Models via Entropy-Based Guidance
Jingwei Zhang, Haoyu Lei, Zijin Feng, Jiacheng Sun, Farzan Farnia · 4. August 2026 · Generative Adversarial Networks and Image Synthesis
Although diffusion models have revolutionized continuous domains like image synthesis through high quality generations and controllable guidance mechanisms, bringing this controllability to the discrete, sequential nature of text remains an open challenge. Meanwhile, current sampling strategies and …
- A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples
Zixuan Fu, Chong Wang, Lanqing Guo, Kailai Zhou, Jiahao Nie, Bihan Wen · 3. August 2026 · Generative Adversarial Networks and Image Synthesis
Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging because a single model must capture both global structure and local texture in the same high-dimensional space. While recent work improves pixel diffusion through alternative prediction ta…
- In-situ Autoguidance: Eliciting Self-Correction in Diffusion Models
Enhao Gu, Haolin Hou · 3. August 2026 · Model Reduction and Neural Networks
The generation of high-quality, diverse, and prompt-aligned images is a central goal in image-generating diffusion models. The popular classifier-free guidance (CFG) approach improves quality and alignment at the cost of reduced variation, creating an inherent entanglement of these effects. Recent w…
- MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation
Rongchao Xu, Lin Jiang, Dahai Yu, Ximiao Li, Guang Wang · 3. August 2026 · Generative Adversarial Networks and Image Synthesis
Utility data (e.g., electricity, water, and gas consumption), collected by ubiquitous sensors and embedded devices, often contains substantial missing values due to various factors such as device failures and data transmission issues. The data missingness can severely impact utility billing accuracy…
- DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models
Omid Ahmadieh, Nima Karimian · 3. August 2026 · Face recognition and analysis
Facial biometric identification relies on the distinctiveness of user attributes within a high-dimensional embedding space. However, the decision boundaries of deep face recognition (FR) systems are often sufficiently narrow that they can be conflated, rendering the models vulnerable to adversarial …
- FreeShadow: Training-Free Shadow Removal via Illumination Transfer and Selective Content Preservation in Diffusion Models
Yinan Wang, Yan Huang, Yong Xu, Patrick Le Callet · 30. Juli 2026 · Generative Adversarial Networks and Image Synthesis
Existing supervised and unsupervised shadow removal methods often suffer from limited generalization due to the insufficient diversity of available training datasets, while zero-shot methods tend to produce artifacts and require time-consuming test-time optimization. To address these issues, we prop…
- TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models
Taewon Kang, Matthias Zwicker · 30. Juli 2026 · Generative Adversarial Networks and Image Synthesis
Text-to-video diffusion models generate temporally coherent content from natural language, yet when a prompt describes an early scene that persists while a new event emerges on top of it---such as "a tall sandcastle standing on a beach where a wave rushes in and washes it away"---generation frequent…
- LumaGuide: Distribution Shaping for Training-Free HDR Generation in Diffusion Models
Bowen Chen, Shreshth Saini, Balu Adsumilli, Alan C. Bovik · 30. Juli 2026 · Generative Adversarial Networks and Image Synthesis
Pretrained diffusion models generate realistic images but are constrained by the statistical biases of their training data, limiting their ability to produce high dynamic range (HDR) content. In this work, we introduce LumaGuide, a training-free framework for distribution shaping in diffusion models…
Die Suche erfasst nur die Titel, nicht den Text der Abstracts. Um den Inhalt der Paper abzufragen, durchsucht der Forschungsassistent die erfassten Abstracts.
