Search
Search: diffusion models
Words are combined with AND. Use quotes for an exact phrase, a leading dash to exclude a word.
Papers
Page 40 of 40
More than 1,000 papers match: here are the 1,000 most recent, ranked by relevance.
- Tensor-Train Joint Modeling for Few-Step Discrete Diffusion
Byoungkwon Kim, Minhyuk Sung · 7 July 2026 · Tensor decomposition and applications
Discrete diffusion promises orders-of-magnitude faster generation than autoregressive (AR) models for sequential discrete data, yet its full potential of few-step generation has remained out of reach due to a fundamental structural limitation. The conditional-independence assumption underlying curre…
- Joint Velocity Slope Diffusion Prior for Structurally Constrained Velocity Model Building
Francesco Brandolin, Tariq Alkhalifah · 7 July 2026 · Seismic Imaging and Inversion Techniques
High-resolution velocity models are crucial for reservoir characterization and subsurface delineation. However, the band limited nature of our surface recorded data limits resolution. Utilizing well measurements to enhance the resolution of our subsurface models is an important objective. To this en…
- Diffusion-warm sampling of the XY model enables fast thermalization at scale
Sehmimul Hoque, Roger Melko, Pooya Ronagh · 1 July 2026 · Quantum many-body systems
We introduce a novel technique for scalable sampling of spin-system states with continuous symmetries using diffusion models. By applying our approach to the XY model, a fundamental continuous-spin model in condensed matter physics, we show that our technique addresses the shortfalls of the Markov c…
- Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS
Sandipan Dhar, Nirmesh J. Shah, Ashishkumar P. Gudmalwar, Pankaj Wasnik · 25 June 2026 · Phonetics and Phonology Research
Diffusion-based text-to-speech (TTS) models have achieved significant improvements in speech quality. However, modeling sharp prosodic transitions and rapid pitch variations in expressive speech remains challenging. Existing diffusion-based TTS decoders commonly utilize periodic nonlinearities such …
- Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models
Kesong Li, Yixuan Xu, Kuo-kun Tseng, Weiyi Lu, Kan Liu, Tao Lan · 21 May 2026 · Multimodal Machine Learning Applications
Direct Preference Optimization (DPO) is successful for alignment in LLMs but still faces challenges in text-to-image generation. Existing studies are confined to denoising diffusion models while overlooking flow-matching, and suffer from an objective mismatch when applying discrete NLP-based DPO to …
- FrequencyBooster: Full-Frequency Modeling for High-Fidelity Pixel Diffusion
Lichen Ma, Zipeng Guo, Yu He, Xiaolong Fu, Luohang Liu, Jingling Fu, Junshi Huang, Yan Li · 19 May 2026 · Generative Adversarial Networks and Image Synthesis
To circumvent the inherent fidelity bottlenecks and optimization misalignment of VAE-based latent diffusion, pixel-space diffusion models have emerged as a compelling end-to-end paradigm. However, existing pixel diffusion models often struggle to balance computational efficiency with the preservatio…
- Heteroscedastic Diffusion for Multi-Agent Trajectory Modeling
Guillem Capellera, Antonio Rubio, Luis Ferraz, Antonio Agudo · 12 May 2026 · Gaussian Processes and Bayesian Inference
Multi-agent trajectory modeling traditionally focuses on forecasting, often neglecting more general tasks like trajectory completion, which is essential for real-world applications such as correcting tracking data. Existing methods also generally predict agents' states without offering any state-wis…
- Esoteric Language Models: A Family of Any-Order Diffusion LLMs
Subham Sekhar Sahoo, Zhihan Yang, Yash Akhauri, Johnna Liu, Deepansha Singh, Zhoujun Cheng, Zhengzhong Liu, Eric Xing, John Thickstun, Arash Vahdat · 1 June 2026 · Media, Religion, Digital Communication
Diffusion-based language models offer a compelling alternative to autoregressive (AR) models by enabling parallel and controllable generation. Within this family, Masked Diffusion Models (MDMs) currently perform best but still underperform AR models in perplexity and lack key inference-time efficien…
- Algebraic Language Models for Inverse Design of Metamaterials via Diffusion Transformers
Li Zheng, Siddhant Kumar, Dennis M. Kochmann · 27 April 2026 · Topology Optimization in Engineering
Generative machine learning models have revolutionized material discovery by capturing complex structure-property relationships, yet extending these approaches to the inverse design of three-dimensional metamaterials remains limited by computational complexity and underexplored design spaces due to …
- Diffusion-aided Task-oriented Semantic Communications with Model Inversion Attack
Xuesong Wang, Mo Li, Xingyan Shi, Zhaoqian Liu, Shenghao Yang · 25 September 2026 · Wireless Signal Modulation Classification
Semantic communication enhances transmission efficiency by conveying semantic information rather than raw input symbol sequences. Task-oriented semantic communication further aims to retain only task-specific information, thereby achieving greater bandwidth savings. However, these neural-network-bas…
- Brownian Motion with a Pulse: A Biostatistician's Guide to Diffusions, Bridges, Functional PCA, and First-Passage Models
Eliuvish Han Cui · 23 June 2026 · Statistical Methods in Clinical Trials
Brownian motion is a compact mathematical language for continuous-time uncertainty in biostatistics. This tutorial develops the process from construction and path properties to tools that recur in applied biomedical work: the Markov and strong Markov properties, the Karhunen-Loeve expansion, functio…
- Diffusion-based Cumulative Adversarial Purification for Vision Language Models
Jia Fu, Yongtao Wu, Yihang Chen, Kunyu Peng, Xiao Zhang, Volkan Cevher, Sepideh Pashami, Anders Holst · 11 June 2026 · Adversarial Robustness in Machine Learning
Vision Language Models (VLMs) have shown remarkable capabilities in multimodal understanding, yet their susceptibility to adversarial perturbations poses a significant threat to their reliability in real-world applications. Despite often being imperceptible to humans, these perturbations can drastic…
- MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models
Yifan Xu, Chao Zhang, Ruifei Ma, Fei Gao, Zhifei Yang, Jiaxing Qi, Zhipeng Chen · 8 June 2026 · Multimodal Machine Learning Applications
The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding. While current VLMs excel at event- or story-level understanding, their ability to capture fine-grained motion details remains limited, primarily due to their focus on …
- A Unified Measure-Theoretic View of Diffusion, Score-Based, and Flow Matching Generative Models
Aditya Ranganath, Mukesh Singhal · 11 May 2026 · Generative Adversarial Networks and Image Synthesis
We survey continuous-time generative modeling methods based on transporting a simple reference distribution to a data distribution via stochastic or deterministic dynamics. We present a unified framework in which diffusion models, score-based generative models, and flow matching are instances of lea…
- Leveraging Pretrained Language Models as Energy Functions for Glauber Dynamics Text Diffusion
Tarun Kathuria, Sachin Kumar · 7 May 2026 · Language and cultural evolution
We present a discrete diffusion-based language model using Glauber dynamics from statistical physics. Our main insight is that instead of trying to train a discrete state space diffusion model using Glauber dynamics with a uniform transition kernel as the forward process, one can set up an ``energy …
- CasLayout: Cascaded 3D Layout Diffusion for Indoor Scene Synthesis with Implicit Relation Modeling
Yingrui Wu, Youkang Kong, Mingyang Zhao, Weize Quan, Dong-Ming Yan, Yang Liu · 1 May 2026 · 3D Shape Modeling and Analysis
Synthesizing realistic 3D indoor scenes remains challenging due to data scarcity and the difficulty of simultaneously enforcing global architectural constraints and local semantic consistency. Existing approaches often overlook structural boundaries or rely on fully connected relation graphs that in…
- Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models
Kaiwen Zheng, Guande He, Min Zhao, Jintao Zhang, Huayu Chen, Jianfei Chen, Chen-Hsuan Lin, Ming-Yu Liu, Jun Zhu, Qianli Ma · 25 June 2026 · Generative Adversarial Networks and Image Synthesis
Autoregressive video diffusion with causal diffusion transformers has emerged as a major paradigm for real-time streaming video generation and action-conditioned interactive world models. In this work, we extend rCM, an advanced diffusion distillation framework, to autoregressive video diffusion. Th…
- Modelling Gas-Phase Reaction Kinetics with Guided Particle Diffusion Sampling
Andrew Millard, Zheng Zhao, Henrik Pedersen · 21 April 2026 · Model Reduction and Neural Networks
Physics-guided sampling with diffusion priors has recently shown strong performance in solving complex systems of partial differential equations (PDEs) from sparse observations. However, these methods are typically evaluated on benchmark problems that do not fully demonstrate their ability to genera…
- Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models
Junjian Li, Xiaolong Liu, Peng Sun, Liantao Wu, Linghan Chen, Yudong Gao, Honglong Chen · 1 October 2026 · Generative Adversarial Networks and Image Synthesis
Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack …
- MixDiffusion: Mixing Diffusion-based Uni-condition Text-to-Image Generation Models for Multi-condition Image Synthesis
Pengcheng Wan, Liang Han, Lin Xu, Bowen Xiao, Liqiang Nie · 21 July 2026 · Generative Adversarial Networks and Image Synthesis
Recent advances in text-to-image (T2I) generation have enabled controllable image synthesis by incorporating conditions beyond text. However, most existing diffusion-based methods are limited to a single type of control condition (e.g., bounding boxes or keypoints), which restricts their flexibility…
- ST-DiffEye: Diffusion-based Continuous Gaze Generation via Joint Scanpath-Trajectory Modeling
Brian Nlong Zhao, Ozgur Kara, Junho Kim, James M. Rehg · 16 June 2026 · Gaze Tracking and Assistive Technology
We study the problem of human gaze modeling, which aims to generate the gaze patterns a viewer produces while observing a visual stimulus. Gaze is primarily captured through two modalities: continuous eye-tracking trajectories, which describe fine-grained motion dynamics, and discrete scanpaths, whi…
- Hiding in Plain Sight: A Diffusion-based Mitigation of Geolocation Privacy Leakage in Vision-Language Models
Yining Wang, Xi Li, Mi Zhang, Xiaohan Zhang, Xiaoyu You, Zhenxing Qian, Mi Wen · 21 September 2026 · Geographic Information Systems Studies
Multimodal large reasoning models (MLRMs) have demonstrated remarkable capabilities in complex visual understanding. However, this very power introduces a critical yet underexplored privacy threat: adversaries can exploit MLRMs to precisely infer users' geographic locations from casually shared phot…
- 2D Versus 3D Diffusion for In Silico Training of Interventional X-ray AI Models
Sampath Rapuri, Jeremy Ko, Benjamin D. Killeen, Russell H. Taylor, Mathias Unberath · 23 June 2026 · Surgical Simulation and Training
The ability to synthesize realistic X-ray images has catalyzed the development of AI models for X-ray image-guided procedures, which otherwise suffer from a lack of available annotated data. Prior work has demonstrated the effectiveness of mechanistic simulation of digitally reconstructed radiograph…
- AirCast-SR: A Foundation Model for Kilometer-Scale Atmospheric Super-Resolution via Latent Consistency Diffusion
Somnath Luitel, Manmeet Singh, Joshua Durkee, Abdullah Al Fahad, Naveen Sudharsan, Prabhjot Singh, Cenlin He, Harsh Kamath, Zong-Liang Yang, Krishnagopal Halder, Sandeep Juneja, Parthasarathi Mukhopadhyay, Saptarishi Dhanuka, Amit Kumar Srivastava · 27 May 2026 · Meteorological Phenomena and Simulations
Operational weather prediction at kilometer scales remains computationally prohibitive for traditional numerical weather prediction (NWP) models, limiting forecast access for applications in energy, agriculture, and disaster management that require fine-grained spatiotemporal detail. Here we introdu…
- Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models
Abrar Alotaibi, Moataz Ahmed · 26 June 2026 · Adversarial Robustness in Machine Learning
Adversarial evaluation of AI systems has matured along four largely disconnected tracks: diffusion-based attacks on text and large language models (LLMs), diffusion-based attacks on image classifiers, jailbreak pipelines against vision-language models, and diffusion-based input purification defenses…
The search covers titles only, not the text of the abstracts. To query the content of the papers, the research assistant searches the indexed abstracts.
