Search
Search: diffusion models
Words are combined with AND. Use quotes for an exact phrase, a leading dash to exclude a word.
Papers
Page 33 of 40
More than 1,000 papers match: here are the 1,000 most recent, ranked by relevance.
- DLM-SWAI: Steering Diffusion Language Models Before They Unmask
Hyeseon An, Yo-Sub Han · 29 May 2026 · Large Language Models
Steering language model generation toward desired textual properties is essential for practical deployment, and inference-time methods are particularly appealing because they enable controllable generation without retraining. Recent work has also highlighted diffusion language models as an emerging …
- GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models
Xiaohang Tang, Keyue Jiang, Che Liu, Qifang Zhao, Xiaoxiao Xu, Sangwoong Yoon, Ilija Bogunovic · 29 May 2026 · Domain Adaptation and Few-Shot Learning
Reinforcement learning (RL) can be used to improve the policy (denoiser) of diffusion large language models (dLLMs), while being hindered by the intractability of the policy likelihood. A dominant and efficient family of methods replaces the likelihood in standard RL with its evidence lower bound (E…
- BlockBatch: Multi-Scale Consensus Decoding for Efficient Diffusion Language Model Inference
Xiaoyou Wu (Celine), Cheng-Jhih Shih (Celine), Binfei Ji (Celine), Yong Liu (Celine), Yingyan (Celine), Lin · 29 May 2026 · Large Language Models
Diffusion language models (dLLMs) generate text by iteratively denoising multiple token positions in parallel, offering an attractive alternative to strictly autoregressive decoding. In practice, however, block-wise dLLM inference exposes a difficult granularity trade-off: small blocks preserve loca…
- When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models
Jungwon Park, Jimyeong Kim, Jungmin Ko, Nojun Kwak, Wonjong Rhee · 28 May 2026 · Large Language Models
Diffusion language models generate text by iteratively selecting and denoising masked positions, making position selection a central inference-time decision. Most training-free methods rely on model confidence, assuming that high-confidence positions are ready to be decoded. However, this assumption…
- Targeted Remasking: Replacing Token Editing with Token-to-Mask Refinement in Discrete Diffusion Language Models
Lin Yao · 27 May 2026 · Large Language Models
Discrete masked diffusion language models such as LLaDA generate text through iterative denoising, where mask tokens are progressively replaced with predicted tokens. LLaDA2.1 introduced a Token-to-Token (T2T) editing mechanism that accelerates generation by directly replacing committed tokens suspe…
- Looped Diffusion Language Models
Sanghyun Lee, Chunsan Hong, Seungryong Kim, Jonghyun Lee, Jongho Park, Dongmin Park · 26 May 2026 · Large Language Models
Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models for language modeling, yet the effective design of transformer architectures for MDMs remains underexplored. In this paper, we show that selectively looping the early-middle transformer layers significant…
- Missing Pattern Recognized Diffusion Imputation Model for Missing Not At Random
Gyuwon Sim, Sumin Lee, Heesun Bae, Byeonghu Na, Doyun Kwon, Ju-Hee Hwang, Jae-Young Lim, Il-Chul Moon · 26 May 2026 · Bayesian Methods and Mixture Models
Missing data frequently arises across diverse domains, including time-series and image domains. In the real world, missing occurrences often depend on the unobservable values themselves, which are referred to as Missing Not at Random (MNAR). In this work, we introduce the Missing Pattern Recognized …
- TUBE: Tangent Upper Bound on Evidence for Discrete Diffusion Language Models
Arseny Ivanov, Sergei Kholkin, Vladislav Gromadskii, Grigoriy Ksenofontov, Ivan Oseledets, Alexander Korotin · 26 May 2026 · Generative Adversarial Networks and Image Synthesis
Log-likelihood is a standard metric for evaluating generative models. Unfortunately, in contrast to autoregressive models (ARMs), discrete diffusion models generally do not admit exact computation of this quantity. Existing evaluations, therefore, rely on the evidence lower bound (ELBO), leaving unc…
- The Path Matters: Learning a Token-Commitment Policy for Diffusion Language Models
Bohang Sun, Max Zhu, Francesco Caso, Jindong Gu, Junchi Yu, Philip Torr, Pietro Li\`o, Jialin Yu · 26 May 2026 · Large Language Models
Diffusion large language models promise faster generation by refining many token positions in parallel, but this parallelism introduces a hidden control problem: which proposed tokens should be transferred into the partially decoded sequence at each step? We refer to this decision as token commitmen…
- Filtered Posterior Mean Collections: A Unified Framework for Analytical Models of Diffusion Generalization
Matthew Niedoba, Berend Zwartsenberg, Frank Wood · 26 May 2026 · Advanced Neuroimaging Techniques and Applications
The neural-network denoising functions which form the backbone of image diffusion models are remarkably consistent in their generalization behaviour across a wide variety of network architectures and training procedure hyperparameters. A recent line of research has sought to model the outputs of the…
- Extracting Training Data from Diffusion Language Models via Infilling
Yihan Wang, N. Asokan · 26 May 2026 · Large Language Models
Memorization in large language models has been studied almost exclusively through prefix-conditioned extraction, a natural choice for autoregressive models. However, diffusion language models (DLMs) can denoise masked tokens at arbitrary positions. Thus, prefix-only probing reveals only one facet of…
- Efficient One-Step Diffusion Restoration Model with Compact Token Compression and Linear Attention
Bingtian Qiao, Yue Shi, Yingjie Zhou, Yong Guo, Guangtao Zhai, Jiezhang Cao · 25 May 2026 · Advanced Image Processing Techniques
Real-world image super-resolution aims to recover high-quality images from complex and unknown real-world degradations. However, existing generative Real-ISR methods largely inherit the dense latent representations and quadratic-cost global modeling paradigm developed for high-resolution image synth…
- Learnability-Informed Fine-Tuning of Diffusion Language Models
Shubham Parashar, Atharv Chagi, Jacob Helwig, Lakshmi Jotsna, Sushil Vemuri, James Caverlee, Dileep Kalathil, Shuiwang Ji · 25 May 2026 · Large Language Models
We aim to improve the reasoning capabilities of diffusion language models (DLMs). While SFT is a popular post-training recipe for autoregressive models, its use in DLMs faces challenges and can even hurt performance, though the underlying causes remain understudied. Our analysis reveals that vanilla…
- PulseCol: Periodically Refreshed Column-Sparse Attention for Accelerating Diffusion Language Models
Yanyi Lyu, Letian Chen, Futing Sun, Miao Zhang, Weili Guan, Liqiang Nie · 21 May 2026 · Large Language Models
Inference in diffusion large language models (dLLMs) is computationally expensive, as full self-attention must be repeatedly executed at each step of the denoising process without KV cache. Recent sparse attention methods for dLLMs mitigate this cost via block-sparse computation, which is applied on…
- Drifting Objectives for Refining Discrete Diffusion Language Models
Daisuke Oba, Hiroki Furuta, Naoaki Okazaki · 20 May 2026 · Large Language Models
Discrete diffusion language models (DDLMs) generate text by iteratively denoising categorical token sequences, while recent drifting methods for continuous generators suggest that part of this sampling-time correction can instead be absorbed into training through an anti-symmetric fixed-point object…
- Tweedie's Formulae and Diffusion Generative Models Beyond Gaussian
Wenpin Tang, Nizar Touzi, Zikun Zhang, Xun Yu Zhou · 20 May 2026 · Generative Adversarial Networks and Image Synthesis
Diffusion models have achieved remarkable success in generating samples from unknown data distributions. Most popular stochastic differential equation-based diffusion models perturb the target distribution by adding Gaussian noise, transforming it into a simple prior, and then use denoising score ma…
- Backdooring Masked Diffusion Language Models
Daniel Yiming Cao, Chengzhong Wang, Sheng-Yen Chou, Chengyu Huang, Pin-Yu Chen, Shengwei An · 20 May 2026 · Adversarial Robustness in Machine Learning
Masked diffusion language models (MDLMs) are emerging as a compelling new paradigm for text generation, but their training-time security remains largely unexplored. Existing backdoor attacks on Gaussian diffusion models or autoregressive language models do not directly apply to MDLMs because MDLMs r…
- Stitched Value Model for Diffusion Alignment
Hyojun Go, Hyungjin Chung, Prune Truong, Goutam Bhat, Li Mi, Zhaochong An, Zixiang Zhao, Dominik Narnhofer, Serge Belongie, Federico Tombari, Konrad Schindler · 20 May 2026 · Generative Adversarial Networks and Image Synthesis
For practical use, diffusion- or flow-based generative models must be aligned with task-specific rewards, such as prompt fidelity or aesthetic preference. That alignment is challenging because the reward is defined for clean output images, but the alignment procedure requires value function estimate…
- Composition of Memory Experts for Diffusion World Models
Sebastian Stapf, Pablo Acuaviva Huertos, Aram Davtyan, Paolo Favaro · 20 May 2026 · Reinforcement Learning in Robotics
World models aim to predict plausible futures consistent with past observations, a capability central to planning and decision-making in reinforcement learning. Yet, existing architectures face a fundamental memory trade-off: transformers preserve local detail but are bottlenecked by quadratic atten…
- Machine Unlearning for Masked Diffusion Language Models
Georu Lee, Seungwon Jeong, Hoki Kim, Jinseong Park, Woojin Lee · 19 May 2026 · Large Language Models
Recent masked diffusion language models (MDLMs), such as LLaDA and Dream, have achieved performance comparable to autoregressive large language models. Unlike autoregressive models, which generate text sequentially, MDLMs generate text by iteratively denoising masked positions in parallel. During fi…
- Wasserstein bounds for denoising diffusion probabilistic models via the F\"ollmer process
Yuta Koike · 19 May 2026 · Markov Chains and Monte Carlo Methods
This paper studies sampling error bounds for denoising diffusion probabilistic models (DDPMs) in the 2-Wasserstein distance. Our contributions are threefold. (i) Under general Lipschitz-type conditions on the score function and for a broad class of variance schedules, including the cosine schedule, …
- A note on connections between the F\"ollmer process and the denoising diffusion probabilistic model
Yuta Koike · 19 May 2026 · Stochastic processes and financial applications
The F\"ollmer process is a Brownian motion conditioned to have a pre-specified distribution at time 1. This process can be interpreted as an "augmented" time-compressed version of the reverse stochastic differential equation (SDE) for the denoising diffusion probabilistic model (DDPM). While this fa…
- Membership Inference Attacks on Discrete Diffusion Language Models
Shailesh Kasivelrajan · 19 May 2026 · Adversarial Robustness in Machine Learning
Masked Diffusion Language Models MDLMs replace autoregressive generation with iterative demasking and their privacy properties are largely unstudied. We study membership inference attacks MIA on fine tuned MDLMs and show they are significantly more vulnerable than current grey box baselines suggest.…
- DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
Amin Karimi Monsefi, Dominic Culver, Nikhil Bhendawade, Lokesh Boominathan, Manuel R. Ciosici, Yizhe Zhang, Irina Belousova · 19 May 2026 · Reinforcement Learning in Robotics
Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. We identify two fundamental weaknesses: the absence of temporal credit a…
- Dynamic Chunking for Diffusion Language Models
Yichen Zhu, Xiaoming Shi, Peng Zhao, Weiyu Chen, Debing Zhang, James Kwok · 18 May 2026 · Language and cultural evolution
Block discrete diffusion language models factorize a sequence autoregressively over fixed-size positional blocks, decoupling within-block parallel denoising from across-block conditioning. We argue that this rigid partition wastes structure already present in the sequence: blocks defined by position…
The search covers titles only, not the text of the abstracts. To query the content of the papers, the research assistant searches the indexed abstracts.
