Search
Search: diffusion models
Words are combined with AND. Use quotes for an exact phrase, a leading dash to exclude a word.
Papers
Page 26 of 40
More than 1,000 papers match: here are the 1,000 most recent, ranked by relevance.
- Cross-Resolution Diffusion Models via Network Pruning
Jiaxuan Ren, Junhan Zhu, Huan Wang · 8 April 2026 · Generative Adversarial Networks and Image Synthesis
Diffusion models have demonstrated impressive image synthesis performance, yet many UNet-based models are trained at certain fixed resolutions. Their quality tends to degrade when generating images at out-of-training resolutions. We trace this issue to resolution-dependent parameter behaviors, where…
- Self-Distilled Trajectory-Aware Boltzmann Modeling: Bridging the Training-Inference Discrepancy in Diffusion Language Models
Kecheng Chen, Ziru Liu, Xijia Tao, Hui Liu, Yibing Liu, Xinyu Fu, Shi Wu, Suiyun Zhang, Dandan Tu, Lingpeng Kong, Rui Liu, Haoliang Li · 13 May 2026 · Large Language Models
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive language models, offering stronger global awareness and highly parallel generation. However, post-training DLMs with standard Negative Evidence Lower Bound (NELBO)-based supervised fine-tuning remains…
- Representation-Space MMD for Diffusion Language Models
Ilya Drobyshevskiy, Ilia Sudakov, Maksim Semenov, Denis Kuznedelev, Maksim Ignatov, Pavel Temirchev, Nikita Balagansky, Viacheslav Meshchaninov, Nikita Gushchin, Dmitry Baranchuk · 6 October 2026 · Natural Language Processing Techniques
We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features at individual token positions, obtainin…
- Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering
Sarim Hashmi, Mukul Ranjan, Abdelrahman Elsayed, Muhammad Umer Sheikh, Fahad Shamshad, Nils Lukas · 6 October 2026 · Adversarial Robustness in Machine Learning
Masked diffusion language models (dLLMs) generate text by iteratively denoising masked positions, re-predicting each token multiple times before it is committed. An autoregressive decoder exposes an answer's distribution once, at the step that commits it; a dLLM exposes it at every denoising step be…
- Towards Unbiased On-Policy Distillation for Block Diffusion Language Models
Zaiquan Yang, Fei Wei, Yong Wang, Yudong Han, Yiyu Li, Zhuofan Zong, Gerhard Petrus Hancke, Xiangxiang Chu, Rynson WH Lau · 6 October 2026 · Large Language Models
On-policy distillation (OPD) has emerged as an effective post-training paradigm for language models, with recent efforts extending it to block diffusion language models (BDLMs). However, existing studies focus almost exclusively on small block sizes, leaving distillation into student models with lar…
- Bayesian Entropy-based Reordering for Calibrated Diffusion Language Models
Zhejun Jiang, Mijung Park · 6 October 2026 · Large Language Models
Masked Diffusion Language Models (MDLMs) generate sequences by iteratively replacing masked tokens with model predictions. At each denoising step, the decoder chooses which positions are sufficiently confident to commit. Existing decoding methods typically rely on softmax confidence, which can be mi…
- SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models
Chung-En Ho (Celine), Weiyu Sun (Celine), Cheng-Jhih Shih (Celine), He Li (Celine), Yong Liu (Celine), Yingyan (Celine), Lin · 6 October 2026 · Power Systems and Technologies
Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit t…
- ALoDLM: Adaptively Looped Diffusion Language Models
Liancheng Fang, Zhuowei Li, Youngeun Kim, Tianchen Zhao, Rajat Koner, Jiaye Wu, Linghan Xu, Xuanbai Chen, Xiang Xu, Zheng Zhang, Jakub Zablocki, Nishant Sankaran, Yifan Xing · 6 October 2026 · Natural Language Processing Techniques
Diffusion language models (DLMs) enable fast generation by predicting multiple tokens in parallel, but their practical adoption remains limited by a persistent quality gap relative to comparably sized autoregressive (AR) models. We attribute this gap to a computation-difficulty mismatch: within a pa…
- Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao, Se-Young Yun, Rahul G. Krishnan · 5 October 2026 · Large Language Models
Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and sh…
- Preserving Mathematical Reasoning in Compressed Diffusion Language Models via Trajectory-Aware Low-Rank Approximation
Tian Liang, Zishan Shao, Yiran Chen · 5 October 2026 · Reservoir Engineering and Simulation Methods
Diffusion language model (dLLM) compression faces a known challenge because calibration is typically performed on clean, fully visible activations, whereas inference traverses partially masked intermediate states. For low-rank compression, this raises two questions. First, can low-rank optimality st…
- Know When to Hold 'em: Correct-Token Retention in Uniform-State Diffusion Language Models
Mojtaba Nafez, James Henderson · 2 October 2026 · Generative Adversarial Networks and Image Synthesis
Uniform-state diffusion models (USDMs) can revise any token at any denoising step, which lets them correct their own mistakes, a key advantage over masked diffusion. Self-correction, however, requires both revising incorrect tokens and retaining correct ones, and we show that current USDMs lack the …
- Hierarchical Continuous Diffusion Language Models
Hui Ren, Zihan Li, Chang Liu, Huidong Liu, Alexander Schwing · 2 October 2026 · Advanced Mathematical Modeling in Engineering
Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, sev…
- Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models
Darpan Aswal, C\'eline Hudelot · 2 October 2026 · Large Language Models
This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-ge…
- Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models
Yue YU, Bowen Zuo, David Crandall, Yinglun Zhu, Dongruo Zhou · 2 October 2026 · Reinforcement Learning in Robotics
Diffusion large language models (dLLMs) generate text by denoising a sequence or successive blocks, allowing several tokens to be revealed in parallel. Reinforcement learning with verifiable rewards (RLVR) reuses terminal feedback across these decisions, even as their conditioning context changes. W…
- ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models
Lianjun Liu, Shipeng Li, You Huang, Weiqi Yan, Mingte Qiu, Huazhong Liu, Xiaofeng Zhu, Yunshan Zhong · 2 October 2026 · Advanced Data Compression Techniques
Mixture-of-Experts (MoE) Diffusion Language Models (DLMs) offer flexible parallel decoding and increased model capacity, but their large number of expert parameters incurs substantial computation and storage costs. Existing low-rank MoE compression methods largely rely on static factorization and fi…
- Time-Anchored Diffusion Language Models: Latent-Space Caching for Fast Generation
Joel Anto Paul, Litu Rout, Aditya Akella, Sanjay Shakkottai · 1 October 2026 · Large Language Models
Recent work on anchored diffusion language models improves denoising by shaping an intermediate latent space with supervised important-token targets. In this work, we introduce time-based (self-supervised) anchoring, which learns and reuses latent anchors without requiring such targets. Our key obse…
- E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models
Arseny Ivanov, Alexander Kolesov, Alexander Korotin, Ivan Oseledets, Mikhail Goncharov · 1 October 2026 · Generative Adversarial Networks and Image Synthesis
Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most…
- Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents
Jiacheng Qiu, Christopher E. Mower, Jan Peters, Haitham Bou-Ammar, Matthieu Zimmer · 1 October 2026 · Large Language Models
Diffusion-based large language models (dLLMs) promise to break the sequential latency bottleneck of autoregressive agents through parallel decoding, but recent evaluations show this efficiency does not transfer to embodied agentic competence: dLLM-backed agents repeatedly fall into retry loops, re-i…
- Acceleration of Diffusion Language Model through Discrete Average Generator
Yidong Ouyang, Zhengyan Wan, Themis Haris, Tian Tan, Liqian Peng, Henry Li, Ziqian Lin, Jianhang Chen, Maryam Karimzadehgan, Alec Go, George Michailidis · 1 October 2026 · Advanced Mathematical Modeling in Engineering
Discrete diffusion models and flow matching have emerged as powerful frameworks for generative modeling over discrete state spaces, yet efficient few-step generation remains a fundamental challenge. In this work, we introduce the Discrete Average Generator, a principled extension of MeanFlow to Cont…
- How Should Diffusion Language Models Edit Code?
Xijia Tao, Ziru Liu, Shansan Gong, Jiacheng Ye, Kecheng Chen, Zirui Wu, Lin Zheng, Xinyu Fu, Rui Liu, Lingpeng Kong · 1 October 2026 · Model-Driven Software Engineering Techniques
Code editing requires a model to decide where to make changes, generate the new content, and preserve everything else. We study how masked diffusion language models divide these responsibilities across four editing interfaces: whole-file rewriting, search-and-replace, locate-then-infill, and token-l…
- Fine-Tuning Diffusion Language Models with Context Selection and Target Weighting
Loay Mualem, Llu\'is Pastor-P\'erez, Vinh Tong, Andrei Manolache, Tanja Bien, Steffen Staab, Mathias Niepert · 1 October 2026 · Large Language Models
Supervised fine-tuning of discrete diffusion language models masks some response tokens and trains the model to recover their original values from the visible context. The masking pattern therefore determines both the context available to the model and the tokens it learns to predict. Uniform random…
- TANGO: Watermarking Masked Diffusion Language Models in Token Pairs
Kasra Arabi, Nir Weinberger, Micah Goldblum, Niv Cohen · 30 September 2026 · Adversarial Robustness in Machine Learning
Masked-diffusion language models fill in masked positions in parallel and in no fixed order. Most practical text watermarks assume left-to-right generation. They key each token to the tokens before it, and in a diffusion model those tokens may still be masked. A fixed green list needs no such contex…
- DreamingGoose: Staged Distillation from Autoregressive Transformers to Bidirectional Recurrent Diffusion Language Models
Julian Boesch, Andrew Wee, Alexander Stranzl · 30 September 2026 · Large Language Models
Pretrained autoregressive Transformers represent a large sunk investment in compute. Existing conversion methods reuse that investment by changing either the architecture (attention to recurrence) or the objective (next-token prediction to denoising), never both. We convert Qwen3 teachers at 1.7B an…
- Simple Diffusion Language Models Are More Effective Few-Step Generators Than Reported
Hasan Amin, Ming Yin, Rajiv Khanna · 30 September 2026 · Natural Language Processing Techniques
Diffusion language models (DLMs) promise fast parallel generation, yet high-quality samples often require large number of refinement steps, which diminishes their advantage in practice. This has led to massive interest in and rapid development of new methods for effective few-step generation. We sho…
- On Trajectory-Aware Training for Masked Diffusion Language Models
Manuel Madeira, Amitis Shidani, Alice Bizeul, Victor Turrisi, Louis B\'ethune, Bhavika Devnani, Dan Busbridge, Pierre Ablin, Jo\~ao Monteiro · 30 September 2026 · Generative Adversarial Networks and Image Synthesis
Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions. The model is trained on randomly masked sequences, whereas inference follows a trajectory shaped by the model's own predictions. Additionally, each step has…
The search covers titles only, not the text of the abstracts. To query the content of the papers, the research assistant searches the indexed abstracts.
