Search
Search: diffusion models
Words are combined with AND. Use quotes for an exact phrase, a leading dash to exclude a word.
Papers
Page 30 of 40
More than 1,000 papers match: here are the 1,000 most recent, ranked by relevance.
- Rethinking the Generation Order of Block Diffusion Language Models
Kai Syun Hou, James Kwok · 28 July 2026 · Generative Adversarial Networks and Image Synthesis
Diffusion language models enable flexible arbitrary-order generation, but existing sampling methods are mostly designed for early masked diffusion models (MDMs). In this work, we study sampling for recent block diffusion language models (BDLMs). We show empirically and analytically that these models…
- Multi-Mask Diffusion Language Models for Few-Step Generation
Sijin Chen, Yinuo Ren, Heyang Zhao, Ziheng Cheng, Quanquan Gu, Lexing Ying · 23 July 2026 · Large Language Models
Masked diffusion models (MDMs) are a promising family of language generators, but achieving high-quality few-step generation remains challenging. In MDMs, all forward trajectories collapse to a single fully masked state, leaving no terminal entropy for consistency-style few-step generation. While re…
- Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models
Brian K Chen, Chong Wu, Kenji Kawaguchi · 22 July 2026 · Large Language Models
Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by producing text block by block. We study a simple plug-and-play inference pattern: first generate a complete draft, then refine the full response using …
- Addressing Limited Data in Auditory Attention Decoding with Diffusion Generative Models
David Rannaleet, Victor Gunnarsson, Bo Bernhardsson, Martin A. Skoglund, Emina Alickovic · 22 July 2026 · Hearing Loss and Rehabilitation
Limited training data constrains deep learning models for Auditory Attention Decoding (AAD) in hearing aids (HAs). AAD uses electroencephalogram (EEG) data to decode listener's attention, enabling real-time tracking of specific sound sources. However, achieving high AAD performance with short time w…
- Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents
Kaixuan Liu, Guojun Xiong, Weinan Zhang, Shengpu Tang · 21 July 2026 · Large Language Models
Evaluating large language model (LLM) agents in multi-turn interactive environments is expensive and risky, as it requires online environment interaction. We propose ADWM (Autoregressive Diffusion World Model), an evaluation framework that estimates the performance of a new LLM agent policy purely f…
- Apeliotes: A Diffusion-Based Modeling Framework for km-scale Multi-Level Atmospheric Fields
Evangelia Rafaela Frastali, Achyut Paudel, Maryam Golbazi, Frank Liu · 21 July 2026 · Meteorological Phenomena and Simulations
High-resolution atmospheric data are required to resolve mesoscale and localized meteorological structures, however such datasets remain limited in many regions of the world. Existing high-resolution weather products are typically produced through dynamical downscaling, which is computationally expe…
- Trace-Based On-Policy Distillation for Masked Diffusion Language Models
Haolin Ren, Ziyang Huang, Chenhao Yuan, Jun Zhao, Kang Liu · 21 July 2026 · Large Language Models
Diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. However, reasoning-oriented post-training for dLLMs remains challenging. Supervised fine-tuning (SFT) for dLLMs requires dense but often off-policy masked states, while reinforcement learning (RL) relie…
- FlowBlock: Wavefront-Parallel Decoding for Self-Correcting Diffusion Language Models
Bing Tian, Haikun Liu, Xiaocheng Zhong, Zhuohui Duan, Zhaokai Luo, Huayi Jin, Zhiyong Wang, Xiaofei Liao · 21 July 2026 · Natural Language Processing Techniques
Block-wise diffusion large language models (dLLMs) decode sequentially at the block level, enabling effective KV-cache reuse across blocks but making inter-block decoding strictly serial. Prior work has attempted to unlock inter-block parallelism through post-training methods, but achieves only mode…
- JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models
Yeachan Jun, Albert No · 21 July 2026 · Large Language Models
Membership inference attacks (MIAs) test whether a candidate example appeared in a model's training data. We study MIAs for fine-tuned discrete diffusion language models (dLLMs), where membership means inclusion in the target model's fine-tuning set. Unlike autoregressive language models, dLLMs allo…
- Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
Darshan Deshpande · 21 July 2026 · Reinforcement Learning in Robotics
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective signals as model performance improves, and sparse rewards over long horizons induce mode collapse on s…
- Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models
Yingqian Cui, Wei Deng, Lantao Mei, Hang Li, Charu C. Aggarwal, Hui Liu, Yue Xing · 20 July 2026 · Large Language Models
Masked diffusion language models (DLMs) enable parallel text generation by iteratively refining masked tokens, offering a promising alternative to autoregressive decoding. Recent lookahead-based decoding methods improve the accuracy--efficiency trade-off by exploring future decoding states before co…
- Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models
Andy Catruna, Emilian Radoi · 20 July 2026 · Large Language Models
While the internal mechanisms of autoregressive (AR) transformers have been studied extensively, much less is known about diffusion language models (DLMs), an emerging alternative that generates text by iterative denoising. In this work, we study how DLMs implement induction, a mechanism behind in-c…
- Mask-Aware Policy Gradients for Diffusion Language Models
Haran Raajesh, Kulin Shah, Adam Klivans, Philipp Kr\"ahenb\"uhl · 17 July 2026 · Large Language Models
Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Masked Diffusion Language Models (MDLMs) remains challenging due to the intractability of the log-likelihood estimation. Existing approaches approximate this log-likelihood by modeling o…
- Constrained Decoding for Diffusion Language Models via Efficient Inference over Finite Automata
Meihua Dang, Stefano Ermon · 9 July 2026 · Machine Learning and Algorithms
Constrained decoding is essential for serving LLMs, ensuring that generated outputs follow specific structures such as JSON schema-formatted function calls. Existing systems are designed for autoregressive models and assume left-to-right generation, masking out invalid next tokens at each step. Diff…
- Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding
Yonggan Fu, Lexington Whalen, Abhinav Garg, Chengyue Wu, Maksim Khadkevich, Nicolai Oswald, Enze Xie, Daniel Egert, Sharath Turuvekere Sreenivas, Shizhe Diao, Chenhan Yu, Ye Yu, Weijia Chen, Sajad Norouzi, Jingyu Liu, Shiyi Lan, Ligeng Zhu, Jin Wang, Jindong Jiang, Morteza Mardani, Mehran Maghoumi, Song Han, Ante Jukić, Nima Tajbakhsh, Jan Kautz, Pavlo Molchanov · 8 July 2026 · Domain Adaptation and Few-Shot Learning
We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architecture. Trained with a joint AR-diffusion objective, Nemotron-Labs-Diffusion can switch modes to sustain high throughput across deployment settings and…
- Diffusion Language Model Parallel Decoding via Product-of-Experts Bridge
Juntong Shi, Brian L. Trippe, Jure Leskovec, Stefano Ermon, Minkai Xu · 8 July 2026 · Large Language Models
Diffusion language models (DLMs) offer substantial speed advantages through parallel decoding, but the lack of token dependencies limits generation quality compared to autoregressive (AR) models. Recent progress attempts to bridge the gap via importance sampling, with DLM being the proposal and AR b…
- dOPSD: On-Policy Self-Distillation for Diffusion Language Models
Phuong Tuan Dat, Qi Li, Xinchao Wang · 7 July 2026 · Large Language Models
Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autoregressive models, but eliciting strong reasoning through post-training remains difficult: supervised fine-tuning is off-policy and suffers from exposure bias, whi…
- Training Hybrid Block Diffusion Language Models with Partial Bidirectionality
Pranshu Chaturvedi, Parth Shroff, Tarun Suresh, Hangoo Kang, Kaiyue Wen · 7 July 2026 · Large Language Models
High-throughput long-context generation is one of the central challenges for large language models. Generation is typically memory-bandwidth-bound rather than compute-bound: each decoding step must stream the accumulated key/value (KV) cache from memory, so bandwidth demand grows with context length…
- TACG: Trajectory-Aware Commit Gating for Diffusion Language Model Decoding
Chengcheng Wang, Tingzhang Luo, Wenhao Li, Jianyuan Guo, Chang Xu · 6 July 2026 · Large Language Models
Diffusion language models (DLLMs) generate text by iteratively denoising masked positions, exposing a trajectory of predictive distributions rather than a single instantaneous belief. Most existing decoders ignore this trajectory and commit tokens from the current snapshot alone, conflating confiden…
- Discrete Diffusion Language Models for Interactive Radiology Report Drafting
Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert · 3 July 2026 · Artificial Intelligence in Healthcare and Education
Diffusion language models, which generate text by denoising a token canvas bidirectionally instead of emitting tokens left to right, have become competitive with autoregressive (AR) generation. Medical foundation models, however, remain almost entirely autoregressive. We adapt a mixture-of-experts d…
- Valdi: Value Diffusion World Models
Christopher Lindenberg, Kashyap Chitta · 2 July 2026 · Reinforcement Learning in Robotics
World models can enable Model Predictive Control (MPC), but this requires dynamics prediction that is both fast enough for online use and expressive enough to represent uncertain futures. Diffusion models offer a natural mechanism for modeling uncertain dynamics, yet their iterative inference proced…
- TAG-DLM: Diffusion Language Models for Text-Attributed Graph Learning
Lingjie Chen, Yuanchen Bei, Haobo Xu, Yanjun Zhao, Yuzhong Chen, Hanghang Tong · 1 July 2026 · Advanced Graph Neural Networks
Text-attributed graphs (TAGs), where each node carries a natural language description, require models to jointly reason over text and graph topology. Existing approaches often handle the two modalities separately: graph neural networks operate on shallow text features, while hybrids of LLMs and grap…
- $x$-Prediction Flow: Efficient Continuous Decoding for Masked Diffusion Language Models
Weitian Wang, Lianlei Shan, Shubham Rai, Cecilia De La Parra, Akash Kumar · 30 June 2026 · Generative Adversarial Networks and Image Synthesis
Masked diffusion language models (MDLMs) generate text by iteratively unmasking tokens, but their standard decoder reduces each step to a binary action: a position is either committed to a single token or left fully masked, discarding rich predictive information rather than carrying it forward, and …
- Adaptive Block Diffusion: Resolving Training-Inference Mismatch in Diffusion Language Models
Gagan Jain · 30 June 2026 · Natural Language Processing Techniques
Diffusion Language Models (DLMs) are typically trained under fixed context structures, restricting denoising to predetermined token subsets. This creates a mismatch between training and inference, where models must operate over arbitrary configurations, leading to degradation off the training grid. …
- Multi-Block Diffusion Language Models
Yijie Jin, Jiajun Xu, Yuxuan Liu, Chenkai Xu, Yi Tu, Jiajun Li, Dandan Tu, Xiaohui Yan, Kai Yu, Pengfei Liu, Zhijie Deng · 30 June 2026 · Caching and Content Delivery
Block Diffusion Language Models (BD-LMs) improve diffusion-based text generation with KV caching and flexible-length generation. A natural next step is to extend them from Single-Block Diffusion (SingleBD) to Multi-Block Diffusion (MultiBD), where a \textit{running-set} of consecutive blocks is deco…
The search covers titles only, not the text of the abstracts. To query the content of the papers, the research assistant searches the indexed abstracts.
