Search
Search: diffusion models
Words are combined with AND. Use quotes for an exact phrase, a leading dash to exclude a word.
Papers
Page 27 of 40
More than 1,000 papers match: here are the 1,000 most recent, ranked by relevance.
- Reliable Parallel Decoding in Masked Diffusion Language Models
Zhenghao He, Bohan Liu, Guangzhi Xiong, Aidong Zhang · 30 September 2026 · Large Language Models
Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidenc…
- FairDiff: Mitigating the Self-Reinforcing Matthew Effect in Diffusion Recommender Models
Song-Li Wu, Xianquan Wang, Zhaocheng Du, Weinan Gan, Jingyi Wang · 30 September 2026 · Recommender Systems and Techniques
While the "Matthew Effect" and filter bubbles are widely recognized outcome-level biases in recommender systems, we reveal that Diffusion Recommender Models (DRMs) uniquely compound this issue through their generative dynamics. Rather than merely inheriting data imbalances, DRMs trigger a self-reinf…
- Diffusion Reward Models
Xiangyang Wang, Bingxiang He, Zeyuan Liu, Jiaze WangZiqing Qiao, Yuxin Zuo, Huan-ang Gao, Cheng Qian, Wenbin Zhang, Ran Li, Youbang Sun, Ning Ding, Yuanchun Shi, Zhiyuan Liu, Chaojun Xiao, Chun Yu · 29 September 2026 · Statistical and Computational Modeling
Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonabl…
- From Position Risks to Block Survival: Faster Generation for Diffusion Language Models
Siwei Chen, Yuxiang Wan, Yifan Yu, Fan Lai · 29 September 2026 · Natural Language Processing Techniques
Diffusion language models (DLMs) can accelerate generation by predicting multiple tokens in parallel, but there is a mismatch between how these tokens are predicted and how they ultimately contribute to generation. Parallel predictions can hardly condition on the tokens selected earlier within the s…
- Hesitation-Aware On-Policy Distillation for Diffusion Language Models
Jianguo Huang, Lipeng Wan, Yanchen Deng, Bo An · 29 September 2026 · Large Language Models
Diffusion large language models (dLLMs) generate text by iterative unmasking. At each denoising step, a dLLM proposes a token at every masked position, but the decoder commits only a confident subset of these proposals. Trace-based on-policy distillation (TOPD) builds on this process by matching the…
- What does FFN compression change downstream? Same-state causal restoration in diffusion language models
Shaurya Omar · 29 September 2026 · Advanced Data Storage Technologies
Diffusion language models (DLMs) enable flexible, parallel generation, but their iterative denoising remains computationally expensive, motivating increasingly aggressive compression. Existing compression objectives largely measure how well compressed computation approximates the original locally, b…
- Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models
Injin Kong, Sunghwan Choi, Yohan Jo · 29 September 2026 · Large Language Models
Masked diffusion language models (MDMs) admit flexible generation orders, making the unmasking strategy an inference decision. Existing methods vary in how they prioritize positions, control parallelism, restrict selection regions, revise predictions, or plan future denoising, yet it remains unclear…
- Constraints Are Graphs, Not Chains: Exact Decoding for Diffusion Language Models
Jianchang Su, Wei Zhang · 29 September 2026 · Constraint Satisfaction and Optimization
Diffusion language models (dLLMs) predict masked positions in arbitrary order, but their exact constrained decoders still encode constraints as sequential languages, whose state must track every unresolved dependency between positions. For relational constraints this encoding grows exponentially: fo…
- Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis
Thong Bach, Dung Nguyen, Thao Minh Le, Truyen Tran · 28 September 2026 · Adversarial Robustness in Machine Learning
Existing attacks and defenses for diffusion-based large language models (dLLMs) target specific vulnerabilities but lack a shared framework explaining why attacks succeed. We propose one by interpreting safety alignment as shaping the denoising energy landscape: a well-aligned model routes harmful q…
- ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks
Zeyu Michael Li, William Xingxu Chen, Bingshuo Qian, Jiayin Liu, Xiang Cheng · 25 September 2026 · Model-Driven Software Engineering Techniques
Fully continuous diffusion language models (dLMs) denoise continuous representations without intermediate discretization, then decode all response tokens in parallel at the final step. Their performance on challenging reasoning tasks remains less established than that of autoregressive (AR) LLMs and…
- BiGraph-Diffuse: A Bidirectional Diffusion Language Model with Graph-Structured Retrieval For Mental Health Counseling
Yuxiang Cheng, Quanwei Tang, Lvhui Lu, Dong Zhang, Shoushan Li, Erik Cambria · 25 September 2026 · Mental Health via Writing
Mental health disorders affect hundreds of millions of people around the world, yet access to professional counseling remains severely limited. AI-powered dialogue systems offer a scalable alternative, but existing models face two fundamental challenges. First, they lack the bidirectional understand…
- Field-of-View Extension in Dental Cone-Beam CT via Implicit Neural Representations and Diffusion Model-Based Refinement
Susanne Schaub, Florentin Bieder, Matheus L. Oliveira, Yulan Wang, Buyanbileg Sodnom-ish, Dorothea Dagassan-Berndt, Michael M. Bornstein, Philippe C. Cattin · 24 September 2026 · Medical Imaging Techniques and Applications
Dental cone-beam computed tomography (CBCT) systems often employ detector configurations that provide a truncated field of view (FOV) that only captures a small part of the patient's anatomy. In this work, we aim to reconstruct an extended FOV using projections of truncated FOV scans. To this end, w…
- Towards Efficient Reasoning: Learning Causal Shortcuts for Diffusion Language Models
Dian Jin, Kairong Han, Baohong Li, Xinpeng Dong, Zijing Hu, Nuanqiao Shan, Fei Wu, Kun Kuang · 24 September 2026 · Large Language Models
Diffusion Language Models (DLMs) have attracted significant attention for their strong reasoning ability. However, under a bidirectional attention mechanism, DLMs operate over an exponentially large exploration space compared to autoregressive models (ARMs), making it challenging to focus on reasoni…
- LOCKR: A Hidden-State Trajectory-Guided Planner for Detecting and Repairing Stable-but-Wrong Lock-In in Diffusion Language Models
Guoshenghui Zhao, Tan Yu, Weijie Zhao · 24 September 2026 · Large Language Models
Diffusion language models generate text through iterative denoising, exposing intermediate trajectories before final answers are produced. We identify a recurring reasoning failure, stable-but-wrong lock-in, where an answer stabilizes early around an incorrect value while substantial denoising remai…
- PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models
Xiaocheng Lu, Shuhan Guo, Ziyue Ma, Jie Zhang, Jian Liu, Jingcai Guo, Haoxuan Che, Song Guo · 23 September 2026 · Power Systems and Technologies
Diffusion language models (dLLMs), such as LLaDA and Dream, have become competitive with autoregressive (AR) LLMs in generation quality while supporting native parallel decoding. A standard acceleration strategy is block-wise decoding, where each forward pass predicts a block of length B and commits…
- Circuit Hypernetworks for Quantum-Augmented Diffusion Language Models
Xiaoqiang Wang, Mengyang Xiong, Jun Dai, Bang Liu · 22 September 2026 · Quantum Computing Algorithms and Architecture
Language models can be adapted by changing the computations applied to individual tokens. Quantum circuits offer one such approach, but evaluating wider circuits inside a large model can be computationally demanding. Here we introduce HyperQ, which adds token-conditioned quantum residual branches to…
- A Ticket from Marginals to Joints: Coupled-Noise Distillation for One-Step Block Generation in Diffusion Language Models
Lin Yao · 21 September 2026 · Language and cultural evolution
Diffusion language models (dLLMs) predict all tokens of a block in parallel, but a single forward pass samples each position from its own marginal distribution, so the tokens need not form a coherent block. We ask whether a discrete masked model can commit an entire block in one pass when its mask e…
- Parallelism, critical windows, and separations among diffusion language models
Sitan Chen, Liye Wang · 18 September 2026 · Advanced Mathematical Modeling in Engineering
A popular selling point of diffusion large language models (dLLMs) is their capacity for parallelism: the ability to generate sequences of text far more efficiently than autoregressive models, which require one forward pass per token. Yet among the many competing paradigms for dLLMs, from masked to …
- Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training
Tarun Suresh, Pranshu Chaturvedi, Hangoo Kang, Parth Shroff, Ishan S. Khare, Hermann Kumbong, Azalia Mirhoseini · 18 September 2026 · Parallel Computing and Optimization Techniques
Block diffusion language models (BDLMs) combine autoregressive dependencies across blocks with parallel denoising within blocks, but long-context training is constrained by distributed attention communication and activation memory. Conventional context parallelism (CP) shards the combined clean-plus…
- dQwen3.5: Hybrid-Attention Diffusion Language Models
Anton Xue, Litu Rout, Aditya Akella, Adam Klivans, Sujay Sanghavi, Sanjay Shakkottai · 18 September 2026 · Generative Adversarial Networks and Image Synthesis
Adapting a pretrained autoregressive (AR) model is a cost-efficient route to a diffusion language model (DLM). While nearly all such adaptations start from a full-attention transformer, AR modeling has shifted toward hybrid architectures that interleave attention and RNN layers. This creates an obst…
- Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference
Leonid Sinev, Ilya Koziev, Vladislav Leshchuk · 18 September 2026 · Natural Language Processing Techniques
Autoregressive language models (ARMs) are constrained by sequential, left-to-right generation, while masked diffusion models (MDMs) enable parallel decoding but suffer from high computational overhead due to the inability to reuse Key-Value (KV) cache and from incoherent generation arising from lear…
- Generating Heterogeneous 3D Geological Microstructures from 2D Images via a Stable Diffusion-Adversarial Model
Ali Aouf, Eric Laloy, Bart Rogiers, Christophe De Vleeschouwer · 18 September 2026 · Enhanced Oil Recovery Techniques
Characterizing the physical properties of clay and cementitious materials matters across many fields, from materials science to geological waste disposal. Property simulation typically calls for 3D imaging, which is expensive, not always accessible, and technically limited for certain materials. Rec…
- Register Tokens for Bounded-State Reasoning in Diffusion Language Models
Albert Ge, Chandan Singh, Yufan Zhuang, Xiaodong Liu, Jianfeng Gao, Frederic Sala · 16 September 2026 · Large Language Models
Masked diffusion language models (dLLMs) generate text by iteratively denoising masked tokens with bidirectional attention. Extending reasoning across generation chunks normally requires keeping earlier generated text in context. We ask whether a dLLM can instead continue reasoning after that text i…
- Fast Convergence for High-Order ODE Solvers in Diffusion Probabilistic Models
Daniel Zhengyu Huang, Jiaoyang Huang, Zhengjiang Lin · 14 September 2026 · Blind Source Separation Techniques
Diffusion probabilistic models generate samples by learning to reverse a noise-injection process that transforms data into noise. A key development is the reformulation of the reverse sampling process as a deterministic probability flow ordinary differential equation (ODE), which allows for efficien…
- CanvasAnneal: Curriculum Reinforcement Learning for Diffusion Language Models
Blake Olson, Yuhang Song, Emmett McQuinn, Yuan Shangguan · 14 September 2026 · Generative Adversarial Networks and Image Synthesis
Diffusion Language Models (DLMs) offer promising parallel generation capabilities but lag behind autoregressive models in complex reasoning and tool-use tasks. While Reinforcement Learning (RL) has recently been applied to enhance DLMs, standard RL approaches suffer from an exploration bottleneck. T…
The search covers titles only, not the text of the abstracts. To query the content of the papers, the research assistant searches the indexed abstracts.
