Search
Search: diffusion models
Words are combined with AND. Use quotes for an exact phrase, a leading dash to exclude a word.
Papers
Page 39 of 40
More than 1,000 papers match: here are the 1,000 most recent, ranked by relevance.
- Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models
Sujung Hong, Chanyong Yoon, Seong Jae Hwang · 15 May 2026 · Multimodal Machine Learning Applications
Large diffusion vision-language models (LDVLMs) have recently emerged as a promising alternative to autoregressive models, enabling parallel decoding for efficient inference and leveraging bidirectional attention for global context. Despite these advances, their behavior under long-form generation r…
- Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement
Injin Kong, Hyoungjoon Lee, Yohan Jo · 15 May 2026 · Large Language Models
Continuous diffusion language models lag behind autoregressive transformers, partly because diffusion is applied in spaces poorly suited to language denoising and token recovery. We propose DiHAL, a geometry-guided diffusion-transformer hybrid that asks where diffusion should enter a pretrained tran…
- Differences in Text Generated by Diffusion and Autoregressive Language Models
Zeyang Zhang, Chengwei Liang, Xingyan Chen, Meiqi Gu, Minrui Luo, Jingzhao Zhang, Tianxing He · 14 May 2026 · Large Language Models
Diffusion language models (DLMs) are promising alternatives to autoregressive language models (ARMs), yet the intrinsic differences in their generated text remain underexplored. We first find empirically that off-the-shelf DLMs exhibit lower $n$-gram entropy, higher semantic coherence, and higher se…
- TextLDM: Language Modeling with Continuous Latent Diffusion
Jiaxiu Jiang, Jingjing Ren, Wenbo Li, Bo Wang, Haoze Sun, Yijun Yang, Jianhui Liu, Yanbing Zhang, Shenghe Zheng, Yuan Zhang, Haoyang Huang, Nan Duan, Wangmeng Zuo · 11 May 2026 · Generative Adversarial Networks and Image Synthesis
Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next step toward a single architecture for both generation (visual synthesis) and understanding (text generation) is to apply this framework to language mo…
- TRACE: Transport Alignment Conformal Prediction via Diffusion and Flow Matching Models
Zhenhan Fang, Aixin Tan, Jian Huang · 11 May 2026 · Generative Adversarial Networks and Image Synthesis
Constructing valid and informative conformal prediction regions for multi-dimensional outputs remains a fundamental challenge. While conformal prediction provides finite-sample, distribution-free coverage guarantees, its practical performance critically depends on the choice of nonconformity score. …
- BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation
Baoyou Chen, Hanchen Xia, Peng Tu, Haojun Shi, Shan Mu, Weihao Yuan, Siyu Zhu · 21 April 2026 · Multimodal Machine Learning Applications
Autoregressive vision-language models (VLMs) deliver strong multimodal capability, but their token-by-token decoding imposes a fundamental inference bottleneck. Diffusion VLMs offer a more parallel decoding paradigm, yet directly converting a pretrained autoregressive VLM into a large-block diffusio…
- Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
Hoeun Lee, Jaeik Kim, Jusang Oh, Jinhyeok Kim, Geon Choi, Hyeonggeun Kim, Jaeyoung Do · 14 September 2026 · Robot Manipulation and Learning
Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a representation of action-dependent scene changes. We bring these predictions into action generation and selection through a shared trajectory model. Dynin-Robotics implements this for…
- Representation Learning in Diffusion and Flow-based Model: An Application Aspect
Yanchen Xu, Sida Huang, Zhenyu Gu, Ruishu Zhu, Yilan Gao, Hongyuan Zhang · 26 August 2026 · Generative Adversarial Networks and Image Synthesis
Diffusion models and flow-based models have recently become the dominant paradigms in generative modeling, largely due to their ability to learn rich, multi-level visual representations through large-scale training. This creates a bidirectional relationship between generative models and representati…
- GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech model
Guanrou Yang, Tian Tan, Qian Chen, Ziyang Ma, Yakun Song, Zhikang Niu, Qi Chen, Wenming Tu, Haitao Li, Shan Yang, Xie Chen · 5 August 2026 · Speech Recognition and Synthesis
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE and track per-step likelihood ratios, introducing stochastic perturbations and substantial overhead. We propose GROW, a…
- Sampling the Schwinger Model with Gauge-Equivariant Diffusion
Octavio Vega, Aida X. El-Khadra · 29 June 2026 · Quantum many-body systems
We present a first study of a diffusion-based approach to accelerated sampling of the $N_f = 2$ lattice Schwinger model. Our work is inspired by recent and growing successes in developing such generative models for ensemble generation in LFT to overcome the well-known critical slowing down problem. …
- SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion
Ruoyu Feng, Jinming Liu, Yuqi Wang, Xin Cheng, Boyuan Liu, Shanglin Li, Hanshen Zhu, Wenfeng Lin, Mingyu Guo, Xin Jin · 23 June 2026 · Generative Adversarial Networks and Image Synthesis
Training image generation foundation models consumes substantial resources. Previous methods have attempted to leverage semantic guidance to accelerate the training process, yet their experiments were only conducted on simple datasets such as ImageNet, at low resolutions, and with small-scale models…
- PG-MAP: Joint MAP Optimization for Inference-Time Alignment of Diffusion and Flow-Matching Models
Ruolan Sun, Pawel Polak · 23 June 2026 · Generative Adversarial Networks and Image Synthesis
Inference-time alignment of pretrained text-to-image models is typically performed along a single control axis, such as classifier-free guidance, attention editing, or reward-based latent perturbations. This limitation prevents modeling joint dependencies between conditioning and latent variables an…
- Learning a Maximum Entropy Model for Visual Textures using Diffusion
Xinyuan Zhao, Eero P. Simoncelli · 17 June 2026 · Generative Adversarial Networks and Image Synthesis
Visual textures -- spatially homogeneous image regions containing repeated elements (e.g. a field of grass, the bark of a tree) -- are ubiquitous in visual scenes and provide important cues for recognizing and analyzing materials and objects. A number of existing texture models extract essential sta…
- Diffusion Transformer World-Action Model for AV Scene Prediction
Ruslan Sharifullin, Benjamin Jiang, Kai Xi Chew · 12 June 2026 · Advanced Vision and Imaging
Action-conditioned world models let an autonomous vehicle predict future camera scenes from its own planned controls, enabling planning and simulation without real-world rollouts, but at compact, trainable scale the futures are ambiguous and the field's standard distortion metrics actively mislead: …
- Diffusion Image Generation with Explicit Modeling of Data Manifold Geometry
Duoduo Xue, Zhiyu Zhu, Junhui Hou · 2 June 2026 · Generative Adversarial Networks and Image Synthesis
Image generative models aim to sample data points from the underlying data manifold, a task that requires learning and decoding a dense, low-dimensional, and compact parameterization space. To achieve this, we propose the Data Manifold-aware Image diffusioN moDel (MIND), a novel framework that expli…
- From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons
Xiangyu Ma, Teng Xiao, Zuchao Li, Lefei Zhang · 28 May 2026 · Large Language Models
Diffusion models promise efficient parallel text generation but rely on bidirectional attention, creating a structural mismatch with pre-trained Autoregressive (AR) models. This incompatibility precludes reusing robust AR priors, necessitating prohibitive pre-training from scratch. To bridge this ga…
- Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models
Yulin Yuan, Hongshuo Zhao, Xiangming Meng · 26 May 2026 · Multimodal Machine Learning Applications
Diffusion-based multimodal large language models (dMLLMs) decode by iteratively predicting tokens at multiple masked positions in parallel. This turns each decoding step into a position-selection problem: the model must choose not only which predictions are reliable in isolation, but also which posi…
- Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models
Andreas Bergmeister, Stefanie Jegelka, Nikolas N\"usken, Carles Domingo-Enrich, Jakiw Pidstrigach · 19 May 2026 · Recommender Systems and Techniques
Diffusion and flow-matching models scale because pretraining is supervised regression: a clean sample is noised analytically, and a model regresses against a closed-form target. RL post-training aligns the model with a reward. In image generation, this makes samples compose objects correctly, render…
- Feedback World Model Enables Precise Guidance of Diffusion Policy
Tuo An, Jindou Jia, Gen Li, Jingliang Li, Chuhao Zhou, Pengfei Liu, Bofan Lyu, Jiaqi Bai, Xinying Guo, Geng Li, Jianfei Yang · 18 May 2026 · Reinforcement Learning in Robotics
World models aim to improve robotic decision making by predicting the consequences of actions. However, in practice, their predictions often become unreliable once the robot encounters states outside the training distribution, limiting their effectiveness at deployment. We observe that execution its…
- Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
Weiming Chen, Xitong Ling, Zhenyang Cai, Xidong Wang, Jiawen Li, Tian Guan, Benyou Wang, Yonghong He · 11 May 2026 · AI in cancer detection
Cell-level dense prediction is central to computational pathology, but remains challenging due to fine-grained histological structures, strong domain shifts, and costly dense annotations. Existing ViT-based pathology foundation models rely on patch tokenization, which can disrupt spatial continuity …
- DiffMI: Breaking Face Recognition Privacy via Diffusion-Driven Training-Free Model Inversion
Hanrui Wang, Shuo Wang, Chun-Shien Lu, Isao Echizen · 4 May 2026 · Face recognition and analysis
Face recognition poses serious privacy risks due to its reliance on sensitive and immutable biometric data. While modern systems mitigate privacy risks by mapping facial images to embeddings (commonly regarded as privacy-preserving), model inversion attacks reveal that identity information can still…
- DGSSM: Diffusion guided state-space models for multimodal salient object detection
Suklav Ghosh, Arijit Sur, Pinaki Mitra · 21 April 2026 · Visual Attention and Saliency Detection
Salient object detection (SOD) requires modeling both long-range contextual dependencies and fine-grained structural details, which remains challenging for convolutional, transformer-based, and Mamba-based state space models. While recent Mamba-based state space approaches enable efficient global re…
- Variational Autoencoding Discrete Diffusion with Enhanced Dimensional Correlations Modeling
Tianyu Xie, Shuchen Xue, Zijin Feng, Tianyang Hu, Jiacheng Sun, Zhenguo Li, Cheng Zhang · 15 April 2026 · Generative Adversarial Networks and Image Synthesis
Discrete diffusion models have recently shown great promise for modeling complex discrete data, with masked diffusion models (MDMs) offering a compelling trade-off between quality and generation speed. MDMs denoise by progressively unmasking multiple dimensions from an all-masked input, but their pe…
- LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling
Yuxin Chen, Chumeng Liang, Hangke Sui, Ruihan Guo, Chaoran Cheng, Jiaxuan You, Ge Liu · 14 April 2026 · Large Language Models
Continuous diffusion models have achieved strong performance across domains such as images. However, in language modeling, prior continuous diffusion language models (DLMs) lag behind discrete counterparts. In this work, we close this gap with LangFlow, the first continuous DLM to rival discrete dif…
- WorldDiT: A Unified Diffusion Architecture for World and Action Modeling
Sen Wang, R. Gnana Praveen, Bidhan Roy, Marcos Villagra · 28 July 2026 · Multimodal Machine Learning Applications
Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a l…
The search covers titles only, not the text of the abstracts. To query the content of the papers, the research assistant searches the indexed abstracts.
