Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Multimodal Machine Learning Applications
8,069 papers indexed
The research grouped under this theme explores how artificial intelligence models simultaneously combine and interpret multiple types of data, such as images, text, or scientific diagrams. It analyzes, for example, biases related to colors in Vision Language Models, assesses their ability to reason about complex figures, or to assemble novel structures under semantic constraints. Other studies focus on improving their active perception, their capacity to express epistemic nuances, or their behavior when faced with missing or misleading visual information.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- China52% · 2,985 papers
- United States36% · 2,061 papers
- United Kingdom6.1% · 349 papers
- South Korea5.7% · 326 papers
- Hong Kong SAR China5.6% · 316 papers
- Singapore4.7% · 267 papers
- Germany4.6% · 262 papers
- India4.3% · 245 papers
Across 5,691 papers on this subject with at least one lab located. 98 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null
Genpei Zhang · 2 October 2026
Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is n…
- MMMG: a Comprehensive and Reliable Benchmark for Multitask Multimodal Generation
Jihan Yao, Yushi Hu, Wenyuan Wang, Bin Han, Shangbin Feng, Guang Yang, Yujie Yi, Bingbing Wen, Ranjay Krishna, Lucy Lu Wang, Yulia Tsvetkov, Noah A. Smith, Banghua Zhu · 2 October 2026
Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation, especially for complex tasks that involve multiple modalities. We present MMMG, the first benchmark to bring the verifiable-task paradigm to multimodal…
- Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris · 2 October 2026
On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains larg…
- Task-Adaptive Grounded 3D-Programmers Using 2D VLMs
Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva, Jan-Nico Zaech, Luc Van Gool, Danda Pani Paudel · 2 October 2026
Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerf…
- MoLE: Mixture of Latent Experts for Complementary Visual Reasoning
Yingcheng Liu, Tianyi Jiang, Yujuan Ding, jiangbo Ai, Xun Jiang, Guoqing Wang, Wei Ye, Yi Bin · 2 October 2026
Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through …
- Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference
Xinye Zhao, Yunkai Dang, Yunchen Wu, Wenbin Li · 2 October 2026
Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially…
- Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning
Yuzhou Wang, Emile Anand, Ijay Narang · 2 October 2026
Reliable visual reasoning requires composing multiple visual observations and returning consistent answers to logically equivalent questions. We introduce Hob-VL, a benchmark for visually grounded Boolean reasoning. Hob-VL comprises two tasks: (1) evaluating whether a Boolean rule holds in an image,…
- Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies
Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh · 2 October 2026
Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pix…
- RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation
Minsu Kim, Jaesung Choe, Jiwoo Lee, Yu-Chiang Frank Wang, Seon Joo Kim · 2 October 2026
Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial r…
- VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision
Vu Dinh Xuan, Duc-Hai Nguyen, Minh-Dung Dao, Vu Quynh Giao, Quang Hong Nguyen, Binh-Son Hua, Barry O'Sullivan, David Murphy, Hoang D. Nguyen · 2 October 2026
Qualitative comparison figures are central evidence in computer vision papers, and vision-language models (VLMs) are increasingly used to judge them. Yet existing benchmarks score only scalar quality or overall preference, so a judge can be rewarded for picking the preferred image for the wrong visu…
- MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation
Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov, Alexey K. Kovalev · 2 October 2026
Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends o…
- LEGO-OPD: Factorized Teacher Composition for Multimodal On-Policy Distillation
Jaeyun Shin, Hangeol Chang, Jong Chul Ye · 2 October 2026
Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary supervision. However, directly using a VLM's full predictive distribut…
- A Framework for Egocentric and Exocentric Procedural Understanding via Temporal Segmentation and Semantic Abstraction
Vivek Chavan, J\"org Kr\"uger · 2 October 2026
Long-horizon ego/exo data contains rich procedural evidence, but are redundant, noisy, and costly to process or retain. We propose a compact framework that converts continuous multimodal workplace video into a structured Procedural State Memory, implemented as a Work Environment Model (WEM). Inspire…
- VISTA: A Visual Harness for Reasoning in an Interactive World
Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He · 2 October 2026
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows th…
- LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification
Haochen Zhang, Laura Yao, Zachary Plotkin, Gengwei Zhang, Tianlong Chen · 2 October 2026
Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model's captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were desi…
- CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement
Dongwei Sun, Yujie Zhang, Bowen Yao, Pei Liu, Jing Yao, Xiangyong Cao · 2 October 2026
Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily…
- Towards Reliable Vision-Language Models for Autonomous Driving
Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner · 2 October 2026
Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-wo…
- AiSearch: Interactive Multi-Modal Search with VLMs
Ali Koksal, Mei Chee Leong, Vicky Sintunata, Ching Ling Chin, Wee Teck Fong · 2 October 2026
Modern retrieval systems must both be automated and interactive, allowing users to search and refine results in real time. We present AiSearch, a flexible multimodal retrieval framework that leverages the zero shot capabilities of Vision Language Models (VLMs) for natural language search over images…
- EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
Shulin Tian, Junsu Kim, Shuai Liu, Hao Li, Yujiao Shen, Sihan Li, Zhe Yang, Yeongon Kim, Feiyu Li, Jialin Wu, Yichi Zhang, Wenhui Wang, Runmao Yao, Yuhao Dong, Zhaoxi Chen, Fangzhou Hong, Antonino Furnari, Jingkang Yang, Hongyuan Zhu, Ziwei Liu · 1 October 2026
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding the…
- Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift
Pengzhan Sun, Shiu-hong Kao, Shijie Li, Yongyi Su, Junbin Xiao, Arjun Reddy Akula, Angela Yao · 1 October 2026
This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suff…
- MindWorldBench: Evaluating Mental-State-to-Behavior Reasoning in Image-to-Video Generation
Ruiqi Li, Xuanyi Liu, Sijia Li, Haofeng Wang, Yuxin Liu, Feng Xie, Songchao Tan, Shiqi Wang, Hanwei Zhu, Yizong Wang, Chuanmin Jia, Siwei Ma · 1 October 2026
Current image-to-video models achieve visual realism and physical plausibility, but reasoning about mental states remains unexplored. Actions are driven by belief, desire, and perception, requiring inference beyond explicit instructions. We introduce MindWorldBench to evaluate mental-state-condition…
- Asking the World: Generalist Physical Reasoning through Agentic World Modeling and Probing
Shenxiang Zeng, Chen Yang, Peiyao Chen, Guohui Zhang, Jiansheng Fan, Chen Wang · 1 October 2026
Physical reasoning from video requires inferring latent physical properties and dynamics beyond direct observation. Direct VLM inference remains unreliable on complex physical tasks without explicit modeling and validation, while predefined tool pipelines rely on task- and domain-specific priors tha…
- Frame Differential On-Policy Self-Distillation for Video Reasoning
Haiying He, Xin Zheng, Shaoli Hu, Shijun Xiao, Xuanhe Liu, Bing Li, Harry Yang · 1 October 2026
Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly fine-grainedvisual or temporal credit assignment. In video reasoning, however, current RL methods typically train with a fixed sparse frame budget: i…
- DecoMoE: Decoupling Visual Propagation and Expert Computation for Efficient Multimodal MoE Inference
Xudong Tan, Peng Ye, Ming Xie, Chenyu Huang, Yaoxin Yang, Jiayuan Fan, Tao Chen · 1 October 2026
Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly because long visual-token sequences repeatedly incur attention, routing, dispatch, and expert-MLP computation. Existing methods typically compress either …
- Distill the Visual Evidence, Not Just the Answer: Cross-World On-Policy Distillation for Vision-Language Models
Yuanhao Sun, Huawei Ji, Jiaxin Ding, Luoyi Fu, Xinbing Wang · 1 October 2026
A central goal of vision-language model (VLM) distillation is to transfer both the teacher's language capabilities and its visual understanding. However, existing methods primarily supervise the student's output, leaving visual understanding implicit. Our analysis reveals that a student can match th…
Other topics in Computer vision and pattern recognition
The topics the OpenAlex classification attaches to the same theme, most active first.
- Generative Adversarial Networks and Image Synthesis4,992 papers / 12 months+39%
- Advanced Neural Network Applications2,354 papers / 12 months+48%
- Advanced Vision and Imaging841 papers / 12 months+78%
- Human Pose and Action Recognition836 papers / 12 months+457%
- Face recognition and analysis482 papers / 12 months+88%
- Image Enhancement Techniques463 papers / 12 months+20%
