Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Multimodal Machine Learning Applications
8069 artículos indexados
Las investigaciones agrupadas bajo este tema exploran cómo los modelos de inteligencia artificial combinan e interpretan simultáneamente varios tipos de datos, como imágenes, texto o esquemas científicos. Analizan, por ejemplo, los sesgos relacionados con los colores en los Vision Language Models, evalúan su capacidad para razonar sobre figuras complejas o para ensamblar estructuras inéditas bajo restricciones semánticas. Otros trabajos abordan la mejora de su percepción activa, su aptitud para expresar matices epistémicos o incluso su comportamiento frente a información visual ausente o engañosa.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- China52 % · 2985 artículos
- Estados Unidos36 % · 2061 artículos
- Reino Unido6,1 % · 349 artículos
- Corea del Sur5,7 % · 326 artículos
- RAE de Hong Kong (China)5,6 % · 316 artículos
- Singapur4,7 % · 267 artículos
- Alemania4,6 % · 262 artículos
- India4,3 % · 245 artículos
Sobre 5691 artículos de este tema con al menos un laboratorio localizado. 98 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null
Genpei Zhang · 2 de octubre de 2026
Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is n…
- MMMG: a Comprehensive and Reliable Benchmark for Multitask Multimodal Generation
Jihan Yao, Yushi Hu, Wenyuan Wang, Bin Han, Shangbin Feng, Guang Yang, Yujie Yi, Bingbing Wen, Ranjay Krishna, Lucy Lu Wang, Yulia Tsvetkov, Noah A. Smith, Banghua Zhu · 2 de octubre de 2026
Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation, especially for complex tasks that involve multiple modalities. We present MMMG, the first benchmark to bring the verifiable-task paradigm to multimodal…
- Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris · 2 de octubre de 2026
On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains larg…
- Task-Adaptive Grounded 3D-Programmers Using 2D VLMs
Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva, Jan-Nico Zaech, Luc Van Gool, Danda Pani Paudel · 2 de octubre de 2026
Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerf…
- MoLE: Mixture of Latent Experts for Complementary Visual Reasoning
Yingcheng Liu, Tianyi Jiang, Yujuan Ding, jiangbo Ai, Xun Jiang, Guoqing Wang, Wei Ye, Yi Bin · 2 de octubre de 2026
Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through …
- Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference
Xinye Zhao, Yunkai Dang, Yunchen Wu, Wenbin Li · 2 de octubre de 2026
Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially…
- Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning
Yuzhou Wang, Emile Anand, Ijay Narang · 2 de octubre de 2026
Reliable visual reasoning requires composing multiple visual observations and returning consistent answers to logically equivalent questions. We introduce Hob-VL, a benchmark for visually grounded Boolean reasoning. Hob-VL comprises two tasks: (1) evaluating whether a Boolean rule holds in an image,…
- Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies
Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh · 2 de octubre de 2026
Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pix…
- RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation
Minsu Kim, Jaesung Choe, Jiwoo Lee, Yu-Chiang Frank Wang, Seon Joo Kim · 2 de octubre de 2026
Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial r…
- VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision
Vu Dinh Xuan, Duc-Hai Nguyen, Minh-Dung Dao, Vu Quynh Giao, Quang Hong Nguyen, Binh-Son Hua, Barry O'Sullivan, David Murphy, Hoang D. Nguyen · 2 de octubre de 2026
Qualitative comparison figures are central evidence in computer vision papers, and vision-language models (VLMs) are increasingly used to judge them. Yet existing benchmarks score only scalar quality or overall preference, so a judge can be rewarded for picking the preferred image for the wrong visu…
- MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation
Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov, Alexey K. Kovalev · 2 de octubre de 2026
Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends o…
- LEGO-OPD: Factorized Teacher Composition for Multimodal On-Policy Distillation
Jaeyun Shin, Hangeol Chang, Jong Chul Ye · 2 de octubre de 2026
Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary supervision. However, directly using a VLM's full predictive distribut…
- A Framework for Egocentric and Exocentric Procedural Understanding via Temporal Segmentation and Semantic Abstraction
Vivek Chavan, J\"org Kr\"uger · 2 de octubre de 2026
Long-horizon ego/exo data contains rich procedural evidence, but are redundant, noisy, and costly to process or retain. We propose a compact framework that converts continuous multimodal workplace video into a structured Procedural State Memory, implemented as a Work Environment Model (WEM). Inspire…
- VISTA: A Visual Harness for Reasoning in an Interactive World
Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He · 2 de octubre de 2026
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows th…
- LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification
Haochen Zhang, Laura Yao, Zachary Plotkin, Gengwei Zhang, Tianlong Chen · 2 de octubre de 2026
Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model's captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were desi…
- CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement
Dongwei Sun, Yujie Zhang, Bowen Yao, Pei Liu, Jing Yao, Xiangyong Cao · 2 de octubre de 2026
Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily…
- Towards Reliable Vision-Language Models for Autonomous Driving
Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner · 2 de octubre de 2026
Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-wo…
- AiSearch: Interactive Multi-Modal Search with VLMs
Ali Koksal, Mei Chee Leong, Vicky Sintunata, Ching Ling Chin, Wee Teck Fong · 2 de octubre de 2026
Modern retrieval systems must both be automated and interactive, allowing users to search and refine results in real time. We present AiSearch, a flexible multimodal retrieval framework that leverages the zero shot capabilities of Vision Language Models (VLMs) for natural language search over images…
- EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
Shulin Tian, Junsu Kim, Shuai Liu, Hao Li, Yujiao Shen, Sihan Li, Zhe Yang, Yeongon Kim, Feiyu Li, Jialin Wu, Yichi Zhang, Wenhui Wang, Runmao Yao, Yuhao Dong, Zhaoxi Chen, Fangzhou Hong, Antonino Furnari, Jingkang Yang, Hongyuan Zhu, Ziwei Liu · 1 de octubre de 2026
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding the…
- Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift
Pengzhan Sun, Shiu-hong Kao, Shijie Li, Yongyi Su, Junbin Xiao, Arjun Reddy Akula, Angela Yao · 1 de octubre de 2026
This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suff…
- MindWorldBench: Evaluating Mental-State-to-Behavior Reasoning in Image-to-Video Generation
Ruiqi Li, Xuanyi Liu, Sijia Li, Haofeng Wang, Yuxin Liu, Feng Xie, Songchao Tan, Shiqi Wang, Hanwei Zhu, Yizong Wang, Chuanmin Jia, Siwei Ma · 1 de octubre de 2026
Current image-to-video models achieve visual realism and physical plausibility, but reasoning about mental states remains unexplored. Actions are driven by belief, desire, and perception, requiring inference beyond explicit instructions. We introduce MindWorldBench to evaluate mental-state-condition…
- Asking the World: Generalist Physical Reasoning through Agentic World Modeling and Probing
Shenxiang Zeng, Chen Yang, Peiyao Chen, Guohui Zhang, Jiansheng Fan, Chen Wang · 1 de octubre de 2026
Physical reasoning from video requires inferring latent physical properties and dynamics beyond direct observation. Direct VLM inference remains unreliable on complex physical tasks without explicit modeling and validation, while predefined tool pipelines rely on task- and domain-specific priors tha…
- Frame Differential On-Policy Self-Distillation for Video Reasoning
Haiying He, Xin Zheng, Shaoli Hu, Shijun Xiao, Xuanhe Liu, Bing Li, Harry Yang · 1 de octubre de 2026
Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly fine-grainedvisual or temporal credit assignment. In video reasoning, however, current RL methods typically train with a fixed sparse frame budget: i…
- DecoMoE: Decoupling Visual Propagation and Expert Computation for Efficient Multimodal MoE Inference
Xudong Tan, Peng Ye, Ming Xie, Chenyu Huang, Yaoxin Yang, Jiayuan Fan, Tao Chen · 1 de octubre de 2026
Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly because long visual-token sequences repeatedly incur attention, routing, dispatch, and expert-MLP computation. Existing methods typically compress either …
- Distill the Visual Evidence, Not Just the Answer: Cross-World On-Policy Distillation for Vision-Language Models
Yuanhao Sun, Huawei Ji, Jiaxin Ding, Luoyi Fu, Xinbing Wang · 1 de octubre de 2026
A central goal of vision-language model (VLM) distillation is to transfer both the teacher's language capabilities and its visual understanding. However, existing methods primarily supervise the student's output, leaving visual understanding implicit. Our analysis reveals that a student can match th…
Otros asuntos del tema Visión por computador y reconocimiento de formas
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Generative Adversarial Networks and Image Synthesis4992 artículos / 12 meses+39 %
- Advanced Neural Network Applications2354 artículos / 12 meses+48 %
- Advanced Vision and Imaging841 artículos / 12 meses+78 %
- Human Pose and Action Recognition836 artículos / 12 meses+457 %
- Face recognition and analysis482 artículos / 12 meses+88 %
- Image Enhancement Techniques463 artículos / 12 meses+20 %
