Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Multimodal Machine Learning Applications
8.069 indexierte Paper
Die unter diesem Thema zusammengefassten Forschungen untersuchen, wie Modelle der künstlichen Intelligenz verschiedene Datentypen wie Bilder, Texte oder wissenschaftliche Diagramme gleichzeitig kombinieren und interpretieren. Sie analysieren beispielsweise Farbbias in Vision Language Models, bewerten deren Fähigkeit, über komplexe Abbildungen zu raisonnieren oder neuartige Strukturen unter semantischen Einschränkungen zusammenzusetzen. Weitere Arbeiten befassen sich mit der Verbesserung ihrer aktiven Wahrnehmung, ihrer Fähigkeit, epistemische Nuancen auszudrücken, oder ihrem Verhalten gegenüber fehlenden oder irreführenden visuellen Informationen.
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- China52 % · 2.985 Artikel
- Vereinigte Staaten36 % · 2.061 Artikel
- Vereinigtes Königreich6,1 % · 349 Artikel
- Südkorea5,7 % · 326 Artikel
- Sonderverwaltungsregion Hongkong5,6 % · 316 Artikel
- Singapur4,7 % · 267 Artikel
- Deutschland4,6 % · 262 Artikel
- Indien4,3 % · 245 Artikel
Über 5.691 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 98 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null
Genpei Zhang · 2. Oktober 2026
Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is n…
- MMMG: a Comprehensive and Reliable Benchmark for Multitask Multimodal Generation
Jihan Yao, Yushi Hu, Wenyuan Wang, Bin Han, Shangbin Feng, Guang Yang, Yujie Yi, Bingbing Wen, Ranjay Krishna, Lucy Lu Wang, Yulia Tsvetkov, Noah A. Smith, Banghua Zhu · 2. Oktober 2026
Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation, especially for complex tasks that involve multiple modalities. We present MMMG, the first benchmark to bring the verifiable-task paradigm to multimodal…
- Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris · 2. Oktober 2026
On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains larg…
- Task-Adaptive Grounded 3D-Programmers Using 2D VLMs
Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva, Jan-Nico Zaech, Luc Van Gool, Danda Pani Paudel · 2. Oktober 2026
Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerf…
- MoLE: Mixture of Latent Experts for Complementary Visual Reasoning
Yingcheng Liu, Tianyi Jiang, Yujuan Ding, jiangbo Ai, Xun Jiang, Guoqing Wang, Wei Ye, Yi Bin · 2. Oktober 2026
Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through …
- Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference
Xinye Zhao, Yunkai Dang, Yunchen Wu, Wenbin Li · 2. Oktober 2026
Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially…
- Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning
Yuzhou Wang, Emile Anand, Ijay Narang · 2. Oktober 2026
Reliable visual reasoning requires composing multiple visual observations and returning consistent answers to logically equivalent questions. We introduce Hob-VL, a benchmark for visually grounded Boolean reasoning. Hob-VL comprises two tasks: (1) evaluating whether a Boolean rule holds in an image,…
- Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies
Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh · 2. Oktober 2026
Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pix…
- RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation
Minsu Kim, Jaesung Choe, Jiwoo Lee, Yu-Chiang Frank Wang, Seon Joo Kim · 2. Oktober 2026
Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial r…
- VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision
Vu Dinh Xuan, Duc-Hai Nguyen, Minh-Dung Dao, Vu Quynh Giao, Quang Hong Nguyen, Binh-Son Hua, Barry O'Sullivan, David Murphy, Hoang D. Nguyen · 2. Oktober 2026
Qualitative comparison figures are central evidence in computer vision papers, and vision-language models (VLMs) are increasingly used to judge them. Yet existing benchmarks score only scalar quality or overall preference, so a judge can be rewarded for picking the preferred image for the wrong visu…
- MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation
Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov, Alexey K. Kovalev · 2. Oktober 2026
Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends o…
- LEGO-OPD: Factorized Teacher Composition for Multimodal On-Policy Distillation
Jaeyun Shin, Hangeol Chang, Jong Chul Ye · 2. Oktober 2026
Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary supervision. However, directly using a VLM's full predictive distribut…
- A Framework for Egocentric and Exocentric Procedural Understanding via Temporal Segmentation and Semantic Abstraction
Vivek Chavan, J\"org Kr\"uger · 2. Oktober 2026
Long-horizon ego/exo data contains rich procedural evidence, but are redundant, noisy, and costly to process or retain. We propose a compact framework that converts continuous multimodal workplace video into a structured Procedural State Memory, implemented as a Work Environment Model (WEM). Inspire…
- VISTA: A Visual Harness for Reasoning in an Interactive World
Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He · 2. Oktober 2026
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows th…
- LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification
Haochen Zhang, Laura Yao, Zachary Plotkin, Gengwei Zhang, Tianlong Chen · 2. Oktober 2026
Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model's captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were desi…
- CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement
Dongwei Sun, Yujie Zhang, Bowen Yao, Pei Liu, Jing Yao, Xiangyong Cao · 2. Oktober 2026
Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily…
- Towards Reliable Vision-Language Models for Autonomous Driving
Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner · 2. Oktober 2026
Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-wo…
- AiSearch: Interactive Multi-Modal Search with VLMs
Ali Koksal, Mei Chee Leong, Vicky Sintunata, Ching Ling Chin, Wee Teck Fong · 2. Oktober 2026
Modern retrieval systems must both be automated and interactive, allowing users to search and refine results in real time. We present AiSearch, a flexible multimodal retrieval framework that leverages the zero shot capabilities of Vision Language Models (VLMs) for natural language search over images…
- EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
Shulin Tian, Junsu Kim, Shuai Liu, Hao Li, Yujiao Shen, Sihan Li, Zhe Yang, Yeongon Kim, Feiyu Li, Jialin Wu, Yichi Zhang, Wenhui Wang, Runmao Yao, Yuhao Dong, Zhaoxi Chen, Fangzhou Hong, Antonino Furnari, Jingkang Yang, Hongyuan Zhu, Ziwei Liu · 1. Oktober 2026
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding the…
- Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift
Pengzhan Sun, Shiu-hong Kao, Shijie Li, Yongyi Su, Junbin Xiao, Arjun Reddy Akula, Angela Yao · 1. Oktober 2026
This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suff…
- MindWorldBench: Evaluating Mental-State-to-Behavior Reasoning in Image-to-Video Generation
Ruiqi Li, Xuanyi Liu, Sijia Li, Haofeng Wang, Yuxin Liu, Feng Xie, Songchao Tan, Shiqi Wang, Hanwei Zhu, Yizong Wang, Chuanmin Jia, Siwei Ma · 1. Oktober 2026
Current image-to-video models achieve visual realism and physical plausibility, but reasoning about mental states remains unexplored. Actions are driven by belief, desire, and perception, requiring inference beyond explicit instructions. We introduce MindWorldBench to evaluate mental-state-condition…
- Asking the World: Generalist Physical Reasoning through Agentic World Modeling and Probing
Shenxiang Zeng, Chen Yang, Peiyao Chen, Guohui Zhang, Jiansheng Fan, Chen Wang · 1. Oktober 2026
Physical reasoning from video requires inferring latent physical properties and dynamics beyond direct observation. Direct VLM inference remains unreliable on complex physical tasks without explicit modeling and validation, while predefined tool pipelines rely on task- and domain-specific priors tha…
- Frame Differential On-Policy Self-Distillation for Video Reasoning
Haiying He, Xin Zheng, Shaoli Hu, Shijun Xiao, Xuanhe Liu, Bing Li, Harry Yang · 1. Oktober 2026
Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly fine-grainedvisual or temporal credit assignment. In video reasoning, however, current RL methods typically train with a fixed sparse frame budget: i…
- DecoMoE: Decoupling Visual Propagation and Expert Computation for Efficient Multimodal MoE Inference
Xudong Tan, Peng Ye, Ming Xie, Chenyu Huang, Yaoxin Yang, Jiayuan Fan, Tao Chen · 1. Oktober 2026
Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly because long visual-token sequences repeatedly incur attention, routing, dispatch, and expert-MLP computation. Existing methods typically compress either …
- Distill the Visual Evidence, Not Just the Answer: Cross-World On-Policy Distillation for Vision-Language Models
Yuanhao Sun, Huawei Ji, Jiaxin Ding, Luoyi Fu, Xinbing Wang · 1. Oktober 2026
A central goal of vision-language model (VLM) distillation is to transfer both the teacher's language capabilities and its visual understanding. However, existing methods primarily supervise the student's output, leaving visual understanding implicit. Our analysis reveals that a student can match th…
Weitere Unterthemen aus Bildverarbeitung und Mustererkennung
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Generative Adversarial Networks and Image Synthesis4.992 Papiere / 12 Monate+39 %
- Advanced Neural Network Applications2.354 Papiere / 12 Monate+48 %
- Advanced Vision and Imaging841 Papiere / 12 Monate+78 %
- Human Pose and Action Recognition836 Papiere / 12 Monate+457 %
- Face recognition and analysis482 Papiere / 12 Monate+88 %
- Image Enhancement Techniques463 Papiere / 12 Monate+20 %
