Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Multimodal Machine Learning Applications
8 069 papiers indexés
Les recherches regroupées sous ce thème explorent comment les modèles d’intelligence artificielle combinent et interprètent simultanément plusieurs types de données, comme des images, du texte ou des schémas scientifiques. Elles analysent par exemple les biais liés aux couleurs dans les Vision Language Models, évaluent leur capacité à raisonner sur des figures complexes ou à assembler des structures inédites sous contraintes sémantiques. D’autres travaux portent sur l’amélioration de leur perception active, leur aptitude à exprimer des nuances épistémiques, ou encore leur comportement face à des informations visuelles absentes ou trompeuses.
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Pays des laboratoires
- Chine52 % · 2 985 articles
- États-Unis36 % · 2 061 articles
- Royaume-Uni6,1 % · 349 articles
- Corée du Sud5,7 % · 326 articles
- R.A.S. chinoise de Hong Kong5,6 % · 316 articles
- Singapour4,7 % · 267 articles
- Allemagne4,6 % · 262 articles
- Inde4,3 % · 245 articles
Sur 5 691 articles de ce sujet dont au moins un laboratoire est situé. 98 pays représentés.
Il s'agit du pays du laboratoire, jamais de la nationalité des personnes. Un article signé depuis plusieurs pays compte pour chacun d'eux, les parts dépassent donc 100 % au total. La couverture est partielle et le manque n'est pas aléatoire : un chercheur dont l'institution est inconnue publie en général peu, ce qui sur-représente les laboratoires établis.
Derniers papiers
- Architecture-Dependent Fusion Pathways in MLLMs
Hebao Zhu, Dongxia Wu · 5 octobre 2026
Multimodal Large Language Models (MLLMs) achieve strong performance across vision-language tasks, yet the internal mechanisms by which visual and textual information are fused across layers remain insufficiently understood. We investigate representative MLLMs from two architectural paradigms: concat…
- Retrospective Open-Vocabulary Memory for Long-Term Object Search
Jiaming Wang, Zhiwei Xue, Chen Jizhuo, Peng Shiqi, Harold Soh · 5 octobre 2026
Long-term object search requires learning where objects usually appear from repeated but uneven observations of a changing environment. We formulate retrospective open-vocabulary memory as probabilistic inference from censored observations, where the key idea is to reason with evidence per opportuni…
- Form and Void: Entangled Composition through an Autonomous AI Agent
Shiwen Wang, Jian Yang, Xu Wang, Xincan Wang, Weiming Dong · 5 octobre 2026
Positive and negative space is a fundamental principle in visual composition, supporting visually coherent forms and layered semantic relationships. Generating such compositions is challenging because it requires coordinated control over two semantic concepts that share a common boundary. Although r…
- Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
Luka Ribar, Jeevan Bhoot, Douglas Orr · 5 octobre 2026
Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model it…
- GeoScaffold: Learning Compact Geometric Latents via Reconstruction for Efficient Vision-Language Navigation
Yixuan Jiang, Wentong Li, An Liu, Zihao Xin, Fulin Tang, Cong Leng, Yang Gao, Jian Cheng · 5 octobre 2026
Recent vision-and-language navigation (VLN) systems increasingly adopt streaming Video-LLM policies that map egocentric RGB observations and instructions directly to low-level actions. Yet these policies inherit weak 3D geometric priors from 2D pretraining. Existing geometry-aware extensions charge …
- Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning
Yudong Han, Yong Wang, Zaiquan Yang, Liang Lin, Chongyang Tao, Xiangxiang Chu, Liyuan Pan · 5 octobre 2026
Reinforcement learning with verifiable rewards has substantially advanced multimodal reasoning, yet it remains fundamentally limited by ambiguous token-level credit assignment. While high-entropy token heuristics encourage possibility exploration, naively extending them to video reasoning tends to i…
- From Patching to Pruning Visual Computation in Vision Language Models
Rahul Chowdhury, Timothy A Rupprecht, Xuan Shen, Shaoyi Huang, Pu Zhao, Yanzhi Wang · 5 octobre 2026
Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder layer, even when token-specific visual computation is unnecessary at many depths. We introduce Patch-to-Prune (P2P), inspired by Mechanistic In…
- Behavior Pack Optimization for Video MLLM Post-Training
Zhaolu Kang, Shiyu Liu, Tailong Luo, Wei Zhang, Yingjie He, Lei Wei, Guansu Wang, Liang He, Siheng Wang, Guangyuan Dong, Jiaqi Su, Shuang Chen, Haoyu Ji, Qishi Zhan, Kaiyue Zhou · 5 octobre 2026
Video multimodal large language models (MLLMs) keep climbing video question answering benchmarks, yet shuffling the frames, masking the segment that supports the answer, or occluding the target object barely changes their predictions. The accuracy rests on appearance and language priors, not on the …
- TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows
Shuai Fu, Jing Gu, Jian Zhou, Zicheng Duan, Gengze Zhou, Qi Wu · 5 octobre 2026
Recent text-to-image models have made substantial progress in photorealism, aesthetics, and text-image alignment. Yet visually appealing images can still violate real-world plausibility, exhibiting malformed object structures, impossible anatomy, physically implausible interactions, or inconsistent …
- CHASE-VLA: Post-Training Quantization Framework for Vision-Language-Action Models with Chunk-Aware Scale Estimation
Jin Hyun, Jung Gyu Min, Gyuhyun Jung, Youngjoo Lee · 5 octobre 2026
Vision-Language-Action (VLA) models map visual observations and language instructions to continuous robot actions, but a diffusion-based action expert (AE) poses a key challenge for low-bit post-training quantization (PTQ). The AE is repeatedly invoked across denoising steps and policy queries, wher…
- Imagine the Future, Internalize the Gist: Efficient VLA Reasoning via Internalized Spatiotemporal Imagination
Shenglan Li, Zhendong Mi, Hengyi Zhu, Jingwu Luo, Chun Kit Chan, Geng Yuan, Yanzhi Wang, Pu Zhao, Shaoyi Huang · 5 octobre 2026
Vision-language-action (VLA) models increasingly incorporate intermediate reasoning to improve robotic manipulation, yet existing approaches primarily reason about observed states without explicitly anticipating future scene evolution. Extending such reasoning to explicit future rollouts at every in…
- CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation
Hang Gao, Wujiang Xu, Zhixing Zhang, Kai Mei, Jingyi Yang, Dimitris N. Metaxas · 5 octobre 2026
Multimodal large language models (MLLMs) have shown strong visual reasoning abilities, but knowledge-intensive visual question answering often requires external textual evidence beyond the image and the model's parametric knowledge. Existing multimodal RAG systems commonly rely on Top-$K$ retrieval …
- OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination
Huiqiang Rong, Haoran Luo, Hui Feng, Zhonghong Ou, Kaiwen Xue, Guoxin Zhang, Yifan Zhu · 5 octobre 2026
Omni-modal large language models (OmniLLMs) unify text, images, audio, and video, yet hallucinate when generation relies on the wrong evidence. Existing inference-time methods can reduce hallucinations, but rarely reveal which evidence sustains a generated commitment. We introduce OmniConfess, a tra…
- Evaluating VQA in Vision Language Models using Cooperative Principles
Monika Shah, Sudarshan Balaji, Somdeb Sarkhel, Sanorita Dey, Deepak Venugopal · 5 octobre 2026
We evaluate the performance of Vision Language Models in Visual Question Answering (VQA) when questions violate Grice's maxims. To do this, we use VLMs to generate question modifiers that add non-essential, ambiguous or false information and show that in the presence of such violations, the VLMs tha…
- Text-Centric Post-Training for Omni-Modal Reasoning
Ziyang Cheng, Yuhao Wang, Hongcheng Liu, Qimin Wu, Jingru Fan, Chen Qian, Yanfeng Wang, Yu Wang · 5 octobre 2026
Improving joint audio-visual reasoning in Omni Large Language Models typically incurs substantial data construction and training costs. Our diagnostics reveal multi-hop reasoning difficulties despite correct answers to all corresponding single-hop questions and suggest partial decoupling in the loca…
- ConfAL-WM: Confidence-Guided Active Learning for Action-Conditioned World Models
Xiang Liu, Kunwei Wu, Miao Liu, Sen Cui, Changshui Zhang · 5 octobre 2026
Action-conditioned world models have become an important foundation for embodied prediction, planning, and synthetic data generation, but their errors under new task and scene distributions are often concentrated in localized spatiotemporal regions such as robot arms, manipulated objects, contact ar…
- Multimodal Representation Learning Conditioned on Semantic Relations
Yang Qiao, Yuntong Hu, Bowen Zhu, Hasibul Haque, Liang Zhao · 5 octobre 2026
Multimodal representation learning has been largely driven by contrastive models such as CLIP, which learn a shared embedding space by aligning paired image-text samples. While effective for general-purpose representation learning, such models typically produce a single embedding per sample that is …
- World Action Planner: Generalizable Robot Decision-Making with Action-Conditioned World Models
Xiangcheng Zhang, Runhan Huang, Yilun Du · 5 octobre 2026
Building generalizable robot agents for diverse applications remains a fundamental challenge. While imitation learning-based policies can perform well in familiar training environments, they often struggle to generalize to novel scenes, layouts, and task compositions. To this end, we present World A…
- Foresight: planning future perception in streaming VLMs without retraining
Ashok Prasad Neupane, Dipan Bartaula, Ankit Belbase, Saugat Adhikari, Samip Ghimire, Saroj Poudel, Binod Bhattarai, Danda Pani Paudel · 5 octobre 2026
Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and…
- Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning
Jinghan Zhao, Yiman Hu, Liang Wu, Jian Xu, Bo Zheng · 5 octobre 2026
E-commerce videos are information-dense and frequently compared by consumers evaluating products and merchants assessing marketing strategies. However, existing multimodal models mainly focus on single-video understanding and have limited ability to compare information across videos. We introduce Ad…
- Relevant Evidence Decoding for Audio-Visual Hallucination Mitigation
Hyunjae Ra, Aecheon Jung, Jungin Park, Sungeun Hong · 5 octobre 2026
Audio-Visual Large Language Models (AV-LLMs) remain prone to cross-modal hallucinations, where one modality incorrectly affects predictions about another. Although contrastive decoding reduces hallucinations in vision-language models, its direct extension to AV-LLMs overlooks a key challenge: differ…
- Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null
Genpei Zhang · 2 octobre 2026
Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is n…
- MMMG: a Comprehensive and Reliable Benchmark for Multitask Multimodal Generation
Jihan Yao, Yushi Hu, Wenyuan Wang, Bin Han, Shangbin Feng, Guang Yang, Yujie Yi, Bingbing Wen, Ranjay Krishna, Lucy Lu Wang, Yulia Tsvetkov, Noah A. Smith, Banghua Zhu · 2 octobre 2026
Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation, especially for complex tasks that involve multiple modalities. We present MMMG, the first benchmark to bring the verifiable-task paradigm to multimodal…
- Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris · 2 octobre 2026
On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains larg…
- Task-Adaptive Grounded 3D-Programmers Using 2D VLMs
Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva, Jan-Nico Zaech, Luc Van Gool, Danda Pani Paudel · 2 octobre 2026
Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerf…
Autres sujets du thème Vision par ordinateur et reconnaissance de formes
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Generative Adversarial Networks and Image Synthesis4 992 papiers / 12 mois+39 %
- Advanced Neural Network Applications2 354 papiers / 12 mois+48 %
- Advanced Vision and Imaging841 papiers / 12 mois+78 %
- Human Pose and Action Recognition836 papiers / 12 mois+457 %
- Face recognition and analysis482 papiers / 12 mois+88 %
- Image Enhancement Techniques463 papiers / 12 mois+20 %
