Emergence logoEmergence

Physical Sciences › Computer Science › Computer Vision and Pattern Recognition

Multimodal Machine Learning Applications

8,069 papers indexed

The research grouped under this theme explores how artificial intelligence models simultaneously combine and interpret multiple types of data, such as images, text, or scientific diagrams. It analyzes, for example, biases related to colors in Vision Language Models, assesses their ability to reason about complex figures, or to assemble novel structures under semantic constraints. Other studies focus on improving their active perception, their capacity to express epistemic nuances, or their behavior when faced with missing or misleading visual information.

This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.

Monthly volume - last 12 months

Lab countries

  1. China52% · 2,985 papers
  2. United States36% · 2,061 papers
  3. United Kingdom6.1% · 349 papers
  4. South Korea5.7% · 326 papers
  5. Hong Kong SAR China5.6% · 316 papers
  6. Singapore4.7% · 267 papers
  7. Germany4.6% · 262 papers
  8. India4.3% · 245 papers

Across 5,691 papers on this subject with at least one lab located. 98 countries represented.

This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.

Latest papers

Other topics in Computer vision and pattern recognition

The topics the OpenAlex classification attaches to the same theme, most active first.

All of Vision →