Emergence logoEmergence

Method

How these numbers are computed

Everything shown on this site comes from a computation, and every computation rests on choices. Here they are, in the order they apply.

The papers

The corpus comes from arXiv, restricted to four categories: cs.AI (artificial intelligence), cs.LG (machine learning), cs.CL (natural language processing) and cs.CV (computer vision). New announcements are collected every day and stored as they are: title, authors, English abstract, link. Nothing is rewritten.

arXiv publishes preprints. A paper listed here has not necessarily been peer reviewed, and its abstract announces a contribution without demonstrating it. All the more reason to open the paper before drawing conclusions.

Themes and topics

Neither the themes nor the topics are invented here. Each paper is matched against the OpenAlex classification, the open catalogue of world scientific research, which organises publications into four levels: domain, field, theme, topic. The match is made through the paper's DOI.

When OpenAlex does not know a paper yet, which is common in the first weeks, it is classified by similarity: its vector representation is compared to the centre of gravity of each topic, and the closest one is kept provided the similarity exceeds 0.30. Below that, the paper stays without a topic rather than being filed at random.

Volumes

Counts are precomputed once a day, per week and per topic. Monthly and quarterly views are sums of those weeks: the three scales cannot contradict each other.

The theme river switches automatically from weekly to monthly beyond six months of history. Over a long period, weekly granularity mostly shows noise.

Emerging signals

Phrases are groups of two or three consecutive words, extracted from the English title and abstract, lowercased and unaccented. For each phrase we count the number of distinct papers using it, not the number of times it appears: otherwise a single verbose paper would be enough to create a signal.

The current window covers the last 30 days, the baseline the six months before it, scaled to the same length. The growth shown is the gap between the two, with smoothing that stops a phrase absent from the baseline from mechanically taking first place.

Three filters remove filler. A phrase must appear in at least 16 papers over the window. Phrases containing a stop word, a modal verb or an auxiliary are rejected: “cannot express” is a sentence structure, not a research topic. Finally a hand-kept blocklist removes the protocol wording that still slips through. Nested phrases are deduplicated, and singular and plural count as one.

The semantic map

Each paper is turned into a vector of 1,024 numbers by an embedding model, which places texts with similar meaning close to one another. Those vectors live in a 1,024-dimensional space that cannot be looked at; a UMAP projection brings them down to two dimensions while trying to preserve neighbourhoods.

The map shows a sample of about 2,000 papers, drawn in proportion to the volume of each topic and biased towards the last twelve months. Distances on it are indicative: two nearby points cover related subjects, but the scale has no unit.

The digest

The digest is written by a language model, from the period's figures and a selection of papers, never from the whole corpus. Three rules govern the writing: one angle per issue, no percentage without the matching absolute volume, and three to five named papers with their links.

A topic only enters the digest if it clears an absolute volume floor over the window. Without that floor, a topic going from three to twenty-one papers would show +600% and take first place with nothing having happened.

The English, Spanish and German versions are translations of the French issue, made afterwards. Until a translation is written, the site shows the French text and says so.

What these numbers do not say

arXiv is not all of AI research: unpublished industrial work, conferences without preprints and non-English research are absent.

Publication volume measures activity, not importance. A topic that doubles may be a fashion as much as a breakthrough.

Emerging phrases are extracted from English text. An idea circulating under several names is split across several phrases and can stay invisible.