2D projections of the embeddings
The semantic map of arXiv AI papers
Each dot is a paper. Two papers close together on the map have similar content. Hover to see the title, click to open it on arXiv, and switch projection to see what the reduction method changes about the drawing.
How it is built
A sample of about 2,000 papers is drawn proportionally to the volumes per topic, weighted towards the last twelve months. Each paper is turned into a vector of 1,024 numbers by an embedding model, which places texts with similar meaning close together.
Those vectors live in a 1,024-dimensional space, impossible to look at. A UMAP projection brings them down to two dimensions while trying to preserve neighbourhoods. The four other tabs are different methods applied to the same sample: they do not show other data, they show other trade-offs.
Distances are indicative. Two nearby dots deal with related subjects, but the scale has no unit and absolute position means nothing.
