Emergence logoEmergence

Emergence - arXiv AI observatory

Digest archives

The digest by email

Get the digest

One email on Monday morning: what moved in AI research last week. No spam, one-click unsubscribe.

AllWeekMonthQuarter
Weekly28 September 2026

AI models that understand and generate human speech are gaining in accuracy, but also in methods to evaluate this accuracy. This week, Speech Recognition and Synthesis dominate with 112 papers published over four weeks, compared to 63 four weeks earlier.

The emergence of the term operating characteristic curve (a tool that measures a model's ability to distinguish correct answers from incorrect ones) in 18 papers shows that researchers are no longer just counting errors. They are also analyzing how these errors are distributed across contexts - for example, a model that recognizes male voices better than female ones, or that stumbles on regional accents.

Three recent papers illustrate this trend:

In the same domain, two other recent papers focus on the robustness of models when faced with languages or situations underrepresented in training data:

Link to this issue →

Weekly21 September 2026

Systems that help find experts or answer technical questions are experiencing a resurgence in activity.

Over four weeks, 62 papers focused on expert finding and question-answering systems (Expert finding and Q&A systems), compared to 25 four weeks earlier. This surge is particularly evident in the evaluation of models and the distribution of tasks between humans and machines.

Three recent papers illustrate this trend:

The term training seeds (19 papers over four weeks, compared to 7 previously) appears in work testing the robustness of models against biased or incomplete initial data. It refers to the starting examples used to train a system before it generates its own data.

Link to this issue →

Weekly14 September 2026

Autonomous agents capable of self-improvement without human intervention are becoming a central topic in AI.

Over four weeks, 16 papers used the term recursive self-improvement, compared to 6 four weeks earlier. This technique allows a system to analyze its own errors and modify its operation to progress, without relying on external data or a human operator. Recent work explores formal frameworks to structure this process and concrete applications, such as agents capable of autonomously exploring new environments.

Here are three representative papers on this trend:

Other rapidly growing terms this week, such as language-model agents (29 papers vs. 12) or agent harnesses (23 papers vs. 11), show that research is focusing on the software infrastructures that enable these agents to operate reliably and scalably.

Link to this issue →

Weekly7 September 2026

AI systems capable of self-improvement are gaining attention.

The term recursive self-improvement (where a model refines its own performance without human intervention) appears in 17 papers over four weeks, compared to 6 four weeks earlier. This theme is emerging primarily in work on autonomous agents - programs that chain complex tasks without supervision.

Three recent papers illustrate this trend:

Research is also focusing on agent harnesses (evaluation frameworks for agents, which define their objectives and constraints): 28 papers in four weeks, compared to 10 previously. These frameworks aim to make agents more reliable on long-horizon tasks, such as project management or decision-making over extended periods. Two examples:

Link to this issue →

Weekly31 August 2026

AI systems acting as autonomous assistants - agents - are shifting in status: the question is no longer just what they can do, but how to prevent them from going off the rails.

The term agent harnesses (the software frameworks that constrain these agents) appears in 33 papers over four weeks, up from 9 four weeks earlier. The titles show that the problem is no longer theoretical: attackers can hijack these frameworks to execute malicious behaviors, and agents retain these biases even when transferred to new servers or when their underlying model is changed.

Three papers illustrate this concern:

The topic goes beyond mere cybersecurity: the same papers discuss language-model agents (28 papers over four weeks), agents that replicate human biases like discrimination between groups or fail to verify available evidence before making a decision. The term failed trajectories (sequences of actions leading to failure) appears in 20 papers, often describing methods that supervise key steps of an agent rather than its final outcome.

Link to this issue →

Weekly24 August 2026

AI models that reason by following explicit rules rather than guessing answers are experiencing a resurgence of interest.

The topic Semantic Web and Ontologies has gone from 10 to 60 articles over four weeks, marking the most significant increase of the period. These works explore how to structure knowledge so that systems can manipulate it logically, as a human would with clear definitions and links. Three expressions are emerging in this context:

  • « points above » (17 articles) refers to measured gaps between a prediction and a reference, often used to assess the robustness of reasoning;
  • « evidence supports » (22 articles) appears in titles that verify whether the steps of a reasoning process are backed by verifiable facts;
  • « local evidence » (17 articles) refers to clues extracted from a limited portion of the data, rather than a global analysis.

Among the representative articles:

These methods target fields where errors are costly - medical diagnostics, software verification, regulatory document analysis - by replacing the imprecision of statistical models with traceable rule-based sequences.

Link to this issue →

Weekly17 August 2026

AI models are now learning to improve without human intervention, by analyzing their own mistakes rather than relying on examples provided by experts. This approach, called on-policy self-distillation (a technique where the model refines its responses by training on its own trials), is increasingly central to research this week.

Over four weeks, 37 papers mention this term, compared to 13 four weeks earlier. A closely related variant, on-policy distillation (where multiple models collaborate to improve), appears in 86 papers, marking an increase of 53 papers from the previous period. These methods aim to reduce the deployment cost of AI systems by limiting the need for human-labeled data, an issue also reflected in titles: 23 papers explicitly mention deployment cost, up from 7 previously.

Recent work explores how to avoid biases introduced by privileged information (hidden data or rules that distort learning) - 29 papers address this, 19 more than a month ago. Several studies also test controlled ablations (targeted removals of components to measure their impact), a practice found in 21 papers.

A few recent examples:

Link to this issue →

Weekly10 August 2026

This week, the focus is on an optimization technique for search agents: on-policy self-distillation (OPSD). 18 papers discuss it, up from 6 four weeks ago, and the term appears in three of the most dynamic topics at the moment - Information Retrieval and Search Behavior (88 papers, +60 in one month), Multi-Agent Systems and Negotiation (104 papers, +45), and language-model agents (18 papers, +11).

The titles show an exploration of the method's limits:

Two variants are emerging: the use of privileged information (23 papers, +13) to guide distillation, and token-level supervision (20 papers, +11) to refine the learning signal. These approaches aim to stabilize agent training in environments where rewards are sparse or noisy, such as information retrieval or negotiation between agents.

Link to this issue →

Weekly3 August 2026

The week of 3 August holds 2,008 ingested papers, against 1,385 the week before. Three neighbouring topics rise together: information retrieval and search behaviour (73 papers over four weeks against 26 over the four before), recommender systems (59 against 27) and personal information management (42 against 26).

Behind those three labels lies one movement: search and recommendation systems are being rewritten around agents built on language models. This week's titles bear less on what those agents promise than on how they are assessed. What their representations are worth, what happens when their own history misleads them, how to trace an answer back to its source.

Worth reading: PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories Universal Pathologies, Conditional Consequences: A Triple-Robustness Analysis of RAG for Multi-Hop Traceability A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing

Link to this issue →

Weekly27 July 2026

This week, AI research refocuses its attention on information retrieval and user behavior: 40 papers published over the last four weeks, compared to 30 in the previous period. The topic now surpasses brain-computer interfaces and recommendation systems, which are also on the rise.

Three trends emerge from the titles:

  • Optimizing commercial results, with models that generate derived search intents to improve product discoverability.
  • Semantic search in open-source code bases, where neural embeddings replace traditional textual queries.
  • Guiding retrieval-augmented generation (RAG) systems with explicit semantic constraints to prevent generated response drift.

To read this week: Improving Item Discoverability in e-Commerce Search via Related Intent Generation MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software Entities GuidedRAG: Semantic Steering of Retrieval-Augmented Generation

Link to this issue →

Weekly20 July 2026

This week, only one topic is advancing in AI research: algorithms and architectures for quantum computing. 74 papers were published over the last four weeks, compared to 60 in the previous period.

Three recent works illustrate this trend:

The other monitored domains are declining or stagnating, with no notable movements to report.

Link to this issue →

Weekly13 July 2026

This week, quantum algorithmics confirms its renewed interest: 70 papers published over the last four weeks, compared to 62 in the previous period. The topic remains the only one growing among the tracked themes, with a modest but continuous increase over the past three weeks.

Recent work explores concrete applications, particularly post-quantum cryptography and multimodal data augmentation. Three notable examples:

Link to this issue →

Weekly6 July 2026

This week, the Scientific Computing and Data Management topic confirms its momentum with 102 articles published over four weeks, compared to 91 in the previous period. The growth (+11 articles) is driven by work at the intersection of large language models and scientific data management.

Three themes emerge in the titles:

  • The integration of LLMs into data science workflows, with evaluations of their automatically generated skills.
  • Data traceability, particularly through watermarking methods for LLM-based agent trajectories.
  • Error analysis in text-to-SQL queries, a key challenge for database automation.

A few representative articles:

Link to this issue →

Weekly29 June 2026

This week, speech recognition and synthesis confirm their momentum: 92 papers published over the last four weeks, compared to 58 in the previous period. Recent work explores hybrid architectures and applications in noisy or multilingual contexts.

Three trends stand out in the titles:

Music processing and vocal emotion analysis follow a parallel trajectory, but with volumes twice as low (44 and 48 papers over 30 days).

Link to this issue →

Weekly22 June 2026

This week, sound and speech processing dominate the trends. Two topics show similar growth: 51 papers on music and audio processing (compared to 23 the previous month) and 97 on speech recognition and synthesis (compared to 44). Practical applications are multiplying, particularly for music analysis and voice synthesis.

A few recent works:

Link to this issue →

Weekly15 June 2026

This week, speech synthesis and recognition clearly dominate the trends. The topic Speech Recognition and Synthesis has risen from 36 to 91 articles in four weeks, marking the strongest growth at the moment.

Three key themes emerge from recent publications:

  • Fine-grained evaluation of speech quality, with models specialized in accent or prosody errors.
  • Adapting models to atypical voices, such as dysarthric speech.
  • Repurposing speech classifiers to generate speech via guided diffusion.

A few representative papers: PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation

Link to this issue →