Emergence logoEmergence

Emergence - arXiv AI observatory

Digest archives

The digest by email

Get the digest

One email on Monday morning: what moved in AI research last week. No spam, one-click unsubscribe.

AllWeekMonthQuarter
Weekly28 September 2026

AI models that understand and generate human speech are gaining in accuracy, but also in methods to evaluate this accuracy. This week, Speech Recognition and Synthesis dominate with 112 papers published over four weeks, compared to 63 four weeks earlier.

The emergence of the term operating characteristic curve (a tool that measures a model's ability to distinguish correct answers from incorrect ones) in 18 papers shows that researchers are no longer just counting errors. They are also analyzing how these errors are distributed across contexts - for example, a model that recognizes male voices better than female ones, or that stumbles on regional accents.

Three recent papers illustrate this trend:

In the same domain, two other recent papers focus on the robustness of models when faced with languages or situations underrepresented in training data:

Link to this issue →

Weekly21 September 2026

Systems that help find experts or answer technical questions are experiencing a resurgence in activity.

Over four weeks, 62 papers focused on expert finding and question-answering systems (Expert finding and Q&A systems), compared to 25 four weeks earlier. This surge is particularly evident in the evaluation of models and the distribution of tasks between humans and machines.

Three recent papers illustrate this trend:

The term training seeds (19 papers over four weeks, compared to 7 previously) appears in work testing the robustness of models against biased or incomplete initial data. It refers to the starting examples used to train a system before it generates its own data.

Link to this issue →

Weekly14 September 2026

Autonomous agents capable of self-improvement without human intervention are becoming a central topic in AI.

Over four weeks, 16 papers used the term recursive self-improvement, compared to 6 four weeks earlier. This technique allows a system to analyze its own errors and modify its operation to progress, without relying on external data or a human operator. Recent work explores formal frameworks to structure this process and concrete applications, such as agents capable of autonomously exploring new environments.

Here are three representative papers on this trend:

Other rapidly growing terms this week, such as language-model agents (29 papers vs. 12) or agent harnesses (23 papers vs. 11), show that research is focusing on the software infrastructures that enable these agents to operate reliably and scalably.

Link to this issue →

Weekly7 September 2026

AI systems capable of self-improvement are gaining attention.

The term recursive self-improvement (where a model refines its own performance without human intervention) appears in 17 papers over four weeks, compared to 6 four weeks earlier. This theme is emerging primarily in work on autonomous agents - programs that chain complex tasks without supervision.

Three recent papers illustrate this trend:

Research is also focusing on agent harnesses (evaluation frameworks for agents, which define their objectives and constraints): 28 papers in four weeks, compared to 10 previously. These frameworks aim to make agents more reliable on long-horizon tasks, such as project management or decision-making over extended periods. Two examples:

Link to this issue →

Monthly1 September 2026

AI systems that improve without human intervention are losing ground, but one area holds steady.

Over four weeks, publications on multi-agent systems and negotiation dropped from 204 to 111 articles, a decline of nearly half. Yet, one subset of this work - those exploring recursive self-improvement - remains stable. The term appears in 33 papers this month, a volume close to the 31 recorded in August. This research focuses on agents capable of correcting their own errors and refining their performance without relying on external data or humans.

Three recent papers illustrate this persistence:

Meanwhile, agent harnesses (software frameworks that define agents' goals and constraints) account for 51 papers this month, up from 21 in August. These tools aim to structure complex, long-running tasks, such as project management or decision-making over time. Two examples:

Elsewhere, most major topics are in decline. Large language models fell from 587 to 271 articles, and multimodal applications (combining text, image, and sound) from 575 to 244. Only speech recognition and synthesis is advancing, with 112 papers compared to 78.

Link to this issue →

Weekly31 August 2026

AI systems acting as autonomous assistants - agents - are shifting in status: the question is no longer just what they can do, but how to prevent them from going off the rails.

The term agent harnesses (the software frameworks that constrain these agents) appears in 33 papers over four weeks, up from 9 four weeks earlier. The titles show that the problem is no longer theoretical: attackers can hijack these frameworks to execute malicious behaviors, and agents retain these biases even when transferred to new servers or when their underlying model is changed.

Three papers illustrate this concern:

The topic goes beyond mere cybersecurity: the same papers discuss language-model agents (28 papers over four weeks), agents that replicate human biases like discrimination between groups or fail to verify available evidence before making a decision. The term failed trajectories (sequences of actions leading to failure) appears in 20 papers, often describing methods that supervise key steps of an agent rather than its final outcome.

Link to this issue →

Weekly24 August 2026

AI models that reason by following explicit rules rather than guessing answers are experiencing a resurgence of interest.

The topic Semantic Web and Ontologies has gone from 10 to 60 articles over four weeks, marking the most significant increase of the period. These works explore how to structure knowledge so that systems can manipulate it logically, as a human would with clear definitions and links. Three expressions are emerging in this context:

  • « points above » (17 articles) refers to measured gaps between a prediction and a reference, often used to assess the robustness of reasoning;
  • « evidence supports » (22 articles) appears in titles that verify whether the steps of a reasoning process are backed by verifiable facts;
  • « local evidence » (17 articles) refers to clues extracted from a limited portion of the data, rather than a global analysis.

Among the representative articles:

These methods target fields where errors are costly - medical diagnostics, software verification, regulatory document analysis - by replacing the imprecision of statistical models with traceable rule-based sequences.

Link to this issue →

Weekly17 August 2026

AI models are now learning to improve without human intervention, by analyzing their own mistakes rather than relying on examples provided by experts. This approach, called on-policy self-distillation (a technique where the model refines its responses by training on its own trials), is increasingly central to research this week.

Over four weeks, 37 papers mention this term, compared to 13 four weeks earlier. A closely related variant, on-policy distillation (where multiple models collaborate to improve), appears in 86 papers, marking an increase of 53 papers from the previous period. These methods aim to reduce the deployment cost of AI systems by limiting the need for human-labeled data, an issue also reflected in titles: 23 papers explicitly mention deployment cost, up from 7 previously.

Recent work explores how to avoid biases introduced by privileged information (hidden data or rules that distort learning) - 29 papers address this, 19 more than a month ago. Several studies also test controlled ablations (targeted removals of components to measure their impact), a practice found in 21 papers.

A few recent examples:

Link to this issue →

Weekly10 August 2026

This week, the focus is on an optimization technique for search agents: on-policy self-distillation (OPSD). 18 papers discuss it, up from 6 four weeks ago, and the term appears in three of the most dynamic topics at the moment - Information Retrieval and Search Behavior (88 papers, +60 in one month), Multi-Agent Systems and Negotiation (104 papers, +45), and language-model agents (18 papers, +11).

The titles show an exploration of the method's limits:

Two variants are emerging: the use of privileged information (23 papers, +13) to guide distillation, and token-level supervision (20 papers, +11) to refine the learning signal. These approaches aim to stabilize agent training in environments where rewards are sparse or noisy, such as information retrieval or negotiation between agents.

Link to this issue →

Weekly3 August 2026

The week of 3 August holds 2,008 ingested papers, against 1,385 the week before. Three neighbouring topics rise together: information retrieval and search behaviour (73 papers over four weeks against 26 over the four before), recommender systems (59 against 27) and personal information management (42 against 26).

Behind those three labels lies one movement: search and recommendation systems are being rewritten around agents built on language models. This week's titles bear less on what those agents promise than on how they are assessed. What their representations are worth, what happens when their own history misleads them, how to trace an answer back to its source.

Worth reading: PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories Universal Pathologies, Conditional Consequences: A Triple-Robustness Analysis of RAG for Multi-Hop Traceability A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing

Link to this issue →

Monthly1 August 2026

AI systems that learn to improve without human data now occupy a central place in research. This month, two topics related to this approach are experiencing unprecedented growth: multi-agent systems and negotiation (172 papers over four weeks, compared to 53 the previous month) and information retrieval (149 papers, compared to 47). These two fields share a common technique, on-policy self-distillation (OPSD), where models refine their responses by training on their own trials rather than on examples provided by humans.

OPSD is not new, but its massive adoption in concrete applications marks a turning point. Recent work no longer merely touts its promises: it tests its limits. Three questions recur in this month’s titles:

  • How to prevent the model from misleading itself with privileged information (hidden data that biases its learning)? 29 papers address this, 19 more than in July.
  • How to stabilize training when rewards are scarce or noisy? 23 papers explore the use of privileged information to guide distillation, and 20 others focus on token-level supervision (a learning signal refined at the level of each word or symbol).
  • How to trace a response back to its source to verify its robustness? The phrases « evidence supports » (22 papers) and « local evidence » (17 papers) appear in work assessing the reliability of step-by-step reasoning.

A few recent examples:

This trend is accompanied by a marked decline in traditional approaches. Generative adversarial networks (268 papers, -23% in one month) and advanced neural networks (135 papers, -24%) are losing ground, as is explainability (107 papers, -38%). Research now seems to favor systems capable of improving on their own, even at the cost of some transparency.

Link to this issue →

Weekly27 July 2026

This week, AI research refocuses its attention on information retrieval and user behavior: 40 papers published over the last four weeks, compared to 30 in the previous period. The topic now surpasses brain-computer interfaces and recommendation systems, which are also on the rise.

Three trends emerge from the titles:

  • Optimizing commercial results, with models that generate derived search intents to improve product discoverability.
  • Semantic search in open-source code bases, where neural embeddings replace traditional textual queries.
  • Guiding retrieval-augmented generation (RAG) systems with explicit semantic constraints to prevent generated response drift.

To read this week: Improving Item Discoverability in e-Commerce Search via Related Intent Generation MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software Entities GuidedRAG: Semantic Steering of Retrieval-Augmented Generation

Link to this issue →

Weekly20 July 2026

This week, only one topic is advancing in AI research: algorithms and architectures for quantum computing. 74 papers were published over the last four weeks, compared to 60 in the previous period.

Three recent works illustrate this trend:

The other monitored domains are declining or stagnating, with no notable movements to report.

Link to this issue →

Weekly13 July 2026

This week, quantum algorithmics confirms its renewed interest: 70 papers published over the last four weeks, compared to 62 in the previous period. The topic remains the only one growing among the tracked themes, with a modest but continuous increase over the past three weeks.

Recent work explores concrete applications, particularly post-quantum cryptography and multimodal data augmentation. Three notable examples:

Link to this issue →

Weekly6 July 2026

This week, the Scientific Computing and Data Management topic confirms its momentum with 102 articles published over four weeks, compared to 91 in the previous period. The growth (+11 articles) is driven by work at the intersection of large language models and scientific data management.

Three themes emerge in the titles:

  • The integration of LLMs into data science workflows, with evaluations of their automatically generated skills.
  • Data traceability, particularly through watermarking methods for LLM-based agent trajectories.
  • Error analysis in text-to-SQL queries, a key challenge for database automation.

A few representative articles:

Link to this issue →

Monthly1 July 2026

July's corpus holds 6,479 papers, against 12,021 in June. The decline touches all twenty tracked topics without exception: it says something about collection, not about the state of research. What remains readable is the order in which topics fall back.

Only one comes through the month nearly intact: quantum computing algorithms and architecture, 69 papers in July against 79 in June, while the whole corpus loses close to half its volume. It is also the only theme the weekly issues saw growing, two weeks in a row.

At the other end, the largest topics are the ones giving up the most ground: large language models (344 papers against 773), multimodal machine learning (311 against 684), adversarial robustness (211 against 464) and image synthesis (171 against 378). All fall back further than the corpus average.

Three quantum papers noted over the month: Cautious optimism for deep parameterized quantum circuits Towards quantum machine learning for assessing the resilience of post-quantum cryptography Approximate Quantum State Preparation Through Proximal Policy Optimization

Link to this issue →

Quarterly1 July 2026

The past quarter marks a structural shift in AI research. Historical topics are collapsing: Large Language Models dropped from 2,054 to 544 papers, Multimodal Machine Learning Applications from 1,610 to 433, and Adversarial Robustness from 1,252 to 277. This widespread contraction - with all topics above the 120-paper threshold declining by 38% to 78% - is not a cyclical adjustment but an exhaustion of the dominant paradigms of the past five years.

Three areas resist this trend, though they do not offset the losses. Scientific Computing and Data Management (175 papers, -39%) and Multi-Agent Systems (133 papers, -48%) maintain significant activity, while Software Engineering Research (120 papers, -57%) stabilizes at the floor. The titles of representative works reveal a common reorientation: AI is no longer studied as an autonomous object but as a component integrated into larger systems.

Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models ATLAS: Discovering Agent Strategies through LLM-Guided Abstraction and Automata Learning Agentic Transaction: Towards ACID-Compliant Agent Systems Building AI-Intensive Software with AI: Early Results and a Cautionary Tale on Measuring Development Cost

Papers no longer focus on improving models but on integrating them into software architectures, transactional protocols, or scientific production pipelines. The vocabulary in the titles - reliable, operating the system around, ACID-compliant, development cost - signals a shift in priorities: the challenge is no longer raw performance but operational reliability in constrained environments. This transition is accompanied by a refocusing on concrete use cases, as suggested by the absence of theoretical work in the representative samples.

The quarter’s overall volume (10,745 papers) confirms that the decline is not a statistical artifact: the field is producing less but differently. The resilient topics are those that articulate AI with critical infrastructures - scientific computing, multi-agent systems, software engineering - rather than treating it as an end in itself.

Link to this issue →

Weekly29 June 2026

This week, speech recognition and synthesis confirm their momentum: 92 papers published over the last four weeks, compared to 58 in the previous period. Recent work explores hybrid architectures and applications in noisy or multilingual contexts.

Three trends stand out in the titles:

Music processing and vocal emotion analysis follow a parallel trajectory, but with volumes twice as low (44 and 48 papers over 30 days).

Link to this issue →

Weekly22 June 2026

This week, sound and speech processing dominate the trends. Two topics show similar growth: 51 papers on music and audio processing (compared to 23 the previous month) and 97 on speech recognition and synthesis (compared to 44). Practical applications are multiplying, particularly for music analysis and voice synthesis.

A few recent works:

Link to this issue →

Weekly15 June 2026

This week, speech synthesis and recognition clearly dominate the trends. The topic Speech Recognition and Synthesis has risen from 36 to 91 articles in four weeks, marking the strongest growth at the moment.

Three key themes emerge from recent publications:

  • Fine-grained evaluation of speech quality, with models specialized in accent or prosody errors.
  • Adapting models to atypical voices, such as dysarthric speech.
  • Repurposing speech classifiers to generate speech via guided diffusion.

A few representative papers: PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation

Link to this issue →

Monthly1 June 2026

Ethics and Multimodal Applications on the Rise

June confirms two major trends in AI research, with volumes significantly exceeding seasonal variations.

Ethics and social impacts: 259 papers published this month, compared to 170 in May (+52%). Recent work addresses governance, bias, and fairness issues, often in connection with concrete applications. Three notable examples:

Multimodal applications: 684 papers, compared to 527 in May (+30%). Models integrating text, image, audio, and video are gaining maturity, with architectures optimized for specific tasks. Among recent publications:

Conversely, topics such as Generative Adversarial Networks (378 papers, -7%) or Graph Neural Networks (274, -8%) are losing ground, though the data does not allow for identifying a single factor.

Link to this issue →

Monthly1 May 2026

Generative Synthesis Emerges as a Research Field

In May 2026, the volume of publications on Generative Adversarial Networks and Image Synthesis surged to 408 articles, compared to 171 the previous month (+138%). This doubling in a single month places the topic at the forefront of growth among the tracked themes, ahead of advances in Stochastic Gradient Optimization Techniques (288 articles, +112%) or Reinforcement Learning in Robotics (442 articles, +87%).

The rise of GANs is not limited to quantity: recent titles explore hybrid architectures and novel applications. Three articles illustrate this diversification:

Work on stochastic optimization (288 articles) and graph neural networks (297 articles, +88%) appears to fuel this momentum, suggesting a convergence between training methods and content generation. Notably, applications in medical imaging (118 articles in Machine Learning in Healthcare, +44%) directly benefit from these advances.

Link to this issue →

Quarterly1 April 2026

Quarterly Review: How AI Transformed Between April and June 2024

Key Takeaways in 3 Points

  • Generative AI enters classrooms: Research on its use in schools surged from 12 to 68 articles in three months, with tools to personalize learning.
  • Robots become more autonomous: Studies on machines capable of acting without human intervention (Autonomous Agents) doubled, rising from 45 to 92 articles.
  • AI steps into human shoes: Work on simulating human behavior (Human Behavior Simulation) exploded (+320%) to test social or medical scenarios.

AI in Schools: A Growing Presence, Not Without Debate

Researchers are exploring how AI tools, such as chatbots or exercise generators, can assist teachers. Some articles show how these technologies adapt lessons to each student’s pace by identifying their difficulties in real time. Others highlight the risks: screen dependency, biases in responses, or the loss of critical thinking. Studies are also multiplying on teacher training, ensuring they can use these tools without being replaced by them. Why does this matter in practice? Because it could reduce inequalities among students—provided these technologies are properly regulated to avoid widening gaps.


Robots That (Almost) Make Their Own Decisions

Machines capable of acting without step-by-step instructions (Autonomous Agents) are increasingly fascinating scientists. This quarter, articles on the topic nearly doubled. Some describe robots that plan their tasks, like a mechanical arm tidying a room without detailed guidance. Others explore systems that negotiate among themselves, such as autonomous cars coordinating to avoid traffic jams. Researchers are also testing limits: What happens if these agents make decisions contrary to human expectations? Why does this matter in practice? Because it could revolutionize factories, transportation, or even emergency response, making machines more responsive—but also more unpredictable.


AI Pretends to Be Human, to Better Understand Us

Simulating human behavior (Human Behavior Simulation) has become a hot topic: publications on the subject quadrupled. Some work uses AI to replicate realistic conversations, like a virtual patient responding to a trainee doctor’s questions. Others model crowds to anticipate panic movements during a concert or sports event. Researchers also use it to test public policies, such as the impact of a tax on consumption habits. Why does this matter in practice? Because it allows risk-free experimentation, whether for training professionals, designing safer cities, or predicting the effects of a reform.


Further Reading

  • Generative AI in Education: A Systematic Review → An overview of AI tools transforming schools, with their promises and pitfalls.
  • Autonomous Agents for Real-World Task Planning → How robots learn to organize their actions without human help, with concrete examples.
  • Simulating Human Behavior in Crisis Situations → A study on AI replicating crowd reactions in emergencies to better prepare responders.
  • Ethical Risks of Human-Like AI in Education → An article warning about the dangers of overly "human" chatbots in classrooms.

Next quarter, we’ll be watching whether debates on AI in schools lead to concrete recommendations and if autonomous robots move from labs to real-world testing.

Link to this issue →

Quarterly1 January 2026

Speech Recognition and Synthesis are becoming a priority focus for AI.

Over three months, 232 papers were published on the subject, compared to 83 in the previous quarter - a nearly threefold increase. Research is no longer limited to dominant languages: several teams are extending models to low-resource languages, such as Nepali, or adapting systems to the constraints of live streaming and vocal style variations.

Research in Information Retrieval and Search Behavior is following the same trajectory, with 139 publications compared to 50 previously. Both topics share a common concern: making models more accurate in real-world contexts, where data is noisy, incomplete, or biased.

A few recent examples:

This surge reflects a shift: after years of refining large language models, researchers are now turning their attention to the interfaces that connect them to the world - voice, queries, and interactions. The targeted applications are less about technical demonstrations and more about concrete tools, such as adaptive tutoring systems (165 papers, +170%) or multilingual voice assistants.

Link to this issue →