Emergence logoEmergence
Weekly

AI research - week of 31 August 2026

AI systems acting as autonomous assistants - agents - are shifting in status: the question is no longer just what they can do, but how to prevent them from going off the rails.

The term agent harnesses (the software frameworks that constrain these agents) appears in 33 papers over four weeks, up from 9 four weeks earlier. The titles show that the problem is no longer theoretical: attackers can hijack these frameworks to execute malicious behaviors, and agents retain these biases even when transferred to new servers or when their underlying model is changed.

Three papers illustrate this concern:

The topic goes beyond mere cybersecurity: the same papers discuss language-model agents (28 papers over four weeks), agents that replicate human biases like discrimination between groups or fail to verify available evidence before making a decision. The term failed trajectories (sequences of actions leading to failure) appears in 20 papers, often describing methods that supervise key steps of an agent rather than its final outcome.

The digest by email

Get the digest

One email on Monday morning: what moved in AI research last week. No spam, one-click unsubscribe.

Browse the topics →