Physical Sciences › Computer Science › Artificial Intelligence
AI-based Problem Solving and Planning
343 papers indexed
Problem-solving and planning in artificial intelligence explore how autonomous systems can develop strategies to achieve goals in complex environments. This research particularly examines the use of world models - internal representations enabling agents to anticipate the consequences of their actions - as well as methods for evaluating and optimizing their decisions over long temporal horizons. Approaches often combine architectures such as transformers or graphs with learning and formal verification techniques to enhance the robustness and adaptability of agents in diverse tasks, ranging from autonomous navigation to academic or operational path planning.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States44% · 93 papers
- China35% · 74 papers
- United Kingdom8.5% · 18 papers
- Germany8% · 17 papers
- Canada7.5% · 16 papers
- Italy4.7% · 10 papers
- France3.3% · 7 papers
- Hong Kong SAR China3.3% · 7 papers
Across 213 papers on this subject with at least one lab located. 40 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- DeepJEPA: Scaling World Models from Within
Zijian Jin, Yunbei Zhang, Yuanzhe Liu, Ming Liu, Baian Chen, Weirui Ye, Shilong Liu, Marco Pavone · 2 October 2026
World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition. We show that making every transition uniformly deeper wastes computation and can degrade planning because useful refin…
- Probabilistic Plan Legibility with Off-the-shelf Planners
Michele Persiani, Thomas Hellstr\"om · 2 October 2026
Legible planning is the creation of plans that best disambiguate their goals from a set of other candidates from an observer's perspective. In this paper we propose a method for legible planning for arbitrary PDDL domains, by extending previous research on legibility to classical planning without re…
- Consistent Plan-Act for Long-Horizon Agentic Tasks
Heng-Zhuang Li, Yi-Kai Zhang, Yu Wang, Yueqing Sun, Jiayuan Zhang, Qi Gu, Han-Jia Ye · 1 October 2026
Long-horizon agentic tasks demand strong reasoning and efficient execution across successive interactions with dynamic environments. A common approach decouples high-level planning from low-level execution through separate planner and actor roles. To investigate coordination failures in these tasks,…
- Risk-Aware Semantic Grounding for Trustworthy LLM-Based Robot Planning
{\L}ukasz Sobczak, Nur Kele\c{s}o\u{g}lu, S{\l}awomir Piotr Nowak · 30 September 2026
Large language models (LLMs) are increasingly used as high-level planners in robot navigation, but their outputs may become unreliable when instructions are ambiguous, unsupported by the environment, or semantically inconsistent. This paper presents a Risk-Aware Semantic Grounding framework for trus…
- ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning
Ke Fang, Yupu Yao, Lu Cheng · 30 September 2026
Latent world models rely on representation geometry for planning, yet regularizing the latent marginal alone does not determine the state-to-state relationships used for action selection. We show that this can cause planning-relevant novelty structure to be weakened as representations are transforme…
- Lookahead-R: Budget-Aware Tool Retrieval via Execution-Centric Planning
Zongze Wu, Yani Guo, Runnan Li · 30 September 2026
Tool retrieval is a critical bottleneck for LLM-based agents operating over large, heterogeneous API ecosystems. Existing approaches face an inherent trade-off: semantic retrievers are fast but suffer from the semantic-functional gap, while execution-based validation improves precision at the cost o…
- Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution
Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera, Marcos L\'opez de Prado, Shadab Khan · 30 September 2026
Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor sys…
- Beyond a single latent space: a dual-latent world model for long-horizon planning
Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang · 30 September 2026
Latent world models often struggle with long-horizon planning despite accurate short-term predictions. Recursive rollouts accumulate errors, while distance concentration in high-dimensional latent spaces can weaken goal discrimination. We introduce the Dual-Latent World Model (Dual-WM), which separa…
- Commitment Hierarchies under Intent Revision: A Belief-Revision Account of Salvage in Tool-Use Agents
Spandan Ghose Chowdhury · 30 September 2026
When a user changes their mind partway through a task, an agent that has already split the task into sub-goals and paid for tool calls must decide, per cached sub-result, whether to keep, patch, or discard it (salvage), restarting wastes valid work and continuing unchanged answers the old question. …
- Distinguish or Homogenize: Last-Chance Policy Identification and Risk-Budgeted Recovery under Irreversible Resource Depletion
Yibo Guo, Xiaodan Wang · 30 September 2026
Under irreversible resource depletion, an agent can spend resources to distinguish among latent fault models, or to change the system state so that the remaining models admit a common acceptable continuation--at which point further diagnosis becomes unnecessary. This distinguish-or-homogenize princi…
- Bayesian Active Learning for Intent Disambiguation in Interactive Robot Planning
Huao Li, Carson Sobolewski, Augustinos Saravanos, William Tan, John Karigiannis, Chuchu Fan · 29 September 2026
Interactive robot planning requires robots to infer and execute human intentions from natural language instructions that are often ambiguous, incomplete, or underspecified. Although large language models (LLMs) provide a powerful interface for clarification, relying on the generative model to drive …
- SAGE: Symbolic Action-Gating and Editing for LLM Task Planners
Trung Minh Bui, JongSul Moon, YoungOuk Kim, Quang-Ngoc Phung, Se-Woong Jun, Dongin Shin · 29 September 2026
Large language models (LLMs) are now the default cognitive core of embodied household agents, yet the plans they emit are rarely checked against a grounded model of the environment before execution, and the task-success they report is often measured on benchmarks so saturated that no method can be s…
- Representation Alignment as a Bottleneck in LLM-Based Retrosynthesis Planning
Hyunwoo Yoo, Cassie Huang, Haebin Shin, Li Zhang, Gail L. Rosen · 29 September 2026
While LLMs show promise in general reasoning, symbolic planning in chemistry remains a bottleneck. Direct ''SMILES-to-PDDL'' attempts fail because they force models to juggle chemical analysis and planning-language structuring simultaneously. We hypothesize that this failure stems from a lack of int…
- FONDANT: Strong and Best-Effort Planning via Antichains
Benjamin Aminof, Tuan Khai Nguyen, Sasha Rubin · 29 September 2026
A classical solution concept in fully observable nondeterministic (FOND) planning, is the strong policy (aka winning strategy in the closely related area of reactive synthesis), i.e., such a policy ensures that the goal is reached in an adversarial environment. When strong policies are not available…
- PDEU-Bench: Benchmarking the Personalized Planning Lifecycle of Tool-Calling LLM Agents
Huayi Lai, Shichao Song, Qingchen Yu, Simin Niu, Mengwei Wang, Hanyu Wang, Xun Liang · 29 September 2026
Large language model (LLM) agents are evolving from tool-calling systems that execute isolated instructions into task-oriented agents that pursue user goals through sustained, multi-step interactions. However, existing benchmarks for personalized tool use largely assess isolated calls or reactive ex…
- SGG-ReflAct: Sub-Goal Guided ReflAct with Structured Planning for Reliable Long-Horizon Reasoning
Jaeho Jung, Sung Hoon Jung · 29 September 2026
Recent advances in reasoning backbones have empowered large language model (LLM)agentstotackle complex, multi-step tasks. However, as reasoning horizons grow, inconsistent internal beliefs induce intermediate errors that cause agents to drift from their goals. This limitation also persists in REFLAC…
- HyperMCTS: Hypergraph-Augmented MCTS for Long-Horizon LLM Agents
Tingsong Xiao, Nithish Balachandar Moudhgalya, Chandrayee Basu, Lichao Wang, Luyang Kong, Benjamin Z. Yao, Zhe Jiang, Jie Hao · 29 September 2026
Long-horizon tasks require large language model (LLM) agents to coordinate decisions under constraints that span an entire solution. Monte Carlo Tree Search (MCTS) offers a promising approach to test-time scaling by exploring alternative action trajectories, but model computation and environment int…
- HTN Planning as a Coordination Layer for Multi-Server MCP Tool Orchestration
Eliott Jacopin, \'Eric Jacopin, Koichi Takahashi · 29 September 2026
The Model Context Protocol (MCP) isolates servers by design: only the host can orchestrate cross-server workflows. When the host is a large language model, the resulting orchestrations are non-deterministic, non-reproducible, and pay one inference round-trip per tool call. We present a coordination …
- ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments
Shengbin Yue, Hongru Wang, Siyuan Wang, Xiaoxin Chen, Wei Chen, Zhongyu Wei · 29 September 2026
Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them withou…
- Planner-as-Router: Joint Plan-Time Model Routing for Cost-Efficient Multi-Agent Workflows
Vivek Kumar Singh, Preeti Priyam, Gautam Bhowmick · 29 September 2026
Running large language model (LLM) agents in production gets expensive fast. A frontier model (the largest, most capable tier) is accurate but can cost 25 times what a small model costs per token, and the gap compounds once a workflow chains several calls together. Planner-as-Router (PaR) attacks th…
- PlanGuard: A Guardrail for Multi-Step Plan Safety in Embodied Agents
Junchi Chen, Changtao Miao, Yuxiao Xiang, Zhenchao Jin, Haojie Yuan, Qi Chu, Tao Gong, He Liu, Bo Zhang, Jiansheng Cai, Zhe Li, Nenghai Yu · 29 September 2026
Embodied task planners may produce multi-step plans whose subtask dependencies and interactions with the environment create physical risks during execution. Yet existing safeguards overlook such compositional risks, as general-purpose guardrails focus on semantic harm and embodied safety detectors a…
- Metro-WM: Long-Horizon Latent Planning with Realisable Sub-Goals
Royson Lee, Fady Rezk, Titouan Parcollet, Timothy Hospedales, Cristina Cornelio · 29 September 2026
Model-predictive control with Joint-Embedding Predictive Architectures (JEPAs) provides a strong zero-shot goal-reaching planner, but it is only effective over short planning horizons. Hierarchical extensions attempt to bridge this gap by learning a macro planner to predict intermediate latent sub-g…
- TemplateCraft: Agentic Visual Template Generation
Hongjie Yu, Zhiyuan Fan, Yuzhe Zhang, Jiangcun Du, Zhicheng Gao, Yuhong Zhang, Xiaokai Zhan, Zongshi Xie · 28 September 2026
The growing popularity of short videos has driven demand for one-click content creation. Visual templates turn uploaded images into personalized content with preset effects, but reusable template generation still requires substantial manual effort in asset preparation and tool orchestration. We prop…
- PTC-Decoder: Towards Intelligent SLMs on Offline Resource-Constrained Edge Devices
Minghui Yu, Ke Mu, Gang Wu · 28 September 2026
Deploying small language models (SLMs) on offline, resource-constrained edge devices such as remote sensing satellites presents a fundamental challenge: their limited reasoning capacity hinders reliable execution of multi-step agent tasks requiring complex tool orchestration. Existing plan-solve par…
- Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think
Xvyuan Liu, Jianjie Fang, Chen Gao, Yong Li · 25 September 2026
Planners built on visual world models commonly score each predicted outcome by its distance to the encoded goal image. We show that this target can limit control even with exact dynamics and globally optimal short-horizon search: reaching a goal may require actions that initially move away from it. …
Other topics in Artificial intelligence
The topics the OpenAlex classification attaches to the same theme, most active first.
- Large Language Models7,407 papers / 12 months+247%
- Adversarial Robustness in Machine Learning3,552 papers / 12 months+118%
- Reinforcement Learning in Robotics2,519 papers / 12 months+117%
- Explainable Artificial Intelligence (XAI)2,319 papers / 12 months+200%
- Domain Adaptation and Few-Shot Learning2,059 papers / 12 months+67%
- Advanced Graph Neural Networks1,926 papers / 12 months+38%
