Social Sciences › Decision Sciences › Information Systems and Management
Scientific Computing and Data Management
987 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos49 % · 219 artículos
- China31 % · 140 artículos
- Reino Unido9,8 % · 44 artículos
- Alemania9,8 % · 44 artículos
- Canadá6,7 % · 30 artículos
- RAE de Hong Kong (China)4,5 % · 20 artículos
- Japón4 % · 18 artículos
- India3,8 % · 17 artículos
Sobre 447 artículos de este tema con al menos un laboratorio localizado. 58 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- ArrivalBench: Agent-Generated Data Pipelines Are Correct Once and Wrong Under Time
Pranay Kothari · 5 de octubre de 2026
Benchmarks for agent-generated data work grade a pipeline by running it once against a fixed snapshot. ArrivalBench instead re-executes the pipeline an agent leaves behind under adversarial but replayable delivery schedules (late, duplicated, out-of-order and retried records) and requires its final …
- From Behavior to Provenance: Attributing Tabular Foundation Models to Synthetic Pretraining Data
Mohamed Bouadi, Nassim Bouarour, Shivam Dubey, Aditya Tanna, Vinay Kumar Sankarapu · 5 de octubre de 2026
Training-data attribution aims to identify which training examples shape model behavior, yet validating such claims is difficult because causal training influence is rarely observable. We argue that controlled synthetic pretraining makes attribution experimentally testable. Using O'PRIOR, a provenan…
- Discriminating Fixture Coverage in Agent-Infrastructure Verification Suites
Xin Xu, Siru Tao · 5 de octubre de 2026
Invariant suites and runtime monitors increasingly gate agent deployment decisions, and the evidence offered for any particular suite is almost always a single observation: it passes an implementation believed correct and fails one believed broken. We measure what that observation is worth. Applying…
- Self-Supervised Scaling of Terminal Environments for Scientific Domains
Zhongzhi Li, Yucheng Shi, Zongxia Li, Junyao Yang, Ruhan Wang, Yu Wang, Jingyuan Huang, Jichao Yu, Ninghao Liu, Haitao Mi, Leowei Liang · 5 de octubre de 2026
Terminal agents are increasingly deployed beyond software engineering in science and other specialized domains. Constructing training environments requires executable reference behavior and a domain-specific verifier that distinguishes semantic correctness from superficially plausible artifacts. Aut…
- Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution Inheritance
Songtao Wei, Yi Li, Zhichun Guo, Bingzhe Li · 5 de octubre de 2026
Multi-agent systems (MAS) built from large language models coordinate specialized agents to tackle complex tasks, but effective workflows are difficult to design in advance. Test-time evolution refines workflows using execution feedback, yet broad revisions can disturb useful components, while re-ex…
- DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents
A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys, Peter W. J. Staar · 5 de octubre de 2026
Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We intr…
- HazardWeaver: Scientific Route Selection for Hazard Analysis Agents
Wangshu Zhu, Xueqi Cheng, Liang Wu, Yushun Dong · 5 de octubre de 2026
Understanding and assessing natural hazards is essential for disaster preparedness and risk reduction. Recent advances in large language models have spurred growing interest in AI agents for hazard analysis, particularly their ability to integrate scientific data, models, and tools into automated wo…
- Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
Zongxia Li, Yucheng Shi, Zhongzhi Li, Junyao Yang, Ruhan Wang, Chengsong Huang, Fuxiao Liu, Haitao Mi, Jordan Boyd-Graber, LeoweiLiang · 5 de octubre de 2026
Successful trajectories on difficult tasks provide valuable supervision for model improvement, but specialized harnesses introduce interventions that may be unavailable during deployment. We propose Recursive Self-Rewrite (RSR), a framework that uses one base model, Qwen-3.8-27B, to discover success…
- VERSE: Verified Self-Evolving Optimizer for Agent Harnesses
Zekai Wang, Yingqiang Ge, Zekun Wang, Hai Wang, Yuhui Xu, Joshua Frandsen, Shancong Fu, Ashia C. Wilson, Chandan K. Reddy · 5 de octubre de 2026
Harness evolution improves an LLM agent's prompts, tools, and workflow, while the optimizer's own tools and procedures often remain fixed. We study whether an optimizer can improve another agent more effectively by also improving how it diagnoses failures, develops edits, and tests their effects. Tw…
- Open-Endedness Bench: Measuring Epistemic Process from Agent Records
Chengyang Shi, Xianglin Ji, Jintao Huang, Jicheng Wang, Yifeng He, Jiachen Liu · 5 de octubre de 2026
Agents are increasingly given open-ended research tasks: discovering an empirical law from self-designed experiments, improving a heuristic whose optimum nobody knows, or beating a standing record. Their execution logs record every step of this research, yet the runs are still judged by their outcom…
- Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents
Ankur Samanta, Yonathan Efroni, Paul Sajda, Kaveh Hassani, Anirudh Goyal · 5 de octubre de 2026
Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are sparse in long execution traces, and their consequences may emerge several investigations later. We introduce Meta-reasoning for Ite…
- When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge
Xi Qin, Isabel Kurth, Xin Cui, Elin Park, Alexander Schaefer, Yaad Oren · 5 de octubre de 2026
Using a frontier model like Claude Opus as a meta-agent to generate terminal tasks and verifiers for RL training is increasingly common. Yet a runnable Docker image and executable test suite do not guarantee a faithful end-to-end pipeline for terminal agent training. We present a meta-agent pipeline…
- Harness Annealing: Learning to Act with Less External Control
Yingxuan Yang, Huacan Chai, Ying Wen · 2 de octubre de 2026
Language agents rely on external harnesses to track state, organize workflows, and verify answers. Beyond providing tools and information, these harnesses supply control decisions about what to investigate, whether to revise, and when to stop. Training on successful harness-supported trajectories ca…
- AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation
Giulio Zeloni, Enrico Lo Conte, Salvatore Rionero, Giuseppe Santoro, Alessandro Rastelli, Fabio Sorrentino · 2 de octubre de 2026
Enterprises adopting retrieval-augmented generation (RAG) face a recurring operational decision: promote, revise, or block a system version. The evidence is incomplete and the metrics come from fallible LLM judges. We report on AGO AI Quality Gate (AGO), an evidence-first quality-gate framework depl…
- It Takes Workflows to Evolve Better Workflows
Xuehang Guo, Haoyu Wang, Haifeng Chen, Yangyi Chen, Zhenhailong Wang, Qingyun Wang · 2 de octubre de 2026
Tackling complex real-world tasks can exceed the capabilities of a single large language model (LLM), motivating the use of multi-agent workflows that coordinate specialized agents to work together on these tasks. Recent methods train LLMs to construct better workflows from execution outcomes, but t…
- Can AI Scientists Coordinate at Runtime?
Zijian Liu, Yangzhixin Luo, Junyu Lu, Yi Li, Yu Chen, David Xu, William F. Shen, Xinchi Qiu, Xisen Wang · 2 de octubre de 2026
Multi-agent AI scientists have shown improving performance across a diverse range of tasks. Yet a common approach is design-time agentic orchestration, which typically relies on fixed workflows. In contrast, human scientists coordinate and adjust their division of labor at runtime. We therefore ask:…
- GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution
Geyi Yang, Zikun Qu, Xiang Li, Zhiyong Wang, Min Zhang, Shipei Zeng, Zhongxiang Dai · 2 de octubre de 2026
The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: rec…
- EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig · 2 de octubre de 2026
When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the start…
- How AI Agents Discover Scientific Equations: From Hydrotope Rediscovery to New Water-Wave Amplitudes
Zihan Zhou, Digvijay Wadekar, Matias Zaldarriaga · 2 de octubre de 2026
We study how AI agents discover and validate scientific formulas using a controlled case study of the hydrotope, a recently discovered geometric formula that combines the different polynomial pieces of nonlinear surface-wave scattering into one global expression. This problem is deceptively difficul…
- Actions with Receipts: Jointly Binding Claims, Evidence, and Execution for Replayable Tool-Agent Auditing
Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Yina Sa, Daren Zha, Jun Xiao · 2 de octubre de 2026
Tool-using agents can expose citations and execution logs while leaving a critical association unaudited: whether the claim shown to a user is the claim emitted by the committed execution and supported by the cited source. A valid citation and a valid trace can therefore remain individually well for…
- The Cognitive Continuity Test: Verifying Governed State Transitions in Persistent AI Agents
Jun He, Deying Yu · 2 de octubre de 2026
Persistent AI agents revise beliefs, consolidate memory, and replace execution substrates. Similar successor states can accompany differently authorized transition claims, while legitimate development can change state substantially. We introduce the Cognitive Continuity Test (CCT), a policy-relative…
- Safety in Self-Evolving Agents: A Survey
Jiahao Chen, Zhou Feng, Oubo Ma, Yichen Yan, Ruixiao Lin, Hangtao Zhang, Linkang Du, Hengyu An, Yong Yang, Jun Liu, Junhao Li, Naen Xu, Chunyi Zhou, Yuan Su, Zehao Jin, Qianli Ma, Leyi Qi, Yiming Wang, Zhe Ma, Yuwen Pu, Mengyao Du, Yuanyi Song, Enhao Huang, Zhihui Fu, Jun Wang, Jinfeng Li, Yuefeng Chen, Hui Xue, Yiming Li, Tianyu Du, Shouling Ji · 2 de octubre de 2026
Large language models (LLMs) exhibit strong general capabilities, yet their parameters typically remain fixed after deployment, limiting learning from new interactions. In open-ended environments, this motivates self-evolving agents that continually update reusable state-including model parameters, …
- The Delegation Danger Band: Why Mid-Capability Sub-Agents Over-Trust Inherited Stale State
Jundong Hu, Shekar Ramachandran · 2 de octubre de 2026
Agent frameworks increasingly delegate work by forking sub-agents; a common default makes the child inherit the parent's full working context. We measure how the effect of inherited state changes with capability, where $C_m$ denotes clean fork-fresh accuracy. We compare 3 inheritance policies: Reset…
- ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha Chowdhery, Akari Asai, Omar Khattab, Yejin Choi, Gunhee Kim, Chelsea Finn · 2 de octubre de 2026
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which ear…
- Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks
Hao Wang, Ting Huang · 2 de octubre de 2026
Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are silently abandoned. We present evidence, from a controlled single-mach…
