Physical Sciences › Computer Science › Information Systems
Software Engineering Research
853 artículos indexados
El estudio de las publicaciones en ingeniería de software en el ámbito de la inteligencia artificial explora cómo los modelos de código interactúan con las prácticas de desarrollo. Examina métodos para evaluar la calidad de las generaciones de código, como el análisis de los hidden states de los modelos preentrenados o el impacto de las influence functions en la detección de anomalías en los prompts. Las investigaciones abordan también los desafíos relacionados con la traducción automática entre lenguajes, el alineamiento de instrucciones para el aprendizaje de representaciones binarias, o incluso los efectos de las tácticas de prompt framing en el rendimiento de los grandes modelos lingüísticos.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos44 % · 229 artículos
- China30 % · 157 artículos
- Canadá8,5 % · 44 artículos
- Reino Unido6,4 % · 33 artículos
- Alemania5,4 % · 28 artículos
- India5 % · 26 artículos
- Singapur4,8 % · 25 artículos
- Australia3,3 % · 17 artículos
Sobre 519 artículos de este tema con al menos un laboratorio localizado. 63 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code
Siyu Wang, Yifan Wang, Yuecheng He · 5 de octubre de 2026
Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test fu…
- Are you Synthesizing or Recalling? Evaluating LLMs on Algorithmic Code Retrieval
Nickil Maveli, Antonio Vergari, Shay B. Cohen · 5 de octubre de 2026
Large language models (LLMs) have demonstrated strong performance in code generation, where success depends on both recalling relevant algorithmic knowledge and reasoning about how to apply it. However, existing LLM pipelines are opaque, with no explicit separation between these two components. We a…
- CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation
Zheng Fang, Yongmin Li, Yichang Zhang, Dongming Jin, Haoyu Wang, Shuai Wang, Zhi Jin, Ge Li · 5 de octubre de 2026
Coding agents can generate code that appears correct but implements behavior the user never intended. This mismatch can arise when an agent silently resolves underspecified requirements through its own assumptions. As subsequent development builds on these assumptions, correcting the resulting behav…
- ClarifyCodeBench: Evaluating LLMs on Clarifying Ambiguous Requirements for Code Generation
Zheng Fang, Dongming Jin, Yihong dong, Yongmin Li, Kechi Zhang, Zhi Jin, Ge Li · 5 de octubre de 2026
Large Language Models have emerged as programming assistants. However, the efficacy of code generation is constrained by the quality of input requirements, which are frequently ambiguous, incomplete, or underspecified. While LLMs excel at one-shot code synthesis, their ability to proactively clarify…
- AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
Xuan Zhang, Longtao Zheng, Cunxiao Du, Bo An, Xin Dong · 2 de octubre de 2026
Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working …
- Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models
Juan S. Santillana · 2 de octubre de 2026
Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 661.6M parameter model (approx. 65% code/technical text; no d…
- When Does a Second Model Help? Cross-Model Review in LLM Verification
Tae-Eun Song · 2 de octubre de 2026
Large language models now generate code, documentation, and analyses, and are increasingly used to review such output. We ask when a second review by a different model helps. Building on the author's earlier preprints, which varied context, repetition, and role structure within one model, we test mo…
- Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents
Junyu Guo, Shangding Gu, Ming Jin, Javad Lavaei · 2 de octubre de 2026
Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-label…
- Cross-Benchmark Transfer from RL on Agentic Coding Tasks
Sushant Mehta, Logan Ritchie, Edwin Chen · 2 de octubre de 2026
Coding agents often fail in the last mile: they build most of a feature but drop a requirement, test only the cases their implementation already handles, break behavior that was supposed to stay intact, or validate against an unchecked assumption. We ask whether reinforcement learning (RL) on expert…
- Self-Evolving Coding Rules for AI Coding Agents
Zhengyuan Jiang, Reachal Wang, Yuepeng Hu, Yupu Wang, Yuqi Jia, Neil Zhenqiang Gong · 2 de octubre de 2026
The performance of AI coding agents is highly dependent on their underlying coding rules. However, existing coding rules are typically hand-crafted and fixed, making the process labor-intensive and often suboptimal. In this work, we propose RuleEvolve, a self-evolving framework for coding rules. Rul…
- Code That Works, Environments That Don't: Measuring Environment Reproducibility in AI-Generated Software
Bhanu Prakash Vangala, Tanu Malik · 2 de octubre de 2026
Code generation has emerged as a central capability of large language models, with coding agents now able to produce functionally correct software projects from natural language prompts. However, functional correctness alone does not capture a critical dimension of generation quality: environment sp…
- Rules to Tools: Executable Checks for LLM Agents in Scientific Computing
Jingjie Ning, Guojiang Zhao, Chen Xu, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Guolin Ke · 2 de octubre de 2026
Scientific coding agents receive equations, boundary conditions, and output requirements in writing, then must assess the programs they revise. Rules to Tools (R2T) supplies prepared executable checks of public scientific requirements. Matched SciCode repair groups share written checks, starting pro…
- LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems
Yun Peng, Zihan Wu, Zeyang Zhuang, Xin Zhou, Rui Shu, Xu Han, Chun Yong Chong, Yuan Wang, Jiakun Liu · 1 de octubre de 2026
Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits from detailed speci…
- Learning When and How to Intervene: A Hindsight-Distilled Sentinel for Coding Agents
Jiangrui Zhao, Chenglong Li, Meng Zhang, Xiaoting Du · 1 de octubre de 2026
Coding agents solve repository-level tasks through sequences of actions, where a single erroneous action can misdirect subsequent decisions and increase recovery costs. Existing approaches use execution feedback for recovery or specialized checks to block errors, but deciding before execution whethe…
- DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation?
Frances Liu, Manny Silva, Paige Calvert, Ayu Adiati, Sarah Sanders · 1 de octubre de 2026
We introduce DoGBENCH (Documentation Generation Benchmark), to our knowledge, the first benchmark for generating and maintaining real user-facing software documentation. It asks whether an agent can produce documentation that experienced technical writers would accept in review. The benchmark contai…
- Self-Spec Verifiable Code Generation
Jiaru Qian, Yihong Dong, Yongmin Li, Hao Zhu, Bin Gu, Ge Li · 1 de octubre de 2026
Large language models (LLMs) may generate unreliable code on corner cases missed by testing, while formal verification can provide machine-checkable guarantees. Recently, researchers have proposed several benchmarks to evaluate the capabilities of LLMs in generating formally verifiable code, where L…
- EngramBench: A Capability-Grounded Benchmark for Skill-Evolution Harnesses
Zhixuan Tan, Pengjie Gu, Zhao Li, Yihan Hu, Xu He, Dong Li, Jianye Hao · 1 de octubre de 2026
While large language models have achieved remarkable success in isolated code generation, authentic software engineering requires sustained reasoning, complex state management, and continuous cross-domain abstraction. However, current evaluations of skill evolution in autonomous agents suffer from a…
- Trustworthy Runtime Error Healing in Real-World Repositories: A Benchmark and Guardrail
Gou Tan, Pengfei Chen, Zhensu Sun, Jieke Shi, Junkai Chen, Ting Zhang, Weifeng Sun, Junda He, Shuai Liang, Chuanfu Zhang, Lwin Khin Shar, David Lo · 1 de octubre de 2026
Runtime error healing lets a crashed program continue by generating code that repairs its live runtime state. Recent work shows that LLMs can generate such healing code, but it is evaluated only on small competition programs, and executing LLM-generated code inside a live process raises safety conce…
- E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch
Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang, Matt Deitke, Pengcheng Yin, Zijian Wang, Rui Hou · 1 de octubre de 2026
Coding agents powered by large language models (LLMs) are evolving from making localized code changes to developing complete software repositories. However, evaluating repository-scale generation remains challenging: tasks must demand system-level reasoning while ensuring that all evaluated behavior…
- Zero2Repo: Can Coding Agents Build Repositories from Scratch?
Pei Yang, Tianyu Shi, Yuhang Yao, Wanyi Chen, Tongyun Yang, Dun Pei, Haonan Wang, Pengbin Feng, Guanxu Yu, Jingchun Huang, Zeyu Zhang, Shuhan Sun, Hao Li, Xiang Li, Jie Xiao, Xinyu Wang, Hanxin Chen, Daqi Li, Qi Jia, Hongshan Lin, Zhizhou Gu, Zijun Tian, Weizhi Du, Lynn Ai, Eric Yang · 1 de octubre de 2026
Coding agents are increasingly asked to build software rather than patch it, yet benchmarks for from-scratch repository construction are mostly limited to a single language and depend on manually curated tasks. We introduce Zero2Repo, a benchmark in which an agent receives a product requirements doc…
- How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
Kirill Brilliantov, Alejandro Hern\'andez-Cano, Emmanuel Abb\'e · 1 de octubre de 2026
Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machin…
- Papers Without Code: Availability of GitHub Repositories Linked in *CL Publications
Selina Meyer, Michael Roth · 30 de septiembre de 2026
Source code and data published at computational linguistics (*CL) venues are increasingly being shared via GitHub. While this generally is a favourable development for the accessibility and potential reusability of research artifacts in natural language processing (NLP), the long-term availability o…
- Opera: A Verbal Critic Framework for Long-horizon Coding Agents
Kai Mei, Zhiyuan Hu, Yutong Dai, Juntao Tan, Yifan Zhang, Dingjie Song, Dimitris N. Metaxas, Silvio Savarese, Ran Xu, Zeyuan Chen · 30 de septiembre de 2026
Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is de…
- Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo · 30 de septiembre de 2026
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and …
- Retrieve, Reproduce, Reveal: Dissecting Retrieval-Augmented Software Vulnerability Detection
Sabrina Kaniewski, Tim Kr\"amer, Julius B\"achle, Markus Enzweiler, Michael Menth, Tobias Heer · 30 de septiembre de 2026
Retrieval-Augmented Generation (RAG) is increasingly used to enhance Large Language Model (LLM)-based software vulnerability detection by grounding predictions in retrieved vulnerability knowledge, such as vulnerability reports. However, existing RAG-based software vulnerability detection (RAG4SVD) …
Otros asuntos del tema Sistemas de información
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Information Retrieval and Search Behavior727 artículos / 12 meses+506 %
- Recommender Systems and Techniques600 artículos / 12 meses+154 %
- Expert finding and Q&A systems205 artículos / 12 meses+1175 %
- Information and Cyber Security169 artículos / 12 meses+1500 %
- Blockchain Technology Applications and Security146 artículos / 12 meses+650 %
- Big Data and Digital Economy127 artículos / 12 meses+14 %
