Physical Sciences › Computer Science › Software
Model-Driven Software Engineering Techniques
112 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States42% · 14 papers
- China30% · 10 papers
- Hong Kong SAR China9.1% · 3 papers
- Germany9.1% · 3 papers
- United Kingdom9.1% · 3 papers
- Türkiye9.1% · 3 papers
- Canada6.1% · 2 papers
- Netherlands6.1% · 2 papers
Across 33 papers on this subject with at least one lab located. 20 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix
Zirui He, Haiyan Zhao, Jingyu Hu, Yinghao Wu, Chenxi Yuan, Yingcong Li, Yandong Bai, Mengnan Du · 2 October 2026
Alignment does not eliminate behavioral errors in language models. Models may still refuse benign requests, call unnecessary tools, or yield to false user claims. Current methods mitigate such errors as a computation problem, and rarely explore if the desired behavior is already encoded in the model…
- Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning
Zichen Xie, Mrigank Pawagi, Lize Shao, Yang Hu, Wenxi Wang · 2 October 2026
Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable principles may prescribe incompatible behavior when applied to the sam…
- MatLoom: Layered Text-to-Material Generation in a Compact Program Space
Anson Y. Lam, Shuqing Li, Michael R. Lyu · 1 October 2026
Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. Each program composes alpha-masked layers whose shared spatial expressions define…
- From Search to Signal: Online Post-Training in Automatic Heuristic Design
Yilun Yuan, Tianyu Zhou, Zhenzhou Tang · 1 October 2026
Large language model (LLM)-based automatic heuristic design (AHD) iteratively proposes and refines heuristics, pairing design rationales with executable code. Task-specific evaluators assess programs; execution outcomes and performance scores guide search. Many AHD systems keep the generator frozen;…
- Right Answers, Costly Models: The Efficiency Gap in LLM-based Optimization Modeling
Zhong Li, Xin Huang, Jinhui Wan, Xiangyi Wang, Shenkai Zhang, Ruiqi Chen, Wenyu Liu, Zaiwen Wen, Ziyan Luo · 1 October 2026
Optimization modeling formulates real-world decision problems as mathematical programs that solvers can use to find optimal decisions. Large language models (LLMs) can automate this process, but the resulting correct formulations can require substantial time and memory to construct and solve, limiti…
- How Should Diffusion Language Models Edit Code?
Xijia Tao, Ziru Liu, Shansan Gong, Jiacheng Ye, Kecheng Chen, Zirui Wu, Lin Zheng, Xinyu Fu, Rui Liu, Lingpeng Kong · 1 October 2026
Code editing requires a model to decide where to make changes, generate the new content, and preserve everything else. We study how masked diffusion language models divide these responsibilities across four editing interfaces: whole-file rewriting, search-and-replace, locate-then-infill, and token-l…
- Code to Control: Synthesizing Parameterized Reactive Controllers
Zergham Ahmed, Joshua B. Tenenbaum, Chris Bates, Samuel J. Gershman · 1 October 2026
Recent LLM-based approaches to control either invoke a language model to select actions or synthesize world models that require planning at every decision, introducing latency that can limit real-time use. We introduce Code to Control, an approach that synthesizes Python controllers which execute di…
- ReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code Generation
Huifei Wang, Xinying Huang, Yiheng Sun, Yifan Yuan · 30 September 2026
Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search fr…
- ManiEdit: Sequential Unstructured Knowledge Editing for Language Models from a Manifold Perspective
Rui Liu, Chenheng Zhang, Haoxuan Li, Zhouchen Lin · 30 September 2026
Large language models (LLMs) inevitably generate some incorrect or outdated content, necessitating efficient and precise mechanisms for continual knowledge updates. However, existing model editing methods struggle to sequentially edit unstructured long-form knowledge, suffering from severe edit forg…
- OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets?
Wenjun Peng, Xinyu Wang · 30 September 2026
Static QA and code-generation benchmarks only partially capture the role that large language models (LLMs) now play as coding agents and research tools. We introduce OptiArena, a budget-controlled testbed for studying whether LLMs can improve executable game-playing algorithms through five rounds of…
- Is manual software optimization a thing of the past?
Pavlin G. Poli\v{c}ar, Martin \v{S}pendl, Toma\v{z} Ho\v{c}evar · 30 September 2026
Scientific software is increasingly required to process larger datasets while maintaining acceptable execution times. Software optimization traditionally requires substantial expertise in programming, algorithms, and numerical methods. Recent advances in large language models (LLMs) offer the possib…
- Compiling Learning Problems into Adaptation Programs for Language Models
Rebecca Ramnauth, Brian Scassellati · 30 September 2026
Model adaptation is typically governed by a fixed recipe, even though different update programs can produce substantially different behavioral outcomes. We introduce adaptation compilation, which reframes where, how, and to what extent a model should adapt as a joint prediction and decision problem.…
- Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes
Xiaokang Ye, Siddhant Hitesh Mantri, Zimeng Chen, Edward Zhang, Zhaoxu Zheng, Yuanheng Li, Yizhao Chen, Tianyang Huang, Lianhui Qin · 30 September 2026
Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably understand 3D structure and precisely control scene state remains unclear. The generated 3D scene is a persistent, executable artifact: a convincing render can hide incorrect spatial relatio…
- From Hand-Crafted to LLM-Based Variation Operators in Metaheuristics: A Tutorial
Camilo Chac\'on Sartori, Guillem Rodr\'iguez-Corominas, Christian Blum · 29 September 2026
Large language models (LLMs) are increasingly being employed as variation operators in metaheuristics, generating or modifying candidate solutions, heuristics, or programs inside iterative search loops. This shift reframes variation as a model call conditioned on different types of information. We i…
- MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses
Xuanjun Chen, Hua-Hsuan Chen, Wei-Chung Lu, Yinghao Ma, Jyh-Shing Roger Jang, Hung-yi Lee · 29 September 2026
Rapid progress in Large Language Models (LLMs) is saturating static benchmarks faster than they can be designed. While existing automated evolution frameworks attempt to generate harder questions by perturbing individual tasks, they remain constrained by rigid, hard-coded generation rules. Moving be…
- HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases
Aditya Kumaran, Rahul Singhal, Karime Maamari, Amine Mhedhbi, Pradyumna Tambwekar · 28 September 2026
Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants while keeping their expected outputs fi…
- ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks
Zeyu Michael Li, William Xingxu Chen, Bingshuo Qian, Jiayin Liu, Xiang Cheng · 25 September 2026
Fully continuous diffusion language models (dLMs) denoise continuous representations without intermediate discretization, then decode all response tokens in parallel at the final step. Their performance on challenging reasoning tasks remains less established than that of autoregressive (AR) LLMs and…
- BigO(Bench): Can LLMs Generate Code with Controlled Time and Space Complexity?
Pierre Chambon, Baptiste Roziere, Benoit Sagot, Gabriel Synnaeve · 23 September 2026
We introduce BigO(Bench), a novel coding benchmark designed to evaluate the capabilities of generative language models in understanding and generating code with specified time and space complexities. This benchmark addresses the gap in current evaluations that often overlook the ability of models to…
- Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training
Jacob Beck, Philip V. Ogren, Ari Kobren · 23 September 2026
Large language models (LLMs) can improve solutions to verifiable scientific and algorithmic problems by spending additional computation at test time. Recent systems achieve strong results with increasingly elaborate evolutionary search harnesses or by updating model parameters during test-time train…
- Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents
Laizhen Li, Jiarui Li, Juanjuan Zhao, Kejiang Ye, Ye Li, Cheng-zhong Xu, Xitong Gao · 23 September 2026
Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserv…
- Toolcompass: Guiding Tool Trialing, Not Suppressing It
Junlin Fang, Chong Zhang, Do Nguyen-Thanh, Xiaogang Xu, Zhen Fang, Sean Du · 23 September 2026
Large language model (LLM) agents must generalize from tools seen during training to unseen tools at deployment. A key challenge is tool trialing, i.e., excessive trials waste the interaction budget, whereas selective trials enable exploration of unfamiliar tools. Existing outcome-based post-trainin…
- AutoGym: Blueprint-First Generation of Verifiable Agent Gyms
Aarati Andrea Noronha, Kavya Ravikumar, Carly Xiaoyu Lin · 22 September 2026
Training agents with reinforcement learning requires a gym, comprising a task, an executable environment in which the task can be attempted, and a verifier that reliably distinguishes success from failure. Constructing such gyms remains manual, expensive, and static. Task sets saturate as models imp…
- H2LooP Telecom Model v1: From Telecom Comprehension to Autonomous Issue and PR Resolution
Amit Singh, Vedant Nipane, Mayank Goel, Pulkit Agrawal, Sairanjan Mishra · 22 September 2026
We present H2LooP Telecom Model v1, a domain-specialized large language models fine-tuned for the telecommunications industry. We release two domain-adapted model variants serving complementary use cases: a comprehension-focused variant for telecom domain question answering and reasoning, and an age…
- GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills
Rui Sun, Zhi Zheng, Zhenkun Wang, Zhichao Lu · 21 September 2026
Skills can improve the performance of Large Language Model (LLM) agents by providing task-specific procedural guidance, while skill optimization further improves their effectiveness through iterative refinement. However, existing skill optimization methods typically represent skills as unstructured …
- The Imitation Game: When LLMs Learn to Reason Like Programs via Code-Centric Reasoning Data Synthesis
Jinyang Zhang, Weibin Liao, Keqin Bao, Sihang Li, Shaobo Wang, Muyang Ye, Hongxin Ding, Yue Fang, Tianyi Tang, Fei Huang, Kexin Yang, Xingzhang Ren, Dayiheng Liu · 16 September 2026
Large Language Models (LLMs) excel at programming tasks but frequently fail at deterministic, fine-grained reasoning in natural language, relying heavily on semantic approximations rather than robust symbolic execution. To bridge this gap, we propose MIMIC, a framework that leverages executable code…
