Physical Sciences › Computer Science › Software
Software Testing and Debugging Techniques
241 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten48 % · 74 Artikel
- China30 % · 46 Artikel
- Kanada9,7 % · 15 Artikel
- Vereinigtes Königreich9,7 % · 15 Artikel
- Deutschland5,8 % · 9 Artikel
- Indien3,9 % · 6 Artikel
- Sonderverwaltungsregion Hongkong3,2 % · 5 Artikel
- Südkorea3,2 % · 5 Artikel
Über 154 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 39 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- When Correct Memory Goes Wrong: Fuzzing Persistent Memory Use in LLM Agents
Yuqiao Meng, Luoxi Tang, Yingxue Zhang, Yuchen Yang, Zhaohan Xi · 1. Oktober 2026
Persistent memory helps LLM agents carry information across long interactions, but correct memory can still be used incorrectly when queries change or memory states evolve. Existing work mainly studies memory content errors or evaluates fixed test cases, leaving memory-use failures hard to discover …
- AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems
Yiming Cheng (The University of Chicago), Alfin Wijaya Rahardja (Fudan University), Mengshi Zhang (TensorBlock, Inc), Zihao Chen (TensorBlock, Inc), Zhenpeng Chen (Tsinghua University), Yiling Lou (University of Illinois Urbana-Champaign) · 30. September 2026
Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number of executable harness bugs while requiring hundreds of human hours to…
- WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses
Haomin Qi, Xiangzhe Xu, Yiming Huang, Jingbo Shang, Chengpeng Wang · 30. September 2026
Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse pub…
- LatentSift: Policy-State Filtering for Token-Efficient Verification of Software Engineering Agents
Yuning Han, Yangchenchen Jin, Tyler Jandreau, Jingwei Sun · 30. September 2026
Test-time scaling improves software engineering agents by generating multiple candidate trajectories and selecting the best one. Verifying and selecting among these long interactions can consume as many tokens as generation itself. Existing hybrid workflows first apply an LLM-based execution-free (E…
- Beyond Rule-Based Mutation Testing: Test-Aware Mutant Generation Using Large Language Models
Nils Kiele, Zainab Saad, Zirui Wang, Steve Drew, Samira Ebrahimi Kahou · 30. September 2026
Mutation testing evaluates test-suite adequacy by injecting synthetic faults into program code. However, traditional rule-based tools often generate large numbers of trivial, redundant, or equivalent mutants that limit their practical use for identifying gaps in a test suite. While recent large lang…
- Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation
Jiayuxuan Yang, Jie M. Zhang, Yiling Lou, Zhenpeng Chen · 30. September 2026
Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guid…
- More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses
Ziyang Xu, Haitian Zhong, Hao Zhou, Hao Qin, Chenhan Jin, Te Qi, Shengze Xu, Tieyong Zeng · 30. September 2026
Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, rep…
- FairTest: Search-Based Fairness Testing for Multi-Agent Reinforcement Learning Systems
Xiaotong Wang, Xuan Xie · 24. September 2026
Multi-agent Reinforcement Learning (MARL) trains a team of agents that share one environment and learn their policies together. Training maximizes the team return, and a high return does not imply that the rewards are shared fairly among the agents in every episode. Testing is an established way to …
- Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
Mingzhe Du, Anh Tuan Luu, Dong Huang, See-Kiong Ng · 22. September 2026
Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and a loose floating-point tolerance, and their verdicts now feed leaderboards and reinforcement-learning rewards. Recent work agrees these checkers are weak and patches them by hand---extra input distributions, fuz…
- Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation
Ana Nunez, Peyman Najafirad · 21. September 2026
Self-play methods that co-train a single language model as both coder and test author promise to move code-generation RL beyond fixed test suites, but they suffer from two coupled pathologies: permissiveness collapse, where pass-rate rewards are maximised by trivial, non-discriminative tests, and co…
- The Output-Space Hypothesis: Enumerative Equivalence Checking for Tensor Programs
Paul Biberstein, Joseph Devietti, Mayur Naik · 18. September 2026
Tensor programs, as used in deep learning models, are a prime target for optimization, as small performance improvements can have a large impact across training or inference workloads. However, such optimizations are complicated and can produce subtle bugs. Traditionally, correctness is assumed when…
- A Study of the Reliability of Agentic AI-Generated Programs
Ayesha Shafique, Barton P. MIller, Elisa R. Heymann · 17. September 2026
Agentic-AI based software development offers the promise of faster completion of the software, greater programmer efficiency, and more reliable code. The question is how can we verify these claims in an objective way? In this project, we attempted to answer this question based on three practices. Fi…
- IWC-Bench: Evaluating Web Application Generation from a Software Testing Perspective
Chenxu Liu, Zilu Zou, Peizhong Gao, Jiawen Tao, Zhexin Zhang, Guang Chen, Haowei Lin, Ying Zhou, Tianyi Bai, Dolly Deng, Suncong Zheng, Maxm Pan · 16. September 2026
Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evaluation remains challenging. Static benchmarks can credit functionality that exists in source code but is unreachable at runtime. Interactive benchmarks …
- TyPatch: Transforming Patches into Typestate Rules for Kernel Bug Detection
Ruoyu Wang, Tuo Li, Jia Li · 15. September 2026
Historical Linux kernel patches capture defect knowledge that applies beyond their original repair sites. Recent work has shown that large language models (LLMs) can generate static-analysis checkers from historical patches and use them to uncover new kernel bugs. However, complete-checker generatio…
- Confidence-Gated Transductive Test Generation for Code Reranking
Sungjae Lee, Youngsik Yoon, Seockbean Song, Siwei Wang, Wei Chen, Jungseul Ok · 14. September 2026
Test case synthesis is crucial for evaluating and ranking programs generated by large language models (LLMs). However, constructing high-quality test cases remains challenging because reliable expected outputs are often difficult to obtain. We propose Confidence-Gated Transductive Test Generation (C…
- SWE-Test: Benchmarking LLM Vulnerability Discovery via Input Prediction
Yuanxiang Shi, Jiayi Lin, Xuanyong Lin, Liangcai Su, Yeheng Duan, Wei Wang, Qi Han, Bing Zhao, Wei Hu, Xander Xu, Chenxiong Qian · 9. September 2026
Vulnerability discovery is becoming an important ability of large language model (LLM) agents: agents that silently miss real defects leave critical software exposed. Rigorously measuring this ability is therefore urgent, but existing benchmarks are gameable through data contamination, score recall …
- ExecCritic: Learn to Test, Test to Improve for Coding Agents
Leitian Tao, Baolin Peng, Haorui Wang, Hang Wang, Hao Cheng, Wenlin Yao, Qianhui Wu, Tao Ge, Sharon Li, Jianfeng Gao · 9. September 2026
Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode incomplete or incorrect behavioral targets; when the same trajectory writes both the patch and the test, their errors ca…
- ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems
Ant\'onio Azevedo, Bruno Lima, Jo\~ao Pascoal Faria · 7. September 2026
Automotive infotainment validation still relies on manual testing, slow, costly, and incompatible with agile releases and OTA updates. Scripted automation only partly helps: it couples test logic to implementation, yielding brittle, high-maintenance suites. Existing LLM-driven frameworks mostly targ…
- Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair
Xuemeng Cai, Jiakun Liu, Linhan Yang, Wei Ma, Lingxiao Jiang · 7. September 2026
Large language models (LLMs) have significantly advanced automated program repair (APR), yet existing evaluations remain largely result-centric and provide limited insight into hallucination during repair. In APR, hallucination may arise not only in final patches but also in the intermediate artifac…
- Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs
Jiacheng Xu, Wentao Zhang, Zhiyi Lyu, Fuxiang Zhang, Chaojie Wang, Yang Liu, Bo An · 4. September 2026
Reinforcement learning (RL) has substantially advanced code generation with large language models (LLMs) through executable feedback. The feedback for coding problems mainly comes from specific test cases, where high-quality test cases are often scarce since they should be both sound and discriminat…
- Counterexamples as Feedback for Agent Self-Correction
Sidhesh Badrinarayan, Adithya Parthasarathy · 4. September 2026
Single-turn code-generation metrics understate a central property of deployed agents: whether they can repair a wrong artifact after receiving concrete feedback. This paper presents A-CEGIS, a lightweight framework that uses counterexamples as feedback for evaluating multi-turn refinement in natural…
- Does Fault Localization Beat a Fresh Attempt? A Placebo-Controlled Study of Test-Guided Code Repair
Anik Jha · 2. September 2026
Fault localization can focus a code model's repair on the statements a failing test implicates, but a targeted edit may succeed merely because it is small, and a second model call may succeed without using the failure at all. We separate these explanations with three arms applied to the same failed …
- RiskBlend: A Multi-Signal Framework for Test Input Prioritization in Machine Learning Regression Testing
Madhusudan Srinivasan, Namith Nishal Raphae · 31. August 2026
When machine learning classifiers are retrained, inputs correctly classified by the previous model version may be misclassified by the updated version, creating regression faults that are costly to detect because verifying predictions against ground truth may require human annotation, expert review,…
- FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs
Zeming Liu, Hang Lyu, Jingtao Zhang · 28. August 2026
Generated operational programs are often validated with either a few hand-written examples or exhaustive regression suites. The former can miss sparse boundary and interaction faults, while the latter can be unnecessarily expensive. We introduce FaultLens, a method for learning compact behavioral te…
- FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs
Ze Sheng, Aleksandar Kezic, Zhicheng Chen, Jeff Huang · 27. August 2026
Evaluating the ability of large language models (LLMs) to discover software bugs is increasingly important. Existing benchmarks typically evaluate this capability by asking the model to generate a proof-of-concept input that triggers a predefined target vulnerability. However, this setup may overloo…
