Alsay main site ↗

Perspective & research agenda · Working paper

AI-assisted programming: from generated code to validated repository changes

Abstract

Success at programme synthesis does not by itself establish dependable repair of an existing repository. This paper synthesizes 12 selected studies spanning code generation, test adequacy and repository-level evaluation. It proposes bounded patches, independent acceptance tests and explicit regression checks, while distinguishing error detection from effective final repair. A paired protocol compares human, one-shot and simplified repair workflows at fixed revisions and budgets. Outcomes include valid repairs, regressions, review costs and uncovered environments. No original experiments or current model ranking are reported. The argument is limited by purposive literature selection and the adequacy of operational tests; passing a finite test suite is not proof of correctness or verification of production deployment.

1 Problem and central argument

Does success at generating a function transfer to reliable repair of an existing system? This paper distinguishes short programmes, competition tasks and repository maintenance instead of generalizing a single benchmark score to all software engineering.

2 Literature scope and research context

This paper purposively selects 12 research documents for a question-led narrative synthesis, not a systematic review or meta-analysis. The selected works were first submitted in 2017–2024 and do not cover all developments through 2026. Titles, authors, submission dates, public abstracts, source pages and PDF access were checked on 8 September 2026. Original research, technical reports and surveys are distinguished rather than treated as equally strong experimental evidence. Selected methodological passages in FActScore, ALCE, EvalPlus and τ-bench were previously examined; we did not conduct full-text appraisal or experimental replication of every work.

The Codex study evaluates functional correctness using tasks including HumanEval and examines repeated sampling and model limitations.[1] The study introducing MBPP evaluates synthesis of short Python programs from natural-language descriptions.[2]

AlphaCode combines large-scale candidate sampling with behaviour-based filtering for competitive programming, a different setting from maintaining an existing software system.[3] Pearce and colleagues find vulnerable code generated by an early version of Copilot in selected security-relevant scenarios; these historical observations do not estimate current product vulnerability rates.[4]

CodeGen studies multi-turn programme synthesis by decomposing requirements into successive subproblems.[5] StarCoder describes an openly released code-model effort, including licensing, personal-information handling and source-attribution arrangements.[6]

Code Llama offers variants for generation, completion and instruction following, highlighting differences in training and use conditions.[7] EvalPlus expands test inputs and identifies errors missed by earlier tests, demonstrating that test adequacy changes measured performance.[8]

SWE-bench evaluates issue resolution in real repositories, requiring coordinated understanding and changes across functions, classes or files.[9] SWE-agent examines how an agent–computer interface for navigation, editing and testing affects software-repair performance.[10]

Agentless uses a simplified localization, repair and patch-validation pipeline, motivating comparisons against simpler workflows.[11] LiveCodeBench collects new competition problems over time and includes self-repair, execution and test-output prediction to address contamination and evaluation coverage.[12]

3 Synthesis and proposed method

The cited studies use different tasks, models, sampling budgets and tests, so their scores do not form a valid league table. We propose separating local behaviour, related-module regressions and target-environment verification. Simple workflows remain necessary controls when assessing whether a more complex agent architecture justifies its cost.

Preserve a minimal reproduction and expected behaviour before limiting the files and invariants a patch may change. Test normal, exceptional and boundary inputs; review dependencies, access control and error handling. Record the revision, tests actually executed, untested environments and a recovery procedure. Code volume is not a quality measure.

4 Proposition and comparisons

At fixed patch scope and cost, does independent testing improve the proportion of valid final repairs, rather than merely reveal more errors? A higher detection rate without more usable final patches must be reported separately, not as better repair effectiveness.

Compare human repair, one-shot generation and a simple localization–repair–validation pipeline. Add independent tests and regression checks in the full workflow. Freeze benchmark revisions, isolate held-out tests from patch generation, and record dependencies and environment hashes.

5 Operational definitions and decision boundaries

Count a code task as a valid repair only if the predefined issue is resolved, independent acceptance checks pass and required regression checks do not fail. Report security review and target-environment coverage separately. This operational definition is bounded by test adequacy, not proof of correctness for every input. Errors detected, repairs completed and review time are distinct outcomes.

Compare human and AI-assisted repair at fixed repository revisions in isolated environments. Record successful repairs, regressions, review time and tool cost. Keep failures, timeouts and human takeovers in the sample. Passing tests only establishes behaviour on executed cases; it neither verifies production deployment nor establishes a current product ranking.

6 A testable evaluation protocol

Treat tasks, not repeated outputs from one task, as sampling units. Use a separate pilot to assess annotation feasibility, cost and plausible effects, then determine the main sample size from precision or power targets. Exclude pilot tasks from held-out evaluation. Stratify Chinese and English tasks and treat translated counterparts as paired rather than independent samples. Publish inclusion and exclusion criteria and freeze inputs, versions, permissions and budgets.

Randomize condition order and compare conditions within tasks; blind adjudicators to condition where feasible. Report raw counts, task-level effect differences and uncertainty intervals without treating repeat runs as extra independent tasks. Report task families separately and test sensitivity to missing outcomes and timeouts. Retain and distinguish timeouts, abstentions and human takeovers. Document when and why a planned procedure changes.

A reproducibility package should identify tasks, data permissions, model and tool versions, prompt templates, configurable random seeds, tests, adjudication rules, failure logs and analysis scripts. Reproducibility does not justify releasing personal or restricted information; redacted structures and controlled-access terms may be appropriate. These are proposed research outputs: this site does not currently provide a completed experimental dataset or analysis code.

7 Interpretive boundaries and alternatives

EvalPlus shows that expanded tests can reveal errors missed by earlier evaluation and also discusses defective reference implementations. More tests do not prove correctness for every input; test validity and final-patch correctness require separate scrutiny.[8]

This paper contains no original experiments or current product comparison. Extra tools, human review or unequal budgets could explain apparent gains; selecting only successful tasks would introduce bias. Findings that contradict expectations should retain failures and negative results, with explicit boundaries for model, language, task family and environment. Historical benchmarks cannot establish guarantees for current systems.

8 Conclusion

Generating a patch, passing executed tests and working in production are different conclusions. Connecting them through revisions, diffs, independent tests and runtime evidence is a proposed delivery framework, not a demonstrated universal safety guarantee.

This companion paper develops the AI+ overview using a shared literature base and framework; it is not an independent empirical study. Read the overview.

9 Declarations and availability

Data and code: No original experimental data were generated. Sources are individually linked. Experimental and analysis code has not been implemented; website code is not an experimental replication package.

Authorship, contributions, funding and competing interests: Alsay is the website publisher, which does not establish scholarly authorship. Author identities, contributions, funding and competing interests require confirmation by accountable contributors. This draft does not assert absence of competing interests or invent affiliations, degrees or ethical approvals.

AI assistance and versions: AI assisted retrieval, organization, editing and Chinese–English translation. The two versions represent one working paper, not two independent studies; scholarly approval by accountable authors remains pending. No peer review, journal acceptance or completed preregistration is claimed.

References

Cited as public arXiv versions; peer-review status was not individually verified. Source pages and PDF access were checked on 2026-09-08; this does not guarantee enduring access or valid full-text conclusions.

  1. [1] Chen, Mark; Tworek, Jerry; Jun, Heewoo; et al.. Evaluating Large Language Models Trained on Code. arXiv:2107.03374 (2021). doi:10.48550/arXiv.2107.03374. PDF. ↩1
  2. [2] Austin, Jacob; Odena, Augustus; Nye, Maxwell; et al.. Program Synthesis with Large Language Models. arXiv:2108.07732 (2021). doi:10.48550/arXiv.2108.07732. PDF. ↩1
  3. [3] Li, Yujia; Choi, David; Chung, Junyoung; et al.. Competition-Level Code Generation with AlphaCode. arXiv:2203.07814 (2022). doi:10.48550/arXiv.2203.07814. PDF. ↩1
  4. [4] Pearce, Hammond; Ahmad, Baleegh; Tan, Benjamin; et al.. Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions. arXiv:2108.09293 (2021). doi:10.48550/arXiv.2108.09293. PDF. ↩1
  5. [5] Nijkamp, Erik; Pang, Bo; Hayashi, Hiroaki; et al.. CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. arXiv:2203.13474 (2022). doi:10.48550/arXiv.2203.13474. PDF. ↩1
  6. [6] Li, Raymond; Allal, Loubna Ben; Zi, Yangtian; et al.. StarCoder: may the source be with you!. arXiv:2305.06161 (2023). doi:10.48550/arXiv.2305.06161. PDF. ↩1
  7. [7] Rozière, Baptiste; Gehring, Jonas; Gloeckle, Fabian; et al.. Code Llama: Open Foundation Models for Code. arXiv:2308.12950 (2023). doi:10.48550/arXiv.2308.12950. PDF. ↩1
  8. [8] Liu, Jiawei; Xia, Chunqiu Steven; Wang, Yuyao; et al.. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. arXiv:2305.01210 (2023). doi:10.48550/arXiv.2305.01210. PDF. ↩1 ↩2
  9. [9] Jimenez, Carlos E.; Yang, John; Wettig, Alexander; et al.. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. arXiv:2310.06770 (2023). doi:10.48550/arXiv.2310.06770. PDF. ↩1
  10. [10] Yang, John; Jimenez, Carlos E.; Wettig, Alexander; et al.. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. arXiv:2405.15793 (2024). doi:10.48550/arXiv.2405.15793. PDF. ↩1
  11. [11] Xia, Chunqiu Steven; Deng, Yinlin; Dunn, Soren; et al.. Agentless: Demystifying LLM-based Software Engineering Agents. arXiv:2407.01489 (2024). doi:10.48550/arXiv.2407.01489. PDF. ↩1
  12. [12] Jain, Naman; Han, King; Gu, Alex; et al.. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. arXiv:2403.07974 (2024). doi:10.48550/arXiv.2403.07974. PDF. ↩1