AI+ · Working paper
Evaluating AI applications: a verification framework for writing, programming and workflows
Abstract
How can AI-generated content become dependable work? Through a narrative analysis of research on assisted writing, programming and tool-based workflows, this paper examines the gap between generation performance and task completion. It proposes a verification framework centred on factual evidence, independent tests and final task states. In writing, factual support should be evaluated alongside information completeness. In programming, detecting errors should be distinguished from completing effective repairs. In workflows, success on one attempt should be distinguished from reliability across repeated runs. These distinctions inform a proposed controlled evaluation, shifting attention from output volume to the verifiability of work and the cost of verification. This is a conceptual study with no original experiments; the effectiveness of the framework remains to be tested empirically.
1 From generative capability to delivery evidence
Generative systems produce fluent text, executable code and tool actions, but none automatically establishes successful delivery. We argue for shifting evaluation from what a model produces to which claims and states can be externally checked. Evidence support, test coverage and final-state verification are distinct levels of judgement. This is a conceptual synthesis and research agenda, not an experimentally validated new algorithm.
2 Literature scope and lines of evidence
This paper purposively selects 150 research documents for a question-led narrative synthesis, not a systematic review or meta-analysis. The selected works were first submitted in 2016–2025 and do not cover all developments through 2026. Titles, authors, submission dates, public abstracts, source pages and PDF access were checked on 8 September 2026. Original research, technical reports and surveys are distinguished rather than treated as equally strong experimental evidence. Selected methodological passages in FActScore, ALCE, EvalPlus and τ-bench were previously examined; we did not conduct full-text appraisal or experimental replication of every work.
The Transformer uses attention for sequence transduction, providing an architectural basis for subsequent language-generation research.[1] GPT-3 demonstrates few-shot task performance through textual examples, while reporting variation across tasks and methodological issues associated with training data.[2]
T5 frames language tasks as text-to-text problems and compares objectives, architectures and datasets; a common task format does not imply a common evaluation criterion.[3] BART studies generation and comprehension through denoising pre-training, illustrating the importance of the training objective.[4]
Retrieval-augmented generation combines parametric models with external retrieval memory for knowledge-intensive tasks.[5] REALM incorporates a knowledge retriever into language-model pre-training, exploring access to documents beyond facts stored in model parameters.[6]
InstructGPT uses demonstrations and preference feedback to improve instruction following, but its authors also report that the resulting models still make simple mistakes.[7] TruthfulQA evaluates questions that elicit common misconceptions, distinguishing imitation of language from factual truthfulness.[8]
FActScore decomposes long-form output into atomic facts and estimates the proportion supported by a reliable knowledge source.[9] ALCE evaluates fluency, correctness and citation quality separately: the presence of citations does not establish that every claim is supported.[10]
Chain-of-Verification separates drafting, verification questions, independent answers and revision, reducing hallucination in the tasks studied.[11] Lost in the Middle finds position-dependent use of relevant information in the models tested; a larger context window does not guarantee effective use of all supplied material.[12]
The Codex study evaluates functional correctness using tasks including HumanEval and examines repeated sampling and model limitations.[13] The study introducing MBPP evaluates synthesis of short Python programs from natural-language descriptions.[14]
AlphaCode combines large-scale candidate sampling with behaviour-based filtering for competitive programming, a different setting from maintaining an existing software system.[15] Pearce and colleagues find vulnerable code generated by an early version of Copilot in selected security-relevant scenarios; these historical observations do not estimate current product vulnerability rates.[16]
CodeGen studies multi-turn programme synthesis by decomposing requirements into successive subproblems.[17] StarCoder describes an openly released code-model effort, including licensing, personal-information handling and source-attribution arrangements.[18]
Code Llama offers variants for generation, completion and instruction following, highlighting differences in training and use conditions.[19] EvalPlus expands test inputs and identifies errors missed by earlier tests, demonstrating that test adequacy changes measured performance.[20]
SWE-bench evaluates issue resolution in real repositories, requiring coordinated understanding and changes across functions, classes or files.[21] SWE-agent examines how an agent–computer interface for navigation, editing and testing affects software-repair performance.[22]
Agentless uses a simplified localization, repair and patch-validation pipeline, motivating comparisons against simpler workflows.[23] LiveCodeBench collects new competition problems over time and includes self-repair, execution and test-output prediction to address contamination and evaluation coverage.[24]
Chain-of-Thought prompting improves performance on selected reasoning tasks through intermediate examples; generated reasoning is not itself external evidence.[25] Self-Consistency samples multiple reasoning paths and aggregates answers, improving performance in the reasoning tasks studied.[26]
ReAct interleaves reasoning and actions, allowing external knowledge or environmental feedback to inform subsequent decisions.[27] Toolformer studies how a model can learn when to invoke APIs, what arguments to use and how to incorporate their results.[28]
Reflexion stores textual reflections on feedback for subsequent attempts without updating model weights.[29] Self-Refine uses the same model to provide feedback and iteratively revise outputs across several tasks.[30]
Tree of Thoughts organizes intermediate reasoning as a process of exploration, evaluation and backtracking for planning or search tasks.[31] AutoGen provides configurable multi-agent conversations connecting models, tools and human input.[32]
AgentBench evaluates reasoning and decision-making in multiple interactive environments and analyses failures in long-horizon reasoning and instruction following.[33] WebArena provides reproducible web environments and long-horizon tasks evaluated by functional task completion.[34]
ToolLLM develops data, training and evaluation for API use, including tool retrieval and search over invocation paths.[35] τ-bench combines simulated users, domain policies and API tools, checking resulting database states and reliability across repeated trials.[36]
3 Extended foundations and counterevidence
The extension covers training and data conditions, retrieval and factual evaluation, software and tool environments, and privacy, bias and uncertainty. Each entry summarizes a research question or method supported by the public abstract, without importing unverified performance figures or treating a larger bibliography as validation of our framework. The cited arXiv version is not necessarily the final journal version; source existence does not guarantee correct conclusions.
3.1 Model and representation foundations
BERT learns bidirectional representations from left and right context and studies transfer to language-understanding tasks.[37] RoBERTa revisits BERT training to examine data and hyperparameter effects, motivating control of training conditions in model comparisons.[38]
XLNet learns bidirectional context through an autoregressive objective over factorization permutations, offering a different route from masked reconstruction.[39] The GPT-4 report describes image and text inputs, text outputs and benchmark evaluation, while acknowledging weaker-than-human performance in many real-world settings.[40]
LLaMA studies foundation models trained on publicly available data, providing a case for examining data accessibility alongside model capability.[41] Llama 2 distinguishes pretrained models from dialogue-tuned variants, making variant identification necessary when citing results.[42]
3.2 Scale, data and training conditions
Scaling Laws studies empirical relationships between cross-entropy loss, parameters, data and training compute, rather than dependable production delivery.[43] The Chinchilla study jointly allocates model size and training tokens under a compute budget, challenging explanations based on parameter count alone.[44]
PaLM examines scaling and few-shot performance alongside bias, toxicity and memorization.[45] OPT releases pretrained models at multiple scales together with training records to support research and inspection.[46]
BLOOM describes an open-access model trained on natural and programming languages, making training-language composition relevant to evaluation.[47] Pythia provides controlled data ordering and multiple checkpoints for studying capability, memorization and bias over training.[48]
3.3 Task adaptation and preference learning
LoRA freezes pretrained weights and trains low-rank update matrices to reduce trainable parameters during adaptation.[49] QLoRA combines a quantized frozen model with low-rank adapters, principally addressing fine-tuning memory requirements.[50]
Prefix-Tuning learns continuous task prefixes on a frozen model for lightweight adaptation to generation tasks.[51] Prompt Tuning learns soft prompts through backpropagation and studies how their effectiveness changes with model scale.[52]
DPO reformulates preference learning through a classification loss; preference alignment should still be evaluated separately from factual support and task completion.[53] Constitutional AI trains assistants using principles, self-critique and AI preference feedback, targeting behavioral constraints rather than universal factual guarantees.[54]
3.4 Data documentation, deduplication and provenance
Datasheets for Datasets proposes documenting motivation, composition, collection and intended uses, informing documentation of source materials.[55] Model Cards proposes reporting intended uses, evaluation procedures and performance across conditions rather than a single score.[56]
Research on training-data deduplication examines repeated text, memorized output and train–test overlap.[57] Kandpal and colleagues study duplicated training sequences and privacy attacks, finding that deduplication reduces attack effectiveness in their evaluated settings.[58]
Data Portraits uses compact records to support training-data membership queries and investigation of test leakage and provenance.[59] Dolma releases an English pretraining corpus and construction documentation, supporting investigation of data-curation effects.[60]
3.5 Benchmark coverage and evaluation methods
MMLU evaluates knowledge and problem solving across academic and professional tasks; its mean score is not accuracy across all applications.[61] BIG-bench assembles diverse tasks and compares performance across model scales, supporting evaluation beyond a single task.[62]
HELM evaluates accuracy, calibration, robustness, fairness, bias, toxicity and efficiency across multiple scenarios.[63] CheckList organizes behavioral tests by capabilities and test types, supplementing weaknesses obscured by average held-out accuracy.[64]
Dynabench collects challenge examples through human-and-model-in-the-loop data creation for dynamic evaluation.[65] Research on neural-network calibration tests correspondence between confidence and correctness and examines temperature scaling; its classification experiments do not guarantee calibrated text generation.[66]
3.6 Reasoning strategies and evidential limits
Zero-shot-CoT studies a short stepwise-reasoning prompt without task exemplars on selected reasoning benchmarks.[67] Least-to-Most decomposes complex problems into sequential subproblems that reuse earlier answers.[68]
PAL generates intermediate programs and delegates execution and computation to an interpreter.[69] Program of Thoughts likewise separates reasoning representation from external computation and evaluates numerical reasoning tasks.[70]
Research on unfaithful explanations finds that input biases can influence predictions without being disclosed in the generated chain of thought.[71] A study of intrinsic self-correction finds that revision without external feedback can fail or degrade performance in the reasoning settings tested.[72]
3.7 Retrievers and external knowledge
DPR uses a dual encoder to learn dense question and passage representations for evidence retrieval in open-domain question answering.[73] ColBERT uses late interaction between query and document representations to study retrieval effectiveness and computational efficiency.[74]
Research combining passage retrieval and generation examines aggregation of information from multiple passages for open-domain answers.[75] Atlas studies few-shot retrieval augmentation and examines the contents and updating of the document index.[76]
RETRO conditions token prediction on retrieved document chunks, studying explicit external memory alongside model parameters.[77] BEIR compares retrieval systems across heterogeneous tasks, cautioning against assuming cross-domain generalization from one dataset.[78]
3.8 Control and evaluation of retrieval augmentation
Self-RAG learns on-demand retrieval and reflection tokens to coordinate retrieval, generation and self-evaluation.[79] CRAG evaluates retrieved-document quality to trigger different retrieval actions, addressing failures of retrieval.[80]
IRCoT interleaves retrieval with multi-step reasoning so that later retrieval uses information already derived.[81] FLARE retrieves using an anticipated next sentence and regenerates it when low-confidence tokens indicate a need for evidence.[82]
RAPTOR recursively clusters and summarizes text chunks to build a retrieval tree at multiple abstraction levels.[83] HyDE generates a hypothetical document as a retrieval intermediary before searching a real corpus; that hypothetical document is not factual evidence to cite.[84]
3.9 Fact verification and summary faithfulness
FEVER labels claims as supported, refuted or lacking sufficient information and records evidence sentences where applicable.[85] The FactCC study jointly assesses summary consistency and locates supporting and conflicting spans.[86]
Maynez and colleagues examine hallucinations through human evaluation of summaries, distinguishing source faithfulness from general text-similarity metrics.[87] SummEval compares automatic metrics, summarization models, and expert and crowd annotations under a common evaluation setup.[88]
QAFactEval examines question generation and answerability in QA-based factual evaluation and their complementarity with entailment-based assessment.[89] TRUE standardizes data from multiple tasks for example-level factual-consistency evaluation, beyond correlations between system averages.[90]
3.10 Hallucination detection and semantic support
SelfCheckGPT uses agreement across sampled answers to identify suspect content; a shared repeated error does not thereby become an externally established fact.[91] HaluEval uses generated and human-annotated samples to test hallucination recognition; its observed rates cannot be extrapolated to all answers.[92]
Ji and colleagues survey definitions, evaluation and mitigation of hallucination in language generation; here it provides conceptual background, not an additional independent experiment.[93] Research on semantic uncertainty measures entropy over meanings rather than string differences alone, accounting for equivalent answer formulations.[94]
AlignScore learns information alignment between two text passages for factual-consistency assessment across settings.[95] Ragas separates retrieved context, faithfulness and output quality to support RAG evaluation without human reference answers.[96]
3.11 Instruction boundaries and safety evaluation
Research on indirect prompt injection shows that malicious instructions in retrieved content can interfere with applications; source content must not automatically acquire operational authority.[97] Research on transferable adversarial attacks studies input suffixes that bypass alignment constraints in tested models, motivating explicit threat models for safety evaluation.[98]
Instruction Hierarchy studies training models to distinguish instruction priorities when lower-trust content conflicts with higher-priority requirements.[99] HarmBench provides a standardized framework for comparing automated red teaming and evaluating attacks and defenses.[100]
JailbreakBench specifies threat models, prompt configurations, scoring and attack artifacts to improve comparability of jailbreak research.[101] SORRY-Bench evaluates safety refusal across fine-grained unsafe topics and linguistic variants rather than a narrow selection of topics.[102]
3.12 Code representation and program generation
CodeBERT learns representations of natural and programming languages and evaluates tasks including code search and documentation generation.[103] GraphCodeBERT incorporates data flow between variables during pretraining rather than treating code solely as a token sequence.[104]
CodeT5 uses an encoder–decoder architecture and identifier-aware objectives to study both code understanding and generation.[105] CodeT selects candidate programs using generated tests and execution agreement; the correctness of generated tests still requires separate scrutiny.[106]
InCoder studies code infilling with bidirectional context alongside left-to-right synthesis, distinguishing editing from continuation.[107] The study introducing PolyCoder compares code models across programming languages, highlighting code corpora and language distribution in evaluation.[108]
3.13 Cross-file context and execution benchmarks
RepoCoder uses iterative retrieval and generation to incorporate cross-file repository information into code completion.[109] RepoBench separates retrieval, code completion and combined-pipeline tasks, distinguishing context acquisition from context use.[110]
CrossCodeEval uses static analysis to construct multilingual completion tasks requiring cross-file context.[111] DS-1000 builds programming tasks around data-science libraries and evaluates execution results alongside constraints such as API usage.[112]
APPS evaluates program generation through natural-language coding problems of varying difficulty and associated test cases.[113] CRUXEval evaluates program input and output prediction separately, showing that generation scores and execution reasoning are distinct capabilities.[114]
3.14 Execution feedback, training and model versions
CodeRL predicts functional correctness with a critic and studies code generation using test feedback and reinforcement learning.[115] Self-Debugging studies code revision using program explanations and execution results; its feedback conditions differ from intrinsic correction without external information.[116]
LEVER learns a verifier from the natural-language request, candidate program and execution result to rerank candidates.[117] DeepSeek-Coder studies project-level code data and infilling objectives, identifying repository context and completion objectives as distinct training factors.[118]
StarCoder2 and The Stack v2 document code provenance and training-data construction alongside evaluation across languages and tasks.[119] WizardCoder applies Evol-Instruct to instruction tuning for code and evaluates code-generation benchmarks.[120]
3.15 Tool selection and interactive memory
Gorilla combines API-document retrieval with call generation and examines changes to documentation at test time.[121] API-Bank provides runnable tools and dialogue tasks to distinguish errors in API planning, retrieval and invocation.[122]
HuggingGPT uses a language model to coordinate task planning, model selection, execution and result summarization.[123] MRKL proposes a system architecture connecting language models, external knowledge and discrete reasoning modules.[124]
Voyager combines a curriculum, executable skill library and environmental feedback in Minecraft; game exploration does not directly establish reliability in real workflows.[125] Generative Agents uses memory, reflection and planning to simulate believable human behavior; behavioral believability differs from business-task correctness.[126]
3.16 Observable environments and task outcomes
Mind2Web collects tasks and action sequences from real websites to study generalization across sites and domains.[127] WorkArena defines knowledge-work tasks in enterprise software and provides an environment for interactive agent evaluation.[128]
OSWorld evaluates cross-application tasks with real computer environments, initial-state configurations and execution-based checking scripts.[129] VisualWebArena studies web tasks requiring both images and text, extending evaluation beyond text-only observations.[130]
AndroidWorld evaluates mobile agents with parameterized tasks and explicit initialization, success checks and teardown logic.[131] AgentDojo tests prompt injection in untrusted tool data, evaluating ordinary tasks alongside attacked conditions.[132]
3.17 Privacy, bias and harmful outputs
Training-data extraction research demonstrates recovery of training text from GPT-2 outputs, showing that publicly sourced inputs can still involve privacy risks.[133] Research on differentially private deep learning uses training mechanisms and privacy-cost analysis to bound leakage risk under the corresponding assumptions.[134]
StereoSet evaluates gender, profession, race and religion stereotypes in English; these categories are not treated here as an exhaustive Chinese-language evaluation.[135] CrowS-Pairs uses sentence pairs to examine masked language models' preferences for social stereotypes in a US context.[136]
RealToxicityPrompts investigates whether apparently innocuous prompts elicit toxic output and compares generation-control methods.[137] BBQ contrasts under-informative and informative question-answering contexts to examine how social bias affects answer selection.[138]
3.18 Uncertainty and selective answering
Research on model self-knowledge distinguishes estimating answer correctness from estimating whether a question is answerable and reports cross-task calibration difficulties.[139] Just Ask for Calibration compares calibration of verbalized confidence and conditional probabilities on selected question-answering benchmarks.[140]
Linguistic Calibration studies confidence in long-form text by asking whether users can form calibrated probabilistic judgments from the output.[141] Conformal Language Modeling calibrates sampling-stop and rejection rules for answer sets with statistical guarantees, not correctness of every single response.[142]
Selective-classification research trades rejection of some instances for risk control, informing joint reporting of coverage and error.[143] Research on risk-controlling prediction sets uses held-out data to calibrate set size and expected loss; its guarantees cannot be detached from the specified sampling and calibration conditions.[144]
3.19 Subsequent models and reasoning training
DeepSeek-R1 studies reinforcement learning for reasoning and performance on verifiable tasks; it does not establish that open-ended factual statements are necessarily correct.[145] Qwen3 reports thinking and non-thinking modes with inference-budget control, reinforcing the need to record both mode and compute budget.[146]
DeepSeek-V3 describes mixture-of-experts design, load balancing and multi-token prediction; active and total parameter counts should be recorded separately.[147] Llama 3 reports multilingual, coding, reasoning and tool-use evaluation while distinguishing pretraining, post-training and multimodal experiments.[148]
OLMo 2 provides weights, training data, logs and intermediate checkpoints, distinguishing open weights from availability of the broader training artifacts.[149] Qwen2.5 describes pretraining and post-training changes and multiple released variants, making model identity necessary when citing its results.[150]
4 An evidence framework across tasks
Rather than combine heterogeneous benchmark scores into a ranking, we distinguish three acceptance objects: claims in writing, patches in programming and state transitions in workflows. All can be recorded as task–constraint–artefact–check–final state, but this shared structure does not justify an unvalidated aggregate score. Good citations cannot replace code tests, and successful receipts cannot replace final-state checks.
| Task | Evidence object | Evaluation dimensions |
|---|---|---|
| Writing | Claims and source locations | Support, coverage and verification cost |
| Programming | Revision diffs and independent tests | Valid repairs, regressions and review cost |
| Workflows | Permission records and observable final states | Repeated reliability, takeover and recovery cost |
5 Falsifiable research propositions
H1: Does a claim–source ledger improve supported-claim precision without worsening required-fact coverage or abstention relative to baseline? Gains achieved only by omitting difficult claims would not support the hypothesis.
Use human writing, direct generation and retrieval-augmented generation as baselines. Add an evidence ledger and independent review in the full workflow; ablate each separately while fixing materials, model version, retrieval corpus and total time budget.
H2: At fixed patch scope and cost, does independent testing improve the proportion of valid final repairs, rather than merely reveal more errors? A higher detection rate without more usable final patches must be reported separately, not as better repair effectiveness.
Compare human repair, one-shot generation and a simple localization–repair–validation pipeline. Add independent tests and regression checks in the full workflow. Freeze benchmark revisions, isolate held-out tests from patch generation, and record dependencies and environment hashes.
H3: Can state verification and idempotent retries improve repeated reliability without broader permissions or increased human takeover? Benefits depending on extra intervention or authority cannot be attributed to workflow structure alone.
Compare a single workflow, multi-agent coordination and a state-verified workflow at matched models, tool permissions, budgets and tasks. Ablate state checks and retry protection separately; inject timeouts, duplicate responses and partial failures only in recoverable environments.
6 Operational definitions and decision boundaries
Let Sᵢ be supported checkable claims in task i and Aᵢ all checkable claims. Let Kᵢ be facts required in advance and Dᵢ required facts correctly retained in the output. Report supported-claim precision Pᵢ and coverage Cᵢ separately so that omission or silence cannot inflate the apparent result. A zero denominator is not applicable; abstentions and omissions are reported separately.
Count a code task as a valid repair only if the predefined issue is resolved, independent acceptance checks pass and required regression checks do not fail. Report security review and target-environment coverage separately. This operational definition is bounded by test adequacy, not proof of correctness for every input. Errors detected, repairs completed and review time are distinct outcomes.
Let yᵢⱼ∈{0,1} indicate that task i in run j satisfies both final-state and policy requirements, with N tasks and a fixed k runs per task. R_all,k is the fraction succeeding in every run; R_any,k is the fraction succeeding at least once. Report both with k, budget and takeover rate. Neither should be replaced by taking the overall mean success rate to the kth power.
7 A testable evaluation protocol
Treat tasks, not repeated outputs from one task, as sampling units. Use a separate pilot to assess annotation feasibility, cost and plausible effects, then determine the main sample size from precision or power targets. Exclude pilot tasks from held-out evaluation. Stratify Chinese and English tasks and treat translated counterparts as paired rather than independent samples. Publish inclusion and exclusion criteria and freeze inputs, versions, permissions and budgets.
Randomize condition order and compare conditions within tasks; blind adjudicators to condition where feasible. Report raw counts, task-level effect differences and uncertainty intervals without treating repeat runs as extra independent tasks. Report task families separately and test sensitivity to missing outcomes and timeouts. Retain and distinguish timeouts, abstentions and human takeovers. Document when and why a planned procedure changes.
A reproducibility package should identify tasks, data permissions, model and tool versions, prompt templates, configurable random seeds, tests, adjudication rules, failure logs and analysis scripts. Reproducibility does not justify releasing personal or restricted information; redacted structures and controlled-access terms may be appropriate. These are proposed research outputs: this site does not currently provide a completed experimental dataset or analysis code.
8 Counterarguments, limitations and scope
More verification is not necessarily better: a reliable ledger can increase review costs, defective tests can reject correct patches, and final-state checks can miss implicit requirements. Lightweight or human-only workflows may be more economical for low-risk, reversible tasks. If additional checks increase cost without improving predefined outcomes, the proposed workflow should be reduced or rejected rather than defended through post-hoc changes to success criteria.
FActScore primarily evaluates biographies against a corresponding knowledge source; it does not directly establish the truth of complex causal arguments. ALCE’s automated entailment judgements can also fail. The proportions proposed below are operational definitions for this protocol, not full reproductions of either benchmark.[9][10]
EvalPlus shows that expanded tests can reveal errors missed by earlier evaluation and also discusses defective reference implementations. More tests do not prove correctness for every input; test validity and final-patch correctness require separate scrutiny.[20]
τ-bench checks final database states together with required responses. We adopt its distinction between at-least-once and all-trial success, but the equations below are direct empirical statistics over a fixed number of trials, not the paper’s combinatorial estimator. Descriptive computation does not require independent and identically distributed trials.[36]
Purposive selection, predominantly abstract-level checking and the absence of original data limit this paper. The present reference set cannot establish a research gap through 2026 or support priority claims. Before submission, searches need updating, key methods require full-text appraisal, version and publication information need checking, and domain researchers must assess the argument. Claims of effectiveness additionally require the proposed studies.
9 Conclusion
Dependability in AI applications should not be defined by generation performance alone. This paper proposes separate acceptance criteria for evidence support, independent tests and task completion, alongside falsifiable hypotheses and bounded evaluation measures. Its contribution is to specify what must be tested, not to declare the proposed approach effective.
10 Companion papers
- Evidence-constrained AI writing: from fluent text to verifiable claims
- AI-assisted programming: from generated code to validated repository changes
- AI workflows: tool coordination and repeatable task completion
11 Declarations and availability
Data and code: No original experimental data were generated. Sources are individually linked. Experimental and analysis code has not been implemented; website code is not an experimental replication package.
Authorship, contributions, funding and competing interests: Alsay is the website publisher, which does not establish scholarly authorship. Author identities, contributions, funding and competing interests require confirmation by accountable contributors. This draft does not assert absence of competing interests or invent affiliations, degrees or ethical approvals.
AI assistance and versions: AI assisted retrieval, organization, editing and Chinese–English translation. The two versions represent one working paper, not two independent studies; scholarly approval by accountable authors remains pending. No peer review, journal acceptance or completed preregistration is claimed.
References
Cited as public arXiv versions; peer-review status was not individually verified. Source pages and PDF access were checked on 2026-09-08; this does not guarantee enduring access or valid full-text conclusions.
- [1] Vaswani, Ashish; Shazeer, Noam; Parmar, Niki; et al.. Attention Is All You Need. arXiv:1706.03762 (2017). doi:10.48550/arXiv.1706.03762. PDF. ↩1
- [2] Brown, Tom B.; Mann, Benjamin; Ryder, Nick; et al.. Language Models are Few-Shot Learners. arXiv:2005.14165 (2020). doi:10.48550/arXiv.2005.14165. PDF. ↩1
- [3] Raffel, Colin; Shazeer, Noam; Roberts, Adam; et al.. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 (2019). doi:10.48550/arXiv.1910.10683. PDF. ↩1
- [4] Lewis, Mike; Liu, Yinhan; Goyal, Naman; et al.. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. arXiv:1910.13461 (2019). doi:10.48550/arXiv.1910.13461. PDF. ↩1
- [5] Lewis, Patrick; Perez, Ethan; Piktus, Aleksandra; et al.. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv:2005.11401 (2020). doi:10.48550/arXiv.2005.11401. PDF. ↩1
- [6] Guu, Kelvin; Lee, Kenton; Tung, Zora; et al.. REALM: Retrieval-Augmented Language Model Pre-Training. arXiv:2002.08909 (2020). doi:10.48550/arXiv.2002.08909. PDF. ↩1
- [7] Ouyang, Long; Wu, Jeff; Jiang, Xu; et al.. Training language models to follow instructions with human feedback. arXiv:2203.02155 (2022). doi:10.48550/arXiv.2203.02155. PDF. ↩1
- [8] Lin, Stephanie; Hilton, Jacob; Evans, Owain. TruthfulQA: Measuring How Models Mimic Human Falsehoods. arXiv:2109.07958 (2021). doi:10.48550/arXiv.2109.07958. PDF. ↩1
- [9] Min, Sewon; Krishna, Kalpesh; Lyu, Xinxi; et al.. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. arXiv:2305.14251 (2023). doi:10.48550/arXiv.2305.14251. PDF. ↩1 ↩2
- [10] Gao, Tianyu; Yen, Howard; Yu, Jiatong; et al.. Enabling Large Language Models to Generate Text with Citations. arXiv:2305.14627 (2023). doi:10.48550/arXiv.2305.14627. PDF. ↩1 ↩2
- [11] Dhuliawala, Shehzaad; Komeili, Mojtaba; Xu, Jing; et al.. Chain-of-Verification Reduces Hallucination in Large Language Models. arXiv:2309.11495 (2023). doi:10.48550/arXiv.2309.11495. PDF. ↩1
- [12] Liu, Nelson F.; Lin, Kevin; Hewitt, John; et al.. Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172 (2023). doi:10.48550/arXiv.2307.03172. PDF. ↩1
- [13] Chen, Mark; Tworek, Jerry; Jun, Heewoo; et al.. Evaluating Large Language Models Trained on Code. arXiv:2107.03374 (2021). doi:10.48550/arXiv.2107.03374. PDF. ↩1
- [14] Austin, Jacob; Odena, Augustus; Nye, Maxwell; et al.. Program Synthesis with Large Language Models. arXiv:2108.07732 (2021). doi:10.48550/arXiv.2108.07732. PDF. ↩1
- [15] Li, Yujia; Choi, David; Chung, Junyoung; et al.. Competition-Level Code Generation with AlphaCode. arXiv:2203.07814 (2022). doi:10.48550/arXiv.2203.07814. PDF. ↩1
- [16] Pearce, Hammond; Ahmad, Baleegh; Tan, Benjamin; et al.. Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions. arXiv:2108.09293 (2021). doi:10.48550/arXiv.2108.09293. PDF. ↩1
- [17] Nijkamp, Erik; Pang, Bo; Hayashi, Hiroaki; et al.. CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. arXiv:2203.13474 (2022). doi:10.48550/arXiv.2203.13474. PDF. ↩1
- [18] Li, Raymond; Allal, Loubna Ben; Zi, Yangtian; et al.. StarCoder: may the source be with you!. arXiv:2305.06161 (2023). doi:10.48550/arXiv.2305.06161. PDF. ↩1
- [19] Rozière, Baptiste; Gehring, Jonas; Gloeckle, Fabian; et al.. Code Llama: Open Foundation Models for Code. arXiv:2308.12950 (2023). doi:10.48550/arXiv.2308.12950. PDF. ↩1
- [20] Liu, Jiawei; Xia, Chunqiu Steven; Wang, Yuyao; et al.. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. arXiv:2305.01210 (2023). doi:10.48550/arXiv.2305.01210. PDF. ↩1 ↩2
- [21] Jimenez, Carlos E.; Yang, John; Wettig, Alexander; et al.. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. arXiv:2310.06770 (2023). doi:10.48550/arXiv.2310.06770. PDF. ↩1
- [22] Yang, John; Jimenez, Carlos E.; Wettig, Alexander; et al.. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. arXiv:2405.15793 (2024). doi:10.48550/arXiv.2405.15793. PDF. ↩1
- [23] Xia, Chunqiu Steven; Deng, Yinlin; Dunn, Soren; et al.. Agentless: Demystifying LLM-based Software Engineering Agents. arXiv:2407.01489 (2024). doi:10.48550/arXiv.2407.01489. PDF. ↩1
- [24] Jain, Naman; Han, King; Gu, Alex; et al.. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. arXiv:2403.07974 (2024). doi:10.48550/arXiv.2403.07974. PDF. ↩1
- [25] Wei, Jason; Wang, Xuezhi; Schuurmans, Dale; et al.. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv:2201.11903 (2022). doi:10.48550/arXiv.2201.11903. PDF. ↩1
- [26] Wang, Xuezhi; Wei, Jason; Schuurmans, Dale; et al.. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171 (2022). doi:10.48550/arXiv.2203.11171. PDF. ↩1
- [27] Yao, Shunyu; Zhao, Jeffrey; Yu, Dian; et al.. ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629 (2022). doi:10.48550/arXiv.2210.03629. PDF. ↩1
- [28] Schick, Timo; Dwivedi-Yu, Jane; Dessì, Roberto; et al.. Toolformer: Language Models Can Teach Themselves to Use Tools. arXiv:2302.04761 (2023). doi:10.48550/arXiv.2302.04761. PDF. ↩1
- [29] Shinn, Noah; Cassano, Federico; Berman, Edward; et al.. Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366 (2023). doi:10.48550/arXiv.2303.11366. PDF. ↩1
- [30] Madaan, Aman; Tandon, Niket; Gupta, Prakhar; et al.. Self-Refine: Iterative Refinement with Self-Feedback. arXiv:2303.17651 (2023). doi:10.48550/arXiv.2303.17651. PDF. ↩1
- [31] Yao, Shunyu; Yu, Dian; Zhao, Jeffrey; et al.. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv:2305.10601 (2023). doi:10.48550/arXiv.2305.10601. PDF. ↩1
- [32] Wu, Qingyun; Bansal, Gagan; Zhang, Jieyu; et al.. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv:2308.08155 (2023). doi:10.48550/arXiv.2308.08155. PDF. ↩1
- [33] Liu, Xiao; Yu, Hao; Zhang, Hanchen; et al.. AgentBench: Evaluating LLMs as Agents. arXiv:2308.03688 (2023). doi:10.48550/arXiv.2308.03688. PDF. ↩1
- [34] Zhou, Shuyan; Xu, Frank F.; Zhu, Hao; et al.. WebArena: A Realistic Web Environment for Building Autonomous Agents. arXiv:2307.13854 (2023). doi:10.48550/arXiv.2307.13854. PDF. ↩1
- [35] Qin, Yujia; Liang, Shihao; Ye, Yining; et al.. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. arXiv:2307.16789 (2023). doi:10.48550/arXiv.2307.16789. PDF. ↩1
- [36] Yao, Shunyu; Shinn, Noah; Razavi, Pedram; et al.. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045 (2024). doi:10.48550/arXiv.2406.12045. PDF. ↩1 ↩2
- [37] Devlin, Jacob; Chang, Ming-Wei; Lee, Kenton; et al.. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 (2018). doi:10.48550/arXiv.1810.04805. PDF. ↩1
- [38] Liu, Yinhan; Ott, Myle; Goyal, Naman; et al.. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 (2019). doi:10.48550/arXiv.1907.11692. PDF. ↩1
- [39] Yang, Zhilin; Dai, Zihang; Yang, Yiming; et al.. XLNet: Generalized Autoregressive Pretraining for Language Understanding. arXiv:1906.08237 (2019). doi:10.48550/arXiv.1906.08237. PDF. ↩1
- [40] OpenAI; Achiam, Josh; Adler, Steven; et al.. GPT-4 Technical Report. arXiv:2303.08774 (2023). doi:10.48550/arXiv.2303.08774. PDF. ↩1
- [41] Touvron, Hugo; Lavril, Thibaut; Izacard, Gautier; et al.. LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971 (2023). doi:10.48550/arXiv.2302.13971. PDF. ↩1
- [42] Touvron, Hugo; Martin, Louis; Stone, Kevin; et al.. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv:2307.09288 (2023). doi:10.48550/arXiv.2307.09288. PDF. ↩1
- [43] Kaplan, Jared; McCandlish, Sam; Henighan, Tom; et al.. Scaling Laws for Neural Language Models. arXiv:2001.08361 (2020). doi:10.48550/arXiv.2001.08361. PDF. ↩1
- [44] Hoffmann, Jordan; Borgeaud, Sebastian; Mensch, Arthur; et al.. Training Compute-Optimal Large Language Models. arXiv:2203.15556 (2022). doi:10.48550/arXiv.2203.15556. PDF. ↩1
- [45] Chowdhery, Aakanksha; Narang, Sharan; Devlin, Jacob; et al.. PaLM: Scaling Language Modeling with Pathways. arXiv:2204.02311 (2022). doi:10.48550/arXiv.2204.02311. PDF. ↩1
- [46] Zhang, Susan; Roller, Stephen; Goyal, Naman; et al.. OPT: Open Pre-trained Transformer Language Models. arXiv:2205.01068 (2022). doi:10.48550/arXiv.2205.01068. PDF. ↩1
- [47] Workshop, BigScience; :; Scao, Teven Le; et al.. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model. arXiv:2211.05100 (2022). doi:10.48550/arXiv.2211.05100. PDF. ↩1
- [48] Biderman, Stella; Schoelkopf, Hailey; Anthony, Quentin; et al.. Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling. arXiv:2304.01373 (2023). doi:10.48550/arXiv.2304.01373. PDF. ↩1
- [49] Hu, Edward J.; Shen, Yelong; Wallis, Phillip; et al.. LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685 (2021). doi:10.48550/arXiv.2106.09685. PDF. ↩1
- [50] Dettmers, Tim; Pagnoni, Artidoro; Holtzman, Ari; et al.. QLoRA: Efficient Finetuning of Quantized LLMs. arXiv:2305.14314 (2023). doi:10.48550/arXiv.2305.14314. PDF. ↩1
- [51] Li, Xiang Lisa; Liang, Percy. Prefix-Tuning: Optimizing Continuous Prompts for Generation. arXiv:2101.00190 (2021). doi:10.48550/arXiv.2101.00190. PDF. ↩1
- [52] Lester, Brian; Al-Rfou, Rami; Constant, Noah. The Power of Scale for Parameter-Efficient Prompt Tuning. arXiv:2104.08691 (2021). doi:10.48550/arXiv.2104.08691. PDF. ↩1
- [53] Rafailov, Rafael; Sharma, Archit; Mitchell, Eric; et al.. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv:2305.18290 (2023). doi:10.48550/arXiv.2305.18290. PDF. ↩1
- [54] Bai, Yuntao; Kadavath, Saurav; Kundu, Sandipan; et al.. Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073 (2022). doi:10.48550/arXiv.2212.08073. PDF. ↩1
- [55] Gebru, Timnit; Morgenstern, Jamie; Vecchione, Briana; et al.. Datasheets for Datasets. arXiv:1803.09010 (2018). doi:10.48550/arXiv.1803.09010. PDF. ↩1
- [56] Mitchell, Margaret; Wu, Simone; Zaldivar, Andrew; et al.. Model Cards for Model Reporting. arXiv:1810.03993 (2018). doi:10.48550/arXiv.1810.03993. PDF. ↩1
- [57] Lee, Katherine; Ippolito, Daphne; Nystrom, Andrew; et al.. Deduplicating Training Data Makes Language Models Better. arXiv:2107.06499 (2021). doi:10.48550/arXiv.2107.06499. PDF. ↩1
- [58] Kandpal, Nikhil; Wallace, Eric; Raffel, Colin. Deduplicating Training Data Mitigates Privacy Risks in Language Models. arXiv:2202.06539 (2022). doi:10.48550/arXiv.2202.06539. PDF. ↩1
- [59] Marone, Marc; Van Durme, Benjamin. Data Portraits: Recording Foundation Model Training Data. arXiv:2303.03919 (2023). doi:10.48550/arXiv.2303.03919. PDF. ↩1
- [60] Soldaini, Luca; Kinney, Rodney; Bhagia, Akshita; et al.. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research. arXiv:2402.00159 (2024). doi:10.48550/arXiv.2402.00159. PDF. ↩1
- [61] Hendrycks, Dan; Burns, Collin; Basart, Steven; et al.. Measuring Massive Multitask Language Understanding. arXiv:2009.03300 (2020). doi:10.48550/arXiv.2009.03300. PDF. ↩1
- [62] Srivastava, Aarohi; Rastogi, Abhinav; Rao, Abhishek; et al.. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models. arXiv:2206.04615 (2022). doi:10.48550/arXiv.2206.04615. PDF. ↩1
- [63] Liang, Percy; Bommasani, Rishi; Lee, Tony; et al.. Holistic Evaluation of Language Models. arXiv:2211.09110 (2022). doi:10.48550/arXiv.2211.09110. PDF. ↩1
- [64] Ribeiro, Marco Tulio; Wu, Tongshuang; Guestrin, Carlos; et al.. Beyond Accuracy: Behavioral Testing of NLP models with CheckList. arXiv:2005.04118 (2020). doi:10.48550/arXiv.2005.04118. PDF. ↩1
- [65] Kiela, Douwe; Bartolo, Max; Nie, Yixin; et al.. Dynabench: Rethinking Benchmarking in NLP. arXiv:2104.14337 (2021). doi:10.48550/arXiv.2104.14337. PDF. ↩1
- [66] Guo, Chuan; Pleiss, Geoff; Sun, Yu; et al.. On Calibration of Modern Neural Networks. arXiv:1706.04599 (2017). doi:10.48550/arXiv.1706.04599. PDF. ↩1
- [67] Kojima, Takeshi; Gu, Shixiang Shane; Reid, Machel; et al.. Large Language Models are Zero-Shot Reasoners. arXiv:2205.11916 (2022). doi:10.48550/arXiv.2205.11916. PDF. ↩1
- [68] Zhou, Denny; Schärli, Nathanael; Hou, Le; et al.. Least-to-Most Prompting Enables Complex Reasoning in Large Language Models. arXiv:2205.10625 (2022). doi:10.48550/arXiv.2205.10625. PDF. ↩1
- [69] Gao, Luyu; Madaan, Aman; Zhou, Shuyan; et al.. PAL: Program-aided Language Models. arXiv:2211.10435 (2022). doi:10.48550/arXiv.2211.10435. PDF. ↩1
- [70] Chen, Wenhu; Ma, Xueguang; Wang, Xinyi; et al.. Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks. arXiv:2211.12588 (2022). doi:10.48550/arXiv.2211.12588. PDF. ↩1
- [71] Turpin, Miles; Michael, Julian; Perez, Ethan; et al.. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. arXiv:2305.04388 (2023). doi:10.48550/arXiv.2305.04388. PDF. ↩1
- [72] Huang, Jie; Chen, Xinyun; Mishra, Swaroop; et al.. Large Language Models Cannot Self-Correct Reasoning Yet. arXiv:2310.01798 (2023). doi:10.48550/arXiv.2310.01798. PDF. ↩1
- [73] Karpukhin, Vladimir; Oğuz, Barlas; Min, Sewon; et al.. Dense Passage Retrieval for Open-Domain Question Answering. arXiv:2004.04906 (2020). doi:10.48550/arXiv.2004.04906. PDF. ↩1
- [74] Khattab, Omar; Zaharia, Matei. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. arXiv:2004.12832 (2020). doi:10.48550/arXiv.2004.12832. PDF. ↩1
- [75] Izacard, Gautier; Grave, Edouard. Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. arXiv:2007.01282 (2020). doi:10.48550/arXiv.2007.01282. PDF. ↩1
- [76] Izacard, Gautier; Lewis, Patrick; Lomeli, Maria; et al.. Atlas: Few-shot Learning with Retrieval Augmented Language Models. arXiv:2208.03299 (2022). doi:10.48550/arXiv.2208.03299. PDF. ↩1
- [77] Borgeaud, Sebastian; Mensch, Arthur; Hoffmann, Jordan; et al.. Improving language models by retrieving from trillions of tokens. arXiv:2112.04426 (2021). doi:10.48550/arXiv.2112.04426. PDF. ↩1
- [78] Thakur, Nandan; Reimers, Nils; Rücklé, Andreas; et al.. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models. arXiv:2104.08663 (2021). doi:10.48550/arXiv.2104.08663. PDF. ↩1
- [79] Asai, Akari; Wu, Zeqiu; Wang, Yizhong; et al.. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. arXiv:2310.11511 (2023). doi:10.48550/arXiv.2310.11511. PDF. ↩1
- [80] Yan, Shi-Qi; Gu, Jia-Chen; Zhu, Yun; et al.. Corrective Retrieval Augmented Generation. arXiv:2401.15884 (2024). doi:10.48550/arXiv.2401.15884. PDF. ↩1
- [81] Trivedi, Harsh; Balasubramanian, Niranjan; Khot, Tushar; et al.. Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions. arXiv:2212.10509 (2022). doi:10.48550/arXiv.2212.10509. PDF. ↩1
- [82] Jiang, Zhengbao; Xu, Frank F.; Gao, Luyu; et al.. Active Retrieval Augmented Generation. arXiv:2305.06983 (2023). doi:10.48550/arXiv.2305.06983. PDF. ↩1
- [83] Sarthi, Parth; Abdullah, Salman; Tuli, Aditi; et al.. RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval. arXiv:2401.18059 (2024). doi:10.48550/arXiv.2401.18059. PDF. ↩1
- [84] Gao, Luyu; Ma, Xueguang; Lin, Jimmy; et al.. Precise Zero-Shot Dense Retrieval without Relevance Labels. arXiv:2212.10496 (2022). doi:10.48550/arXiv.2212.10496. PDF. ↩1
- [85] Thorne, James; Vlachos, Andreas; Christodoulopoulos, Christos; et al.. FEVER: a large-scale dataset for Fact Extraction and VERification. arXiv:1803.05355 (2018). doi:10.48550/arXiv.1803.05355. PDF. ↩1
- [86] Kryściński, Wojciech; McCann, Bryan; Xiong, Caiming; et al.. Evaluating the Factual Consistency of Abstractive Text Summarization. arXiv:1910.12840 (2019). doi:10.48550/arXiv.1910.12840. PDF. ↩1
- [87] Maynez, Joshua; Narayan, Shashi; Bohnet, Bernd; et al.. On Faithfulness and Factuality in Abstractive Summarization. arXiv:2005.00661 (2020). doi:10.48550/arXiv.2005.00661. PDF. ↩1
- [88] Fabbri, Alexander R.; Kryściński, Wojciech; McCann, Bryan; et al.. SummEval: Re-evaluating Summarization Evaluation. arXiv:2007.12626 (2020). doi:10.48550/arXiv.2007.12626. PDF. ↩1
- [89] Fabbri, Alexander R.; Wu, Chien-Sheng; Liu, Wenhao; et al.. QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization. arXiv:2112.08542 (2021). doi:10.48550/arXiv.2112.08542. PDF. ↩1
- [90] Honovich, Or; Aharoni, Roee; Herzig, Jonathan; et al.. TRUE: Re-evaluating Factual Consistency Evaluation. arXiv:2204.04991 (2022). doi:10.48550/arXiv.2204.04991. PDF. ↩1
- [91] Manakul, Potsawee; Liusie, Adian; Gales, Mark J. F.. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. arXiv:2303.08896 (2023). doi:10.48550/arXiv.2303.08896. PDF. ↩1
- [92] Li, Junyi; Cheng, Xiaoxue; Zhao, Wayne Xin; et al.. HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. arXiv:2305.11747 (2023). doi:10.48550/arXiv.2305.11747. PDF. ↩1
- [93] Ji, Ziwei; Lee, Nayeon; Frieske, Rita; et al.. Survey of Hallucination in Natural Language Generation. arXiv:2202.03629 (2022). doi:10.48550/arXiv.2202.03629. PDF. ↩1
- [94] Kuhn, Lorenz; Gal, Yarin; Farquhar, Sebastian. Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation. arXiv:2302.09664 (2023). doi:10.48550/arXiv.2302.09664. PDF. ↩1
- [95] Zha, Yuheng; Yang, Yichi; Li, Ruichen; et al.. AlignScore: Evaluating Factual Consistency with a Unified Alignment Function. arXiv:2305.16739 (2023). doi:10.48550/arXiv.2305.16739. PDF. ↩1
- [96] Es, Shahul; James, Jithin; Espinosa-Anke, Luis; et al.. Ragas: Automated Evaluation of Retrieval Augmented Generation. arXiv:2309.15217 (2023). doi:10.48550/arXiv.2309.15217. PDF. ↩1
- [97] Greshake, Kai; Abdelnabi, Sahar; Mishra, Shailesh; et al.. Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. arXiv:2302.12173 (2023). doi:10.48550/arXiv.2302.12173. PDF. ↩1
- [98] Zou, Andy; Wang, Zifan; Carlini, Nicholas; et al.. Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv:2307.15043 (2023). doi:10.48550/arXiv.2307.15043. PDF. ↩1
- [99] Wallace, Eric; Xiao, Kai; Leike, Reimar; et al.. The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions. arXiv:2404.13208 (2024). doi:10.48550/arXiv.2404.13208. PDF. ↩1
- [100] Mazeika, Mantas; Phan, Long; Yin, Xuwang; et al.. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. arXiv:2402.04249 (2024). doi:10.48550/arXiv.2402.04249. PDF. ↩1
- [101] Chao, Patrick; Debenedetti, Edoardo; Robey, Alexander; et al.. JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models. arXiv:2404.01318 (2024). doi:10.48550/arXiv.2404.01318. PDF. ↩1
- [102] Xie, Tinghao; Qi, Xiangyu; Zeng, Yi; et al.. SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal. arXiv:2406.14598 (2024). doi:10.48550/arXiv.2406.14598. PDF. ↩1
- [103] Feng, Zhangyin; Guo, Daya; Tang, Duyu; et al.. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. arXiv:2002.08155 (2020). doi:10.48550/arXiv.2002.08155. PDF. ↩1
- [104] Guo, Daya; Ren, Shuo; Lu, Shuai; et al.. GraphCodeBERT: Pre-training Code Representations with Data Flow. arXiv:2009.08366 (2020). doi:10.48550/arXiv.2009.08366. PDF. ↩1
- [105] Wang, Yue; Wang, Weishi; Joty, Shafiq; et al.. CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. arXiv:2109.00859 (2021). doi:10.48550/arXiv.2109.00859. PDF. ↩1
- [106] Chen, Bei; Zhang, Fengji; Nguyen, Anh; et al.. CodeT: Code Generation with Generated Tests. arXiv:2207.10397 (2022). doi:10.48550/arXiv.2207.10397. PDF. ↩1
- [107] Fried, Daniel; Aghajanyan, Armen; Lin, Jessy; et al.. InCoder: A Generative Model for Code Infilling and Synthesis. arXiv:2204.05999 (2022). doi:10.48550/arXiv.2204.05999. PDF. ↩1
- [108] Xu, Frank F.; Alon, Uri; Neubig, Graham; et al.. A Systematic Evaluation of Large Language Models of Code. arXiv:2202.13169 (2022). doi:10.48550/arXiv.2202.13169. PDF. ↩1
- [109] Zhang, Fengji; Chen, Bei; Zhang, Yue; et al.. RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation. arXiv:2303.12570 (2023). doi:10.48550/arXiv.2303.12570. PDF. ↩1
- [110] Liu, Tianyang; Xu, Canwen; McAuley, Julian. RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems. arXiv:2306.03091 (2023). doi:10.48550/arXiv.2306.03091. PDF. ↩1
- [111] Ding, Yangruibo; Wang, Zijian; Ahmad, Wasi Uddin; et al.. CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion. arXiv:2310.11248 (2023). doi:10.48550/arXiv.2310.11248. PDF. ↩1
- [112] Lai, Yuhang; Li, Chengxi; Wang, Yiming; et al.. DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation. arXiv:2211.11501 (2022). doi:10.48550/arXiv.2211.11501. PDF. ↩1
- [113] Hendrycks, Dan; Basart, Steven; Kadavath, Saurav; et al.. Measuring Coding Challenge Competence With APPS. arXiv:2105.09938 (2021). doi:10.48550/arXiv.2105.09938. PDF. ↩1
- [114] Gu, Alex; Rozière, Baptiste; Leather, Hugh; et al.. CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution. arXiv:2401.03065 (2024). doi:10.48550/arXiv.2401.03065. PDF. ↩1
- [115] Le, Hung; Wang, Yue; Gotmare, Akhilesh Deepak; et al.. CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning. arXiv:2207.01780 (2022). doi:10.48550/arXiv.2207.01780. PDF. ↩1
- [116] Chen, Xinyun; Lin, Maxwell; Schärli, Nathanael; et al.. Teaching Large Language Models to Self-Debug. arXiv:2304.05128 (2023). doi:10.48550/arXiv.2304.05128. PDF. ↩1
- [117] Ni, Ansong; Iyer, Srini; Radev, Dragomir; et al.. LEVER: Learning to Verify Language-to-Code Generation with Execution. arXiv:2302.08468 (2023). doi:10.48550/arXiv.2302.08468. PDF. ↩1
- [118] Guo, Daya; Zhu, Qihao; Yang, Dejian; et al.. DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence. arXiv:2401.14196 (2024). doi:10.48550/arXiv.2401.14196. PDF. ↩1
- [119] Lozhkov, Anton; Li, Raymond; Allal, Loubna Ben; et al.. StarCoder 2 and The Stack v2: The Next Generation. arXiv:2402.19173 (2024). doi:10.48550/arXiv.2402.19173. PDF. ↩1
- [120] Luo, Ziyang; Xu, Can; Zhao, Pu; et al.. WizardCoder: Empowering Code Large Language Models with Evol-Instruct. arXiv:2306.08568 (2023). doi:10.48550/arXiv.2306.08568. PDF. ↩1
- [121] Patil, Shishir G.; Zhang, Tianjun; Wang, Xin; et al.. Gorilla: Large Language Model Connected with Massive APIs. arXiv:2305.15334 (2023). doi:10.48550/arXiv.2305.15334. PDF. ↩1
- [122] Li, Minghao; Zhao, Yingxiu; Yu, Bowen; et al.. API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs. arXiv:2304.08244 (2023). doi:10.48550/arXiv.2304.08244. PDF. ↩1
- [123] Shen, Yongliang; Song, Kaitao; Tan, Xu; et al.. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. arXiv:2303.17580 (2023). doi:10.48550/arXiv.2303.17580. PDF. ↩1
- [124] Karpas, Ehud; Abend, Omri; Belinkov, Yonatan; et al.. MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning. arXiv:2205.00445 (2022). doi:10.48550/arXiv.2205.00445. PDF. ↩1
- [125] Wang, Guanzhi; Xie, Yuqi; Jiang, Yunfan; et al.. Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv:2305.16291 (2023). doi:10.48550/arXiv.2305.16291. PDF. ↩1
- [126] Park, Joon Sung; O'Brien, Joseph C.; Cai, Carrie J.; et al.. Generative Agents: Interactive Simulacra of Human Behavior. arXiv:2304.03442 (2023). doi:10.48550/arXiv.2304.03442. PDF. ↩1
- [127] Deng, Xiang; Gu, Yu; Zheng, Boyuan; et al.. Mind2Web: Towards a Generalist Agent for the Web. arXiv:2306.06070 (2023). doi:10.48550/arXiv.2306.06070. PDF. ↩1
- [128] Drouin, Alexandre; Gasse, Maxime; Caccia, Massimo; et al.. WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?. arXiv:2403.07718 (2024). doi:10.48550/arXiv.2403.07718. PDF. ↩1
- [129] Xie, Tianbao; Zhang, Danyang; Chen, Jixuan; et al.. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. arXiv:2404.07972 (2024). doi:10.48550/arXiv.2404.07972. PDF. ↩1
- [130] Koh, Jing Yu; Lo, Robert; Jang, Lawrence; et al.. VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks. arXiv:2401.13649 (2024). doi:10.48550/arXiv.2401.13649. PDF. ↩1
- [131] Rawles, Christopher; Clinckemaillie, Sarah; Chang, Yifan; et al.. AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents. arXiv:2405.14573 (2024). doi:10.48550/arXiv.2405.14573. PDF. ↩1
- [132] Debenedetti, Edoardo; Zhang, Jie; Balunović, Mislav; et al.. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. arXiv:2406.13352 (2024). doi:10.48550/arXiv.2406.13352. PDF. ↩1
- [133] Carlini, Nicholas; Tramer, Florian; Wallace, Eric; et al.. Extracting Training Data from Large Language Models. arXiv:2012.07805 (2020). doi:10.48550/arXiv.2012.07805. PDF. ↩1
- [134] Abadi, Martín; Chu, Andy; Goodfellow, Ian; et al.. Deep Learning with Differential Privacy. arXiv:1607.00133 (2016). doi:10.48550/arXiv.1607.00133. PDF. ↩1
- [135] Nadeem, Moin; Bethke, Anna; Reddy, Siva. StereoSet: Measuring stereotypical bias in pretrained language models. arXiv:2004.09456 (2020). doi:10.48550/arXiv.2004.09456. PDF. ↩1
- [136] Nangia, Nikita; Vania, Clara; Bhalerao, Rasika; et al.. CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models. arXiv:2010.00133 (2020). doi:10.48550/arXiv.2010.00133. PDF. ↩1
- [137] Gehman, Samuel; Gururangan, Suchin; Sap, Maarten; et al.. RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. arXiv:2009.11462 (2020). doi:10.48550/arXiv.2009.11462. PDF. ↩1
- [138] Parrish, Alicia; Chen, Angelica; Nangia, Nikita; et al.. BBQ: A Hand-Built Bias Benchmark for Question Answering. arXiv:2110.08193 (2021). doi:10.48550/arXiv.2110.08193. PDF. ↩1
- [139] Kadavath, Saurav; Conerly, Tom; Askell, Amanda; et al.. Language Models (Mostly) Know What They Know. arXiv:2207.05221 (2022). doi:10.48550/arXiv.2207.05221. PDF. ↩1
- [140] Tian, Katherine; Mitchell, Eric; Zhou, Allan; et al.. Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback. arXiv:2305.14975 (2023). doi:10.48550/arXiv.2305.14975. PDF. ↩1
- [141] Band, Neil; Li, Xuechen; Ma, Tengyu; et al.. Linguistic Calibration of Long-Form Generations. arXiv:2404.00474 (2024). doi:10.48550/arXiv.2404.00474. PDF. ↩1
- [142] Quach, Victor; Fisch, Adam; Schuster, Tal; et al.. Conformal Language Modeling. arXiv:2306.10193 (2023). doi:10.48550/arXiv.2306.10193. PDF. ↩1
- [143] Geifman, Yonatan; El-Yaniv, Ran. Selective Classification for Deep Neural Networks. arXiv:1705.08500 (2017). doi:10.48550/arXiv.1705.08500. PDF. ↩1
- [144] Bates, Stephen; Angelopoulos, Anastasios; Lei, Lihua; et al.. Distribution-Free, Risk-Controlling Prediction Sets. arXiv:2101.02703 (2021). doi:10.48550/arXiv.2101.02703. PDF. ↩1
- [145] DeepSeek-AI; Guo, Daya; Yang, Dejian; et al.. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12948 (2025). doi:10.48550/arXiv.2501.12948. PDF. ↩1
- [146] Yang, An; Li, Anfeng; Yang, Baosong; et al.. Qwen3 Technical Report. arXiv:2505.09388 (2025). doi:10.48550/arXiv.2505.09388. PDF. ↩1
- [147] DeepSeek-AI; Liu, Aixin; Feng, Bei; et al.. DeepSeek-V3 Technical Report. arXiv:2412.19437 (2024). doi:10.48550/arXiv.2412.19437. PDF. ↩1
- [148] Grattafiori, Aaron; Dubey, Abhimanyu; Jauhri, Abhinav; et al.. The Llama 3 Herd of Models. arXiv:2407.21783 (2024). doi:10.48550/arXiv.2407.21783. PDF. ↩1
- [149] OLMo, Team; Walsh, Pete; Soldaini, Luca; et al.. 2 OLMo 2 Furious. arXiv:2501.00656 (2024). doi:10.48550/arXiv.2501.00656. PDF. ↩1
- [150] Qwen; :; Yang, An; et al.. Qwen2.5 Technical Report. arXiv:2412.15115 (2024). doi:10.48550/arXiv.2412.15115. PDF. ↩1