Perspective & research agenda · Working paper
Evidence-constrained AI writing: from fluent text to verifiable claims
Abstract
AI-assisted writing must reconcile readability with traceable factual support. This paper synthesizes 12 selected studies on generation, retrieval and factual evaluation, then proposes a claim–source ledger with independent review. The central hypothesis is that evidence support can improve without reducing required-fact coverage or increasing abstention. A paired protocol compares human writing, direct generation and retrieval-based workflows under matched budgets. Support and coverage are measured separately, with explicit treatment of omissions, unresolved claims and verification costs. No original experiment is reported; the selected historical literature neither establishes current product performance nor guarantees correctness in specialist domains. The proposal remains subject to empirical testing and accountable scholarly review.
1 Problem and central argument
How can a writing workflow preserve checkable facts, qualifications and sources without confusing readability with truth? This paper addresses public or authorized materials, not a guarantee of correctness in specialist practice.
2 Literature scope and research context
This paper purposively selects 12 research documents for a question-led narrative synthesis, not a systematic review or meta-analysis. The selected works were first submitted in 2017–2024 and do not cover all developments through 2026. Titles, authors, submission dates, public abstracts, source pages and PDF access were checked on 8 September 2026. Original research, technical reports and surveys are distinguished rather than treated as equally strong experimental evidence. Selected methodological passages in FActScore, ALCE, EvalPlus and τ-bench were previously examined; we did not conduct full-text appraisal or experimental replication of every work.
The Transformer uses attention for sequence transduction, providing an architectural basis for subsequent language-generation research.[1] GPT-3 demonstrates few-shot task performance through textual examples, while reporting variation across tasks and methodological issues associated with training data.[2]
T5 frames language tasks as text-to-text problems and compares objectives, architectures and datasets; a common task format does not imply a common evaluation criterion.[3] BART studies generation and comprehension through denoising pre-training, illustrating the importance of the training objective.[4]
Retrieval-augmented generation combines parametric models with external retrieval memory for knowledge-intensive tasks.[5] REALM incorporates a knowledge retriever into language-model pre-training, exploring access to documents beyond facts stored in model parameters.[6]
InstructGPT uses demonstrations and preference feedback to improve instruction following, but its authors also report that the resulting models still make simple mistakes.[7] TruthfulQA evaluates questions that elicit common misconceptions, distinguishing imitation of language from factual truthfulness.[8]
FActScore decomposes long-form output into atomic facts and estimates the proportion supported by a reliable knowledge source.[9] ALCE evaluates fluency, correctness and citation quality separately: the presence of citations does not establish that every claim is supported.[10]
Chain-of-Verification separates drafting, verification questions, independent answers and revision, reducing hallucination in the tasks studied.[11] Lost in the Middle finds position-dependent use of relevant information in the models tested; a larger context window does not guarantee effective use of all supplied material.[12]
3 Synthesis and proposed method
Training, retrieval and evaluation intervene at different points and cannot be combined into one writing-accuracy score. Our synthesis is to establish claim–source links before composing paragraphs and to retain unresolved claims as explicit gaps. Organizing traceable excerpts by topic can be more defensible than merely increasing input length.
For each checkable claim, record its wording, source, page or paragraph, access date and verification status. Review numbers, names, causal language and distinctions such as planned versus completed. Model self-checks can suggest questions, but confirmation must return to evidence and publication requires an accountable human reviewer.
4 Proposition and comparisons
Does a claim–source ledger improve supported-claim precision without worsening required-fact coverage or abstention relative to baseline? Gains achieved only by omitting difficult claims would not support the hypothesis.
Use human writing, direct generation and retrieval-augmented generation as baselines. Add an evidence ledger and independent review in the full workflow; ablate each separately while fixing materials, model version, retrieval corpus and total time budget.
5 Operational definitions and decision boundaries
Let Sᵢ be supported checkable claims in task i and Aᵢ all checkable claims. Let Kᵢ be facts required in advance and Dᵢ required facts correctly retained in the output. Report supported-claim precision Pᵢ and coverage Cᵢ separately so that omission or silence cannot inflate the apparent result. A zero denominator is not applicable; abstentions and omissions are reported separately.
Compare human-only and AI-assisted workflows on the same non-sensitive materials. Predefine claim segmentation and supported, contradicted and unresolved labels. Report evidence support together with required-fact coverage, omissions, verification time and reader comprehension. An output with no checkable claims has an undefined support proportion, not perfect factuality.
6 A testable evaluation protocol
Treat tasks, not repeated outputs from one task, as sampling units. Use a separate pilot to assess annotation feasibility, cost and plausible effects, then determine the main sample size from precision or power targets. Exclude pilot tasks from held-out evaluation. Stratify Chinese and English tasks and treat translated counterparts as paired rather than independent samples. Publish inclusion and exclusion criteria and freeze inputs, versions, permissions and budgets.
Randomize condition order and compare conditions within tasks; blind adjudicators to condition where feasible. Report raw counts, task-level effect differences and uncertainty intervals without treating repeat runs as extra independent tasks. Report task families separately and test sensitivity to missing outcomes and timeouts. Retain and distinguish timeouts, abstentions and human takeovers. Document when and why a planned procedure changes.
A reproducibility package should identify tasks, data permissions, model and tool versions, prompt templates, configurable random seeds, tests, adjudication rules, failure logs and analysis scripts. Reproducibility does not justify releasing personal or restricted information; redacted structures and controlled-access terms may be appropriate. These are proposed research outputs: this site does not currently provide a completed experimental dataset or analysis code.
7 Interpretive boundaries and alternatives
FActScore primarily evaluates biographies against a corresponding knowledge source; it does not directly establish the truth of complex causal arguments. ALCE’s automated entailment judgements can also fail. The proportions proposed below are operational definitions for this protocol, not full reproductions of either benchmark.[9][10]
This paper contains no original experiments or current product comparison. Extra tools, human review or unequal budgets could explain apparent gains; selecting only successful tasks would introduce bias. Findings that contradict expectations should retain failures and negative results, with explicit boundaries for model, language, task family and environment. Historical benchmarks cannot establish guarantees for current systems.
8 Conclusion
Verifiable writing depends on an auditable relationship between claims, sources and qualifications, not the number of citations. Whether an evidence ledger justifies its additional cost remains a question for paired evaluation.
This companion paper develops the AI+ overview using a shared literature base and framework; it is not an independent empirical study. Read the overview.
9 Declarations and availability
Data and code: No original experimental data were generated. Sources are individually linked. Experimental and analysis code has not been implemented; website code is not an experimental replication package.
Authorship, contributions, funding and competing interests: Alsay is the website publisher, which does not establish scholarly authorship. Author identities, contributions, funding and competing interests require confirmation by accountable contributors. This draft does not assert absence of competing interests or invent affiliations, degrees or ethical approvals.
AI assistance and versions: AI assisted retrieval, organization, editing and Chinese–English translation. The two versions represent one working paper, not two independent studies; scholarly approval by accountable authors remains pending. No peer review, journal acceptance or completed preregistration is claimed.
References
Cited as public arXiv versions; peer-review status was not individually verified. Source pages and PDF access were checked on 2026-09-08; this does not guarantee enduring access or valid full-text conclusions.
- [1] Vaswani, Ashish; Shazeer, Noam; Parmar, Niki; et al.. Attention Is All You Need. arXiv:1706.03762 (2017). doi:10.48550/arXiv.1706.03762. PDF. ↩1
- [2] Brown, Tom B.; Mann, Benjamin; Ryder, Nick; et al.. Language Models are Few-Shot Learners. arXiv:2005.14165 (2020). doi:10.48550/arXiv.2005.14165. PDF. ↩1
- [3] Raffel, Colin; Shazeer, Noam; Roberts, Adam; et al.. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 (2019). doi:10.48550/arXiv.1910.10683. PDF. ↩1
- [4] Lewis, Mike; Liu, Yinhan; Goyal, Naman; et al.. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. arXiv:1910.13461 (2019). doi:10.48550/arXiv.1910.13461. PDF. ↩1
- [5] Lewis, Patrick; Perez, Ethan; Piktus, Aleksandra; et al.. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv:2005.11401 (2020). doi:10.48550/arXiv.2005.11401. PDF. ↩1
- [6] Guu, Kelvin; Lee, Kenton; Tung, Zora; et al.. REALM: Retrieval-Augmented Language Model Pre-Training. arXiv:2002.08909 (2020). doi:10.48550/arXiv.2002.08909. PDF. ↩1
- [7] Ouyang, Long; Wu, Jeff; Jiang, Xu; et al.. Training language models to follow instructions with human feedback. arXiv:2203.02155 (2022). doi:10.48550/arXiv.2203.02155. PDF. ↩1
- [8] Lin, Stephanie; Hilton, Jacob; Evans, Owain. TruthfulQA: Measuring How Models Mimic Human Falsehoods. arXiv:2109.07958 (2021). doi:10.48550/arXiv.2109.07958. PDF. ↩1
- [9] Min, Sewon; Krishna, Kalpesh; Lyu, Xinxi; et al.. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. arXiv:2305.14251 (2023). doi:10.48550/arXiv.2305.14251. PDF. ↩1 ↩2
- [10] Gao, Tianyu; Yen, Howard; Yu, Jiatong; et al.. Enabling Large Language Models to Generate Text with Citations. arXiv:2305.14627 (2023). doi:10.48550/arXiv.2305.14627. PDF. ↩1 ↩2
- [11] Dhuliawala, Shehzaad; Komeili, Mojtaba; Xu, Jing; et al.. Chain-of-Verification Reduces Hallucination in Large Language Models. arXiv:2309.11495 (2023). doi:10.48550/arXiv.2309.11495. PDF. ↩1
- [12] Liu, Nelson F.; Lin, Kevin; Hewitt, John; et al.. Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172 (2023). doi:10.48550/arXiv.2307.03172. PDF. ↩1