Perspective & research agenda · Working paper
AI workflows: tool coordination and repeatable task completion
Abstract
Tool invocation and a model’s report of success do not establish completed work. This paper synthesizes 12 selected studies on reasoning, tools, feedback and interactive evaluation to propose workflows organized around observable states, explicit permissions and recoverable failures. The central hypothesis is that state checks and idempotent retries can improve repeated reliability without broader authority or more human intervention. A matched-budget protocol compares simple and coordinated workflows under controlled fault injection. All-trial and at-least-once success are reported separately alongside takeovers and recovery costs. No original experiments are presented. The proposal is a testable research agenda, not evidence that additional agents or checks necessarily improve outcomes.
1 Problem and central argument
Once a model can invoke tools, what establishes completion beyond its own success narrative? This paper considers controlled tasks using public materials; methodological discussion does not authorize publication, payment or permission changes on a user’s behalf.
2 Literature scope and research context
This paper purposively selects 12 research documents for a question-led narrative synthesis, not a systematic review or meta-analysis. The selected works were first submitted in 2017–2024 and do not cover all developments through 2026. Titles, authors, submission dates, public abstracts, source pages and PDF access were checked on 8 September 2026. Original research, technical reports and surveys are distinguished rather than treated as equally strong experimental evidence. Selected methodological passages in FActScore, ALCE, EvalPlus and τ-bench were previously examined; we did not conduct full-text appraisal or experimental replication of every work.
Chain-of-Thought prompting improves performance on selected reasoning tasks through intermediate examples; generated reasoning is not itself external evidence.[1] Self-Consistency samples multiple reasoning paths and aggregates answers, improving performance in the reasoning tasks studied.[2]
ReAct interleaves reasoning and actions, allowing external knowledge or environmental feedback to inform subsequent decisions.[3] Toolformer studies how a model can learn when to invoke APIs, what arguments to use and how to incorporate their results.[4]
Reflexion stores textual reflections on feedback for subsequent attempts without updating model weights.[5] Self-Refine uses the same model to provide feedback and iteratively revise outputs across several tasks.[6]
Tree of Thoughts organizes intermediate reasoning as a process of exploration, evaluation and backtracking for planning or search tasks.[7] AutoGen provides configurable multi-agent conversations connecting models, tools and human input.[8]
AgentBench evaluates reasoning and decision-making in multiple interactive environments and analyses failures in long-horizon reasoning and instruction following.[9] WebArena provides reproducible web environments and long-horizon tasks evaluated by functional task completion.[10]
ToolLLM develops data, training and evaluation for API use, including tool retrieval and search over invocation paths.[11] τ-bench combines simulated users, domain policies and API tools, checking resulting database states and reliability across repeated trials.[12]
3 Synthesis and proposed method
Reasoning traces, retrieved documents and execution receipts provide different kinds of evidence. We propose evaluating observable task states instead of model self-reports. A single success does not establish dependable operation; architecture comparisons should hold the model, inputs and invocation budget fixed and include a single-workflow baseline.
Specify input, processing, checking and delivery stages, retaining input identifiers, artefacts, timestamps and states. Mark failed reads as missing. Before retrying, check whether an action already took effect. Bound tools, writable objects, runtime and retries; missing information or a need for additional authority requires human resolution.
4 Proposition and comparisons
Can state verification and idempotent retries improve repeated reliability without broader permissions or increased human takeover? Benefits depending on extra intervention or authority cannot be attributed to workflow structure alone.
Compare a single workflow, multi-agent coordination and a state-verified workflow at matched models, tool permissions, budgets and tasks. Ablate state checks and retry protection separately; inject timeouts, duplicate responses and partial failures only in recoverable environments.
5 Operational definitions and decision boundaries
Let yᵢⱼ∈{0,1} indicate that task i in run j satisfies both final-state and policy requirements, with N tasks and a fixed k runs per task. R_all,k is the fraction succeeding in every run; R_any,k is the fraction succeeding at least once. Report both with k, budget and takeover rate. Neither should be replaced by taking the overall mean success rate to the kth power.
Repeat the same task set in a recoverable environment and separately record completion, policy compliance, human takeover and recovery cost. Define success before injecting network faults, duplicate responses or missing inputs. All-trial success and success in at least one trial answer different questions and must not be conflated.
6 A testable evaluation protocol
Treat tasks, not repeated outputs from one task, as sampling units. Use a separate pilot to assess annotation feasibility, cost and plausible effects, then determine the main sample size from precision or power targets. Exclude pilot tasks from held-out evaluation. Stratify Chinese and English tasks and treat translated counterparts as paired rather than independent samples. Publish inclusion and exclusion criteria and freeze inputs, versions, permissions and budgets.
Randomize condition order and compare conditions within tasks; blind adjudicators to condition where feasible. Report raw counts, task-level effect differences and uncertainty intervals without treating repeat runs as extra independent tasks. Report task families separately and test sensitivity to missing outcomes and timeouts. Retain and distinguish timeouts, abstentions and human takeovers. Document when and why a planned procedure changes.
A reproducibility package should identify tasks, data permissions, model and tool versions, prompt templates, configurable random seeds, tests, adjudication rules, failure logs and analysis scripts. Reproducibility does not justify releasing personal or restricted information; redacted structures and controlled-access terms may be appropriate. These are proposed research outputs: this site does not currently provide a completed experimental dataset or analysis code.
7 Interpretive boundaries and alternatives
τ-bench checks final database states together with required responses. We adopt its distinction between at-least-once and all-trial success, but the equations below are direct empirical statistics over a fixed number of trials, not the paper’s combinatorial estimator. Descriptive computation does not require independent and identically distributed trials.[12]
This paper contains no original experiments or current product comparison. Extra tools, human review or unequal budgets could explain apparent gains; selecting only successful tasks would introduce bias. Findings that contradict expectations should retain failures and negative results, with explicit boundaries for model, language, task family and environment. Historical benchmarks cannot establish guarantees for current systems.
8 Conclusion
Dependable coordination requires observable final states, explicit authority and recoverable failures. Agent counts and reasoning length are not completion evidence; the proposed workflow still requires repeated comparison with simpler baselines.
This companion paper develops the AI+ overview using a shared literature base and framework; it is not an independent empirical study. Read the overview.
9 Declarations and availability
Data and code: No original experimental data were generated. Sources are individually linked. Experimental and analysis code has not been implemented; website code is not an experimental replication package.
Authorship, contributions, funding and competing interests: Alsay is the website publisher, which does not establish scholarly authorship. Author identities, contributions, funding and competing interests require confirmation by accountable contributors. This draft does not assert absence of competing interests or invent affiliations, degrees or ethical approvals.
AI assistance and versions: AI assisted retrieval, organization, editing and Chinese–English translation. The two versions represent one working paper, not two independent studies; scholarly approval by accountable authors remains pending. No peer review, journal acceptance or completed preregistration is claimed.
References
Cited as public arXiv versions; peer-review status was not individually verified. Source pages and PDF access were checked on 2026-09-08; this does not guarantee enduring access or valid full-text conclusions.
- [1] Wei, Jason; Wang, Xuezhi; Schuurmans, Dale; et al.. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv:2201.11903 (2022). doi:10.48550/arXiv.2201.11903. PDF. ↩1
- [2] Wang, Xuezhi; Wei, Jason; Schuurmans, Dale; et al.. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171 (2022). doi:10.48550/arXiv.2203.11171. PDF. ↩1
- [3] Yao, Shunyu; Zhao, Jeffrey; Yu, Dian; et al.. ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629 (2022). doi:10.48550/arXiv.2210.03629. PDF. ↩1
- [4] Schick, Timo; Dwivedi-Yu, Jane; Dessì, Roberto; et al.. Toolformer: Language Models Can Teach Themselves to Use Tools. arXiv:2302.04761 (2023). doi:10.48550/arXiv.2302.04761. PDF. ↩1
- [5] Shinn, Noah; Cassano, Federico; Berman, Edward; et al.. Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366 (2023). doi:10.48550/arXiv.2303.11366. PDF. ↩1
- [6] Madaan, Aman; Tandon, Niket; Gupta, Prakhar; et al.. Self-Refine: Iterative Refinement with Self-Feedback. arXiv:2303.17651 (2023). doi:10.48550/arXiv.2303.17651. PDF. ↩1
- [7] Yao, Shunyu; Yu, Dian; Zhao, Jeffrey; et al.. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv:2305.10601 (2023). doi:10.48550/arXiv.2305.10601. PDF. ↩1
- [8] Wu, Qingyun; Bansal, Gagan; Zhang, Jieyu; et al.. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv:2308.08155 (2023). doi:10.48550/arXiv.2308.08155. PDF. ↩1
- [9] Liu, Xiao; Yu, Hao; Zhang, Hanchen; et al.. AgentBench: Evaluating LLMs as Agents. arXiv:2308.03688 (2023). doi:10.48550/arXiv.2308.03688. PDF. ↩1
- [10] Zhou, Shuyan; Xu, Frank F.; Zhu, Hao; et al.. WebArena: A Realistic Web Environment for Building Autonomous Agents. arXiv:2307.13854 (2023). doi:10.48550/arXiv.2307.13854. PDF. ↩1
- [11] Qin, Yujia; Liang, Shihao; Ye, Yining; et al.. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. arXiv:2307.16789 (2023). doi:10.48550/arXiv.2307.16789. PDF. ↩1
- [12] Yao, Shunyu; Shinn, Noah; Razavi, Pedram; et al.. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045 (2024). doi:10.48550/arXiv.2406.12045. PDF. ↩1 ↩2