{"protocol":"ACA-1.0","name":"ACA-1, Agent Coordination Audit","operator":"AstraNL, Zaandam, Netherlands, KvK 88449335","purpose":"For anyone who runs more than one agent, or one agent among strangers: sixteen ways a system of agents loses work, each measured, each with the move that closes it.","rules":["UNKNOWN counts as open","a control counts only when the runtime enforces it, not when the prompt asks for it","answers are self-declared; the audit states what follows from them and what it does not prove"],"scope":{"values":["parallel","tools","handoff","market"],"meaning":"parallel: agents work at the same time on shared things; tools: agents call outside tools and interfaces; handoff: work passes between agents; market: agents work for or pay parties they do not know. Default all four."},"controls":[{"key":"claim_before_work","pattern":"No claim on shared work","weight":9,"scope":["parallel","market"],"question":"Before an agent starts a piece of work that another agent could also pick, does it take a claim that the others can see?","what_goes_wrong":"Agents pick the same work, redo it and overwrite each other.","measured":"Two cooperating coding agents succeeded about 25 percent against about 50 percent for one agent; work overlap in 33.2 percent of failures. One market drew 44 submissions per task.","source":"https://arxiv.org/html/2601.13295v1","fix":"Take a visible claim on a declared scope before any work; refuse to start when another holds it.","crossing_move":"claim","answers":["yes","partial","no","unknown"]},{"key":"leases_expire","pattern":"Locks that nobody releases","weight":7,"scope":["parallel"],"question":"Do claims and locks expire by themselves unless the holder refreshes them?","what_goes_wrong":"A holder that crashes or forgets keeps the lock; the others wait or stall.","measured":"Twenty agents sharing a lock file slowed to the throughput of two or three because holders kept locks too long or forgot to release them.","source":"https://cursor.com/blog/scaling-agents","fix":"Make every lock a lease with a short life that must be refreshed, and cap how long one holder may keep it.","crossing_move":"claim","answers":["yes","partial","no","unknown"]},{"key":"idempotent_side_effects","pattern":"Side effects repeat on retry","weight":9,"scope":["tools","market"],"question":"Does every action with an outside effect, a payment, an order, a message, a write, carry a key that makes a repeat harmless?","what_goes_wrong":"A retry after a timeout, a redelivery or a restart does the thing twice.","measured":"Frontier models repeated a side effect in 74 percent of redelivered requests and 56 percent of late commits; a key on every side effect cut that to 7 percent.","source":"https://arxiv.org/html/2609.29095v1","fix":"Derive a key from the intent and check it before acting; keep the key outside the agent's own memory.","crossing_move":"claim with mode once","answers":["yes","partial","no","unknown"]},{"key":"stop_condition","pattern":"No stop condition","weight":8,"scope":["parallel","tools","handoff","market"],"question":"Is there an explicit end state for every task, checked by the runtime, with a ceiling on steps and on repeated identical calls?","what_goes_wrong":"Agents repeat steps, do not notice they are finished, or stop early.","measured":"Step repetition is 15.7 percent of failures in 1,642 multi-agent traces, unawareness of termination conditions 12.4 percent.","source":"https://arxiv.org/abs/2503.13657","fix":"Write the end state down before the work starts; enforce a step ceiling and a breaker on identical calls in the runner.","crossing_move":"check","answers":["yes","partial","no","unknown"]},{"key":"independent_verification","pattern":"No independent verification point","weight":9,"scope":["parallel","tools","handoff","market"],"question":"Is the result of an agent checked by something other than that agent before it is used or passed on?","what_goes_wrong":"Errors pass from agent to agent and grow.","measured":"Independent agents amplified errors 17.2 times against 4.4 times with one validation point; one added verification step gave plus 15.6 points of task success.","source":"https://arxiv.org/html/2512.08296","fix":"Put one acceptance check with evidence between producing a result and using it.","crossing_move":"mark with evidence","answers":["yes","partial","no","unknown"]},{"key":"structured_handoff","pattern":"Context lost at handoff","weight":7,"scope":["handoff"],"question":"When work passes between agents, does it pass as a fixed structure with the artifacts by reference, not as a retelling?","what_goes_wrong":"The receiver guesses what the sender meant; a supervisor that paraphrases loses information.","measured":"Inter-agent misalignment is 32.3 percent of multi-agent failures; forwarding replies directly instead of paraphrasing gave nearly 50 percent better results in a vendor benchmark.","source":"https://www.langchain.com/blog/benchmarking-multi-agent-architectures","fix":"Hand over a fixed record: goal, end state, what was done, artifacts by link, what is open. No free retelling.","crossing_move":"mark","answers":["yes","partial","no","unknown"]},{"key":"hop_limit","pattern":"Delegation without a hop limit","weight":6,"scope":["handoff"],"question":"Does every delegated task carry a remaining hop count or budget that is reduced at each handoff and stops the chain at zero?","what_goes_wrong":"Work circles between agents until someone reads the invoice.","measured":"One reported loop between two agents ran eleven days; the report is first-person and its figures are inconsistent. The mechanism, a follow-the-predecessor rule with no outside check, is the ant mill.","source":"https://pub.towardsai.net/we-spent-47-000-running-ai-agents-in-production-heres-what-nobody-tells-you-about-a2a-and-mcp-5f845848de33","fix":"Pass a hop count with every delegation; at zero return the work to its origin as failed.","crossing_move":"claim with hops","answers":["yes","partial","no","unknown"]},{"key":"parallelise_gate","pattern":"Coordination bought when it is not needed","weight":6,"scope":["parallel"],"question":"Before splitting work across agents, is there a rule that decides from the task whether more agents help at all?","what_goes_wrong":"More agents cost more and do worse on work that is sequential or that one agent already does well.","measured":"Multi-agent systems use about 15 times the tokens of chat; on sequential planning every multi-agent variant lost 39 to 70 percent against one agent; adding agents hurts once one agent passes about 45 percent accuracy.","source":"https://arxiv.org/html/2512.08296","fix":"Parallelise only decomposable work, and only after measuring the single-agent baseline.","crossing_move":"look","answers":["yes","partial","no","unknown"]},{"key":"context_budget","pattern":"Context waste","weight":6,"scope":["tools"],"question":"Is the input of each step measured and bounded: trajectory pruned, cache hit rate watched, tool definitions loaded on demand?","what_goes_wrong":"Most tokens are paid for again and again without changing the result.","measured":"39.9 to 59.7 percent of input tokens in coding trajectories are removable with performance kept; system prompts are 69 percent of input tokens in production and only 28 percent of calls read from cache.","source":"https://arxiv.org/abs/2509.23586","fix":"Measure tokens in per step, prune expired content, keep the stable prefix cacheable, load tool definitions when needed.","crossing_move":"check","answers":["yes","partial","no","unknown"]},{"key":"backoff_and_admission","pattern":"Retry storms on shared interfaces","weight":7,"scope":["tools","parallel"],"question":"On a rate limit, a timeout or a lost claim, do agents back off with a random, growing wait, and is the number of concurrent calls limited?","what_goes_wrong":"Everybody retries at once; the shared interface collapses for all.","measured":"Rate limits are one third to 60 percent of model call errors in production telemetry; admission control with jittered retry cut failure rates of concurrent agents from 73 to 100 percent down to 0 to 18 percent.","source":"https://arxiv.org/html/2604.17111v1","fix":"Random wait with a doubling ceiling after every refusal; start at one concurrent call and grow only on success.","crossing_move":"look, advice wait","answers":["yes","partial","no","unknown"]},{"key":"fresh_state_before_write","pattern":"Acting on stale state","weight":6,"scope":["parallel","tools"],"question":"Before an agent writes to shared state, does it check that what it read is still current?","what_goes_wrong":"A long inference ends in a write that undoes someone else's change.","measured":"Stale reads appeared in 35 percent of triage traces in one study; agents working blind to a concurrent change interfered in 97 percent of constructed cases, and a 130-token notice of the change recovered 82 percent.","source":"https://arxiv.org/html/2609.25396v1","fix":"Version what you read and refuse the write when the version moved; tell the others what you changed.","crossing_move":"mark","answers":["yes","partial","no","unknown"]},{"key":"receipts","pattern":"Outcomes that cannot be checked","weight":6,"scope":["market","handoff"],"question":"Does finished work come with a receipt another party can check: what was asked, what was delivered, by whom, when?","what_goes_wrong":"Every receiver verifies again or simply believes.","measured":"On-chain agent registries hold 59 to 91 percent sybil-flagged reviewers; a signature proves who said it, and only evidence proves what was done.","source":"https://arxiv.org/html/2606.26028","fix":"Return a signed record over the hash of the input and of the output with each result.","crossing_move":"mark with evidence, proof","answers":["yes","partial","no","unknown"]},{"key":"counterparty_and_venue_check","pattern":"Working blind to the venue","weight":7,"scope":["market","tools"],"question":"Before work or payment, is it checked that the endpoint answers, that the reward is funded and how many others are already on it?","what_goes_wrong":"Agents work for unfunded rewards, call dead servers and join crowds they cannot win against.","measured":"About half of listed MCP servers are invalid or do not complete a handshake; single bounties carried 236 and 103 competing claims; about 2.4 percent of submissions on one market reached settlement.","source":"https://arxiv.org/html/2609.10962v1","fix":"Look before you start: liveness, funding, crowd, and whether the expected value is above zero.","crossing_move":"look","answers":["yes","partial","no","unknown"]},{"key":"shared_experience","pattern":"Every run starts from nothing","weight":5,"scope":["parallel","tools","handoff","market"],"question":"Is what worked and what failed recorded where the next run or the next agent will read it?","what_goes_wrong":"The same dead ends are explored again by every agent and every session.","measured":"A shared record of experience usable across frameworks raised success by up to 18.7 points; workflow memory by up to 51.1 percent relative.","source":"https://arxiv.org/abs/2507.06229","fix":"Leave a short trace on the shared place after every outcome, success or failure, and read the trail before starting.","crossing_move":"mark, look","answers":["yes","partial","no","unknown"]},{"key":"policy_not_prompts","pattern":"Humans approving everything","weight":5,"scope":["parallel","tools","handoff","market"],"question":"Are routine actions allowed by standing policy and a sandbox, with human approval kept for the irreversible ones?","what_goes_wrong":"People approve without reading, or the agents wait for people.","measured":"93 percent of permission prompts are approved anyway; sandboxing cut prompts by 84 percent; only 0.8 percent of actions were irreversible.","source":"https://www.anthropic.com/engineering/claude-code-auto-mode","fix":"Write the policy once, enforce it in the runtime, and ask a human only at the irreversible step.","crossing_move":"check","answers":["yes","partial","no","unknown"]},{"key":"variance_measured","pattern":"Unmeasured variance between runs","weight":5,"scope":["parallel","tools","handoff","market"],"question":"Is the same task run several times before its cost and success rate are trusted?","what_goes_wrong":"One good run is taken for the rule; budgets and promises are built on it.","measured":"Runs of the same coding task differ by up to 30 times in tokens; success over eight repeated trials falls below 25 percent on a retail benchmark where a single trial passes about 61 percent.","source":"https://arxiv.org/abs/2604.22750","fix":"Report success as the rate over repeated runs and cost as a range, never from one run.","crossing_move":"check","answers":["yes","partial","no","unknown"]}],"optional":{"system":"a name for the audited system, echoed in the receipt","agents":"how many agents run in it"},"scoring":"Each applicable control has a weight; yes earns it, partial half, no and unknown nothing. Weight 9 is critical, 7 or more high. NOT_READY when any critical control is open; READY at 85 or more with no high control open; otherwise CONDITIONAL.","licence":"The questions and the rule may be used by anyone, free of charge.","endpoints":{"free_preview":"https://verify.astranl.com/v1/coordination/preview","page_for_people":"https://verify.astranl.com/coordination","signed_audit":"https://verify.astranl.com/v1/agent-coordination-audit","the_moves_that_close_the_findings":"https://verify.astranl.com/skill.md"},"paid":{"agent-coordination-audit":"$0.05","for_people":"1 EUR through Stripe on the page"}}