Counterfactual Generation from Audit Trails in Multi-Agent Negotiation Systems
Preprint, 2026
Abstract. When autonomous agents negotiate on behalf of human principals, adverse outcomes demand an answer to the question: which facts, disclosures, or contract clauses were decision-critical, and what changes would have altered the result? We formalise this as a constrained optimisation over a structural causal model (SCM) of the multi-agent coordination pipeline, organise the resulting counterfactual interventions into three interpretable classes — evidence, clause, and protocol counterfactuals — and prove that the lowest-cost single-class intervention upper-bounds the counterfactual margin, with exactness under an explicit layer-local pivotality condition. Building on the established correspondence between counterfactual explanations and adversarial perturbations, we derive closed-form margin formulae for linear decision boundaries under weighted ℓ∞ cost functions via Lagrange duality and the Fenchel conjugate. The layered structure of the causal graph enables per-layer verification of counterfactual claims, connecting to recent work on verifiable causality analysis. We extend the framework to adversarial settings by modelling the counterparty’s disclosure strategy as an endogenous variable, distinguishing between factual, disclosure, and policy counterfactuals.
