Multi-agent systems are increasingly proposed to automate multi-step, safety-critical clinical workflows, yet benchmark accuracy is an insufficient proxy for deployment readiness: the coordination failures that matter for patient safety surface across the steps of the process, not in any single accuracy metric. In a simulated emergency-room onboarding workflow run entirely by LLM agents, we diagnose these failures across two studies that vary contextual knowledge, communication structure, and model reasoning. Our preliminary results suggest that team structure, rather than knowledge availability, may be the primary bottleneck: several failure modes persist even when an exhaustive knowledge base is supplied. We also surface a deployment risk that benchmark scores alone would miss: a stronger reasoning model (o3) plans more capably than a non-reasoning model (GPT-4o) yet appears to introduce a wider variety of failure modes. These early findings suggest that process-level evaluation, auditable coordination, and clinician-in-the-loop oversight are prerequisites for safe deployment.
@misc{bai2025masmarscoordinationfailures,
title={From MAS to MARS: Coordination Failures and Reasoning Trade-offs in Hierarchical Multi-Agent Robotic Systems within a Healthcare Scenario},
author={Yuanchen Bai and Zijian Ding and Shaoyue Wen and Xiang Chang and Angelique Taylor},
year={2025},
eprint={2508.04691},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2508.04691},
}