This is a real, unedited ProSearch output, generated from academic sources in minutes. Each [Ref N] maps to a source in your library.
[Strategic Research Advisory: Multi-Agent AI]
Abstract
Multi-Agent AI (often “LLM-based multi-agent systems” or “agentic multi-agent systems”) is rapidly moving from conceptual prototypes to production deployments, motivated by the limits of single-agent designs under long-horizon tasks, complex tool use, and enterprise constraints (governance, sovereignty, modularity). The provided sources span: (i) architectural frameworks for multi-agent AI and its socio-technical pathways [Ref 4, Ref 12]; (ii) domain implementations in design, engineering, AEC inspection, offshore surveillance, materials discovery, and healthcare [Ref 1, Ref 19, Ref 111, Ref 37, Ref 23, Ref 61, Ref 98]; (iii) reliability tooling such as interactive debugging and steering [Ref 32, Ref 39]; (iv) safety/security perspectives including risk taxonomies and security field-building [Ref 18, Ref 36, Ref 88]; and (v) evaluation/benchmarking and automation of evidence synthesis [Ref 63, Ref 83, Ref 106]. Collectively, they indicate a shift from “agents as prompts” toward engineered systems with orchestration layers, structured communication, and runtime governance.
The most consequential gaps, however, are methodological: (1) weak causal evidence for when multi-agent systems outperform strong single-agent baselines (and at what cost), despite explicit calls to answer this [Ref 15]; (2) immature, non-standard evaluation protocols for multi-agent reliability, safety, and alignment in interactive contexts [Ref 4, Ref 15, Ref 34, Ref 41]; (3) underdeveloped multi-agent security engineering for emergent threats (collusion, miscoordination, tool abuse) [Ref 18, Ref 36, Ref 88]; and (4) limited developer tooling and observability for debugging long, branching agent interactions in real deployments [Ref 32, Ref 39].
Top recommendations: (i) develop a rigorous, cost-aware evaluation framework and benchmark suite that compares multi-agent vs single-agent performance under matched tool access, latency budgets, and governance constraints, directly addressing effectiveness and safety together [Ref 15, Ref 41, Ref 72]; (ii) create “protocolized coordination” and “negative-feedback” reliability patterns that are auditable and fail-closed, then evaluate them empirically on realistic tasks [Ref 2, Ref 114, Ref 32]; and (iii) build and test security threat models and mitigations for multi-agent interactions (including collusion detection and policy enforcement layers) with reproducible red-team harnesses [Ref 36, Ref 88, Ref 80, Ref 84]. These directions are publishable because they target foundational questions the field itself flags as unresolved, and they naturally produce artifacts (benchmarks, protocols, datasets, tools) valued by high-impact venues.
Keywords
Multi-agent AI; LLM agents; orchestration; evaluation; benchmarking; coordination protocols; AI safety; multi-agent security; interactive debugging; governance; retrieval-augmented generation; socio-technical systems
1. Introduction
Multi-agent AI systems replace monolithic “do-everything” assistants with teams of specialized agents that coordinate through structured interactions, shared state, tool use, and supervisory control. The appeal is practical: complex work is parallelizable, requires diverse competencies (retrieval, planning, coding, verification), and benefits from cross-checking. Several provided sources emphasize the shift from static workflows to adaptive systems of interacting agents that coordinate in real time for knowledge work automation [Ref 4, Ref 12].
Research activity now spans at least two intertwined traditions. First is the classical multi-agent systems and multi-agent reinforcement learning lineage, focusing on decentralized decision-making, coordination, incentives, and learning dynamics (with longstanding critiques about unclear problem formulations in MARL) [Ref 29]. Second is the newer wave of LLM-powered tool-using agents, where “agent teams” are built using orchestration frameworks and communication protocols, with growing industrial attention to deployment reliability, observability, and governance,,. The provided references include both a modern conceptual architecture for “multi-agent artificial intelligence” as layered systems [Ref 4] and older but rigorous work integrating planning with multi-agent environments, including soundness/completeness conditions [Ref 31], showing the continuity between symbolic planning and today’s agentic systems.
From a publication perspective, the field’s center of gravity is moving toward measurable claims: when multi-agent designs *actually* outperform single-agent baselines, what safety risks emerge uniquely from interaction effects, and how we can evaluate and govern such systems before they are widely deployed. These questions are explicitly foregrounded in an “Outlook” paper that frames effectiveness and safety as core axes and asks when MAS are more effective than single agents [Ref 15]. Meanwhile, security-focused sources argue that interacting agents create novel, under-studied threats (including collusion and coordinated attacks) and call for a distinct “multi-agent security” research agenda [Ref 36].
This advisory document synthesizes the provided sources and broader established knowledge to (i) assess what is currently strong vs fragile in multi-agent AI research, (ii) identify high-impact gaps, and (iii) propose concrete, publishable research directions with methodology and positioning guidance.
2. Critical Analysis of Reviewed Sources
2.1 Multi-agent AI: Collaborative Design…Interior Design Workflow
This paper proposes a four-agent collaborative design framework for interior design: data analysis, requirement guidance, scheme generation, and design optimization [Ref 1]. Two agents (data analysis, requirement guidance) are implemented and evaluated in real-world projects with reported workload reduction and improved requirement understanding; two others remain conceptual and integrate parametric modeling, graph networks, style transfer, reinforcement learning, and VR interaction [Ref 1]. A key strength is its anchoring in real projects and explicit workflow decomposition. A limitation is that half the system is conceptual, and the evaluation evidence (as described in the abstract) is not detailed enough to judge rigor (metrics, baselines, sample size), leaving generalizability open.
2.2 Distributed Negative Feedback Optimization for Multi-Agent AI
This work argues that single-agent guardrails have a “blind spot” because an agent cannot reliably detect errors in its own reasoning, proposing Distributed Negative Feedback Optimization (DNFO) as a control-engineering-inspired closed-loop approach [Ref 2]. It introduces design principles: redundant intelligence architecture, complementary agent chains, and a triple-layer failsafe (including emergency stop, liveness monitoring, mutual exclusion) with inspiration from ISO 26262 [Ref 2]. Its key contribution is an engineering-oriented reliability blueprint that is naturally testable. The limitation (based on provided text) is that details of the “controlled before/after evaluation” and what “Buddys” entails are truncated here, making it hard to assess evidence quality and scope.
2.3 What Are Multi-Agent AI Systems
This source provides an accessible definition and emphasizes role specialization, parallel work, tool use, and routing GPU inference only when needed [Ref 3]. It outlines agent components and deployment considerations (as indicated), and it stresses scalability and reliability. Its value is conceptual clarity for system builders. Its limitation is that it is descriptive rather than evidentiary; it does not provide a research-grade evaluation or formal claims.
2.4 Multi-agent AI (five-component layered architecture)
This paper frames multi-agent artificial intelligence (MAAI) as a “foundational shift” in knowledge work automation and proposes a structured five-component layered architecture: foundation model; data-centric perception/action; dynamic orchestration; agent-integrated workflow; interaction interface [Ref 4]. It explicitly structures research pathways: technical capabilities, organizational integration, and socio-technical implications such as fairness, accountability, and labor transformation [Ref 4]. The strength is a clear conceptual model that disentangles technical and organizational dimensions. The limitation is that, from the abstract, it appears primarily conceptual; empirical validation is not described.
2.5 Multi-Agent Systems: Architecture, Applications & Real-World Impact
This source describes MAS as distributed intelligence with benefits (scalability, resilience, coordination) and promises “real-world examples” [Ref 5]. Its strength is framing multi-agent systems as a foundation for “next generation” AI and discussing orchestration patterns and governance at a high level. As provided, it repeats key takeaways and reads as overview material; it lacks methodological detail or verifiable comparative evaluation.
2.6 Multi-Agent Systems: How AI Agents Work Together (2026 Guide)
This guide identifies five architecture patterns (hierarchical, sequential pipeline, collaborative, competitive, swarm) and names “CrewAI, AutoGen, and LangGraph” as leading frameworks in 2026 [Ref 6]. It asserts a “sweet spot” of 3–10 agents due to coordination overhead [Ref 6]. Its strength is actionable taxonomy for system design. Its limitation is that these are assertions without experimental backing in the excerpt; a publishable paper would need empirical substantiation and careful definition of “best.”
2.7 Multi-Agent AI Systems and How Multiple AI Agents Work Together
This article emphasizes specialized agents, orchestrator coordination, and communication via structured data such as JSON [Ref 7]. It suggests multi-agent works best for parallelizable and “read-heavy workloads,” and claims companies like Tesla rely on such systems for real-time decision-making [Ref 7]. The strength is the practical deployment orientation. The limitation is that claims about industrial usage are not evidenced in the excerpt; it also mixes explanation with service marketing, which reduces its value as an academic anchor.
2.8 Multi-agent Embodied AI: Advances and Future Directions
This arXiv survey positions embodied AI as sensor/actuator systems learning from real-world feedback, and argues most embodied AI research remains single-agent and assumes static/closed environments, while real environments are dynamic/open and require collaboration [Ref 8]. It explicitly notes existing multi-agent embodied research is “narrow” and relies on simplified models that fail to capture open-world complexity [Ref 8]. Strength: clearly articulated gap statement and motivation for multi-agent embodied benchmarks. Limitation: as an abstract-only view here, specifics of taxonomy, datasets, and evaluation recommendations are not visible.
2.9 Multi-Agent AI Systems Complete Guide 2026 - Calmops
This guide claims multi-agent systems moved “from research to production” by 2026 and covers architecture, patterns, memory/context management, best practices, and mentions an “Agent-to-Agent (A2A) protocol” being developed by Google, Anthropic, and other labs [Ref 9]. Strength: highlights interoperability and protocol standardization as practical challenges. Limitation: it is a secondary summary; academic work would need primary protocol specs and empirical tests.
2.10 AI Agent Orchestration: A 2026 Guide to Multi-Agent Systems
This guide frames orchestration as coordination of specialized agents and cites adoption statistics (e.g., AI adoption and rapid agentic adoption) and “measurable improvements” [Ref 10]. It names LangGraph, CrewAI, AutoGen as frameworks [Ref 10]. Strength: focuses on orchestration as an engineering discipline. Limitation: metrics and sources for the claimed improvements are not provided in the excerpt; for publication, treat as motivation, not evidence.
2.11 Multi-Agent Systems: The Complete Deep Dive into Collaborative AI
This source explains MAS concepts including emergence and mixed cooperative/competitive dynamics [Ref 11]. Strength: basic conceptual overview. Limitation: no research methodology or evaluation.
2.12 Designing Multi-Agent Intelligence
This piece argues enterprises are pivoting from single “do-everything” agents to multi-agent systems due to domain breadth, data sovereignty, and modularity needs [Ref 12]. It describes agents as coupling LLM/SLM cores, domain toolsets, and memories, coordinated by an orchestrator, and stresses emergent behavior as the “breakthrough” [Ref 12]. Strength: clear articulation of enterprise constraints shaping architectures. Limitation: not an academic evaluation; it should be used to derive requirements and hypotheses rather than as evidence of performance.
2.13 Multi-Agent AI Systems in Healthcare…Systematic Review
Only minimal metadata is provided (title and DOI) with no abstract content to assess contribution details [Ref 13]. It is likely valuable as a healthcare-focused synthesis, but based on provided text, its scope, methods, and findings cannot be evaluated here. This is a gap in your source pack: include full text/abstract in your working notes before relying on it.
2.14 Multi agent AI for tactical maneuvering
This paper models a one-vs-one aerial dogfighting scenario as a two-person zero-sum perfect information game and applies simultaneous-move Monte Carlo Tree Search (MCTS) online, using self-play and demonstrating numerically in a simulated 2D application [Ref 14]. Strength: concrete formalization, algorithmic clarity, and simulation-based evaluation. Limitation: it is “multi-agent” in the sense of adversarial interaction rather than LLM-agent orchestration; direct relevance to LLM-based MAS is mainly methodological (game-theoretic decision-making, evaluation in simulation).
2.15 An Outlook on the Opportunities and Challenges of Multi-Agent AI Systems
This paper explicitly asks: when are MAS more effective than single-agent systems, what new safety risks arise from interactions, and how to evaluate reliability and structure [Ref 15]. It proposes a formal framework focused on effectiveness and safety and draws analogies to distributed estimation and sensor fusion, with experiments on data science automation [Ref 15]. Strength: directly identifies the field’s central open questions and sets a publishable agenda. Limitation: abstract-level detail only here; you will need the full paper to see the formalism and experimental setup.
2.16 Multi Agent Systems AI: What It Is & How It Works
This is a general introduction emphasizing decentralized load, parallelism, and structured protocols; it mentions context degradation and hallucinations in single-agent long workflows [Ref 16]. Strength: useful for background framing and terminology. Limitation: non-research overview without rigorous evaluation.
2.17 Advanced Game-Theoretic Frameworks…A 2025 Outlook
This paper discusses dynamic coalition formation, language-based utilities, sabotage risks, partial observability, and includes repeated games, Bayesian updates for adversarial detection, and moral framing in payoff structures [Ref 17]. Strength: brings formal game theory to agent interaction, including adversarial contexts. Limitation: “outlook” framing suggests it may be more conceptual/forecasting; empirical grounding (beyond “simulations and coding schemes”) is unclear from the excerpt.
2.18 Multi-Agent Risks from Advanced AI
This report provides a taxonomy of multi-agent risks with three failure modes (miscoordination, conflict, collusion) and seven risk factors (information asymmetries, network effects, selection pressures, destabilising dynamics, commitment problems, emergent agency, multi-agent security) [Ref 18]. Strength: clear, structured vocabulary that can be operationalized into threat models and benchmarks. Limitation: taxonomies need translation into measurable tests; the paper highlights directions but does not itself guarantee mitigations.
2.19 AI Agents in Engineering Design: A Multi-Agent Framework for…Car Design
This work proposes “Design Agents” that automate and accelerate tasks in automotive design (conceptual sketching, styling, 3D retrieval/generative modeling, CFD meshing, aerodynamic simulations) using VLMs, LLMs, and geometric deep learning, with industry-standard benchmarks and high-fidelity aerodynamic simulations [Ref 19]. Strength: strong domain grounding, tool integration, and claims of large cycle-time reductions. Limitation: without full paper details, it is unclear how the baseline comparisons were controlled and whether human-in-the-loop costs are accounted for.
2.20 AI for Explaining Decisions in Multi-Agent Environments
This paper argues explanation is especially important in multi-agent environments where goals may depend on other agents’ preferences, and proposes “Explainable decisions in Multi-Agent Environments (xMASE)” as a research direction, with emphasis on user satisfaction, fairness, envy, privacy, and environment properties [Ref 20]. Strength: establishes explanation as multi-factor and preference-dependent. Limitation: it frames a direction more than delivering a concrete algorithmic solution (based on excerpt).
2.21 Analysis: Capabilities, Limitations, and Premature Patterns of AI Agents
This meta-analysis-like document aims to classify promising but structurally limited patterns in LLM-based agent systems, explicitly excluding speculative AGI and marketing claims, and emphasizing observable/reproducible cases [Ref 21]. Strength: a falsifiability-oriented stance that aligns with publishable evaluation culture. Limitation: without detailed results in the excerpt, its key classifications and conclusions are not visible.
2.22 Guide to multi-agent systems (MAS)
This source defines MAS, contrasts with single-agent systems, and highlights distributed control and scalability to large numbers of agents [Ref 22]. Strength: clear definitions. Limitation: general overview, not research evidence.
2.23 Rapid and automated alloy design with GNN-powered LLM-driven multi-agent AI
This paper presents a multi-agent AI for alloy discovery combining LLMs, role-specialized agents, and a newly developed GNN model to predict properties (Peierls barrier; solute/screw dislocation interaction energy) for a ternary NbMoTa alloy system, reducing reliance on costly calculations and enabling autonomous exploration [Ref 23]. Strength: exemplary hybrid architecture where non-LLM models carry the “physics bottleneck,” improving feasibility. Limitation: generalizability beyond the specific alloy system and property targets must be established; also the reliability of autonomous trend identification requires careful validation.
2.24 Agent AI: Surveying the Horizons of Multimodal Interaction
This survey defines “Agent AI” as interactive systems that perceive multimodal inputs and produce embodied actions, arguing grounded environments can mitigate hallucinations and improve context awareness [Ref 24]. Strength: connects embodiment/multimodality to reliability. Limitation: survey-level; needs operational benchmarks for claims.
2.25 AI agents and agentic systems: A multi-expert analysis
This multi-expert perspective argues agents reshape industries via decentralized decision-making, organizational restructuring, and cross-functional collaboration, citing applications in healthcare, supply chain, etc. [Ref 25]. Strength: socio-technical framing and multi-domain orientation. Limitation: expert perspective is valuable for hypotheses but not a substitute for empirical evaluation.
2.26 Beyond Single Systems: How Multi-Agent AI Is Reshaping Ethics in Radiology
This paper argues multi-agent radiological AI amplifies the “black box” problem into “compound opacity,” where traditional XAI for isolated predictions fails for distributed multi-step reasoning [Ref 26]. It frames an “autonomy–transparency paradox” in radiology with regulatory and trust implications [Ref 26]. Strength: precise articulation of explainability failure modes in multi-agent workflows. Limitation: needs concrete measurement tools for “compound opacity” and practical evaluation designs.
2.27 AI Agent for Education: von Neumann Multi-Agent System Framework
This paper proposes a von Neumann-inspired architecture decomposing agents into control, logic, storage, and I/O modules, defining operations such as task deconstruction, self-reflection, memory processing, tool invocation, and referencing techniques like Chain-of-Thought and multi-agent debate [Ref 27]. Strength: modular conceptualization with education focus. Limitation: educational effectiveness and learning outcomes require controlled studies; abstract does not specify evaluations.
2.28 AI Agents Meet Blockchain: A Survey on Secure and Scalable Collaboration for Multi-Agents
This survey examines synergy between agents and blockchain for secure/scalable collaboration, and identifies challenges: coordination complexity, interoperability, privacy, and future directions including governance and interpretability [Ref 28]. Strength: highlights decentralized infrastructure as enabler and source of constraints. Limitation: blockchain is often unnecessary overhead; publishable work must demonstrate concrete threat models and measurable benefits.
2.29 On the agenda(s) of research on multi-agent learning
This survey critiques MARL for lack of clarity about the problems being addressed, proposes five well-defined problems, and singles out one not adequately addressed [Ref 29]. Strength: methodological critique and problem-definition discipline—highly relevant because LLM-based MAS risks repeating “unclear problem” patterns. Limitation: it addresses MARL, so mapping to LLM-based MAS requires careful conceptual translation.
2.30 Best Multi-Agent AI Frameworks…CrewAI vs AutoGen vs LangGraph
This source claims benchmarking and tiers for frameworks and advises budgeting context/token costs [Ref 30]. Strength: practical cost awareness. Limitation: not a peer-reviewed benchmark; details of methodology not provided here.
2.31 IMPACTing SHOP: Putting an AI Planner Into a Multi-Agent Environment
This paper integrates the SHOP HTN planning system into the IMPACT multi-agent environment, defining A-SHOP and showing soundness and completeness under certain conditions [Ref 31]. Strength: strong formal properties and a clear integration formalism between planning and multi-agent environments. Limitation: the work predates LLMs; relevance is as a blueprint for formal guarantees and interfaces between planners and agents, which modern LLM-MAS research largely lacks.
2.32 Interactive Debugging and Steering of Multi-Agent AI Systems
This paper identifies developer challenges (reviewing long conversations, localizing errors, lack of interactive debugging tools, iterating on configurations) via interviews with five developers, then introduces AGDebugger with browsing/messaging UI, message edit/reset, and overview visualization [Ref 32]. It includes a two-part user study (14 participants) and finds interactive message resets are important for debugging [Ref 32]. Strength: empirically grounded HCI contribution with concrete tooling. Limitation: sample sizes are modest; broader generalization across domains and agent architectures needs replication.
2.33 8 Best Multi-Agent AI Frameworks for 2026
This source emphasizes enterprise needs: auditability, confidence scoring, human oversight, and alignment between framework choice and compliance-heavy workflows [Ref 33]. Strength: highlights evaluation and governance requirements. Limitation: descriptive list format; not research evidence.
2.34 The Coming Crisis of Multi-Agent Misalignment…
This position paper argues alignment in MAS is dynamic and interaction-dependent, shaped by social environment (collaborative/cooperative/competitive), and calls for simulation environments, benchmarks, and evaluation frameworks for alignment in interactive multi-agent contexts [Ref 34]. Strength: strong problem statement aligned with emerging risks. Limitation: it is a position paper; your publishable contribution would be to instantiate benchmarks/metrics and test interventions.
2.35 An AI Agent for Fully Automated Multi-Omic Analyses
This paper introduces AutoBA, an autonomous LLM-based agent for automated multi-omic analyses with minimal user input, step-by-step plans, multiple LLM backends (online/local), and an automated code repair mechanism to improve stability [Ref 35]. Strength: attention to privacy/security (local usage) and stability via code repair. Limitation: it is framed as a single agent rather than multi-agent; it is nonetheless a strong baseline/foil for multi-agent claims in bioinformatics.
2.36 Open Challenges in Multi-Agent Security…
This paper introduces “multi-agent security” as a field, arguing free-form protocols enable threats like secret collusion and coordinated swarm attacks; network effects spread privacy breaches, disinformation, jailbreaks, and data poisoning; and research is fragmented across many fields [Ref 36]. Strength: explicit articulation of new threat surface and trade-offs (security–utility; security–security). Limitation: being “open challenges,” it needs follow-on work converting taxonomies into test suites and mitigations.
2.37 Offshore Production Surveillance and Intervention Using Multi Agent AI
This paper describes an agentic AI framework for offshore production surveillance/intervention with multiple specialized agents (Data-QC, Well-Surveillance, Asset-Surveillance, Well-Screening, Model-Management, Log-Interpreter, plus others for knowledge and alarms) and rich datasets (production history, coordinates, interventions, petrophysical, completion) [Ref 37]. Strength: real industrial workflow decomposition and explicit agent roles. Limitation: evaluation details, reliability/safety constraints, and human oversight mechanisms are not provided in the abstract.
2.38 Grid-Wise Control for Multi-Agent Reinforcement Learning in Video Game AI
This work proposes a convolutional encoder-decoder policy that outputs grid-wise actions controlling varying numbers of agents, promoting communication via receptive fields and enabling fast parallel exploration; evaluated in StarCraft II tasks [Ref 38]. Strength: concrete MARL method addressing variable agent count. Limitation: not LLM-based; relevance is methodological and benchmark design (spatial, variable team sizes).
2.39 Interactive Debugging and Steering… (arXiv version)
This appears to be the arXiv version of the AGDebugger work with the same abstract and contributions [Ref 39]. Strength/limitations match Ref 32; you should avoid double-counting and cite one version consistently for a given claim.
2.40 Multi-agent systems in industry: current trends & future challenges
This 2013 paper introduces MAS in industrial applications (manufacturing, handling, logistics) and discusses road-blockers for adoption and future challenges [Ref 40]. Strength: long-view industrial perspective; useful to compare with current LLM-agent adoption narratives. Limitation: predates LLMs; its “road-blockers” may differ but can be reinterpreted as integration/governance/robustness challenges.
2.41 AI Agent Systems: Architectures, Applications, and Evaluation
This survey offers a taxonomy of agent architectures (reasoning, planning/control, tool calling), orchestration patterns, deployment settings, and stresses evaluation difficulties due to non-determinism, long-horizon credit assignment, tool/environment variability, and hidden costs like retries and context growth [Ref 41]. Strength: evaluation-aware and systems-centric; extremely relevant for publishable methodology. Limitation: survey; needs empirical instantiations.
2.42 Generic Multi-Agent AI Framework for Weighted Dynamic Corridor Price Optimisation
This paper proposes a theoretical taxonomy/ontology for a domain-specific multi-agent AI as an internal price advisor integrated with ERP and other data sources, referencing Nash equilibrium and principal-agent theory in cooperative/semi-cooperative models [Ref 42]. Strength: explicit economic/game-theoretic grounding and enterprise integration orientation. Limitation: it is “theoretical analysis”; empirical evaluation in real pricing settings is needed.
2.43 The Future of Systematic Reviews: AI and Multi-Agent Automation
This article claims AI tools reduce screening time and mentions recall rates and efficiency gains, and describes multi-agent orchestration across the review pipeline [Ref 43]. Strength: identifies SR automation as a compelling application with measurable process metrics. Limitation: specific performance numbers in a blog-like source require careful verification via primary studies before academic use.
2.44 Modeling AI-Human Collaboration as a Multi-Agent Adaptation
This paper formalizes AI-human collaboration via agent-based simulation using an NK model, distinguishing optimization-based AI search vs satisficing human adaptation, and finds sequencing effects (human-first then AI-refine can maximize joint performance in sequenced tasks) [Ref 44]. Strength: provides a formal lens for task architecture and sequencing—directly applicable to orchestrator design decisions. Limitation: simulation-based results may not translate directly to real LLM-agent systems; needs empirical confirmation.
2.45 Scalability in modeling and simulation systems for multi-agent, AI, and machine learning applications
This paper argues current simulation environments are not designed for decentralized intelligent systems at scale, and recommends investment in measuring scalability from cost–benefit perspective, tools for scalability dimensions, and formalizing specs of scalability requirements met by systems [Ref 45]. Strength: focuses on scalability as a formal requirement and measurement problem. Limitation: it discusses modeling/simulation environments generally; needs mapping to LLM-agent infra constraints (token throughput, latency, context growth).
2.46 Conversational AI Multi-Agent Interoperability…Universal Open APIs
This paper analyzes interoperability frameworks and proposes OVON (Open Voice Network) architecture with universal natural-language-based APIs, discovery specs, and manifest publication for service lookup [Ref 46]. Strength: concrete interoperability mechanism (discovery + manifest) suited to multi-agent ecosystems. Limitation: as an abstract, does not show adoption, performance, or security analysis of the API layer.
2.47 Explainable AI Models Applied to the Multi-agent Environment of Financial Markets
This paper uses gradient boosting decision trees to predict large S&P 500 price drops from 150 features and uses Shapley values for feature attribution and local explanation, analyzing March 2020 meltdown [Ref 47]. Strength: rigorous explainability method in a multi-agent real-world domain (markets). Limitation: it is not an agentic MAS paper; “multi-agent environment” here refers to the market, not coordinated AI agents.
2.48 Single Agent vs Multi-Agent: When to Build a Multi-Agent System
This post discusses agent design components, ReAct workflows, and includes a walkthrough of a Multi-Agent RAG system [Ref 48]. Strength: practical design heuristics and baseline selection guidance. Limitation: not research evidence; should not be used for quantitative claims.
2.49 Multi Agent System in AI - GeeksforGeeks
Introductory explanation of MAS components and interactions [Ref 49]. Strength: definitions. Limitation: not scholarly.
2.50 Mastering Multi-Agent Orchestration: Coordination Is the New Scale Frontier
This article argues single-agent systems fail under domain overload, governance complexity, and performance bottlenecks; multi-agent orchestration stabilizes via task decomposition, shared memory, standardized tools, and protocols, and claims high AI initiative failure rate due to governance/integration issues [Ref 50]. Strength: frames orchestration as core engineering challenge. Limitation: quantitative claims need primary sources; treat as motivation.
2.51 What is a Multi-Agent System? | IBM
Definition of MAS and agent components (LLMs, tools, memory, planning) [Ref 51]. Strength: clear conceptualization. Limitation: overview.
2.52 Multi-agent system - Wikipedia
Encyclopedic overview distinguishing MAS vs agent-based modeling and listing related topics [Ref 52]. Strength: terminology map. Limitation: not citable for research claims in academic venues.
2.53 Multi-Agent AI Platform Comparison 2026: Complete Tool Guide
Tool/platform comparison framing; excerpt is incomplete [Ref 53]. Strength: potentially useful market scan. Limitation: not research; incomplete content here.
2.54 CrewAI Review 2026 - Multi-Agent Framework (Pricing)
This review describes CrewAI’s role-based orchestration, open-source + paid platform features (visual building, tracing, guardrails), and cautions about production scaling [Ref 54]. Strength: practical operational concerns (tracing, scaling). Limitation: anecdotal evaluation; not research-grade.
2.55 The Best AI Agents in 2026…Compared
This guide provides market sizing figures and agent component breakdown [Ref 55]. Strength: high-level adoption context. Limitation: market claims require independent verification for academic use; not MAS-specific evidence.
2.56 Multi-Agent Collaboration Mechanisms: A Survey of LLMs
This survey proposes an extensible framework characterizing collaboration mechanisms by actors, types (cooperation/competition/coopetition), structures, strategies, and coordination protocols, and identifies lessons and open challenges [Ref 56]. Strength: strong taxonomy tailored to LLM-based MAS collaboration. Limitation: survey; operational definitions and evaluation metrics need adoption.
2.57 Creativity in LLM-based Multi-Agent Systems: A Survey (GitHub)
This survey argues creativity is overlooked by existing MAS surveys and proposes a roadmap for creative MAS evaluation in multimodal text/image artifacts [Ref 57]. Strength: identifies a distinct evaluation dimension (novelty + value) often missing in MAS benchmarks. Limitation: GitHub-hosted survey; ensure stable versioning and peer-review considerations when citing.
2.58 State of AI Agents (enterprise report)
This report claims multi-agent systems in enterprises grew by 327% in less than four months and that evaluation/governance tool usage correlates with more projects reaching production [Ref 58]. Strength: motivates evaluation and governance as adoption accelerators. Limitation: enterprise reports may not disclose methodology; use for context, not as scientific evidence unless methods are transparent.
2.59 AI Agent Adoption 2026: What the Data Shows | Gartner, IDC
Only title-level content is provided, no data or methods [Ref 59]. You should not rely on it without obtaining the underlying analyst reports or details.
2.60 Generative Multi-Agent Collaboration in Embodied AI: A Systematic Review
This survey provides a taxonomy by architectures and embodiment modalities, analyzes building blocks (perception, planning, communication, feedback), and discusses challenges/future directions [Ref 60]. Strength: systematic review framing; useful for embodied MAS research directions. Limitation: abstract-level only here.
2.61 AI agent in healthcare: applications, evaluations, and future directions
This review traces AI agent evolution in healthcare applications and analyzes evaluation frameworks and metrics, proposing seven future directions including embodied systems, hybrid expert models, expanded evaluation paradigms, safety/controllability, ethical governance, and role guidance for staff [Ref 61]. Strength: explicitly evaluation-focused and future-direction oriented in a high-stakes domain. Limitation: healthcare-specific; generalizing to other domains requires care.
2.62 Unified multimodal GenAI platform integrating GraphRAG multi-agent systems…
This paper proposes a unified platform integrating multi-agent systems with GraphRAG to address relational consistency and multi-document aggregation issues, supporting tasks like QA, entity extraction, text-to-SQL, and fact verification, describing a modular pipeline with 5 conceptual layers [Ref 62]. Strength: bridges multi-agent orchestration with knowledge graph retrieval for multi-task enterprise document reasoning. Limitation: the “conceptual layers” phrasing suggests parts may be architectural rather than fully validated; evaluation details are not in excerpt.
2.63 Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System
This work proposes a multi-agent system for end-to-end automated meta-analysis with tool calls, using strategies like hybrid review, hierarchical extraction, self-proving, and feedback checking to reduce hallucinations in screening and extraction [Ref 63]. It constructs a benchmark with 729 papers across 3 domains and >10,000 data points, reporting improvements over an LLM baseline [Ref 63]. Strength: benchmark creation + concrete hallucination-mitigation strategies + end-to-end evaluation. Limitation: needs careful scrutiny of benchmark labeling, leakage risks, and reproducibility (but the abstract signals seriousness).
2.64 Can ‘Deep Research’ agents…perform systematic review and meta-analysis?
This paper evaluates and compares performances of AI agents autonomously conducting SRMAs, noting end-to-end ability remains untested and setting up a comparative evaluation [Ref 64]. Strength: directly challenges “autonomous research” claims with evaluation. Limitation: excerpt does not show results; still important as a framing reference.
2.65 How we built our multi-agent research system
This engineering write-up describes a multi-agent research feature with a planner agent spawning parallel search agents, and highlights challenges in coordination, evaluation, and reliability learned from prototype to production [Ref 65]. Strength: concrete production lessons and architecture patterns. Limitation: not academic; still useful to derive realistic failure modes and constraints for research hypotheses.
2.66 MetaMind: A Multi-Agent Transformer-Driven Framework for Automated Network Meta-Analyses
Only partial snippet is shown (background line) without methodological details [Ref 66]. Obtain full text before using.
2.67 AI-Driven Research Assistant (GitHub README)
This README describes a multi-agent research assistant using LangChain, OpenAI GPT models, and LangGraph, with supervisor/critic agents and a “Note Taker Agent” to maintain project state as an alternative to transmitting full history [Ref 67]. Strength: highlights practical memory/context management idea (“state note taking”). Limitation: a code repository description is not validated research; still may inspire publishable ablation studies if formalized and evaluated.
2.68 Beyond Single Systems…Ethics in Radiology (PubMed record)
This appears to be the PubMed entry for the radiology ethics paper [Ref 68], duplicating Ref 26 at abstract level. Prefer citing the full PDF for substantive claims.
2.69 Preserving Cultural Identity…Translation Through Multi-Agent AI Systems
This paper proposes a multi-agent translation framework with agents for translation, interpretation, synthesis, and bias evaluation using CrewAI/LangChain, reporting improved performance over GPT-4o and emphasizing low-resource language cultural nuance [Ref 69]. Strength: clear multi-agent role decomposition + bias evaluation role. Limitation: translation quality evaluation must be carefully designed (human evaluation, error taxonomy); claims require inspecting the full methodology.
2.70 AstroAgents: A Multi-Agent AI for Hypothesis Generation from Mass Spectrometry Data
This paper proposes an eight-agent system (data analyst, planner, three domain scientists, accumulator, literature reviewer using Semantic Scholar, critic) to generate hypotheses from mass spectrometry data and papers [Ref 70]. Strength: strong example of multi-agent hypothesis generation pipeline and explicit agent roster. Limitation: hypothesis quality evaluation is difficult; needs expert scoring protocols, novelty checks, and contamination controls.
2.71 Collaboration Promotes Group Resilience in Multi-Agent RL
This paper introduces “group resilience” in MARL and empirically evaluates collaboration protocols, reporting that collaborative approaches achieve higher group resilience than non-collaborative counterparts [Ref 71]. Strength: defines a measurable robustness property and tests it. Limitation: MARL setting; still offers a concept (resilience under perturbation) that LLM-MAS evaluation could adapt.
2.72 TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management…
This review adapts TRiSM to agentic multi-agent systems, proposes a risk taxonomy, and introduces two metrics: Component Synergy Score (CSS) and Tool Utilization Efficacy (TUE) [Ref 72]. Strength: explicitly metric-driven governance and lifecycle framing. Limitation: as a review, metrics need external validation and adoption.
2.73 AI Hiring with LLMs: A Context-Aware and Explainable Multi-Agent Framework…
This paper proposes a four-agent resume screening framework (extractor, evaluator with RAG, summarizer, score formatter) and compares AI scores with HR professional ratings on anonymized resumes [Ref 73]. Strength: concrete human-grounded evaluation setup and explainability orientation. Limitation: fairness/bias evaluation must be central for hiring; abstract does not specify robust fairness auditing.
2.74 Simulating Cooperative Prosocial Behavior with Multi-Agent LLMs…
This work tests whether multi-agent LLM systems replicate human public goods game behavior under multiple treatments and explores “unbounded actions” beyond lab constraints [Ref 74]. Strength: strong bridge between multi-agent LLM simulation and social science experimental paradigms. Limitation: simulation validity and sensitivity to prompting/model choice must be carefully managed.
2.75 Temporal Spectrum Cartography…Multi-Agent Learning
This paper proposes a two-stage GenAI framework with a masked autoencoder for spectrum map reconstruction and a multi-agent diffusion policy to optimize UAV sensor trajectories, validated by numerical experiments [Ref 75]. Strength: combination of generative modeling + multi-agent RL for sensing/control. Limitation: domain-specific; not LLM-MAS.
2.76 IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
This work proposes an open-source framework to generate synthetic benchmarks via policy-driven graph modeling and interactive simulations, offering fine-grained diagnostics compared to static benchmarks [Ref 76]. Strength: evaluation infrastructure as primary contribution. Limitation: synthetic benchmarks risk realism gaps; must validate against real transcripts or user studies.
2.77 MOSAIC: Modeling Social AI for Content Dissemination and Regulation…
This paper provides a social network simulation with LLM agents and evaluates content moderation strategies for misinformation, analyzing reasoning vs engagement patterns and open-sourcing software [Ref 77]. Strength: large-scale simulation as testbed for emergent behaviors and regulation strategies. Limitation: simulation-to-reality validity is a major concern; needs calibration.
2.78 Mapping Student-AI Interaction Dynamics in Multi-Agent Learning Environments…
This study uses an online learning platform with multiple AI agents, includes 305 students and 19,365 dialogue lines, identifies engagement patterns, and reports differential learning gains/motivation by prior knowledge [Ref 78]. Strength: strong sample size, mixed methods (tests + dialogue analysis), and actionable pedagogical insights. Limitation: generalization across subjects/institutions requires replication; also agent design details matter.
2.79 Fairness in Agentic AI: A Unified Framework…
This survey treats fairness as emergent property of agent interactions, integrating constraints, mitigation strategies, and incentives, claiming empirical validation and improved equity [Ref 79]. Strength: fairness framed at system level rather than model level. Limitation: “empirical validation” is unspecified in excerpt; must inspect experiments carefully to avoid overclaiming.
2.80 Prompt Injection Detection and Mitigation via AI Multi-Agent NLP Frameworks
This paper proposes multi-agent layered detection/enforcement for prompt injection and evaluates on 500 engineered prompts, introducing multiple metrics and a composite vulnerability score [Ref 80]. Strength: operational security metrics and explicit evaluation dataset size. Limitation: “engineered prompts” may not reflect real attacker distributions; still valuable as benchmark seed.
2.81 TAMA: Human-AI Collaborative Thematic Analysis…Clinical Interviews
This paper proposes a multi-agent LLM framework for thematic analysis with human expert coordination, reporting improvements (thematic hit rate, coverage, distinctiveness) on clinical interviews [Ref 81]. Strength: clear evaluation metrics and human-in-the-loop design. Limitation: qualitative analysis validity must be ensured (inter-rater agreement, audit trail).
2.82 Multi-Agent Penetration Testing AI for the Web
This paper introduces MAPTA and reports performance on the 104-challenge XBOW benchmark with detailed success rates and cost analysis [Ref 82]. Strength: unusually strong quantitative evaluation including costs and early-stopping heuristics, and end-to-end exploit validation. Limitation: focused on security testing; also it is dated 2508, which may be beyond your current project timeframe, but it is highly informative for evaluation culture.
2.83 AI Hospital: Benchmarking LLMs in a Multi-agent Medical Interaction Simulator
This paper introduces AI Hospital with multi-agent NPCs (Patient, Examiner, Chief Physician) and a benchmark MVME using Chinese medical records; proposes dispute resolution mechanism; finds multi-turn interaction performance gaps persist [Ref 83]. Strength: benchmark + mechanism + finding that multi-turn realism exposes weaknesses. Limitation: language/culture specificity (Chinese records) may limit generalization; still valuable.
2.84 Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance…
This paper proposes a modular policy-driven runtime enforcement layer that intercepts and logs actions, using declarative rules and a Trust Factor scoring mechanism, evaluated in simulations including adversarial agents [Ref 84]. Strength: governance decoupled from agent internals; strong for auditable compliance. Limitation: simulation regimes and rule design may not match real enterprise policy complexity; needs real-world pilots.
2.85 Accelerating Drug Discovery Through Agentic AI…DMTA Cycle
This paper introduces “Tippy,” a five-agent system automating the DMTA cycle with Safety Guardrail oversight, claiming production-ready implementation and workflow efficiency improvements [Ref 85]. Strength: strong workflow mapping to a canonical drug discovery cycle. Limitation: abstract claims “first production-ready” and “significant improvements” need close verification; also safety validation in wet-lab context is critical.
2.86 DrugAgent: Automating AI-aided Drug Discovery Programming…
This paper proposes a planner/instructor multi-agent framework for drug discovery ML programming and reports ROC-AUC improvement over ReAct on DTI tasks [Ref 86]. Strength: provides a quantitative baseline comparison and public availability (per abstract). Limitation: “ReAct” baseline choice and fairness of tool access must be checked.
2.87 Are AI agents the new machine translation frontier?
This paper analyzes single vs multi-agent MT and describes a pilot legal MT study with specialized agents (translation, adequacy review, fluency review, final editing) suggesting improved quality and domain adaptability [Ref 87]. Strength: maps professional translation workflow into agent roles. Limitation: pilot nature; needs robust human evaluation and error typology.
2.88 Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
This paper formalizes secret collusion via steganography in communicating agents, proposes mitigations, and provides an evaluation framework testing collusion capabilities; notes steganographic capabilities limited but suggests monitoring as capabilities grow [Ref 88]. Strength: concrete threat model + evaluation framework—highly publishable direction for follow-up. Limitation: mitigation effectiveness and false positive trade-offs likely complex; needs systematic study.
2.89 AgentMesh: A Cooperative Multi-Agent Generative AI Framework for Software Development Automation
This paper proposes a planner/coder/debugger/reviewer pipeline and discusses limitations such as error propagation and context scaling [Ref 89]. Strength: maps software lifecycle into agent roles and candidly notes limitations. Limitation: case-study evidence may not generalize; needs standardized benchmarks and cost metrics.
2.90 Advancing Geometry with AI: Multi-agent Generation of Polytopes
This paper claims a highly parallel AI system generated novel polytopes surpassing best known bounds in multiple problems, emphasizing massive parallel example generation [Ref 90]. Strength: demonstrates multi-agent/parallel search can yield scientific discoveries. Limitation: details of “multi-agent” vs distributed search need careful interpretation; also verification of mathematical novelty is crucial.
2.91 MA-RAG: Multi-Agent Retrieval-Augmented Generation…
This paper proposes specialized agents for RAG stages and reports improvements on multi-hop QA benchmarks, with ablations showing planner/extractor importance [Ref 91]. Strength: strong benchmark orientation and modular RAG decomposition. Limitation: it mentions chain-of-thought communication; publication strategy must consider how to evaluate without relying on hidden reasoning (and to ensure compliance with model provider policies).
2.92 AI Metropolis: Scaling LLM-based Multi-Agent Simulation with Out-of-order Execution
This paper introduces a simulation engine improving parallelism by tracking real dependencies and achieving 1.3x–4.15x speedups [Ref 92]. Strength: systems contribution with clear performance evaluation. Limitation: simulation engine relevance depends on your target application; more impactful if coupled to benchmark tasks.
2.93 Optimizing Generative AI Networking…MAS and Mixture of Experts
This paper discusses MAS and MoE for AIGC networking and proposes a multi-agent-enabled MoE-PPO framework for 3D object generation and data transfer scenarios with simulation results [Ref 93]. Strength: cross-layer integration (AI + networking). Limitation: may be too domain-specific unless you target networking venues.
2.94 APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation…
This paper proposes an agentic pipeline with committee reviewers and feedback loops to generate verifiable multi-turn agent data, releasing synthetic trajectories and training models outperforming frontier models on benchmarks [Ref 94]. Strength: addresses data scarcity for agent training with “verifiable blueprint” approach. Limitation: synthetic data realism and benchmark alignment must be scrutinized; also claims about outperforming frontier models are strong and require careful replication.
2.95 Can We Trust AI Agents? Case Study…Ethical AI
This paper uses design science research to build a multi-agent debate prototype for AI ethics issues, evaluating via thematic analysis and comparative studies, reporting more extensive code/documentation generation than baseline trials and surfacing compliance terms (GDPR, EU AI Act) [Ref 95]. Strength: methodological transparency (DSR) and evaluation via multiple qualitative/quantitative lenses. Limitation: more code lines is not necessarily better; correctness and compliance must be verified.
2.96 Age of Information Minimization using Multi-agent UAVs…
This paper proposes mean field hybrid PPO with LSTM to reduce AoI in UAV swarms, reporting up to 45% and 57% AoI reductions compared to baselines [Ref 96]. Strength: clear quantitative improvements and baseline comparisons. Limitation: MARL/control domain; not directly LLM-MAS.
2.97 Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures
This survey compares classical MAS and foundation-model-based MAS in a closed-loop coordination framework (perception, communication, decision-making, control), and outlines open challenges and opportunities [Ref 97]. Strength: bridges classical and LFM-based MAS—useful for positioning your work as continuity plus novelty. Limitation: survey.
2.98 Multi-agent artificial intelligence designs novel catalysts for ultrafast water purification
This paper presents ECOMATS, integrating expert-validated knowledge graphs with seven fine-tuned LLMs to design catalysts, using a five-dimensional evaluation framework and a triple-agent blind review with score fusion to improve reliability [Ref 98]. Strength: high-quality exemplar of multi-agent scientific discovery with explicit reliability mechanism (blind review + fusion). Limitation: replication in other scientific domains is needed; evaluation framework design choices could be contested.
2.99 Multi-Agent collaboration patterns with Strands Agents and Amazon Nova
This article describes collaboration patterns (manager-agent, swarms, agent graph, workflow), claims multi-agent collaboration can improve success rates substantially, and emphasizes computational/cost demands (thousands of prompts per request) [Ref 99]. Strength: highlights cost/throughput as first-class constraints and introduces design pattern vocabulary. Limitation: quantitative claim (“up to 70% higher”) is not supported with primary evidence in excerpt.
2.100 LangGraph vs CrewAI vs AutoGen: AI Agent Framework Comparison [2026]
This article compares frameworks and suggests ReAct is prevalent, claiming LangGraph offers fine-grained control and production suitability, CrewAI faster development but less flexibility, AutoGen excels in negotiation but harder to debug [Ref 100]. Strength: practical decision guidance. Limitation: not a controlled evaluation; debugging difficulty is asserted.
2.101 Multi-Agent AI Systems Enterprise Guide 2026
This enterprise guide makes quantitative performance claims (e.g., 3x faster, 60% better accuracy; surge in inquiries) and argues failures are often orchestration/context-transfer issues at handoffs [Ref 101]. Strength: identifies handoff/context transfer as a likely failure locus—good research hypothesis. Limitation: quantitative claims require verification; treat as non-verified context.
2.102 Multi-Agent Systems: A Survey (IEEE)
This survey covers definitions, features, applications, challenges including coordination, security, and task allocation, and provides classification of applications and challenges [Ref 102]. Strength: broad canonical survey anchor (classical MAS). Limitation: not LLM-era specific; use for foundational concepts.
2.103 Large Language Model Based Multi-agents: A Survey of Progress and Challenges
This survey is explicitly about LLM-based multi-agents, listing authors and being part of IJCAI proceedings as given in the excerpt [Ref 103]. Strength: high-value anchor for LLM-MAS progress/challenges. Limitation: the excerpt lacks the abstract; you should read full to avoid mischaracterizing its taxonomy.
2.104 LLM-Based Autonomous Multi-Agent Systems: A Comprehensive Survey
This survey proposes a unified taxonomy across architecture paradigms, coordination mechanisms (including protocol families), planning strategies, and evaluation metrics; it identifies open challenges like scalability to 100+ agents, protocol standardization, and multi-agent alignment [Ref 104]. Strength: evaluation and protocol standardization emphasized—exactly where publishable gaps are. Limitation: “Anonymous Authors” indicates it may be a preprint draft; venue is not confirmed here.
2.105 How Multi-Agent AI Editorial Review Works (9 Specialized Agents Explained)
This article describes multi-agent editorial review as staged specialist passes with labeled outputs and early “gates,” arguing it scales better than single-prompt feedback [Ref 105]. Strength: provides a concrete example of workflow gating and role specialization. Limitation: domain is editorial; not research evidence.
2.106 Latterview: A Multi-Agent Framework for Systematic Review Automation…
This paper introduces LatteReview, a Python framework with configurable LLM agents for screening and abstraction, multi-round workflows, and evaluates on public systematic review datasets with junior agents and a senior agent resolving disagreements, using AUC and accuracy [Ref 106]. Strength: concrete evaluation design (junior/senior arbitration) on multiple datasets and structured outputs. Limitation: excerpt truncates results; still, the methodology is clearly publishable.
2.107 LLM Agents for Smart City Management…
This study tests three hypotheses (routing, RAG effectiveness, integration with urban systems) and reports strong pipeline selection accuracy (94–99%) and RAG-driven response accuracy improvements (17% and 55% on query types), with highest G-Eval scores when combining document DB + service APIs [Ref 107]. Strength: quantified improvements with a defined validation dataset (150 QA pairs) and clear ablation-style comparisons. Limitation: evaluation metrics like G-Eval need careful justification and potential human validation.
2.108 AI hospitals use agent-driven multi-agent systems… (survey on OSF)
This survey analyzes 72 studies on LLM-based medical agents (2023–2025), proposing taxonomy and challenges (roles, interaction patterns, reasoning, memory, tool integration) and future platform development for simulation and evaluation [Ref 108]. Strength: consolidates fragmented domain and identifies component-level challenges. Limitation: healthcare-specific; but evaluation ideas transfer.
2.109 Mesoscale impact of trader psychology…multi-agent AI approach
This paper describes a stock market simulator built around a multi-agent architecture with autonomous investor agents [Ref 109]. Strength: multi-agent simulation with behavioral grounding. Limitation: not LLM-based; still useful as a simulation paradigm.
2.110 PromptBio: A Multi-Agent AI Platform for Bioinformatics Data Analysis
This paper presents PromptBio with multi-agent PromptGenie and other modes, using specialized agents and prevalidated tools for reproducible bioinformatics workflows, discussing extensibility, monitoring, and compliance [Ref 110]. Strength: reproducibility and tool validation emphasis, which is often missing in LLM-agent papers. Limitation: needs rigorous benchmarking against human analysts and single-agent baselines.
2.111 LLM-informed multi-agent AI system for drone-based visual inspection…
This paper proposes a five-agent framework (Router, PathPlanner, Controller, Perceptioner, Retriever) and introduces a pipeline producing semantic point clouds and 3D Scene Graphs (3DSGs) for spatial-semantic reasoning, evaluated via simulations and lab experiments; it notes reliance on commercial API LLMs introduces latency/adaptation gaps [Ref 111]. Strength: strong integration of 3D representations (3DSG) as shared knowledge storage and reasoning engine, plus evaluation beyond simulation. Limitation: latency and domain adaptation are identified but not solved—an excellent research gap.
2.112 Multi-Agent Modeling and Simulation in the AI Age
This survey reviews multi-agent modeling and simulation (MAMS), hybrid modeling with system dynamics, MARL modeling/simulation, and large-scale multi-agent simulation platforms [Ref 112]. Strength: foundational for simulation-based evaluation methodology and platform selection. Limitation: not LLM-specific.
2.113 AstroAgents PDF version
Duplicate of Ref 70 content in PDF form [Ref 113]; treat as same source.
2.114 Frame Handshake Protocol (FHP) v1.0
This work introduces a deterministic coordination layer combining explicit frames/constraints/permissions with distributed systems primitives (shared state vectors, append-only event logs, two-phase commit) to make agent coordination auditable, replayable, and recoverable, addressing drift, frame shifts, and state divergence [Ref 114]. Strength: unusually concrete protocolization and “fail-closed” orientation that directly targets known coordination failure modes. Limitation: requires empirical validation across tasks and comparison with looser, natural-language-based coordination.
2.3 Comparative Summary Table
| Ref | Focus / Scope | Method or Approach | Key Contribution | Main Limitation or Gap |
|---|---|---|---|---|
| [Ref 1] | Interior design workflow | 4-agent workflow; 2 implemented + real projects; 2 conceptual | Demonstrates practical value of implemented agents and proposes advanced conceptual agents | Half conceptual; evaluation details limited in excerpt |
| [Ref 2] | Reliability/safety engineering for MAS | Negative feedback control framing; redundancy + complementary chains + failsafes | Closed-loop verification principles for multi-agent correctness | Evaluation details truncated; needs broader empirical validation |
| [Ref 4] | Conceptual framework for MAAI | Five-component layered architecture; research pathways | Disentangles technical/organizational/human dimensions | Conceptual; limited empirical evidence in abstract |
| [Ref 15] | Effectiveness & safety outlook | Formal framework; experiments on data science automation | Sharp open questions: when MAS beats single-agent; safety via interactions | Needs full details for operationalization |
| [Ref 18] | Multi-agent risk taxonomy | Failure modes + risk factors taxonomy | Clear vocabulary for emergent multi-agent failures | Needs translation into benchmarks/tests |
| [Ref 31] | Planning in multi-agent environments | HTN planning integration; soundness/completeness | Formal guarantees blueprint | Pre-LLM; adaptation needed |
| [Ref 32] | Debugging multi-agent teams | Interviews + AGDebugger tool + user study | Identifies debugging needs; shows value of message resets | Modest sample sizes; domain breadth needed |
| [Ref 36] | Multi-agent security agenda | Taxonomy + trade-offs | Positions security as distinct field for interacting agents | Fragmented; requires concrete test suites |
| [Ref 63] | Automated meta-analysis | Tool-calling multi-agent strategies + new benchmark | End-to-end evaluation; hallucination mitigation | Benchmark validity/reproducibility must be scrutinized |
| [Ref 83] | Medical interaction simulator | Multi-agent simulator + MVME benchmark + dispute resolution | Reveals multi-turn gaps; benchmark infra | Language/domain specificity |
| [Ref 84] | Runtime compliance/governance | Policy-driven enforcement layer + trust scoring | Decoupled, auditable governance mechanism | Simulation-only evidence in abstract |
| [Ref 91] | Multi-agent RAG | Planner/Extractor/etc. for multi-hop QA | Strong benchmark improvements + ablations | CoT-based communication raises evaluation/opacity issues |
| [Ref 98] | Scientific discovery (water catalysts) | KG + 7 fine-tuned LLMs + blind review + score fusion | Reliability-oriented evaluation framework for discovery | Needs replication; domain specificity |
| [Ref 111] | AEC drone inspection | 5-agent framework + 3DSG pipeline; sim + lab | 3DSGs as shared reasoning substrate; practical evaluation | Latency/adaptation gaps from API LLMs remain |
| [Ref 114] | Coordination protocol | Fail-closed protocol + event logs + 2PC | Auditable, replayable coordination against drift | Requires empirical comparisons & usability validation |
2.4 Overall Assessment of the Provided Sources
Your source set is unusually broad and current, spanning conceptual frameworks [Ref 4], engineering protocols [Ref 114], debugging tools with user studies [Ref 32], risk/security agendas [Ref 18, Ref 36, Ref 88], and domain implementations with nontrivial evaluation in healthcare and urban systems [Ref 83, Ref 107]. This diversity is a strength for designing publishable, cross-cutting research.
What is missing is a coherent *evaluation spine* connecting these works: standardized tasks, baselines, and metrics that allow apples-to-apples comparisons of multi-agent vs single-agent, and comparisons among coordination/governance patterns under controlled budgets (latency, token cost, tool-call limits). Several papers call for evaluation frameworks, but only a subset provide strong benchmarking artifacts [Ref 63, Ref 83, Ref 76]. This gap is precisely where high-impact publication opportunities lie.
3. The Broader Research Landscape
Research in this field has established that “multi-agent systems” is a long-standing AI area encompassing distributed problem solving, coordination, negotiation, and learning in stochastic games. Classical MAS emphasizes communication protocols, coordination mechanisms, and formal properties (e.g., convergence or equilibrium concepts), while MARL focuses on non-stationarity, credit assignment, and scalability in cooperative/competitive settings. The provided IEEE MAS survey indicates canonical challenges like coordination, security, and task allocation have been persistent for years [Ref 102], and critiques in MARL emphasize that unclear problem definitions impede progress [Ref 29].
The newer LLM-based MAS wave (often called agentic AI systems) reframes agents as tool-using, language-coordinated components where coordination and verification are implemented via prompts, schemas, and orchestrators rather than fixed symbolic protocols. Surveys now attempt to unify taxonomies across architecture paradigms, coordination mechanisms, and evaluation metrics [Ref 41, Ref 56, Ref 97, Ref 103, Ref 104]. This wave also renews connections to symbolic planning (e.g., HTN planning) and distributed systems engineering, but often without importing the formal guarantees seen in earlier planning-in-MAS work [Ref 31].
Ongoing debates and unresolved questions increasingly center on: (i) whether multi-agent systems genuinely add capability/robustness vs being a repackaging of ensembles or self-consistency methods—explicitly asked in the “Outlook” framing [Ref 15]; (ii) how interaction dynamics introduce novel safety failures (miscoordination, collusion), highlighted by risk taxonomies and alignment position pieces [Ref 18, Ref 34]; and (iii) how to evaluate long-horizon, tool-using, non-deterministic systems with hidden costs and variable environments, emphasized by agent evaluation surveys [Ref 41]. In high-stakes domains (healthcare, hiring, radiology), the challenge expands to governance, explainability, and accountability for multi-step distributed reasoning [Ref 26, Ref 61, Ref 73].
Publication venues and communities are split across AI systems, HCI, security, and domain-specific applied fields. The presence of AAAI-hosted material on explanation in multi-agent environments [Ref 20] and ACM-style work on debugging tools [Ref 32] suggests clear routes for both algorithmic and human-centered contributions. ArXiv surveys and frameworks indicate intense activity, but top-tier acceptance typically requires either (a) a new benchmark or dataset with strong baselines, (b) a demonstrably better coordination/reliability method with ablations and cost accounting, or (c) a real-world deployment study with rigorous evaluation and reproducibility.
4. Strengths of Existing Research
A major strength is the emergence of explicit architectural frameworks that separate concerns: foundation model capabilities, perception/action through data and tools, orchestration, workflow integration, and interfaces [Ref 4]. This layered viewpoint supports better experimental design because it clarifies what variable you are changing (coordination protocol vs memory design vs tool router).
Second, there is growing methodological seriousness around evaluation and benchmarking. Examples include end-to-end automated meta-analysis with a newly constructed benchmark of 729 papers and >10,000 data points [Ref 63], medical multi-agent simulators with dedicated benchmarks and mechanisms like dispute resolution [Ref 83], and benchmark generation frameworks for conversational agents using policy-driven graphs and simulations [Ref 76]. This trend is crucial: publishable work in 2026-era agentic AI increasingly demands reproducible, benchmarkable evidence.
Third, domain case studies demonstrate that multi-agent designs can be meaningfully integrated into complex workflows beyond toy tasks. In AEC drone inspection, the use of 3D Scene Graphs as knowledge storage and reasoning engines suggests a credible pathway to grounded, spatially consistent agent behavior, supported by simulation and lab evaluation [Ref 111]. In materials discovery, combining LLM agents with a GNN property predictor shows a pragmatic hybrid approach that reduces computational burden while retaining domain fidelity [Ref 23], and ECOMATS illustrates reliability mechanisms such as triple-agent blind review and score fusion in scientific discovery [Ref 98]. These are strong templates for publishable applied MAS.
5. Weaknesses and Limitations of Existing Research
First, comparative claims are often under-identified and under-controlled. Many sources assert that multi-agent systems are “more capable” or “more resilient,” but do not specify baselines, budgets, or task distributions. The “Outlook” paper’s question—when are MAS more effective than single-agent systems?—exists because the evidence is not yet settled [Ref 15]. Without matched compute/token cost, tool access parity, and careful single-agent prompt/program baselines, reviewer skepticism will remain high.
Second, coordination and reliability are frequently handled implicitly through natural language rather than explicit protocols, enabling failure modes like drift, frame shifts, and state divergence. The need for deterministic coordination layers is directly addressed by the Frame Handshake Protocol’s fail-closed design with shared state vectors, append-only logs, and two-phase commit [Ref 114]. But most empirical papers do not benchmark “protocolized” coordination against informal messaging; thus we lack evidence about what reliability improvements are achievable and at what usability cost.
Third, developer tooling and observability lag behind system complexity. Debugging studies show developers struggle to review long conversations, localize errors, and iterate on agent configurations, motivating tools like AGDebugger and demonstrating strategies such as message resets [Ref 32]. Yet debugging research remains small-scale and not yet integrated into standardized evaluation pipelines. For publication, this is both a weakness (field immaturity) and an opportunity (tool + benchmark co-design).
Fourth, safety/security is fragmented and under-instrumented. Risk taxonomies describe miscoordination, conflict, and collusion [Ref 18], and security papers argue that free-form agent protocols enable secret collusion and swarm attacks, with network effects spreading failures [Ref 36]. Work on secret collusion via steganography provides an evaluation framework and suggests monitoring is needed as model capabilities evolve [Ref 88]. However, the field lacks standardized red-team harnesses and benchmarks analogous to what exists in classical security testing—except for emerging examples like multi-agent penetration testing evaluated with costs and end-to-end validation [Ref 82]. This gap makes it hard to compare mitigations.
Finally, explainability and accountability are underdeveloped for multi-step, multi-agent reasoning. Radiology ethics work highlights “compound opacity” where conventional XAI breaks down in distributed workflows [Ref 26], and xMASE emphasizes that explanations must incorporate multiple agents’ preferences and fairness/privacy constraints [Ref 20]. But few works propose measurable explainability targets for MAS (e.g., what constitutes a sufficient explanation trace), leaving regulators and practitioners without actionable standards.
6. Critical Research Gaps
1. Causal, cost-aware evidence of when MAS beats strong single-agent baselines Unknown: Under what task structures, tool constraints, and budgets do multi-agent systems outperform single-agent systems in accuracy, robustness, and time-to-completion? Why it matters: Without this, MAS research risks being seen as engineering “folklore” rather than science; reviewers will demand controlled comparisons. Evidence it’s open: Explicitly raised as a key question in the effectiveness/safety outlook [Ref 15] and reinforced by evaluation challenges noted in agent surveys (non-determinism, hidden costs) [Ref 41]. Impact: HIGH.
2. Standardized evaluation metrics for coordination quality and system-level reliability Unknown: How to measure “coordination quality” beyond task success—e.g., error propagation, disagreement handling, verification coverage, and failure recovery. Why it matters: Multi-agent systems can succeed for the wrong reasons (lucky tool calls) or fail silently; we need system-property metrics. Evidence it’s open: TRiSM review proposes metrics like CSS and TUE, implying lack of standard metrics [Ref 72]; protocol work emphasizes drift and divergence issues [Ref 114]. Impact: HIGH.
3. Protocolized, auditable coordination vs natural-language coordination: effectiveness trade-offs Unknown: Do fail-closed protocols (frames, permissions, logs, commits) materially improve correctness and auditability without destroying flexibility and usability? Why it matters: This is the bridge from prototypes to regulated production, and a key differentiator for publishable systems work. Evidence it’s open: FHP proposes mechanisms but does not (in provided text) report broad empirical benchmarks [Ref 114]; debugging studies show coordination complexity is hard to manage [Ref 32]. Impact: HIGH.
4. Multi-agent security benchmarks and red-team harnesses for interaction-born threats Unknown: How to systematically test miscoordination, collusion, and tool-mediated attacks in MAS under realistic constraints. Why it matters: Interaction effects can amplify vulnerabilities; without benchmarks, mitigations cannot be compared. Evidence it’s open: “Open challenges” framing and fragmented research noted in multi-agent security agenda [Ref 36]; collusion/steganography work proposes evaluation but needs broader adoption [Ref 88]. Impact: HIGH.
5. Explainability for multi-agent, multi-step workflows (“compound opacity”) Unknown: What explanation artifacts (traces, rationales, provenance, preference alignment) enable user trust and regulatory oversight in MAS decisions. Why it matters: High-stakes domains (radiology, hiring, healthcare) require explanations that traditional XAI does not provide. Evidence it’s open: xMASE frames explanation as a new research direction and highlights difficulty of maximizing user satisfaction [Ref 20]. Impact: HIGH.
6. Human-in-the-loop interaction design and debugging at scale Unknown: Which interface primitives (reset, branch, replay, causal tracing) reliably help users steer/repair MAS with long interaction histories. Why it matters: Production systems fail without debuggability; also publishable in HCI/AI tooling venues. Evidence it’s open: AGDebugger addresses needs but studies remain limited and domain breadth is open [Ref 32]. Impact: MEDIUM-HIGH.
7. Grounding and shared representations for embodied/spatial MAS Unknown: How to design shared world models (e.g., scene graphs) and verification loops that reduce hallucinations and improve coordination in physical tasks. Why it matters: Embodied MAS are crucial for robotics, inspection, and healthcare; grounding is a path to reliability. Evidence it’s open: Embodied surveys note current work is narrow and simplified [Ref 8]; AEC inspection notes latency/adaptation gaps and provides a 3DSG approach that invites extension [Ref 111]. Impact: MEDIUM-HIGH.
7. Recommended Research Directions
7.1 Priority Research Opportunities
Recommendation 1: A Cost-Constrained Benchmark for “When MAS Wins”
- Research Question: Under matched tool access and fixed budgets (tokens, tool calls, latency), when does multi-agent orchestration outperform a single strong agent across task classes (modular vs sequenced; low vs high interdependence)?
- Why It Matters: Directly answers the field’s central effectiveness question [Ref 15] and aligns with evaluation complexity concerns (hidden retries/context growth) [Ref 41].
- Addresses Gap: Gap 1.
- Suggested Methodology:
- Build a task suite spanning: document reasoning (multi-doc), coding+testing, planning+execution, and fact verification (inspired by tasks in [Ref 62, Ref 91]).
- Define *budget profiles*: e.g., 5/20/50 tool calls; token caps; latency caps.
- Compare: single-agent ReAct-style baseline vs MAS patterns (pipeline, hierarchical, debate/consensus) with ablations on agent count and communication structure.
- Report: success, time, cost, variance across runs, failure typology.
- Expected Contribution: A publishable benchmark + empirical map of regimes where MAS is justified.
- Publication Potential: Strong for agent evaluation venues and AI systems venues; also aligns with surveys calling for expanded evaluation paradigms [Ref 41, Ref 61].
Recommendation 2: Protocolized Coordination Ablations (FHP vs “Free Chat”)
- Research Question: Does fail-closed protocol coordination (frames, constraints, permissions, event logs, 2PC) reduce drift and state divergence compared to natural-language coordination, and what is the performance/usability cost?
- Why It Matters: Auditability and replayability are production-critical; this is a clear systems research contribution [Ref 114].
- Addresses Gap: Gaps 2 & 3.
- Suggested Methodology:
- Implement two coordination modes in the same MAS: (i) free-form messaging, (ii) FHP-style structured coordination [Ref 114].
- Use tasks with known failure modes: long-horizon multi-step workflows and multi-document aggregation (motivated by [Ref 62, Ref 111]).
- Measure: divergence rate (state inconsistency), recovery success, audit completeness (can you replay to reproduce outcome), and developer time to debug (pair with AGDebugger-inspired metrics [Ref 32]).
- Expected Contribution: Evidence-driven protocol engineering results; could define new coordination metrics.
- Publication Potential: AI systems + software engineering + trustworthy AI venues; very publishable because it converts protocol ideas into measurable claims.
Recommendation 3: Negative-Feedback Reliability Patterns as a General MAS Design Primitive
- Research Question: Can control-inspired “distributed negative feedback” (redundancy + complementary evaluation + failsafes) reduce reasoning errors and tool misuse in MAS without excessive cost?
- Why It Matters: Moves beyond ad hoc “critic agents” into a principled closed-loop design [Ref 2].
- Addresses Gap: Gaps 2 & 3.
- Suggested Methodology:
- Implement DNFO-style structures: redundant checkers with engineered diversity; complementary evaluators for different facets; triple-layer failsafe [Ref 2].
- Evaluate on tasks where errors matter (fact verification, numerical reasoning, policy compliance).
- Compare against common baselines: single critic, self-consistency, debate. Track false positives/negatives, overhead, and “error of omission” capture rate.
- Expected Contribution: A generalizable reliability architecture and empirical trade-off curves.
- Publication Potential: Trustworthy AI + agent systems; reviewers value principled engineering tied to measurable improvements.
Recommendation 4: Multi-Agent Security Testbed for Collusion and Tool-Mediated Attacks
- Research Question: How do interaction patterns and communication channels enable collusion, covert channels, and coordinated attacks, and which mitigations measurably reduce risk?
- Why It Matters: Security agenda argues this is under-studied yet critical [Ref 36]; collusion via steganography provides a concrete threat model and evaluation framework [Ref 88].
- Addresses Gap: Gap 4.
- Suggested Methodology:
- Build a red-team harness with scenarios for miscoordination/conflict/collusion per taxonomy [Ref 18].
- Include covert-channel tests based on collusion evaluation frameworks [Ref 88].
- Evaluate mitigations: structured JSON-only channels, content filters, governance enforcement interceptors (inspired by runtime interception in GaaS [Ref 84]), and multi-agent prompt-injection defenses with metrics like ISR/POF/CCS [Ref 80].
- Expected Contribution: Reproducible security benchmark + mitigation comparisons.
- Publication Potential: Security + AI safety venues; likely high impact due to novelty and urgency.
Recommendation 5: Explainability Artifacts for Multi-Agent Decisions (Operationalizing “Compound Opacity”)
- Research Question: What explanation interfaces and artifacts increase user satisfaction and trust in multi-agent decisions when goals/preferences are multi-party, and how do they trade off with privacy and fairness?
- Why It Matters: Radiology ethics identifies compound opacity as a unique MAS problem [Ref 26]; xMASE frames explanation as essential and challenging [Ref 20].
- Addresses Gap: Gap 5.
- Suggested Methodology:
- Define explanation outputs: provenance graph (which agent/tool/source contributed), preference trace (whose constraints mattered), and counterfactual “why not” explanations.
- Conduct controlled user studies in a constrained domain (e.g., resume screening [Ref 73] or clinical triage simulation [Ref 83]) measuring satisfaction, perceived fairness, and ability to detect errors.
- Quantify explanation fidelity via replayability/logs (link to FHP-style event logs [Ref 114]).
- Expected Contribution: A measurable explainability framework for MAS, not just conceptual discussion.
- Publication Potential: HCI + AI ethics + healthcare AI venues.
Recommendation 6: Grounded Shared World Models for Embodied/Spatial MAS
- Research Question: Do shared structured representations (e.g., 3D Scene Graphs) reduce hallucinations and improve coordination efficiency in embodied multi-agent tasks, and what verification loops are needed?
- Why It Matters: Embodied AI surveys argue current multi-agent embodied work is narrow and simplified [Ref 8]; AEC inspection shows a promising 3DSG pipeline and points to latency/adaptation gaps [Ref 111].
- Addresses Gap: Gap 7.
- Suggested Methodology:
- Extend the 3DSG-based architecture: add verification agents that validate spatial claims against the graph; add “perception confidence” propagation.
- Evaluate in simulation + lab replication style like, with metrics: navigation success, defect detection accuracy, correction rate after verifier feedback, and latency/cost.
- Expected Contribution: A replicable grounded MAS design pattern with empirical evidence.
- Publication Potential: Robotics/embodied AI + AEC automation venues.
7.2 Quick-Win Opportunities
1. Agent-count and structure ablation on an existing task suite Use a limited set of tasks (e.g., multi-hop QA + fact verification) and compare 1-agent vs 3-agent vs 5-agent under fixed tool-call budgets, reporting variance and cost. Leverage MA-RAG’s agent roles as a template [Ref 91], but evaluate across multiple orchestration patterns.
2. Developer debugging study replication with your own MAS Replicate AGDebugger-style findings in your domain (e.g., document-processing MAS) by logging failures, measuring time-to-fix with/without message reset and replay tools [Ref 32]. Even a modest study can be publishable if coupled to a released dataset of interaction traces.
3. Prompt-injection multi-agent defense benchmark extension Start from the 500 engineered prompts evaluation concept and add tool-using contexts (RAG + APIs), reporting ISR/POF/CCS-style metrics and comparing governance interceptors [Ref 80, Ref 84]. This can yield a solid security workshop paper within 6–12 months.
8. How to Maximise Your Publication Success
High-impact reviewers in this area increasingly demand *evaluation rigor under constraints*. Treat token cost, tool-call count, latency, and variance across runs as first-class experimental variables, not engineering footnotes—this is explicitly motivated by evaluation challenges like retries and context growth [Ref 41] and by real multi-agent systems issuing many prompts per request [Ref 99]. Papers that show “accuracy improved” without cost curves are now easy to reject.
Position your contribution against the field’s own stated open questions. Use the framing from the effectiveness/safety outlook—“when MAS is more effective than single agents” and “what new safety risks arise”—as your narrative hook [Ref 15]. Then deliver something concrete: a benchmark, protocol, or metric suite. Benchmarks and open-source artifacts (as in AI Hospital and Manalyzer) are especially publishable because they become infrastructure others cite [Ref 83, Ref 63].
Avoid the common rejection pattern: “interesting system, unclear novelty.” The way to prevent this is to isolate one core innovation (e.g., fail-closed coordination [Ref 114] or negative feedback verification [Ref 2]) and run comprehensive ablations: remove components, vary agent diversity, vary communication schemas, and show exactly what drives gains. Also include a *strong single-agent baseline* with comparable tool access; otherwise reviewers will argue your result is just “more compute.”
Collaboration strategy: pair with at least one domain expert if your evaluation is in a regulated area (healthcare, hiring, radiology). Sources in these domains highlight governance, explainability, and staff role impacts as critical future directions [Ref 61, Ref 26]. A domain co-author strengthens dataset validity, evaluation metrics, and ethical review—often decisive for acceptance.
Finally, release reproducible assets whenever possible: interaction logs, red-team harnesses, benchmark tasks, and evaluation scripts. Multi-agent research suffers from irreproducibility due to hidden prompts and proprietary tools; artifact release is a competitive advantage and improves citation potential.
9. Suggested Research Framework
A coherent publishable agenda can be sequenced in three phases. Phase 1 (0–9 months): build a cost-constrained evaluation harness that compares MAS vs single-agent across a small but representative task suite, with rigorous budgets and variance reporting (Section 7.1 Recommendation 1). This yields a first strong paper and sets up your lab’s evaluation “spine.”
Phase 2 (9–18 months): introduce a coordination innovation—either protocolized coordination (FHP-style) [Ref 114] or negative-feedback optimization [Ref 2]—and evaluate it within the harness, producing a second paper focused on reliability/coordination metrics and auditable replay. Phase 3 (18–30 months): expand into security and governance by building a red-team benchmark for collusion/tool attacks [Ref 36, Ref 88] and testing governance-as-a-service enforcement [Ref 84] integrated with your protocol/logging stack. This creates a compelling program: effectiveness → reliability → security/governance, with reusable artifacts across papers.
10. Conclusion
The most important unresolved issue is not “how to build” multi-agent systems, but *how to prove* when they are better—and safer—than strong single-agent baselines under realistic cost and governance constraints [Ref 15, Ref 41]. Without that evidence, the field remains vulnerable to hype cycles and irreproducible claims.
The highest-priority recommendation is to create a cost-constrained, variance-aware evaluation and benchmarking framework, then use it to test protocolized coordination (fail-closed, auditable logs) and negative-feedback reliability patterns [Ref 114, Ref 2]. This combination is highly publishable because it converts widely discussed challenges—coordination failures, opacity, security risk—into measurable, reproducible science.
11. References (IEEE format)
[1] X. Li, P. Wu, Z. He, J. Li, and L. Fan, "Multi-agent AI: Collaborative Design with Multiple AI Tools in Interior Design Workflow | Springer Nature Link," Communications in computer and information science, 2024, doi: 10.1007/978-3-031-78531-3_23.
[2] T. Seki, "Distributed Negative Feedback Optimization for Multi-Agent AI," Zenodo (CERN European Organization for Nuclear Research), 2026, doi: 10.5281/zenodo.19160814.
[3] "What Are Multi-Agent AI Systems," n.d. [Online]. Available: https://www.runpod.io/articles/guides/what-are-multi-agent-ai-systems.
[4] S. Allmendinger, L. Bonenberger, K. Endres, D. Fetzer, H. Gimpel, and N. Kühl, "Multi-agent AI," Electronic Markets, 2026, doi: 10.1007/s12525-025-00862-z.
[5] "Multi-Agent Systems: Architecture, Applications & Real-World Impact," n.d. [Online]. Available: https://www.cognizant.com/us/en/ai-lab/blog/what-are-multi-agent-systems.
[6] "Multi-Agent Systems: How AI Agents Work Together (2026 Guide)," n.d. [Online]. Available: https://whatisagentic.ai/learn/multi-agent-systems/.
[7] "What Is a Multi-Agent AI System and How Does It Work?," n.d. [Online]. Available: https://trigma.com/artificial-intelligence/multi-agent-ai-systems/.
[8] "[2505.05108] Multi-agent Embodied AI: Advances and Future Directions," n.d. [Online]. Available: https://arxiv.org/abs/2505.05108.
[9] "Multi-Agent AI Systems Complete Guide 2026 - Calmops," n.d. [Online]. Available: https://calmops.com/ai/multi-agent-ai-systems-2026-complete-guide/.
[10] "AI Agent Orchestration: A 2026 Guide to Multi-Agent Systems," n.d. [Online]. Available: https://a-listware.com/blog/ai-agent-orchestration.
[11] "Multi-Agent Systems: The Complete Deep Dive into Collaborative AI - SO Development," n.d. [Online]. Available: https://so-development.org/multi-agent-systems-the-complete-deep-dive-into-collaborative-a….
[12] "Designing Multi-Agent Intelligence - Microsoft for Developers," n.d. [Online]. Available: https://developer.microsoft.com/blog/designing-multi-agent-intelligence.
[13] I. P. Nweke, C. O. Ogadah, K. Koshechkin, and P. M. Oluwasegun, "Multi-Agent AI Systems in Healthcare: A Systematic Review Enhancing Clinical Decision-Making," 2025, doi: 10.9734/ajmpcp/2025/v8i1288.
[14] K. Srivastava and A. Surana, "Multi agent AI for tactical maneuvering," 2022, doi: 10.1117/12.2617157.
[15] "[2505.18397] An Outlook on the Opportunities and Challenges of Multi-Agent AI Systems," n.d. [Online]. Available: https://arxiv.org/abs/2505.18397.
[16] "Multi Agent Systems AI: What It Is & How It Works," n.d. [Online]. Available: https://www.scaler.com/topics/multi-agent-systems-ai-what-it-is/.
[17] P. Malinovskiy, "ADVANCED GAME-THEORETIC FRAMEWORKS FOR MULTI-AGENT AI CHALLENGES: A 2025 OUTLOOK," International Research Journal of Modernization in Engineering Technology and Science, 2025, doi: 10.56726/irjmets69135.
[18] "[2502.14143] Multi-Agent Risks from Advanced AI," n.d. [Online]. Available: https://arxiv.org/abs/2502.14143.
[19] "[2503.23315] AI Agents in Engineering Design: A Multi-Agent Framework for Aesthetic and Aerodynamic Car Design," n.d. [Online]. Available: https://arxiv.org/abs/2503.23315.
[20] S. Kraus et al., "AI for Explaining Decisions in Multi-Agent Environments," Proceedings of the AAAI Conference on Artificial Intelligence, 2020, doi: 10.1609/aaai.v34i09.7077.
[21] "AI Agents Meta-Analysis 2025 - Capabilities,... | AiBrain," n.d. [Online]. Available: https://www.askaibrain.com/en/posts/meta-analysis-ai-agents/.
[22] "What is a multi-agent system in AI? | Google Cloud," n.d. [Online]. Available: https://cloud.google.com/discover/what-is-a-multi-agent-system.
[23] A. Ghafarollahi and M. J. Buehler, "Rapid and automated alloy design with graph neural network-powered large language model-driven multi-agent AI," MRS Bulletin, 2025, doi: 10.1557/s43577-025-00953-4.
[24] "[2401.03568] Agent AI: Surveying the Horizons of Multimodal Interaction," n.d. [Online]. Available: https://arxiv.org/abs/2401.03568.
[25] L. Hughes et al., ""AI agents and agentic systems: A multi-expert analysis" by Laurie Hughes, Yogesh K. Dwivedi et al," Journal of Computer Information Systems, 2025, doi: 10.1080/08874417.2025.2483832.
[26] S. Salehi, Y. Singh, P. Habibi, and B. J. Erickson, "Beyond Single Systems: How Multi-Agent AI Is Reshaping Ethics in Radiology," Bioengineering, 2025, doi: 10.3390/bioengineering12101100.
[27] "[2501.00083] AI Agent for Education: von Neumann Multi-Agent System Framework," n.d. [Online]. Available: https://arxiv.org/abs/2501.00083.
[28] M. M. Karim, D. H. Van, S. Khan, Q. Qu, and Y. Kholodov, "AI Agents Meet Blockchain: A Survey on Secure and Scalable Collaboration for Multi-Agents," Future Internet, 2025, doi: 10.3390/fi17020057.
[29] R. Powers, T. Grenager, and Y. Shoham, "On the agenda(s) of research on multi-agent learning," n.d. [Online]. Available: https://core.ac.uk/download/pdf/22875077.pdf.
[30] "Best Multi-Agent AI Frameworks for 2025 & 2026: CrewAI vs AutoGen vs LangGraph | LLM Practical Experience Hub," n.d. [Online]. Available: https://langcopilot.com/posts/2025-11-01-top-multi-agent-ai-frameworks-2024-guide.
[31] J. Dix, H. Muñoz‐Avila, D. Nau, and L. Zhang, "IMPACTing SHOP: Putting an AI Planner Into a Multi-Agent Environment | Annals of Mathematics and Artificial Intelligence | Springer Nature Link," Annals of Mathematics and Artificial Intelligence, 2003, doi: 10.1023/a:1021560510377.
[32] W. Epperson et al., "Interactive Debugging and Steering of Multi-Agent AI Systems," 2025, doi: 10.1145/3706598.3713581.
[33] "8 Best Multi-Agent AI Frameworks for 2026," n.d. [Online]. Available: https://www.multimodal.dev/post/best-multi-agent-ai-frameworks.
[34] "[2506.01080] The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process," n.d. [Online]. Available: https://arxiv.org/abs/2506.01080.
[35] J. Zhou et al., "An AI Agent for Fully Automated Multi‐Omic Analyses," Advanced Science, 2024, doi: 10.1002/advs.202407094.
[36] "[2505.02077] Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents," n.d. [Online]. Available: https://arxiv.org/abs/2505.02077.
[37] D. Shekhawat, J. Barua, K. Bhatia, and S. Saumya, "Offshore Production Surveillance and Intervention Using Multi Agent AI," 2025, doi: 10.2118/226728-ms.
[38] L. Han et al., "Grid-Wise Control for Multi-Agent Reinforcement Learning in Video Game AI," Rare & Special e-Zone (The Hong Kong University of Science and Technology), 2019. [Online]. Available: http://repository.hkust.edu.hk/ir/Record/1783.1-103623.
[39] "[2503.02068] Interactive Debugging and Steering of Multi-Agent AI Systems," n.d. [Online]. Available: https://arxiv.org/abs/2503.02068.
[40] K. Schild, L. Monostori, M. Wooldridge, P. Leitão, and V. Mařík, "Multi-agent systems in industry: current trends & future challenges," 'Springer Science and Business Media LLC', 2013, doi: 10.1007/978-3-642-34422-0_13.
[41] "AI Agent Systems: Architectures, Applications, and Evaluation," n.d. [Online]. Available: https://arxiv.org/html/2601.01743v1.
[42] W. Kurz, "Generic Multi-Agent AI Framework for Weighted Dynamic Corridor Price Optimisation," Journal of Next-Generation Research 5 0, 2025, doi: 10.70792/jngr5.0.v1i2.65.
[43] "The Future of Systematic Reviews: AI and Multi-Agent Automation | SystematicReviewTools.app," n.d. [Online]. Available: https://systematicreviewtools.app/blog/systematic-review-technology.
[44] "[2504.20903] Modeling AI-Human Collaboration as a Multi-Agent Adaptation," n.d. [Online]. Available: https://arxiv.org/abs/2504.20903.
[45] C. Newton, J. S. Singleton, C. Copland, S. Kitchen, and J. Hudack, "Scalability in modeling and simulation systems for multi-agent, AI, and machine learning applications," 2021, doi: 10.1117/12.2585723.
[46] "[2407.19438] Conversational AI Multi-Agent Interoperability, Universal Open APIs for Agentic Natural Language Multimodal Communications," n.d. [Online]. Available: https://arxiv.org/abs/2407.19438.
[47] J. J. Ohana, S. Ohana, E. Benhamou, D. Saltiel, and B. Guez, "Explainable AI (XAI) Models Applied to the Multi-agent Environment of Financial Markets | Springer Nature Link," Lecture notes in computer science, 2021, doi: 10.1007/978-3-030-82017-6_12.
[48] "Single Agent vs Multi-Agent: When to Build a Multi-Agent System | Towards Data Science," n.d. [Online]. Available: https://towardsdatascience.com/single-agent-vs-multi-agent-when-to-build-a-multi-agent-sys….
[49] "Multi Agent System in AI - GeeksforGeeks," n.d. [Online]. Available: https://www.geeksforgeeks.org/artificial-intelligence/multi-agent-system-in-ai/.
[50] "Multi-Agent AI Orchestration Guide & 2026 Updates," n.d. [Online]. Available: https://www.codebridge.tech/articles/mastering-multi-agent-orchestration-coordination-is-t….
[51] "What is a Multi-Agent System? | IBM," n.d. [Online]. Available: https://www.ibm.com/think/topics/multiagent-system.
[52] "Multi-agent system - Wikipedia," n.d. [Online]. Available: https://en.wikipedia.org/wiki/Multi-agent_system.
[53] "Multi-Agent AI Platform Comparison 2026: Complete Tool Guide," n.d. [Online]. Available: https://promethium.ai/guides/multi-agent-ai-platform-comparison-2026/.
[54] "CrewAI Review 2026 - Multi-Agent Framework (Pricing)," n.d. [Online]. Available: https://vibecoding.app/blog/crewai-review.
[55] "The Best AI Agents in 2026: Tools, Frameworks, and Platforms Compared | DataCamp," n.d. [Online]. Available: https://www.datacamp.com/blog/best-ai-agents.
[56] "[2501.06322] Multi-Agent Collaboration Mechanisms: A Survey of LLMs," n.d. [Online]. Available: https://arxiv.org/abs/2501.06322.
[57] "GitHub - MiuLab/MultiAgent-Survey: Survey Paper for Multi-Agent Interaction and Collaboration · GitHub," n.d. [Online]. Available: https://github.com/MiuLab/MultiAgent-Survey.
[58] "2026 State of AI Agents: Enterprise Insights on Building AI | Databricks," n.d. [Online]. Available: https://www.databricks.com/resources/ebook/state-of-ai-agents.
[59] "AI Agent Adoption 2026: What the Data Shows | Gartner, IDC," n.d. [Online]. Available: https://joget.com/ai-agent-adoption-in-2026-what-the-analysts-data-shows/.
[60] "[2502.11518] Generative Multi-Agent Collaboration in Embodied AI: A Systematic Review," n.d. [Online]. Available: https://arxiv.org/abs/2502.11518.
[61] "AI agent in healthcare: applications, evaluations, and future directions | npj Artificial Intelligence," n.d. [Online]. Available: https://www.nature.com/articles/s44387-026-00076-4.
[62] "A unified multimodal GenAI platform integrating GraphRAG multi-agent systems and custom language models for intelligent document processing and knowledge synthesis | Scientific Reports," n.d. [Online]. Available: https://www.nature.com/articles/s41598-026-47145-x.
[63] "[2505.20310] Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System," n.d. [Online]. Available: https://arxiv.org/abs/2505.20310.
[64] "Can ‘Deep Research’ agents and general AI agentic systems autonomously perform systematic review and meta-analysis? | Eye," n.d. [Online]. Available: https://www.nature.com/articles/s41433-025-04138-w.
[65] "How we built our multi-agent research system \ Anthropic," n.d. [Online]. Available: https://www.anthropic.com/engineering/multi-agent-research-system.
[66] "MetaMind: A Multi-Agent Transformer-Driven Framework for Automated Network Meta-Analyses | medRxiv," n.d. [Online]. Available: https://www.medrxiv.org/content/10.1101/2025.08.04.25332893v1.
[67] "ai-data-analysis-MultiAgent/README.md at main · aimped-ai/ai-data-analysis-MultiAgent · GitHub," n.d. [Online]. Available: https://github.com/aimped-ai/ai-data-analysis-MultiAgent/blob/main/README.md.
[68] "Beyond Single Systems: How Multi-Agent AI Is Reshaping Ethics in Radiology," n.d. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/41155099/.
[69] "[2503.04827] Preserving Cultural Identity with Context-Aware Translation Through Multi-Agent AI Systems," n.d. [Online]. Available: https://arxiv.org/abs/2503.04827.
[70] "[2503.23170] AstroAgents: A Multi-Agent AI for Hypothesis Generation from Mass Spectrometry Data," n.d. [Online]. Available: https://arxiv.org/abs/2503.23170.
[71] "[2111.06614] Collaboration Promotes Group Resilience in Multi-Agent RL," n.d. [Online]. Available: https://arxiv.org/abs/2111.06614.
[72] "[2506.04133] TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems," n.d. [Online]. Available: https://arxiv.org/abs/2506.04133.
[73] "[2504.02870] AI Hiring with LLMs: A Context-Aware and Explainable Multi-Agent Framework for Resume Screening," n.d. [Online]. Available: https://arxiv.org/abs/2504.02870.
[74] "[2502.12504] Simulating Cooperative Prosocial Behavior with Multi-Agent LLMs: Evidence and Mechanisms for AI Agents to Inform Policy Decisions," n.d. [Online]. Available: https://arxiv.org/abs/2502.12504.
[75] "[2505.15571] Temporal Spectrum Cartography in Low-Altitude Economy Networks: A Generative AI Framework with Multi-Agent Learning," n.d. [Online]. Available: https://arxiv.org/abs/2505.15571.
[76] "[2501.11067] IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems," n.d. [Online]. Available: https://arxiv.org/abs/2501.11067.
[77] "[2504.07830] MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations," n.d. [Online]. Available: https://arxiv.org/abs/2504.07830.
[78] "[2506.02993] Mapping Student-AI Interaction Dynamics in Multi-Agent Learning Environments: Supporting Personalised Learning and Reducing Performance Gaps," n.d. [Online]. Available: https://arxiv.org/abs/2506.02993.
[79] "[2502.07254] Fairness in Agentic AI: A Unified Framework for Ethical and Equitable Multi-Agent System," n.d. [Online]. Available: https://arxiv.org/abs/2502.07254.
[80] "[2503.11517] Prompt Injection Detection and Mitigation via AI Multi-Agent NLP Frameworks," n.d. [Online]. Available: https://arxiv.org/abs/2503.11517.
[81] "[2503.20666] TAMA: A Human-AI Collaborative Thematic Analysis Framework Using Multi-Agent LLMs for Clinical Interviews," n.d. [Online]. Available: https://arxiv.org/abs/2503.20666.
[82] "[2508.20816] Multi-Agent Penetration Testing AI for the Web," n.d. [Online]. Available: https://arxiv.org/abs/2508.20816.
[83] "[2402.09742] AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator," n.d. [Online]. Available: https://arxiv.org/abs/2402.09742.
[84] "[2508.18765] Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance and Policy Enforcement," n.d. [Online]. Available: https://arxiv.org/abs/2508.18765.
[85] "[2507.09023] Accelerating Drug Discovery Through Agentic AI: A Multi-Agent Approach to Laboratory Automation in the DMTA Cycle," n.d. [Online]. Available: https://arxiv.org/abs/2507.09023.
[86] "[2411.15692] DrugAgent: Automating AI-aided Drug Discovery Programming through LLM Multi-Agent Collaboration," n.d. [Online]. Available: https://arxiv.org/abs/2411.15692.
[87] "[2504.12891] Are AI agents the new machine translation frontier? Challenges and opportunities of single- and multi-agent systems for multilingual digital communication," n.d. [Online]. Available: https://arxiv.org/abs/2504.12891.
[88] "[2402.07510] Secret Collusion among AI Agents: Multi-Agent Deception via Steganography," n.d. [Online]. Available: https://arxiv.org/abs/2402.07510.
[89] "[2507.19902] AgentMesh: A Cooperative Multi-Agent Generative AI Framework for Software Development Automation," n.d. [Online]. Available: https://arxiv.org/abs/2507.19902.
[90] "[2502.05199] Advancing Geometry with AI: Multi-agent Generation of Polytopes," n.d. [Online]. Available: https://arxiv.org/abs/2502.05199.
[91] "[2505.20096] MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning," n.d. [Online]. Available: https://arxiv.org/abs/2505.20096.
[92] "[2411.03519] AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution," n.d. [Online]. Available: https://arxiv.org/abs/2411.03519.
[93] "[2405.12472] Optimizing Generative AI Networking: A Dual Perspective with Multi-Agent Systems and Mixture of Experts," n.d. [Online]. Available: https://arxiv.org/abs/2405.12472.
[94] "[2504.03601] APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay," n.d. [Online]. Available: https://arxiv.org/abs/2504.03601.
[95] "[2411.08881] Can We Trust AI Agents? A Case Study of an LLM-Based Multi-Agent System for Ethical AI," n.d. [Online]. Available: https://arxiv.org/abs/2411.08881.
[96] "[2405.00056] Age of Information Minimization using Multi-agent UAVs based on AI-Enhanced Mean Field Resource Allocation," n.d. [Online]. Available: https://arxiv.org/abs/2405.00056.
[97] "[2604.18133] Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures," n.d. [Online]. Available: https://arxiv.org/abs/2604.18133.
[98] "Multi-agent artificial intelligence designs novel catalysts for ultrafast water purification | Nature Water," n.d. [Online]. Available: https://www.nature.com/articles/s44221-026-00634-9.
[99] "Multi-Agent collaboration patterns with Strands Agents and Amazon Nova | Artificial Intelligence," n.d. [Online]. Available: https://aws.amazon.com/blogs/machine-learning/….
[100] "LangGraph vs CrewAI vs AutoGen: AI Agent Framework Comparison [2026]," n.d. [Online]. Available: https://www.meta-intelligence.tech/en/insight-ai-agent-frameworks.
[101] "Multi-Agent AI Systems Enterprise Guide 2026 - AgileSoftLabs Blog," n.d. [Online]. Available: https://www.agilesoftlabs.com/blog/2026/03/multi-agent-ai-systems-enterprise-guide.
[102] "Multi-Agent Systems: A Survey | IEEE Journals & Magazine | IEEE Xplore," n.d. [Online]. Available: https://ieeexplore.ieee.org/document/8352646.
[103] "Large Language Model Based Multi-agents: A Survey of Progress and Challenges | IJCAI," n.d. [Online]. Available: https://www.ijcai.org/proceedings/2024/890.
[104] "LLM-Based Autonomous Multi-Agent Systems:," n.d. [Online]. Available: https://nanoagentteam.github.io/assets/autonomous-multi-agent-survey.pdf.
[105] "How Multi-Agent AI Editorial Review Works (9 Specialized Agents Explained) | Editorial Conductor," n.d. [Online]. Available: https://editorial-conductorai.com/blog/multi-agent-ai-editorial-review-nine-agents-explain….
[106] "CAIMI 2025 - Call for Papers," n.d. [Online]. Available: https://siim.org/wp-content/uploads/….
[107] A. V. Kalyuzhnaya et al., "LLM Agents for Smart City Management: Enhancing Decision Support Through Multi-Agent AI Systems," Smart Cities, 2025, doi: 10.3390/smartcities8010019.
[108] Z. Yao and H. Yu, "OSF," 2025, doi: 10.31219/osf.io/bv5sg_v1.
[109] J. Lussange, S. Palminteri, S. Bourgeois‐Gironde, and B. Gutkin, "Mesoscale impact of trader psychology on stock markets: a multi-agent AI approach," arXiv (Cornell University), 2019, doi: 10.48550/arxiv.1910.10099.
[110] M. Zhang et al., "PromptBio: A Multi-Agent AI Platform for Bioinformatics Data Analysis," bioRxiv (Cold Spring Harbor Laboratory), 2025, doi: 10.1101/2025.07.05.663295.
[111] J. Liu, H. Li, C. Chai, K. Chen, and D. Wang, "A LLM-informed multi-agent AI system for drone-based visual inspection for infrastructure," Advanced Engineering Informatics, 2025, doi: 10.1016/j.aei.2025.103643.
[112] W. Fan, P. Chen, D. Shi, X. Guo, and L. Kou, "TSINGHUA SCIENCE AND TECHNOLOGY," Tsinghua Science & Technology, 2021, doi: 10.26599/tst.2021.9010005.
[113] D. Saeedi, D. Buckner, J. C. Aponte, and A. Aghazadeh, "ASTRO AGENTS : A M ULTI-AGENT AI FOR HYPOTHESIS," ArXiv.org, 2025, doi: 10.48550/arxiv.2503.23170.
[114] K. Lionis, "Frame Handshake Protocol (FHP) v1.0: A Fail-Closed Coordination Protocol for Agentic and Multi-Agent AI Systems," Zenodo (CERN European Organization for Nuclear Research), 2026, doi: 10.5281/zenodo.18453840.
12. Key Literature From Broader Field (TIER 3 — Verify Before Citing)
- [TRAINING — verify before citing] M. Wooldridge, *An Introduction to MultiAgent Systems*, book, approx. 2002–2009.
- [TRAINING — verify before citing] L. Busoniu, R. Babuska, B. De Schutter, “A Comprehensive Survey of Multiagent Reinforcement Learning,” journal article, approx. 2008.
- [TRAINING — verify before citing] R. Sutton and A. Barto, *Reinforcement Learning: An Introduction*, book, 1998/2018 editions (for RL foundations used in MARL).
- [TRAINING — verify before citing] M. E. J. Newman, *Networks: An Introduction*, book, approx. 2010 (for network effects relevant to multi-agent risk).
- [TRAINING — verify before citing] N. Jennings, K. Sycara, M. Wooldridge, “A Roadmap of Agent Research and Development,” approx. late 1990s (agent research roadmap).
- [TRAINING — verify before citing] M. Shoham and K. Leyton-Brown, *Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations*, book, approx. 2008.
- [TRAINING — verify before citing] Y. Shoham, R. Powers, T. Grenager, “If Multi-Agent Learning is the Answer, What is the Question?” approx. 2007 (related to problem-definition critiques; note you already have a critique source ).
- [TRAINING — verify before citing] FIPA Agent Communication Language / specifications (for classical agent interoperability standards).
Create a report like this on your own topic.
Start your free research