A curated list of resources for multi-agent AI.
- Foundations & surveys
- Orchestration frameworks
- Communication protocols & standards
- Environments & simulation
- Benchmarks & evaluation
- Observability & trace tooling
- Cooperative AI, coordination & game theory
- Safety, security & failure modes
- Datasets
- Key papers
- Researchers & labs
- Funders, programs & RFPs
- Workshops, venues & communities
- Contributing
- License
- An Introduction to MultiAgent Systems (2nd ed.): Michael Wooldridge's standard textbook on classical multi-agent systems, covering agent architectures, interaction, and game-theoretic foundations (Wiley, 2009).
- Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations: Shoham and Leyton-Brown's graduate textbook on the formal foundations of multi-agent systems, available as a free PDF (Cambridge University Press, 2009).
- Open Problems in Cooperative AI: Dafoe et al. set out a research agenda for building AI that can cooperate with humans and other agents (2020).
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges: Guo et al. survey LLM-based multi-agent systems across agent profiling, communication, and capability growth (2024).
- LLM Multi-Agent Systems: Challenges and Open Problems: Han et al. catalog open problems in LLM multi-agent systems, including task allocation, reasoning, context, and memory (2024).
- Multi-Agent Collaboration Mechanisms: A Survey of LLMs: Tran et al. propose a framework characterizing LLM multi-agent collaboration by actors, types, structures, strategies, and coordination protocols (2025).
- LangChain: A Python framework for building LLM applications and agents by chaining together models, tools, and integrations.
- CrewAI: A Python framework for orchestrating role-playing autonomous agents that collaborate through Crews and event-driven Flows.
- LangGraph: A low-level orchestration framework from LangChain for building stateful, long-running multi-agent systems as graphs.
- Microsoft AutoGen: A multi-agent conversation framework from Microsoft, now in maintenance mode with development continuing in the Microsoft Agent Framework.
- Microsoft Agent Framework: Microsoft's multi-language framework for building production agents and multi-agent workflows in .NET and Python, positioned as the successor to AutoGen and Semantic Kernel.
- OpenAI Agents SDK: OpenAI's Python SDK for building multi-agent workflows with handoffs, guardrails, and tracing across many LLM providers.
- Google Agent Development Kit (ADK): Google's code-first Python toolkit for building, evaluating, and deploying agents and multi-agent systems.
- CAMEL: A Python framework for building and studying communicative multi-agent systems and large-scale agent societies.
- MetaGPT: A multi-agent framework that assigns software-company roles to LLM agents to turn a one-line requirement into software artifacts.
- ChatDev: A multi-agent platform that simulates a virtual software company in which LLM agents collaborate to develop software.
- Microsoft Semantic Kernel: A model-agnostic SDK from Microsoft for integrating LLMs and building agents in C#, Python, and Java.
- LlamaIndex: A data framework for LLM applications that includes agent and multi-agent workflow capabilities over user data.
- AG2: A community-governed fork of the original AutoGen for building conversable agents and multi-agent workflows.
- Magentic-One: A generalist multi-agent system built on AutoGen AgentChat that uses an orchestrator plus specialized agents to solve open-ended web and file tasks.
- OpenAI Swarm: An educational, lightweight multi-agent orchestration framework from OpenAI, superseded by the OpenAI Agents SDK.
- Model Context Protocol (MCP): An open standard, originally from Anthropic and now contributed to the Linux Foundation's Agentic AI Foundation, that connects AI applications to external data sources, tools, and prompts.
- Agent2Agent (A2A): An open protocol, originally developed by Google and now hosted by the Linux Foundation, that lets independent AI agents discover one another and exchange tasks and messages across vendors and frameworks.
- AGNTCY / Internet of Agents: A Linux Foundation project, initially open-sourced by Cisco, providing infrastructure for agent discovery, identity, messaging, and observability across multi-agent systems.
- Agent Communication Protocol (ACP): An open, REST-based protocol introduced by IBM's BeeAI for agent interoperability, now developed as part of A2A under the Linux Foundation.
- Agent Network Protocol (ANP): An open-source protocol using W3C Decentralized Identifiers to provide identity, encrypted communication, and capability discovery for agent-to-agent collaboration.
- FIPA ACL Message Structure Specification: The 2002 FIPA standard defining the message structure of the Agent Communication Language, a historical reference for interoperable multi-agent communication.
- KQML: The UMBC resource hub for the Knowledge Query and Manipulation Language, an early-1990s DARPA-sponsored language and protocol for knowledge exchange among software agents.
- PettingZoo: A Python API standard and environment suite for multi-agent reinforcement learning, maintained by the Farama Foundation.
- Melting Pot: A DeepMind suite of multi-agent reinforcement learning substrates and test scenarios for evaluating social generalization.
- Concordia: A DeepMind library for generative agent-based social simulation using a game-master pattern.
- OpenSpiel: A DeepMind collection of environments and algorithms for research in games and multi-agent reinforcement learning.
- Generative Agents (Smallville): The Stanford research codebase for generative agents that simulate human behavior in the Smallville environment.
- AgentVerse: A framework for deploying multiple LLM agents in task-solving and social-simulation environments.
- AI Town: A deployable starter kit for a virtual town where AI characters live, chat, and socialize.
- MAgent2: A Farama Foundation engine and reference environments for gridworld scenarios with very large numbers of agents.
- JaxMARL: A JAX library of GPU-accelerated multi-agent reinforcement learning environments and baseline algorithms.
- VirtualHome: A Python and Unity platform for simulating multi-agent household activities via programs.
- AgentSims: A sandbox infrastructure for evaluating LLM agents through task-based simulations in an interactive town environment.
- MultiAgentBench / MARBLE: An ACL 2025 benchmark and framework that evaluates the collaboration and competition of LLM agents across coordination topologies using milestone-based metrics.
- REALM-Bench: A benchmark for evaluating multi-agent systems on real-world, dynamic planning and scheduling tasks across multiple agent frameworks.
- AgentBench: An ICLR 2024 benchmark that evaluates LLMs as autonomous agents across eight distinct interactive environments.
- AgentBoard: A NeurIPS 2024 benchmark and analytical evaluation board for multi-turn LLM agents, reporting fine-grained progress metrics across nine tasks.
- AgentGym: A framework and benchmark suite for evaluating and evolving LLM-based agents across 14 diverse interactive environments in a unified format.
- τ-bench / tau2-bench: A benchmark from Sierra for tool-agent-user interaction that evaluates agents conversing with a simulated user across domains such as airline, retail, telecom, and banking.
- GAIA: A benchmark of real-world questions for general AI assistants requiring reasoning, multimodal understanding, web browsing, and tool use, with a public Hugging Face leaderboard.
- Inspect: An open-source evaluation framework from the UK AI Security Institute providing composable datasets, agents, tools, and scorers for LLM and agentic evaluations.
- ControlArena: A library from the UK AI Security Institute and Redwood Research for running AI control experiments across settings, model organisms, and protocols, built on Inspect.
- SWE-bench: A benchmark that tasks agents with resolving real-world GitHub issues by generating patches verified through unit tests.
- BrowserGym: A Gym environment from ServiceNow for web-task automation that unifies benchmarks such as MiniWoB, WebArena, and WorkArena.
- WebArena: A self-hostable, realistic web environment with 812 long-horizon tasks for building and evaluating autonomous web agents.
- MLAgentBench: A benchmark of end-to-end machine-learning experimentation tasks for evaluating language agents as AI research agents.
- LangSmith: Framework-agnostic SaaS platform from LangChain for tracing, evaluating, and monitoring LLM and agent applications.
- Langfuse: Open-source (self-hostable) LLM engineering platform offering tracing, evaluation, prompt management, and datasets, with OpenTelemetry integration.
- Arize Phoenix: Open-source AI observability and evaluation tool with native OpenTelemetry support for tracing LLM and agent applications.
- Weights & Biases Weave: Open-source toolkit from Weights & Biases for logging, tracing, and evaluating LLM application inputs, outputs, and calls.
- AgentOps: Open-source Python SDK for monitoring AI agents, with session replay, LLM cost tracking, and integrations for CrewAI, AutoGen, LangChain, and the OpenAI Agents SDK.
- Langtrace: Open-source, OpenTelemetry-based observability tool providing tracing, evaluations, and metrics for LLMs, frameworks, and vector databases.
- Pydantic Logfire: Observability platform built on OpenTelemetry by the Pydantic team, with features for tracing LLM calls and agent behavior.
- Helicone: Open-source LLM observability platform for monitoring, evaluating, and experimenting on traces and sessions.
- OpenLLMetry (Traceloop): Open-source set of OpenTelemetry-based extensions, maintained by Traceloop, for instrumenting GenAI and LLM applications.
- Inspect log viewer: The UK AI Security Institute's evaluation framework, including a web-based log viewer for inspecting agent and evaluation transcripts.
- Transluce Docent: Tool from Transluce for searching, clustering, and analyzing long agent transcripts to surface behaviors and failures.
- OpenTelemetry GenAI semantic conventions: Standard schema under OpenTelemetry defining attributes, spans, metrics, and events for generative AI and agent telemetry.
- OpenInference: Specification and instrumentation libraries from Arize that complement OpenTelemetry to standardize tracing of AI applications.
- Cooperative AI Foundation: A nonprofit foundation that funds and coordinates research aimed at improving the cooperative capabilities of AI systems.
- Cooperative AI: machines must learn to find common ground (Nature 2021): A Nature comment by Dafoe and colleagues arguing that AI research should be reconceived to address problems of cooperation.
- FOCAL: Foundations of Cooperative AI Lab (CMU): A Carnegie Mellon University lab, directed by Vincent Conitzer, developing game-theoretic foundations for cooperation among autonomous AI agents.
- CICERO: Meta AI's agent for the game Diplomacy that combines a language model with strategic reasoning to negotiate with human players.
- Hanabi Learning Environment: A DeepMind research platform for the cooperative card game Hanabi, used as a benchmark for multi-agent coordination under imperfect information.
- Welfare Diplomacy: A general-sum variant of Diplomacy and open-source engine designed to benchmark and incentivize cooperative capabilities in language-model agents.
- GovSim / Cooperate or Collapse: A simulation platform and paper studying whether societies of LLM agents can sustainably manage shared common-pool resources.
- Multi-Agent Risks from Advanced AI: A Cooperative AI Foundation technical report presenting a taxonomy of risks that arise when advanced AI agents interact, along with a research agenda.
- Open Challenges in Multi-Agent Security: A paper defining multi-agent security as a field and presenting a threat taxonomy and research agenda for securing networks of interacting AI agents.
- Why Do Multi-Agent LLM Systems Fail? (MAST): A study introducing the Multi-Agent System Failure Taxonomy, identifying 14 failure modes across popular multi-agent LLM frameworks.
- MAST repository: The code and annotated trace dataset accompanying the "Why Do Multi-Agent LLM Systems Fail?" paper.
- AgentDojo: A dynamic benchmark environment for evaluating prompt-injection attacks and defenses against tool-using LLM agents.
- OWASP Agentic AI: Threats and Mitigations: The OWASP GenAI Security Project's threat-model reference cataloguing security risks and mitigations for agentic AI systems.
- OWASP Top 10 for Agentic Applications (2026): A ranked list of the most critical security risks for autonomous, tool-using AI agents, produced by the OWASP GenAI Security Project.
- Algorithmic Collusion by Large Language Models (Fish et al.): A paper showing that LLM-based pricing agents can autonomously reach supracompetitive prices in oligopoly and auction settings.
- Artificial Intelligence, Algorithmic Pricing, and Collusion (Calvano et al.): An American Economic Review study demonstrating that Q-learning pricing algorithms can learn collusive strategies without communicating.
- ToolBench: Open-source instruction-tuning dataset and platform for tool learning, built from 16,000+ real-world RapidAPI APIs with single- and multi-tool scenarios.
- xLAM function-calling 60k: Salesforce dataset of 60,000 verified function-calling examples generated by the APIGen pipeline across 3,673 executable APIs.
- APIGen-MT-5k: Salesforce dataset of 5,000 synthetic multi-turn agent trajectories produced via simulated agent-human interplay.
- Berkeley Function-Calling Leaderboard: UC Berkeley Gorilla project's question-and-function-documentation dataset used to evaluate LLM function-calling across multiple categories.
- AgentInstruct (THUDM): Dataset of 1,866 multi-turn agent interaction trajectories across six tasks, used to train the AgentLM models.
- Orca-AgentInstruct-1M: Microsoft dataset of 1 million synthetic instruction pairs generated by the AgentInstruct multi-agent workflow, covering skills including tool use and coding.
- ToolACE: Tool-learning dataset generated by a multi-agent pipeline over a pool of 26,507 APIs, covering single, parallel, dependent, and multi-turn function calls.
- Glaive function-calling v2: Dataset of roughly 52,000 samples for training models on function-calling tasks, including no-call, single-call, and multi-call examples.
- Hermes Function-Calling V1: Nous Research dataset of single- and multi-turn function-calling and structured JSON output samples in ShareGPT format.
- Melting Pot: Leibo et al. introduce a multi-agent reinforcement learning evaluation suite of test scenarios measuring generalization to novel social situations (2021).
- CICERO: Human-level play in the game of Diplomacy: Meta's FAIR team report an agent combining a language model with planning and reasoning that reached human-level play in Diplomacy (Science, 2022).
- Generative Agents: Interactive Simulacra of Human Behavior: Park et al. populate a sandbox town with LLM agents that remember, reflect, and plan to produce believable individual and emergent social behavior (2023).
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society: Li et al. propose a role-playing framework using inception prompting to drive autonomous cooperation between two LLM agents (2023).
- Improving Factuality and Reasoning via Multiagent Debate: Du et al. show that having multiple LLM instances debate over several rounds improves factuality and reasoning (2023).
- Debating with More Persuasive LLMs Leads to More Truthful Answers: Khan et al. study multi-agent debate between LLMs as a method for scalable oversight and for eliciting more truthful answers from non-expert judges (ICML 2024).
- ChatDev: Communicative Agents for Software Development: Qian et al. organize LLM agents into a virtual software company with chat-driven design, coding, testing, and documentation stages (2023).
- MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework: Hong et al. encode standardized operating procedures into an assembly-line multi-agent framework for software tasks (2023).
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation: Wu et al. present an open-source framework for building applications from customizable, conversable LLM agents (2023).
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors: Chen et al. propose a framework that dynamically adjusts agent group composition and studies emergent social behaviors during collaboration (2023).
- Concordia: Generative Agent-Based Modeling: Vezhnevets et al. release a library using a Game Master pattern to simulate generative agents grounded in physical, social, or digital space (2023).
- Agent-as-a-Judge: Evaluate Agents with Agents: Zhao et al. extend LLM-as-a-judge to agentic evaluators that give intermediate feedback across a task, with a code-generation benchmark (2024).
- Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks: Fourney et al. present a modular orchestrator-led multi-agent system for open-ended web and file-based tasks (2024).
- Generative Agent Simulations of 1,000 People: Park et al. build agents from qualitative interviews with 1,052 real individuals and measure how accurately they reproduce survey responses and behaviors (2024).
- MultiAgentBench: Evaluating the Collaboration and Competition of LLM Agents: Zhu et al. introduce a benchmark with milestone-based metrics across collaborative and competitive scenarios and coordination topologies (2025).
- How we built our multi-agent research system: Anthropic's engineering write-up describes an orchestrator-worker architecture in which a lead agent spawns subagents for parallel research (2025).
Generative social simulation & cooperative AI
- Joel Z. Leibo: Google DeepMind; creator of Melting Pot and Concordia, and a cultural-evolution framing of multi-agent cooperation.
- Alexander Vezhnevets: Google DeepMind; lead of Concordia and generative agent-based modeling.
- Allan Dafoe: Google DeepMind; founder of the Centre for the Governance of AI and the Cooperative AI Foundation.
- Edgar A. Duéñez-Guzmán: Game theory and multi-agent reinforcement learning; contributor to Melting Pot and Concordia.
Cooperative AI & game theory
- Lewis Hammond: Cooperative AI Foundation; multi-agent risk and safety.
- Jesse Clifton: Cooperative AI Foundation / Center on Long-Term Risk; bargaining and conflict.
- Vincent Conitzer: Carnegie Mellon University; directs FOCAL, on game-theoretic foundations of cooperative AI.
- Michael P. Wellman: University of Michigan; empirical game-theoretic analysis and market games.
- David C. Parkes: Harvard University; mechanism design and economics-and-computation.
Multi-agent reinforcement learning
- Jakob Foerster: University of Oxford (FLAIR) and Meta FAIR; deep multi-agent reinforcement learning.
- Christian Schroeder de Witt: University of Oxford; multi-agent security and covert communication.
- Natasha Jaques: University of Washington / Google DeepMind; social reinforcement learning.
- Max Kleiman-Weiner: University of Washington; computational models of cooperation and social learning.
- Eugene Vinitsky: New York University; multi-agent reinforcement learning at scale.
- Dylan Hadfield-Menell: MIT CSAIL; cooperative inverse reinforcement learning and alignment.
- Stuart Russell: UC Berkeley; directs CHAI, on assistance games and provably beneficial AI.
Self-play & negotiation
- Noam Brown: OpenAI; self-play, search, and strategic reasoning (CICERO).
- Tim Baarslag: CWI and Utrecht University; automated negotiation and the ANAC competition.
Failure modes, collusion & social intelligence
- Mert Cemri: UC Berkeley; multi-agent failure taxonomy (MAST).
- Maarten Sap: Carnegie Mellon University / Allen Institute for AI; social intelligence in language agents (SOTOPIA).
- Sara Fish: Harvard University; algorithmic collusion among LLM agents.
- Emilio Calvano: LUISS / Toulouse School of Economics; algorithmic pricing and collusion.
Computational social science of agents
- Iyad Rahwan: Max Planck Institute for Human Development; machine behavior and cooperative AI.
- Andrea Baronchelli: City St George's, University of London; emergence of social conventions in human and LLM populations.
Cooperative agents & robustness (adjacent)
- Sheila McIlraith: University of Toronto; theory of mind for cooperative agents.
- Adam Gleave: FAR AI; multi-agent robustness and adversarial policies.
- Susmit Jha: SRI International; assured autonomy and trustworthy AI.
- Cooperative AI Foundation: A charitable foundation, backed by a philanthropic commitment from Macroscopic Ventures, that funds research aimed at improving the cooperative intelligence of advanced AI systems.
- Cooperative AI Foundation: Research Grants: The Foundation's open call for research proposals on cooperation-relevant capabilities and propensities in AI systems.
- Schmidt Sciences: Science of Trustworthy AI: A Schmidt Sciences program funding technical research to understand, predict, and control risks from frontier AI systems, including a research aim on multi-agent risks.
- Schmidt Sciences: AI Agents: A pilot program supporting academic and nonprofit researchers studying how multiple AI agents communicate and coordinate with one another.
- Macroscopic Ventures: A nonprofit grantmaker whose focus areas include cooperative AI and reducing risks from advanced AI, and which funds the Cooperative AI Foundation.
- Coefficient Giving: Navigating Transformative AI: The fund (from the grantmaker formerly named Open Philanthropy) supporting technical safety, governance, and field-building work on risks from transformative AI.
- ARIA: Safeguarded AI: A UK Advanced Research and Invention Agency programme developing methods for fleets of AI agents to produce formally verified, quantitatively safe artefacts.
- Foresight Institute: AI Safety Grants: A grant program whose focus areas include decentralized and cooperative AI alongside other AI safety topics.
- NSF: Safe Learning-Enabled Systems: A US National Science Foundation program, run in partnership with Open Philanthropy and Good Ventures, funding foundational research on the safety of learning-enabled systems.
- DICE: Decentralized Artificial Intelligence through Controlled Emergence (DARPA): A DARPA program developing theory and algorithms for decentralized coordination and local inference control among collectives of heterogeneous AI agents.
- MATHBAC: Mathematics of Boosting Agentic Communication (DARPA): A 2026 DARPA program developing mathematical frameworks to design and provide guarantees for communication protocols among collectives of AI agents.
- AAMAS (via IFAAMAS): The International Foundation for Autonomous Agents and Multiagent Systems, which sponsors the annual AAMAS conference for agents and multi-agent systems research.
- AAMAS 2026: The 25th International Conference on Autonomous Agents and Multiagent Systems, held in Paphos, Cyprus, in May 2026.
- Cooperative AI Workshop (NeurIPS): The Cooperative AI Foundation's recurring workshops at NeurIPS and other machine learning conferences.
- Concordia Contest: A NeurIPS competition challenging participants to build language-model agents that exhibit cooperative intelligence in text-based mixed-motive scenarios.
- Melting Pot Contest: A NeurIPS competition, run with Google DeepMind and MIT, evaluating multi-agent reinforcement learning agents on mixed-motive cooperation problems.
- ANAC (Automated Negotiating Agents Competition): An annual international competition, held in conjunction with AAMAS or IJCAI since 2010, in which researchers develop and benchmark automated negotiation agents.
- Cooperative AI Summer School: An annual residential program run by the Cooperative AI Foundation for students and early-career researchers entering the field.
Contributions are welcome. Please open a pull request that:
- adds entries in the existing
**[Name](URL)**: one neutral sentence.format, - places each entry in the most fitting section, ordered by prominence,
- links the official repo, homepage, or paper (arXiv/DOI), and
- keeps descriptions factual and free of marketing language.
Suggestions, corrections, and dead-link reports via issues are equally valued.
To the extent possible under law, contributors have waived all copyright and related or neighboring rights to this work.
