Nicolas99-9/multiagentics

A curated list of resources for multiagent AI.

★ 0Forks 0GitHub ↗Compare

README

Awesome Multiagentics Awesome

A curated list of resources for multi-agent AI.

Contents

Foundations & surveys

Orchestration frameworks

  • LangChain: A Python framework for building LLM applications and agents by chaining together models, tools, and integrations.
  • CrewAI: A Python framework for orchestrating role-playing autonomous agents that collaborate through Crews and event-driven Flows.
  • LangGraph: A low-level orchestration framework from LangChain for building stateful, long-running multi-agent systems as graphs.
  • Microsoft AutoGen: A multi-agent conversation framework from Microsoft, now in maintenance mode with development continuing in the Microsoft Agent Framework.
  • Microsoft Agent Framework: Microsoft's multi-language framework for building production agents and multi-agent workflows in .NET and Python, positioned as the successor to AutoGen and Semantic Kernel.
  • OpenAI Agents SDK: OpenAI's Python SDK for building multi-agent workflows with handoffs, guardrails, and tracing across many LLM providers.
  • Google Agent Development Kit (ADK): Google's code-first Python toolkit for building, evaluating, and deploying agents and multi-agent systems.
  • CAMEL: A Python framework for building and studying communicative multi-agent systems and large-scale agent societies.
  • MetaGPT: A multi-agent framework that assigns software-company roles to LLM agents to turn a one-line requirement into software artifacts.
  • ChatDev: A multi-agent platform that simulates a virtual software company in which LLM agents collaborate to develop software.
  • Microsoft Semantic Kernel: A model-agnostic SDK from Microsoft for integrating LLMs and building agents in C#, Python, and Java.
  • LlamaIndex: A data framework for LLM applications that includes agent and multi-agent workflow capabilities over user data.
  • AG2: A community-governed fork of the original AutoGen for building conversable agents and multi-agent workflows.
  • Magentic-One: A generalist multi-agent system built on AutoGen AgentChat that uses an orchestrator plus specialized agents to solve open-ended web and file tasks.
  • OpenAI Swarm: An educational, lightweight multi-agent orchestration framework from OpenAI, superseded by the OpenAI Agents SDK.

Communication protocols & standards

  • Model Context Protocol (MCP): An open standard, originally from Anthropic and now contributed to the Linux Foundation's Agentic AI Foundation, that connects AI applications to external data sources, tools, and prompts.
  • Agent2Agent (A2A): An open protocol, originally developed by Google and now hosted by the Linux Foundation, that lets independent AI agents discover one another and exchange tasks and messages across vendors and frameworks.
  • AGNTCY / Internet of Agents: A Linux Foundation project, initially open-sourced by Cisco, providing infrastructure for agent discovery, identity, messaging, and observability across multi-agent systems.
  • Agent Communication Protocol (ACP): An open, REST-based protocol introduced by IBM's BeeAI for agent interoperability, now developed as part of A2A under the Linux Foundation.
  • Agent Network Protocol (ANP): An open-source protocol using W3C Decentralized Identifiers to provide identity, encrypted communication, and capability discovery for agent-to-agent collaboration.
  • FIPA ACL Message Structure Specification: The 2002 FIPA standard defining the message structure of the Agent Communication Language, a historical reference for interoperable multi-agent communication.
  • KQML: The UMBC resource hub for the Knowledge Query and Manipulation Language, an early-1990s DARPA-sponsored language and protocol for knowledge exchange among software agents.

Environments & simulation

  • PettingZoo: A Python API standard and environment suite for multi-agent reinforcement learning, maintained by the Farama Foundation.
  • Melting Pot: A DeepMind suite of multi-agent reinforcement learning substrates and test scenarios for evaluating social generalization.
  • Concordia: A DeepMind library for generative agent-based social simulation using a game-master pattern.
  • OpenSpiel: A DeepMind collection of environments and algorithms for research in games and multi-agent reinforcement learning.
  • Generative Agents (Smallville): The Stanford research codebase for generative agents that simulate human behavior in the Smallville environment.
  • AgentVerse: A framework for deploying multiple LLM agents in task-solving and social-simulation environments.
  • AI Town: A deployable starter kit for a virtual town where AI characters live, chat, and socialize.
  • MAgent2: A Farama Foundation engine and reference environments for gridworld scenarios with very large numbers of agents.
  • JaxMARL: A JAX library of GPU-accelerated multi-agent reinforcement learning environments and baseline algorithms.
  • VirtualHome: A Python and Unity platform for simulating multi-agent household activities via programs.
  • AgentSims: A sandbox infrastructure for evaluating LLM agents through task-based simulations in an interactive town environment.

Benchmarks & evaluation

  • MultiAgentBench / MARBLE: An ACL 2025 benchmark and framework that evaluates the collaboration and competition of LLM agents across coordination topologies using milestone-based metrics.
  • REALM-Bench: A benchmark for evaluating multi-agent systems on real-world, dynamic planning and scheduling tasks across multiple agent frameworks.
  • AgentBench: An ICLR 2024 benchmark that evaluates LLMs as autonomous agents across eight distinct interactive environments.
  • AgentBoard: A NeurIPS 2024 benchmark and analytical evaluation board for multi-turn LLM agents, reporting fine-grained progress metrics across nine tasks.
  • AgentGym: A framework and benchmark suite for evaluating and evolving LLM-based agents across 14 diverse interactive environments in a unified format.
  • τ-bench / tau2-bench: A benchmark from Sierra for tool-agent-user interaction that evaluates agents conversing with a simulated user across domains such as airline, retail, telecom, and banking.
  • GAIA: A benchmark of real-world questions for general AI assistants requiring reasoning, multimodal understanding, web browsing, and tool use, with a public Hugging Face leaderboard.
  • Inspect: An open-source evaluation framework from the UK AI Security Institute providing composable datasets, agents, tools, and scorers for LLM and agentic evaluations.
  • ControlArena: A library from the UK AI Security Institute and Redwood Research for running AI control experiments across settings, model organisms, and protocols, built on Inspect.
  • SWE-bench: A benchmark that tasks agents with resolving real-world GitHub issues by generating patches verified through unit tests.
  • BrowserGym: A Gym environment from ServiceNow for web-task automation that unifies benchmarks such as MiniWoB, WebArena, and WorkArena.
  • WebArena: A self-hostable, realistic web environment with 812 long-horizon tasks for building and evaluating autonomous web agents.
  • MLAgentBench: A benchmark of end-to-end machine-learning experimentation tasks for evaluating language agents as AI research agents.

Observability & trace tooling

  • LangSmith: Framework-agnostic SaaS platform from LangChain for tracing, evaluating, and monitoring LLM and agent applications.
  • Langfuse: Open-source (self-hostable) LLM engineering platform offering tracing, evaluation, prompt management, and datasets, with OpenTelemetry integration.
  • Arize Phoenix: Open-source AI observability and evaluation tool with native OpenTelemetry support for tracing LLM and agent applications.
  • Weights & Biases Weave: Open-source toolkit from Weights & Biases for logging, tracing, and evaluating LLM application inputs, outputs, and calls.
  • AgentOps: Open-source Python SDK for monitoring AI agents, with session replay, LLM cost tracking, and integrations for CrewAI, AutoGen, LangChain, and the OpenAI Agents SDK.
  • Langtrace: Open-source, OpenTelemetry-based observability tool providing tracing, evaluations, and metrics for LLMs, frameworks, and vector databases.
  • Pydantic Logfire: Observability platform built on OpenTelemetry by the Pydantic team, with features for tracing LLM calls and agent behavior.
  • Helicone: Open-source LLM observability platform for monitoring, evaluating, and experimenting on traces and sessions.
  • OpenLLMetry (Traceloop): Open-source set of OpenTelemetry-based extensions, maintained by Traceloop, for instrumenting GenAI and LLM applications.
  • Inspect log viewer: The UK AI Security Institute's evaluation framework, including a web-based log viewer for inspecting agent and evaluation transcripts.
  • Transluce Docent: Tool from Transluce for searching, clustering, and analyzing long agent transcripts to surface behaviors and failures.
  • OpenTelemetry GenAI semantic conventions: Standard schema under OpenTelemetry defining attributes, spans, metrics, and events for generative AI and agent telemetry.
  • OpenInference: Specification and instrumentation libraries from Arize that complement OpenTelemetry to standardize tracing of AI applications.

Cooperative AI, coordination & game theory

  • Cooperative AI Foundation: A nonprofit foundation that funds and coordinates research aimed at improving the cooperative capabilities of AI systems.
  • Cooperative AI: machines must learn to find common ground (Nature 2021): A Nature comment by Dafoe and colleagues arguing that AI research should be reconceived to address problems of cooperation.
  • FOCAL: Foundations of Cooperative AI Lab (CMU): A Carnegie Mellon University lab, directed by Vincent Conitzer, developing game-theoretic foundations for cooperation among autonomous AI agents.
  • CICERO: Meta AI's agent for the game Diplomacy that combines a language model with strategic reasoning to negotiate with human players.
  • Hanabi Learning Environment: A DeepMind research platform for the cooperative card game Hanabi, used as a benchmark for multi-agent coordination under imperfect information.
  • Welfare Diplomacy: A general-sum variant of Diplomacy and open-source engine designed to benchmark and incentivize cooperative capabilities in language-model agents.
  • GovSim / Cooperate or Collapse: A simulation platform and paper studying whether societies of LLM agents can sustainably manage shared common-pool resources.

Safety, security & failure modes

Datasets

  • ToolBench: Open-source instruction-tuning dataset and platform for tool learning, built from 16,000+ real-world RapidAPI APIs with single- and multi-tool scenarios.
  • xLAM function-calling 60k: Salesforce dataset of 60,000 verified function-calling examples generated by the APIGen pipeline across 3,673 executable APIs.
  • APIGen-MT-5k: Salesforce dataset of 5,000 synthetic multi-turn agent trajectories produced via simulated agent-human interplay.
  • Berkeley Function-Calling Leaderboard: UC Berkeley Gorilla project's question-and-function-documentation dataset used to evaluate LLM function-calling across multiple categories.
  • AgentInstruct (THUDM): Dataset of 1,866 multi-turn agent interaction trajectories across six tasks, used to train the AgentLM models.
  • Orca-AgentInstruct-1M: Microsoft dataset of 1 million synthetic instruction pairs generated by the AgentInstruct multi-agent workflow, covering skills including tool use and coding.
  • ToolACE: Tool-learning dataset generated by a multi-agent pipeline over a pool of 26,507 APIs, covering single, parallel, dependent, and multi-turn function calls.
  • Glaive function-calling v2: Dataset of roughly 52,000 samples for training models on function-calling tasks, including no-call, single-call, and multi-call examples.
  • Hermes Function-Calling V1: Nous Research dataset of single- and multi-turn function-calling and structured JSON output samples in ShareGPT format.

Key papers

Researchers & labs

Generative social simulation & cooperative AI

  • Joel Z. Leibo: Google DeepMind; creator of Melting Pot and Concordia, and a cultural-evolution framing of multi-agent cooperation.
  • Alexander Vezhnevets: Google DeepMind; lead of Concordia and generative agent-based modeling.
  • Allan Dafoe: Google DeepMind; founder of the Centre for the Governance of AI and the Cooperative AI Foundation.
  • Edgar A. Duéñez-Guzmán: Game theory and multi-agent reinforcement learning; contributor to Melting Pot and Concordia.

Cooperative AI & game theory

  • Lewis Hammond: Cooperative AI Foundation; multi-agent risk and safety.
  • Jesse Clifton: Cooperative AI Foundation / Center on Long-Term Risk; bargaining and conflict.
  • Vincent Conitzer: Carnegie Mellon University; directs FOCAL, on game-theoretic foundations of cooperative AI.
  • Michael P. Wellman: University of Michigan; empirical game-theoretic analysis and market games.
  • David C. Parkes: Harvard University; mechanism design and economics-and-computation.

Multi-agent reinforcement learning

  • Jakob Foerster: University of Oxford (FLAIR) and Meta FAIR; deep multi-agent reinforcement learning.
  • Christian Schroeder de Witt: University of Oxford; multi-agent security and covert communication.
  • Natasha Jaques: University of Washington / Google DeepMind; social reinforcement learning.
  • Max Kleiman-Weiner: University of Washington; computational models of cooperation and social learning.
  • Eugene Vinitsky: New York University; multi-agent reinforcement learning at scale.
  • Dylan Hadfield-Menell: MIT CSAIL; cooperative inverse reinforcement learning and alignment.
  • Stuart Russell: UC Berkeley; directs CHAI, on assistance games and provably beneficial AI.

Self-play & negotiation

  • Noam Brown: OpenAI; self-play, search, and strategic reasoning (CICERO).
  • Tim Baarslag: CWI and Utrecht University; automated negotiation and the ANAC competition.

Failure modes, collusion & social intelligence

  • Mert Cemri: UC Berkeley; multi-agent failure taxonomy (MAST).
  • Maarten Sap: Carnegie Mellon University / Allen Institute for AI; social intelligence in language agents (SOTOPIA).
  • Sara Fish: Harvard University; algorithmic collusion among LLM agents.
  • Emilio Calvano: LUISS / Toulouse School of Economics; algorithmic pricing and collusion.

Computational social science of agents

  • Iyad Rahwan: Max Planck Institute for Human Development; machine behavior and cooperative AI.
  • Andrea Baronchelli: City St George's, University of London; emergence of social conventions in human and LLM populations.

Cooperative agents & robustness (adjacent)

  • Sheila McIlraith: University of Toronto; theory of mind for cooperative agents.
  • Adam Gleave: FAR AI; multi-agent robustness and adversarial policies.
  • Susmit Jha: SRI International; assured autonomy and trustworthy AI.

Funders, programs & RFPs

Workshops, venues & communities

  • AAMAS (via IFAAMAS): The International Foundation for Autonomous Agents and Multiagent Systems, which sponsors the annual AAMAS conference for agents and multi-agent systems research.
  • AAMAS 2026: The 25th International Conference on Autonomous Agents and Multiagent Systems, held in Paphos, Cyprus, in May 2026.
  • Cooperative AI Workshop (NeurIPS): The Cooperative AI Foundation's recurring workshops at NeurIPS and other machine learning conferences.
  • Concordia Contest: A NeurIPS competition challenging participants to build language-model agents that exhibit cooperative intelligence in text-based mixed-motive scenarios.
  • Melting Pot Contest: A NeurIPS competition, run with Google DeepMind and MIT, evaluating multi-agent reinforcement learning agents on mixed-motive cooperation problems.
  • ANAC (Automated Negotiating Agents Competition): An annual international competition, held in conjunction with AAMAS or IJCAI since 2010, in which researchers develop and benchmark automated negotiation agents.
  • Cooperative AI Summer School: An annual residential program run by the Cooperative AI Foundation for students and early-career researchers entering the field.

Contributing

Contributions are welcome. Please open a pull request that:

  • adds entries in the existing **[Name](URL)**: one neutral sentence. format,
  • places each entry in the most fitting section, ordered by prominence,
  • links the official repo, homepage, or paper (arXiv/DOI), and
  • keeps descriptions factual and free of marketing language.

Suggestions, corrections, and dead-link reports via issues are equally valued.

License

CC0

To the extent possible under law, contributors have waived all copyright and related or neighboring rights to this work.

Contributors

suchow

Issues