LLM-Agent-Benchmark-List
LLM-Agent-Benchmark-List is a curated, continuously updated repository of benchmarks and surveys for evaluating large language models (LLMs) and LLM-powered agents. It organizes resources by category (e.g., ToolUse, Reasoning, Agent, Code) to help researchers and practitioners find relevant evaluation tools. The project aims to streamline the process of identifying appropriate benchmarks for LLM and agent assessment.
✨ Key features
- Categorized list of LLM and agent benchmarks
- Includes surveys and papers with links
- Covers ToolUse, Reasoning, Knowledge, Graph, Video, Code, Alignment, Agent, Multimodal
- Continuous updates and community contributions via PRs/issues
- Provides direct links to papers and project pages
🎯 Use cases
- Finding benchmarks for evaluating LLM tool-use capabilities
- Selecting benchmarks for agent reasoning and planning
- Identifying code generation benchmarks for LLMs
- Locating surveys on LLM evaluation methodologies
- Discovering multimodal evaluation benchmarks
⚠️ Good to know
The list is continuously updated and may not be exhaustive; it relies on community contributions.
❓ FAQ
How can I contribute to this list?
You can contribute via PRs, issues, emails, or other methods as stated in the README.
What categories of benchmarks are included?
Categories include Survey, ToolUse, Reasoning, Knowledge, Graph, Video, Code, Alignment, Agent, and Multimodal.
Are there benchmarks for evaluating agents in real-world environments?
Yes, examples include WebArena, OSWorld, and Terminal-Bench for terminal environments.
Does the list include recent benchmarks?
Yes, it includes entries up to 2026, and it is continuously updated.
📊 Repository
🤖 Overview, features, install steps and FAQ were generated from the project's README on Sep 4, 2026. Always check the original source before running commands.