AgentBench
Active·★ 3.8k·Apache-2.0·Updated 2026-02-08
★ Trending★ Essential
AgentBench is a comprehensive benchmark for evaluating Large Language Models (LLMs) as agents across diverse environments, now featuring a function-calling version integrated with AgentRL. It provides a containerized setup for various tasks like OS interaction, database operations, and web shopping, enabling robust and reproducible agent evaluation.
#LLM Evaluation#Agent Benchmarking#Function Calling#Docker#Multi-task Learning
01
Features
01Comprehensive LLM-as-Agent Evaluation across diverse environments.
02Function Calling integration for advanced agent interaction.
03Fully containerized deployment using Docker Compose for reproducibility.
04Multi-task and multi-turn interaction for realistic agent assessment.
05Extensible framework for adding new evaluation tasks.
02
Compatibility
Docker
Native
Verified via docs
Python
Native
Verified via docs
OpenAI API
Supported
Verified via docs
Large Language Models
Supported
Verified via docs
03
Quick start
1
$ pip install -r requirements.txt
04
Use cases
↳Systematically benchmark the performance of various LLM-based agents.
↳Develop and refine advanced LLM agent architectures and strategies.
↳Conduct academic research on the capabilities and limitations of agentic AI.
05
Alternatives
Comments
Log in to leave a comment
No comments yet. Be the first!