AgentBench
Active·★ 3.8k·Apache-2.0·Updated 2026-02-08
★ Trending★ Essential
AgentBench is a comprehensive benchmark for evaluating Large Language Models (LLMs) as agents across diverse environments, now featuring a function-calling version integrated with AgentRL. It provides a containerized setup for various tasks like OS interaction, database operations, and web shopping, enabling robust and reproducible agent evaluation.
#LLM Evaluation#Agent Benchmarking#Function Calling