AgentBench
AgentBench is a comprehensive benchmark for evaluating Large Language Models (LLMs) as agents across diverse environments, now featuring a function-calling version integrated with AgentRL. It provides a containerized setup for various tasks like OS interaction, database operations, and web shopping, enabling robust and reproducible agent evaluation.
AgentBench is currently grouped under Observability, which makes it easier to evaluate through workflow fit instead of isolated features alone. Based on the available data, it leans most heavily toward Comprehensive LLM-as-Agent Evaluation across diverse environments. and Systematically benchmark the performance of various LLM-based agents.. The listed license is Apache-2.0, which is useful when adoption constraints matter. It also shows measurable community traction with 3.6k GitHub stars.
Features
Why choose it
Trade-offs
Compatibility
Quick start
Use cases
How it compares
Alternatives
Related searches
Comments
No comments yet. Be the first!