AgentIndex icon
AgentIndex
ToolsCategoriesTrendingNewCompare
Submit Tool
ToolsCategoriesTrendingNewCompare
Home/
RAG / Knowledge Base/
chinese-llm-benchmark
chinese-llm-benchmark logo

chinese-llm-benchmark

Active·★ 6.3k·Updated 2026-07-12
★ Trending★ Essential

ReLE Benchmark (formerly CLiB) provides a continuously updated evaluation for Chinese AI large language models, covering over 337 commercial and open-source LLMs. It offers multi-dimensional capability assessments across various domains, along with comprehensive rankings and a large defect library for model improvement.

chinese-llm-benchmark is currently grouped under RAG / Knowledge Base, which makes it easier to evaluate through workflow fit instead of isolated features alone. Based on the available data, it leans most heavily toward Extensive coverage of 337+ commercial and open-source Chinese LLMs. and Comparing and selecting the best performing LLMs for specific applications.. It also shows measurable community traction with 6.3k GitHub stars.

#LLM Evaluation#Chinese LLMs#AI Benchmark#Model Ranking#Defect Analysis#Data Analysis#Communication
↗ Visit site★ GitHub
01

Features

01Extensive coverage of 337+ commercial and open-source Chinese LLMs.
02Multi-dimensional evaluation across 7 main domains and ~300 sub-dimensions.
03Provides detailed ranking lists for various capabilities and specific domains.
04Offers a large defect library with over 2 million LLM flaws for research and improvement.
05Supports customized model selection and free evaluation services for private models.
02

Why choose it

+Extensive coverage of 337+ commercial and open-source Chinese LLMs.
+Comparing and selecting the best performing LLMs for specific applications.
+Covers 6 supported environments or platforms, which is helpful for broader deployment needs.
+The latest recorded update is 2026-07-12, which suggests the project is still actively maintained.
03

Trade-offs

!This page does not list a concrete install command, so you may need to verify setup steps in the official docs before adopting it.
!There are at least 8 related tools in the same category, so the best choice is easier to make after side-by-side comparison.
04

Compatibility

OpenAI (GPT series)
Supported
Verified via docs
Google (Gemini series)
Supported
Verified via docs
Anthropic (Claude series)
Supported
Verified via docs
Baidu (ERNIE series)
Supported
Verified via docs
Alibaba (Qwen series)
Supported
Verified via docs
DeepSeek
Supported
Verified via docs
05

Use cases

↳Comparing and selecting the best performing LLMs for specific applications.
↳Identifying weaknesses and improving the capabilities of large language models.
↳Benchmarking private or custom LLMs against public models for performance and cost optimization.
06

How it compares

≈chinese-llm-benchmark sits in the RAG / Knowledge Base category, so it makes more sense to evaluate it alongside tools like mindsdb instead of in isolation.
≈If your main need is closer to "Comparing and selecting the best performing LLMs for specific applications.", that use case is a better lens for comparison than broad feature checklists alone.
≈chinese-llm-benchmark's licensing and community traction are both easier to judge in category context.
07

Alternatives

mindsdb logo
mindsdb★ 39.5k
Federated Query Engine for AI - The only MCP Server you'll ever need
vs →
Brave Search MCP logo
Brave Search MCP★ 88.6k
Allow your AI Agent to search the real-time internet using Brave Search API. Essential for getting up-to-date information.
vs →
Claude Flow logo
Claude Flow★ 65.2k
The leading agent orchestration platform for Claude. Deploy intelligent multi-agent swarms.
vs →
CopilotKit logo
CopilotKit★ 36.2k
React UI + elegant infrastructure for AI Copilots, AI chatbots, and in-app AI agents. The Agentic Frontend.
vs →
awesome-n8n-templates logo
awesome-n8n-templates★ 24.0k
Supercharge your workflow automation with this curated collection of n8n templates! Instantly connect your favorite apps-like Gmail, Telegram, Google Drive, Slack, and more-with ready-to-use, AI-powered automations. Save time, boost productivity, and unlock the true potential of n8n in just a few clicks.
vs →
dagster logo
dagster★ 15.9k
An orchestration platform for the development, production, and observation of data assets.
vs →
genai-toolbox logo
genai-toolbox★ 16.0k
MCP Toolbox for Databases is an open source MCP server for databases.
vs →
mcp-chrome logo
mcp-chrome★ 12.2k
Chrome MCP Server is a Chrome extension-based Model Context Protocol (MCP) server that exposes your Chrome browser functionality to AI assistants like Claude, enabling complex browser automation, content analysis, and semantic search.
vs →
See all alternatives →

Related searches

chinese-llm-benchmark AlternativesBest RAG / Knowledge Base Tools 2026Open Source RAG / Knowledge Basechinese-llm-benchmark Tutorialchinese-llm-benchmark Vs CompetitorsLLM EvaluationChinese LLMsAI Benchmark

Comments

Log in to leave a comment
  • ?
    usr_seed_0647May 22, 2026

    The reliable agent design scales well from prototype to production — 5、minimax-m2、deepseek-v3. Good documentation, reduces onboarding time.

  • ?
    usr_seed_0255May 3, 2026

    The clean approach to agent memory is more reliable than alternatives — rele评测:中文ai大模型能力评测(持续更新):目前已囊括335个大模型,覆盖chatgpt、gpt-5. Would recommend for clean use cases.

  • ?
    usr_seed_0631Mar 29, 2026

    The robust agent design scales well from prototype to production. Runs fine on Python 3.11.

  • ?
    usr_seed_0803Mar 14, 2026

    The solid approach to agent memory is more reliable than alternatives. The maintainers are responsive to issues.

On this page
01Features02Why choose it03Trade-offs04Compatibility05Use cases06How it compares07Alternatives
Stats
GitHub Stars★ 6.3k
Last commit1w ago
StatusActive
License—
CategoryRAG / Knowledge Base
Trend (30d)
+0.2k↑ 4.6%
Links
Documentation↗Discussion↗Issues↗Releases↗

Deploy on DigitalOcean — Get $200 Free Credit

Ad
© 2026 AgentIndex.app|Built by a 10-year iOS Developer.
QYSGitHubBuy me a coffee ☕

Browse by Category

Code AssistantWorkflow AutomationRAG / Knowledge BaseMulti-AgentBrowser AutomationLLM InfraDev ToolingObservability

Not affiliated with Anthropic, OpenAI or Microsoft.