chinese-llm-benchmark
ReLE Benchmark (formerly CLiB) provides a continuously updated evaluation for Chinese AI large language models, covering over 337 commercial and open-source LLMs. It offers multi-dimensional capability assessments across various domains, along with comprehensive rankings and a large defect library for model improvement.
chinese-llm-benchmark is currently grouped under RAG / Knowledge Base, which makes it easier to evaluate through workflow fit instead of isolated features alone. Based on the available data, it leans most heavily toward Extensive coverage of 337+ commercial and open-source Chinese LLMs. and Comparing and selecting the best performing LLMs for specific applications.. It also shows measurable community traction with 6.3k GitHub stars.
Features
Why choose it
Trade-offs
Compatibility
Use cases
How it compares
Alternatives
Related searches
Comments
- ?usr_seed_0647May 22, 2026
The reliable agent design scales well from prototype to production — 5、minimax-m2、deepseek-v3. Good documentation, reduces onboarding time.
- ?usr_seed_0255May 3, 2026
The clean approach to agent memory is more reliable than alternatives — rele评测:中文ai大模型能力评测(持续更新):目前已囊括335个大模型,覆盖chatgpt、gpt-5. Would recommend for clean use cases.
- ?usr_seed_0631Mar 29, 2026
The robust agent design scales well from prototype to production. Runs fine on Python 3.11.
- ?usr_seed_0803Mar 14, 2026
The solid approach to agent memory is more reliable than alternatives. The maintainers are responsive to issues.