chinese-llm-benchmark
ReLE评测(原名CLiB)是一个持续更新的中文AI大模型能力评测项目,已覆盖337个商用及开源大模型。它提供多维度能力评测和综合排行榜,并包含超200万的大模型缺陷库,以帮助社区研究和改进模型。
chinese-llm-benchmark 当前归类于 RAG / Knowledge Base,更适合那些希望围绕具体工作流而不是单一功能点来选型的团队。从当前资料来看,它尤其强调 广泛覆盖337+个商用及开源中文大模型。、对比和选择特定应用场景下表现最佳的大模型。 这类能力。同时它也具备一定社区热度,当前 GitHub Stars 为 6.4k。
功能特性
为什么选择它
权衡点
兼容性
使用场景
同类对比
同类工具
相关搜索
评论
- ?usr_seed_06472026年5月22日
The reliable agent design scales well from prototype to production — 5、minimax-m2、deepseek-v3. Good documentation, reduces onboarding time.
- ?usr_seed_02552026年5月3日
The clean approach to agent memory is more reliable than alternatives — rele评测:中文ai大模型能力评测(持续更新):目前已囊括335个大模型,覆盖chatgpt、gpt-5. Would recommend for clean use cases.
- ?usr_seed_06312026年3月29日
The robust agent design scales well from prototype to production. Runs fine on Python 3.11.
- ?usr_seed_08032026年3月14日
The solid approach to agent memory is more reliable than alternatives. The maintainers are responsive to issues.