OpenCompass
ActiveDescription
OpenCompass is a comprehensive LLM evaluation platform supporting a wide range of models including Llama, Mistral, GPT-4, Qwen, GLM, and Claude across 100+ benchmark datasets.
Key Features
- Comprehensive benchmarking across 100+ datasets covering knowledge, reasoning, math, code and more
- Compatible with major LLMs — Llama, Mistral, GPT-4, Qwen, GLM, Claude, etc.
- CascadeEvaluator for sequential multi-evaluator pipelines on complex assessment scenarios
- Built-in GenericLLMEvaluator (LLM-as-Judge) and MATHVerifyEvaluator for math reasoning
- CompassHub leaderboard and CompassRank model ranking visualizations
- Recommended by Meta AI, integrated into the Llama Get Started workflow
Use Cases
Strengths & Limitations
✅ Strengths
- • Actively maintained, recent updates
- • High community interest (7.4k stars)
- • Permissive open-source license (Apache-2.0)
- • Established track record (3 years in production)
⚠️ Limitations
- • High issue backlog (394 open issues)
Categories
Quick Start
Install with pip install opencompass, configure models and datasets via OpenCompass config files, and run opencompass to start evaluation. Supports on-demand dataset loading from ModelScope, with built-in CompassHub leaderboard for result visualization.