Preparing for system design interviews?  Try bugzed.com →

LM Evaluation Harness

JSON →
library 0.4.11 ·python
verified Jun 28, 2026

LM Evaluation Harness (lm-eval) is a comprehensive framework for evaluating language models on a wide range of benchmarks and tasks. It supports various model backends (HuggingFace, vLLM, SGLang, etc.) and provides a standardized way to compare model performance. The current version is 0.4.11, and it maintains a rapid release cadence with frequent minor updates and occasional breaking changes.

total hits 18
actors 6 distinct systems
last hit 18d ago OAI-SearchBot
GPTBot
3
ByteDance
3
OAI-SearchBot
2
Script
1
Humans
3

top countries 🇺🇸 United States · 🇸🇬 Singapore · 🇨🇦 Canada · 🇩🇪 Germany · 🇬🇧 United Kingdom