← AI Terminology

HELM - Holistic Evaluation of Language Models

HELM is a living evaluation framework from Stanford CRFM that measures LLMs across many scenarios and metrics (accuracy, calibration, robustness, fairness, bias, toxicity, efficiency).

It emphasises multi-metric, transparent reporting.
Why It Matters in AI
Accuracy-only leaderboards hide harms and brittleness. HELM popularised holistic scorecards for foundation models — important for governance, procurement, and responsible release.
Key Points
Aspect Description
Org Stanford CRFM
Use Compare models beyond one number
Idea Many scenarios × many metrics
Related HELM lite/modifications; other scorecards
Audience Researchers, policymakers, enterprises
Transparency Standardised prompts, versioned results
Simple Analogy
A full medical checkup with bloodwork, stress test, and lifestyle review — not only weighing yourself on a scale.
Common Usage Examples
  • Read HELM model scorecards
  • Adopt multi-metric internal evals
  • Track toxicity + accuracy together
  • Cite HELM scenarios in model cards
Summary
In short: HELM holistically evaluates LLMs across accuracy, robustness, fairness, and more — multi-metric scorecards instead of single leaderboards.