← AI Terminology
GPQA - Graduate-Level Google-Proof Q&A
GPQA is a benchmark of extremely hard multiple-choice science questions written so that even experts struggle and web search is of limited help.
It targets graduate-level reasoning beyond trivia.
It targets graduate-level reasoning beyond trivia.
Why It Matters in AI
As MMLU saturates, harder science evals are needed. GPQA (especially Diamond split) differentiates frontier reasoning models and resists simple memorisation/search, making it a key 2024–2026 benchmark.
Key Points
| Aspect | Description |
|---|---|
| Use | Frontier model cards; reasoning-mode evals |
| Caveat | Still possible contamination; interpret carefully |
| Design | “Google-proof” aspirationally; hard for non-experts |
| Splits | Main, Diamond (hardest) |
| Content | Biology, physics, chemistry expert questions |
| Related | MMLU, MATH, Humanity’s last exam-style sets |
Simple Analogy
An oral exam for PhD candidates where looking up the textbook mid-question barely helps — you must reason.
Common Usage Examples
- Report GPQA-Diamond with/without tools
- Compare reasoning models vs chat models
- Use as holdout for overfit MMLU systems
- Pair with expert human baselines
Summary
In short: GPQA is a graduate-level, hard science QA benchmark designed to stress genuine reasoning beyond easy recall.