← AI Terminology

GPQA - Graduate-Level Google-Proof Q&A

GPQA is a benchmark of extremely hard multiple-choice science questions written so that even experts struggle and web search is of limited help.

It targets graduate-level reasoning beyond trivia.
Why It Matters in AI
As MMLU saturates, harder science evals are needed. GPQA (especially Diamond split) differentiates frontier reasoning models and resists simple memorisation/search, making it a key 2024–2026 benchmark.
Key Points
Aspect Description
Use Frontier model cards; reasoning-mode evals
Caveat Still possible contamination; interpret carefully
Design “Google-proof” aspirationally; hard for non-experts
Splits Main, Diamond (hardest)
Content Biology, physics, chemistry expert questions
Related MMLU, MATH, Humanity’s last exam-style sets
Simple Analogy
An oral exam for PhD candidates where looking up the textbook mid-question barely helps — you must reason.
Common Usage Examples
  • Report GPQA-Diamond with/without tools
  • Compare reasoning models vs chat models
  • Use as holdout for overfit MMLU systems
  • Pair with expert human baselines
Summary
In short: GPQA is a graduate-level, hard science QA benchmark designed to stress genuine reasoning beyond easy recall.