GPQA
Learn what the GPQA benchmark measures, how GPQA Diamond scores work, and how to interpret results when evaluating advanced scientific reasoning in AI.
GPQA stands for Graduate-Level Google-Proof Question Answering, a benchmark that evaluates how well AI systems answer difficult scientific questions. It tests specialized knowledge and reasoning in biology, chemistry, and physics through multiple-choice questions written by domain experts. Intuitively, GPQA asks whether a system can apply scientific understanding—not merely recognize familiar facts or retrieve a matching sentence.
For a large language model, which processes and generates language, GPQA provides a focused measure of advanced scientific question-answering ability. It is an evaluation tool, not a model architecture, training technique, or certification of scientific competence.
How GPQA Works#
Each question presents four answer choices, including one correct answer and three distractors: plausible but incorrect alternatives. Solving it may require interpreting assumptions, connecting several principles, or ruling out explanations that initially seem reasonable.
For example, a chemistry-style question might describe competing reaction pathways and ask which product is favored under particular conditions. Knowing a reaction’s name would not necessarily suffice; the system must apply the relevant constraints. This illustrates the task style rather than reproducing a benchmark question.
“Google-proof” describes the benchmark’s design goal: questions should remain challenging for knowledgeable nonspecialists even with web access. It does not mean answers are impossible to find online or that browsing is prohibited in every evaluation. Whether tools are allowed is part of the testing protocol and must be reported.
GPQA Diamond and Related Concepts#
The main GPQA set contains 448 questions. GPQA Diamond is a more selectively filtered, 198-question subset emphasizing expert agreement and difficulty for nonspecialists. The GPQA Diamond evaluation documentation describes its four-choice format and scientific domains. Diamond is not a separate model or a universal intelligence test, and its scores should not be treated as interchangeable with main-set scores.
SuperGPQA is a separate, broader graduate-level evaluation covering many disciplines rather than only GPQA’s three scientific domains. Its similar name does not make its results directly comparable.
GPQA also differs from question answering, the general task of producing answers to questions. GPQA is one specific evaluation of that capability. Unlike visual question answering, it does not test interpreting images to answer questions. Strong scientific text performance therefore cannot establish visual perception accuracy.
What a GPQA Score Means#
The central metric is answer accuracy: the fraction of questions answered correctly. With four choices, uniform random guessing has an expected accuracy of 25%. A hypothetical result of 150 correct answers out of 198 is approximately 75.8%.
However, the percentage alone is incomplete. Comparisons should identify the subset, model version, prompt, tool access, reasoning budget, and number of attempts. A single answer per question measures a different setup from generating several answers and selecting among them.
Small score differences also deserve caution. One additional correct answer on Diamond changes accuracy by about 0.5 percentage points. Binomial confidence intervals can help communicate uncertainty, although they do not resolve dataset bias or protocol differences.
Correct choices do not guarantee correct explanations. A system can select the right option while producing a hallucination—a plausible-sounding but unsupported explanation. GPQA scores should therefore complement, not replace, application-specific testing.
Two Practical Application Examples#
Scientific assistant selection. Consider a laboratory comparing assistants that explain unexpected enzyme activity. GPQA can provide an initial signal of scientific capability. The team should then test its own experimental scenarios and have specialists review explanations. Confusing correlation with mechanism could misdirect follow-up experiments despite a strong benchmark score.
Microscopy assistance. Consider a system combining a cell detector with an assistant that explains changes in cell counts. GPQA can inform evaluation of the language component’s scientific reasoning, but it says nothing about whether the detector missed overlapping cells. Evaluate detection and interpretation separately, then test the complete workflow. This context-specific approach aligns with the NIST AI Risk Management Framework.
Practical Evaluation Guidance#
Protect evaluation integrity through separate training, validation, and test sets. Avoid data leakage, where evaluation information influences development and inflates apparent performance. Public benchmark exposure can likewise complicate interpretation.
For a mixed vision-and-language application, evaluate the visual component with its own metrics. After installing ultralytics with pip install ultralytics, this documented Ultralytics YOLO validation workflow evaluates YOLO26:
from ultralytics import YOLO
if __name__ == "__main__":
# Load a pretrained object detector.
model = YOLO("yolo26n.pt")
# Compare predictions with labeled validation images.
metrics = model.val(data="coco8.yaml")
# Report detection performance, not scientific reasoning.
print("mAP50-95:", metrics.box.map)The COCO8 example dataset demonstrates the API; it is too small for deployment conclusions. The output is mean average precision, which evaluates detection across bounding-box overlap thresholds—not a GPQA score. The essential principle is to match each performance claim to the capability actually measured.









