Humanity's Last Exam (HLE)
Learn what Humanity’s Last Exam (HLE) measures, how its AI benchmark scores expert reasoning and confidence, and what its results reveal—and don’t.
Humanity’s Last Exam (HLE) is an AI benchmark that tests expert-level academic knowledge and reasoning through difficult questions with verifiable answers. Think of it as a demanding exam spanning many specialties rather than a general-knowledge quiz. It helps reveal where an AI system can solve challenging problems—and where fluent, confident responses conceal mistakes. (labs.scale.com)
How Humanity’s Last Exam Works#
HLE is a benchmark dataset: a standardized collection used to compare systems under defined conditions. Developed by the Center for AI Safety and Scale AI, its original finalized public collection contains 2,500 questions across more than 100 subjects, including mathematics, natural sciences, engineering, and humanities. (agi.safe.ai)
Questions include multiple-choice items and short-answer problems. They are closed-ended, meaning a correct answer can be checked, even when reaching it requires substantial reasoning. Expert contributors and reviewers help establish difficulty and answer quality. HLE is multimodal: some questions require interpreting images alongside text rather than processing language alone. (labs.scale.com)
This makes HLE relevant to large language models, which process and generate language, and image-capable systems. Its visual questions overlap with visual question answering, where a system answers questions about an image. However, recognizing a diagram’s components is only part of solving a specialized academic problem. (labs.scale.com)
Understanding Scores and Confidence#
Two measurements are especially important:
- Answer accuracy: The fraction of questions answered correctly. It measures successful answers, not whether every step in an explanation is sound.
- Confidence calibration: How closely confidence matches observed correctness. Among answers assigned 80% confidence, approximately 80% should be correct in a well-calibrated system. (scikit-learn.org)
HLE evaluations can request a final answer and a confidence estimate. These capture different properties: a system may improve accuracy while remaining overconfident when wrong. A stated confidence percentage is therefore not automatically a reliable probability. (agi.safe.ai)
When reading HLE evaluation methodology, check the question version, text-only versus multimodal coverage, prompts, and grading procedure. Also check whether external tools were allowed; tool-assisted and unassisted results describe different systems. (labs.scale.com)
How HLE Differs From Related Concepts#
HLE addresses benchmark saturation: when many systems score near a test’s ceiling, that test becomes less useful for distinguishing their capabilities. Harder questions create more room to measure differences. The name expresses that ambition; it does not mean AI evaluation ends with this exam. (labs.scale.com)
It is not a certificate of artificial general intelligence, meaning broadly applicable intelligence across tasks. Strong academic answers do not establish autonomous discovery, dependable planning, or safe behavior. The NIST AI Risk Management Framework addresses broader concerns such as trustworthiness and context-specific risk. (agi.safe.ai)
HLE also differs from object detection, which identifies and locates objects in images. Academic answer accuracy cannot substitute for measuring whether a detector finds the right objects in deployment images. (docs.ultralytics.com)
Two Practical Application Examples#
Scientific assistant selection. Consider a team choosing an assistant to explain advanced physics problems. HLE can provide an initial signal of academic capability, but the team should also test representative questions from its own work. A confidently incorrect answer could misdirect calculations or experiments, so expert review remains important. Treat this as a screening example, not proof that a particular assistant is suitable. (agi.safe.ai)
Visual engineering assistance. Consider an assistant interpreting equipment diagrams. HLE’s image-based questions illustrate why seeing symbols and understanding their relationships are separate challenges. A system might identify a component correctly yet infer the wrong operating behavior. Evaluation should therefore cover the complete task, including interpretations and consequences—not just visual recognition. (labs.scale.com)
Practical Evaluation Guidance#
Protect evaluation integrity by keeping training, validation, and test sets separate. Avoid data leakage, where evaluation information influences model development and makes performance appear better than it generalizes. HLE’s public questions require similar care around prior exposure. (developers.google.com)
For computer vision, use task-specific evaluation alongside any reasoning benchmark. After installing ultralytics, this documented Ultralytics YOLO validation workflow evaluates YOLO26 on the small COCO8 example dataset: (docs.ultralytics.com)
from ultralytics import YOLO
if __name__ == "__main__":
model = YOLO("yolo26n.pt")
# Compare detections with the dataset's validation labels.
metrics = model.val(data="coco8.yaml")
# Report the detector's localization and classification metric.
print("mAP50-95:", metrics.box.map)The output is mean average precision across several bounding-box overlap thresholds, not an HLE score. COCO8 demonstrates the workflow; representative held-out application data is needed for meaningful deployment evaluation. The central lesson is to match each performance claim to the capability actually tested. (docs.ultralytics.com)









