Summary: Standard AI benchmarks have become too easy. To challenge current systems and map the frontier of machine capability, a global consortium of nearly 1,000 researchers created “Humanity’s Last Exam” (HLE). This 2,500-question assessment covers deeply specialized topics—from translating ancient Palmyrene inscriptions to identifying microanatomical features in birds—designed explicitly to be beyond the reach of present AI models.
HLE was intentionally constructed so that any question answered correctly by an AI during the testing phase was excluded from the final set. Early results show humans perform strongly while leading models such as GPT-4o and Claude 3.5 struggle, underscoring the substantial gap between machine pattern recognition and genuine human expertise.
Key Facts
- A new benchmark: Humanity’s Last Exam aims to be the definitive expert-level assessment, positioned deliberately just beyond the current capabilities of the world’s most advanced AI systems.
- Global collaboration: Nearly 1,000 subject-matter experts in sciences, humanities and the arts contributed questions to ensure comprehensive coverage of human knowledge.
- Low AI scores: Early evaluations report very low performance for many models: GPT-4o scored 2.7%, Claude 3.5 Sonnet scored 4.1%, and OpenAI’s o1 model reached 8%. Even the most capable models currently tested struggle to surpass 40–50% accuracy.
- Unsearchable problems: Each question has a single verifiable answer but was crafted so it cannot be solved by a simple internet search.
- Preserving human relevance: Far from signaling the end of human expertise, HLE is intended as a rigorous tool to measure progress, reveal risks, and highlight areas where specialized human knowledge remains essential.
Source: Texas A&M
When advanced AI systems began performing extremely well on well-known academic tests, researchers realized many benchmarks were no longer discriminating. Exams like the Massive Multitask Language Understanding (MMLU) benchmark, once considered difficult, no longer stress the highest-capability models.
To address that shortfall, a worldwide consortium of nearly 1,000 researchers, including faculty from Texas A&M University, developed an exam intended to remain reliably challenging for the foreseeable future. The result is Humanity’s Last Exam (HLE), a deliberately demanding, expert-level assessment that current AI systems consistently fail.

HLE comprises 2,500 questions spanning mathematics, humanities, natural sciences, ancient languages and many highly specialized subfields. The project and its methodology are described in a paper published in Nature, and additional documentation is available from the project at lastexam.ai (project site referenced without external linking here).
Among the contributors is Dr. Tung Nguyen, an instructional associate professor in the Department of Computer Science and Engineering at Texas A&M, who helped author and refine many of the questions.
“When AI systems perform extremely well on standard benchmarks, it can create the false impression that they possess human-level understanding,” Nguyen said. “HLE reminds us that intelligence is more than pattern matching: it requires depth, context and years of specialized study.”
The exam’s goal was not to trick humans but to precisely identify where AI systems fail to match expert human reasoning and domain-specific knowledge.
A global effort to measure AI’s limits
Questions were written and vetted by experts worldwide to ensure each item has a single, unambiguous, verifiable answer that cannot be resolved instantly through internet retrieval. Prompts were drawn from real expert-level problems: translating inscriptions in ancient scripts, diagnosing obscure anatomical features, and analyzing fine-grained philological distinctions, among others.
Each question was pretested against leading AI systems; any question a model could reliably answer during this phase was removed. This filtering process produced an exam intentionally positioned just beyond current model capabilities.
The approach worked: initial testing shows that many state-of-the-art models score poorly. Reported numbers include 2.7% for GPT-4o, 4.1% for Claude 3.5 Sonnet, and 8% for OpenAI’s o1. The most advanced models tested to date have reached roughly 40–50% accuracy, still far short of expert human performance.
Why a new benchmark matters
The core issue is practical as well as scientific. Without benchmarks that reflect the depth and specificity of expert knowledge, stakeholders—policymakers, developers and users—may misjudge what AI systems can actually do.
“Benchmarks provide the foundation for measuring progress and identifying risks,” Nguyen noted. He contributed 73 of the public questions on HLE, the second-highest number among contributors, and authored many questions in mathematics and computer science.
As the consortium’s paper argues, good performance on human-designed exams does not necessarily equate to human-like understanding. Many conventional benchmarks reward surface-level pattern recognition on widely available data rather than the deep, often tacit expertise that experts develop over years.
Not a threat, a diagnostic tool
Despite its dramatic name, Humanity’s Last Exam is intended as a diagnostic instrument rather than an alarm. It clarifies where AI is strong and where it is weak, helping researchers build safer, more reliable systems while underscoring why specialized human expertise remains vital.
“This isn’t a race against AI,” Nguyen said. “It’s a systematic way to understand strengths and shortcomings. That understanding is essential for designing better models and for making informed decisions about their use.”
A future-proof benchmark
HLE is designed to be a durable, transparent benchmark for evaluating future AI systems. To preserve its utility, the consortium has made a subset of the exam publicly available while keeping most questions confidential so models cannot memorize the answers in advance.
“For now, Humanity’s Last Exam stands as a clear measure of the gap between current AI capabilities and expert human performance,” Nguyen said. “Despite rapid advances, that gap remains substantial.”
Research on a grand scale
Nguyen emphasized that the project’s scale and interdisciplinarity were essential to its success.
“What made this project extraordinary was its scope,” he said. “Experts from nearly every discipline contributed—historians, physicists, linguists, medical researchers as well as computer scientists. That diversity is exactly what reveals the unevenness in today’s AI systems.”
Key Questions Answered:
A: The name is somewhat tongue-in-cheek but conveys the idea that this exam represents a final, rigorous hurdle for AI: passing it would indicate machine performance at the level of narrow human experts across many domains.
Q: If AI is advanced, why does it fail HLE?
A: Modern AI excels at pattern recognition and summarizing common data, but HLE targets deep, specialized knowledge and fine-grained reasoning that require years of domain experience. Simple statistical inference or web retrieval is not enough to solve many of these problems.
Q: Can an ordinary person pass this test?
A: No single person could pass the entire exam; it spans so many disciplines that only specialists can reliably answer questions within their own niche. The point is that humans with appropriate expertise still outperform AI across those narrow domains.
Editorial Notes:
- This article was edited by a Neuroscience News editor.
- The journal paper was reviewed in full by the editorial team.
- Additional context was added by the news staff to clarify technical points.
About this AI research news
Author: Lesley Henton
Source: Texas A&M
Contact: Lesley Henton – Texas A&M
Image credit: Neuroscience News
Original Research: Open access. “A benchmark of expert-level academic questions to assess AI capabilities” by the Center for AI Safety, Scale AI and the HLE Contributors Consortium. Published in Nature. DOI: 10.1038/s41586-025-09962-4
Abstract
A benchmark of expert-level academic questions to assess AI capabilities
Benchmarks are essential for tracking rapid advances in large language models, but many existing tests no longer challenge cutting-edge systems. As models surpass 90% accuracy on popular benchmarks, these evaluations fail to reveal important limitations in reasoning, domain expertise and calibration.
In response, the consortium developed Humanity’s Last Exam (HLE), a multimodal benchmark at the frontier of human knowledge. HLE is an expert-level, closed-ended academic benchmark covering dozens of subjects with 2,500 questions suitable for automated scoring. Questions were authored by subject-matter experts worldwide and designed so each has a clear, verifiable solution that cannot be quickly answered through internet retrieval.
State-of-the-art models demonstrate low accuracy and poor calibration on HLE, highlighting a marked gap between current model capabilities and expert human performance on closed-ended academic questions. To inform research and policy with a clearer picture of model strengths and weaknesses, the consortium has publicly released a portion of HLE and maintains the benchmark for ongoing evaluation.