Skip to main content
Products & Tools

Humanity's Last Exam (Diamond) - Scale AI

Challenging LLMs at the frontier of human knowledge HLE-Diamond, built with the Center for AI Safety (CAIS) , is a refined subset of Humanity’s Last Exam (HLE) resulting from a year-long process of cleaning and refinement with input from research communities.

By Precis Daily Newsroom3 min read595 words
Illustration for: Humanity's Last Exam (Diamond) - Scale AI
Illustration
Key points
  • At release in January 2025, frontier models scored in the single digits and systematically exhibited uncalibrated overconfidence in their answers.
  • It consists of 1,000 questions drawn from the existing HLE pool and its held-out reserve.
  • The strongest model we tested lands at 60.6%, well below the noise ceiling.

Challenging LLMs at the frontier of human knowledge HLE-Diamond, built with the Center for AI Safety (CAIS) , is a refined subset of Humanity’s Last Exam (HLE) resulting from a year-long process of cleaning and refinement with input from research communities. In partnership with the Center for AI Safety, we addressed the problem of benchmark saturation by creating Humanity's Last Exam (HLE): the toughest, most subject-diverse, multi-modal questions we could assemble, designed to be the last academic exam of its kind for AI. HLE tests both depth of reasoning (eg. world-class mathematical problems) and breadth of knowledge across its subject domains, providing a precise measurement of model capability. At release in January 2025, frontier models scored in the single digits and systematically exhibited uncalibrated overconfidence in their answers. HLE-Diamond is where that work lands now. It consists of 1,000 questions drawn from the existing HLE pool and its held-out reserve. High accuracy on HLE-Diamond would demonstrate that AI has achieved expert-level performance on closed-ended cutting-edge scientific knowledge. It would not alone suggest autonomous research capabilities or "artificial general intelligence." Most of these changes were made incrementally over a year of HLE-Rolling updates. HLE-Diamond is where they land as a single release. Size. 1,000 questions, down from 2,500. Every question comes from the existing HLE pool and its held-out reserve. Two partitions. HLE-Diamond consists of 500 reasoning and 500 knowledge questions, measuring reasoning and expert knowledge, respectively. Final HLE-Diamond results are aggregated between the two partitions. A frozen question set. HLE-Diamond is fixed, with stable question IDs, so a score from today is comparable to a score from next quarter. Headroom. The strongest model we tested lands at 60.6%, well below the noise ceiling. We expect HLE-Diamond to carry signal for the next 6 to 12 months. Model Overall Reasoning Knowledge GPT-6 Astra 60.6% 75.6% 45.6% Claude Opus 5.5 55.0% 63.2% 46.8% Claude Fable 5.1 51.3% 62.0% 40.6% GPT-6 Sol 33.8% 44.2% 23.4% Gemini 3.8 Flash 34.3% 38.6% 30.0% Muse Spark 1.3 25.4% 31.6% 19.2% Grok 4.7 23.4% 32.4% 14.4% Overall, we observed a strong correlation between reasoning and knowledge partitions; models that perform better on reasoning consistently score higher on knowledge. We report accuracy on the reasoning and knowledge partitions, plus an overall accuracy across all 1,000 questions. Models are ranked on the leaderboard using overall accuracy. We also use the model's own stated confidence to derive an RMS calibration error , using the implementation from Hendrycks et al., 2022 with the default hyperparameters provided. We want to emphasize calibration error as an important metric alongside accuracy. Evaluation is automatic. Models are prompted to give a final answer and an estimation of confidence using the system prompts (or user prompt when not configurable), following the setup from Wei et al., 2024 . Because HLE-Diamond uses closed-form solutions, we use an LLM judge as an automatic extractor and judge to compare the model response against the ground truth answer. HLE-Diamond questions are designed to be answerable in a closed-book setting, testing both reasoning and expert knowledge. Since HLE is also used to evaluate the capabilities of agentic systems, we outline our recommended settings for evaluating HLE-Diamond with tools here . Humanity's Last Exam was a global collaborative effort developed in partnership with the Center for AI Safety . We extend our deepest gratitude to all participating question contributors and expert reviewers involved in creating and refining the dataset, and to the researchers whose feedback across HLE-Rolling shaped HLE-Diamond. Rank (UB): 1 + the number of models whose lower CI bound exceeds this model’s upper CI bound.

Sources

Summarized from the linked originals.

Related stories

Illustration for: How Genie One reshapes work for finance teams
Products & Tools

• Give finance professionals an AI coworker grounded in their business context to accelerate decision-making • Enable teams to analyze performance, model scenarios, investigate variances, and speed up forecasting and financial reporting • Apply consistent guardrails across data access, actions, and AI usage, so every...

Databricks Blog5 min