explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

  1. Home
  2. /
  3. Dictionary
  4. /
  5. Humanity's Last Exam
Evaluation & Benchmarksaka HLE

Humanity's Last Exam

Humanity's Last Exam is a 2,500-question benchmark of expert-written questions across dozens of academic subjects, built to stay hard after older knowledge tests saturated.

Ask Melo about this← all terms

Released in January 2025, HLE is intentionally adversarial: questions were solicited from subject-matter experts specifically to resist models that had memorized older benchmarks like MMLU. Frontier scores moved quickly after launch, and model cards increasingly report separate no-tools and with-tools figures, since web and code access change the result more than raw model capability does.

Related terms

AI BenchmarkGPQAContamination AuditTool UseExploitGymBrowseComp

Where Humanity's Last Exam comes up

  • OpenRouter Web Search Benchmarks: How to Pick a Search Tool for Agents
  • Inkling-Small: Thinking Machines’ 12B-Active MoE Matches Inkling at 1/4 Size
  • AI Benchmarks in 2026: The Complete Guide to MMLU, GPQA, SWE-bench, and Beyond
  • Gemini 3.8 Flash Is Official: Benchmarks, Flash Cyber, and Pricing