AIME 2025 / Humanity's Last Exam (HLE)

AIME 2025 / Humanity's Last Exam (HLE) is a benchmark for evaluating AI on elite-level math and PhD-general reasoning, featuring the hardest known problems across advanced mathematics and science. It tests an AI's ability to solve complex, multi-step challenges that require deep logical deduction and expert knowledge.