Agents’ Last Exam is a community-built benchmark of over 1,000 economically valuable, long-horizon agent tasks, contributed by 300+ industry practitioners and led by UC Berkeley RDI. The benchmark measures how far frontier AI agents remain from reliably completing realistic, high-stakes work: frontier agents achieve only a 2.6% full-pass rate on the hardest tier. I contributed tasks as one of the 300+ practitioner contributors.