Epoch AI is looking for a researcher to evaluate frontier AI models on hard-to-grade tasks drawn from real-world scenarios. About the role We’re seeking a Researcher to lead a new effort evaluating how well frontier models perform on the kinds of open-ended tasks that make up real office work. You will curate a suite of realistic tasks to serve as a benchmark, design the grading rubrics for AI performance, and run newly-released models through the suite, assessing their performance both quantitatively and qualitatively. The focus is on how models handle messy, real-world work rather than on scientific knowledge or programming ability. The role makes heavy use of AI tools, but strong software engineering experience is not required. Comfort setting up AI-assisted automated workflows is a plus. If this role sounds interesting, we are also looking for researchers on multiple other teams. Applications are rolling.
Epoch AI is an AI evaluation and benchmarking company that provides rigorous, independent insights into key trends in artificial intelligence. The company operates an AI Benchmarking Hub that evaluates frontier AI models to help researchers, developers, and policymakers better understand AI development and capabilities. Their mission is to deliver public, trustworthy evaluations of AI capabilities on challenging benchmarks, empowering stakeholders to make well-informed decisions about AI. Epoch AI focuses on implementing and developing AI benchmarks within their evaluation infrastructure, primarily using the Inspect library, and collaborating with AI providers to track and evaluate new model releases.