Epoch AI is looking for a Software Engineer who will help us evaluate frontier AI models, enabling researchers, developers, and policymakers to better understand AI development. The role will involve running and maintaining our benchmarking infrastructure as well as contributing to the development of brand new benchmarks. About the role Please do not include a cover letter, photograph, or headshot of yourself, or any personal information that is not relevant to the role for which you're applying (including marital status, age, identity traits, etc.). We are looking for a Software Engineer to help us expand and develop our AI Benchmarking Hub. You will work closely with the rest of the benchmarking team to run and maintain benchmarks, integrate with AI providers, set up existing benchmarks to run on our infrastructure, help design and develop brand new benchmarks, and facilitate internal experiments. This role is fully remote, and we are able to hire in many countries. We invite anyone who is interested to apply, regardless of background, experience, or credentials. Applications are rolling.
Epoch AI is an AI evaluation and benchmarking company that provides rigorous, independent insights into key trends in artificial intelligence. The company operates an AI Benchmarking Hub that evaluates frontier AI models to help researchers, developers, and policymakers better understand AI development and capabilities. Their mission is to deliver public, trustworthy evaluations of AI capabilities on challenging benchmarks, empowering stakeholders to make well-informed decisions about AI. Epoch AI focuses on implementing and developing AI benchmarks within their evaluation infrastructure, primarily using the Inspect library, and collaborating with AI providers to track and evaluate new model releases.