Overview In this role you will design and implement evaluation environments for frontier AI models, including LLMs and multimodal systems, to produce auditable accuracy signals guiding model releases. You’ll research novel evaluation methods for emerging model families and scale