Amazon's Artificial General Intelligence (AGI) organization is seeking an Applied Scientist III to advance the science of Responsible AI evaluation for large language models and generative AI. In this role, you will lead the design and development of rigorous evaluation methods, benchmarks, and metrics that measure the safety, fairness, robustness, and trustworthiness of frontier models. You will work with large-scale datasets, modern deep learning frameworks, and world-class scientists and engineers to turn research into evaluation systems that shape model launch decisions at Amazon scale.
Key job responsibilities
- Lead the design and implementation of evaluation frameworks, benchmarks, and metrics for responsible AI, including safety, fairness, robustness, and harmful content.
- Build scalable automated evaluation pipelines for large language models, including model-based and human-in-the-loop evaluation.
- Partner with pretraining, post-training, and product teams to translate evaluation results into model improvements and launch decisions.
- Conduct rigorous experimentation and statistical analysis, and publish research at top venues.
- Mentor junior scientists and help raise the scientific bar of the team.
- Champion responsible AI practices across the model development lifecycle.
About the team
The AGI Responsible AI (RAI) team builds the science and systems that make Amazon's large language models safe, fair, and trustworthy. We work on problems spanning safety evaluation, content moderation, watermarking, bias mitigation, and alignment. Our team values scientific rigor, customer obsession, and rapid iteration, and we collaborate closely with pretraining, post-training, and product teams across AGI.