Join the AI Studios Engineering org within Prime Video and Amazon MGM Studios as a Machine Learning Engineer on CreativeFlux, the ML platform powering Nara, our AI-native content creation platform for professional animation and live action/VFX.
You will own the training and serving infrastructure that puts generative models in front of artists working on Prime Video's animated and live action productions. You'll work daily with Applied Scientists, animators, filmmakers, and storytellers, building the systems that turn new model architectures into production-grade tools artists can rely on under tight deadlines.
This role suits engineers who want their work to expand what storytelling can look like, putting new generative capabilities into the hands of the people building Prime Video's animated and live action shows.
Key job responsibilities
Build and operate ML training and serving infrastructure for the CreativeFlux platform. Strong hands-on experience with EKS / Kubernetes is required.
Design, deploy, and operate ML training and inference workloads on EKS, including GPU node groups, custom schedulers, autoscaling, resource quotas, and gang scheduling for multi-GPU jobs
Integrate generative ML models (image, video, audio) into production inference pipelines, partnering with Applied Scientists to take new architectures from research to scaled deployment
Build training and fine-tuning pipelines on EKS and SageMaker, including data preparation, distributed training, evaluation harnesses, and checkpointing strategies
Improve GPU utilization, throughput, and latency on the inference fleet through batching, quantization, model compilation, serving framework tuning, and Kubernetes-native scaling primitives (HPA, KEDA, custom controllers)
Build automated evaluation pipelines and quality metrics that catch model regressions before they reach Artists, including human-in-the-loop review where needed
Contribute to integrations with Digital Content Creation (DCC) tools such as Adobe Animate, After Effects, Maya, and Storyboard Pro at the points where models are surfaced to creators
Own operational health of model endpoints on EKS, including monitoring, alerting, on-call response, capacity planning, and post-incident analysis
Write clean, well-tested code and participate in design and code reviews
Stay current with generative ML, distributed training, inference optimization, and GPU efficiency techniques, and apply them where they move team metrics
A day in the life
As an MLE on the AI Studios team, you'll spend most of your time deploying and tuning generative ML models, optimizing inference pipelines, and building the EKS infrastructure that lets Applied Scientists move new architectures into production. A typical day might include scaling a video generation endpoint to handle a production deadline, debugging a multi-GPU training job that's checkpointing too slowly, or working with a Scientist to take a new image model from a research notebook to a serving fleet. You'll work with other Nara platform teams who owns other building blocks, making sure ML capabilities land cleanly in the workflows artists actually use. You'll participate in team standups, model review sessions, on-call rotations, and platform-wide design reviews with the other Nara teams. You'll learn from experienced engineers and scientists and gradually take on more complex challenges as you grow.
About the team
We're AI Studios within Prime Video and Amazon MGM Studios, building Nara, our AI-native content creation platform. Our mission is to make Amazon a leader in AI-enabled content creation, giving creators direct control over model outputs while reducing the cost and timeline of animation and live action productions.
Our engineering team spans distributed systems engineers building scalable infrastructure, ML engineers optimizing 1P model training and inference, and full-stack engineers building creative facing agentic tools.