The AWS Neuron Science Core Algorithm team is looking for talented Applied Scientists to push the frontier of hardware-aware machine learning for Trainium and Inferentia, the AWS Machine Learning accelerators. In this rare role at the intersection of LLM modeling, large-scale training systems, and hardware/datatype co-design, you own model and algorithm decisions jointly with AWS custom silicon. You will own solutions end-to-end from research through production, publish at top venues, and work alongside distinguished engineers and scientists in a strategic growth area for AWS.
We actively work on these areas:
- Low-precision training and inference: MXFP8, MXFP4, and sub-4-bit training and inference recipes, stochastic rounding, and Trn4 datatype exploration.
- Trn-friendly architectures: model architectures that exploit hardware strengths without sacrificing quality.
- System-aware optimizers & efficient distributed systems: efficient optimizers and distributed system that gives best accuracy, co-designed with the hardware.
- Foundation-model pre-training accuracy: end-to-end validation across model scales, catching training divergence early, and equivalence-checking tooling.
- GenAI for systems: RL post-training for NKI kernel generation, mitigating reward-hacking and accelerating under low precision on Trn.
Key job responsibilities
- Own scientific problems end-to-end - from research and experimentation through production impact - applying rigorous evaluation to complex, ill-defined problems at large scale.
- Develop production-quality code in PyTorch or JAX and integrate scientific components into large-scale training and inference systems with operational excellence and efficient resource usage.
- Partner with foundation-model, engineering, and hardware-architecture teams so your findings directly inform what gets built into Trainium and shipped in the product stack.
- Mentor fellow scientists and interns, give constructive peer reviews, and help shape team goals, priorities, and the technical roadmap.
- Author and publish research at top peer-reviewed venues (ICLR, NeurIPS, ICML, MLSys) and engage the broader scientific community.
A day in the life
You might start your morning reviewing large-scale training runs — checking accuracy at a new low-precision datatype or debugging a divergence before it costs a run — then join a design discussion with engineering partners on how to land your recipe in the production stack. After lunch you could be whiteboarding a Trn-friendly architecture variant or an RL post-training approach for kernel generation with a teammate, then writing code to prototype it on Trainium. You will regularly present findings to the team and to leadership, review peers' and interns' work, and stay connected with the academic community.
About the team
AWS Neuron is the software of Trainium and Inferentia, the AWS Machine Learning chips. Inferentia delivers best-in-class ML inference performance at the lowest cost in the cloud to our AWS customers. Trainium is designed to deliver the best-in-class ML training performance at the lowest training cost in the cloud, and it's all being enabled by AWS Neuron. Neuron is a software that includes an ML compiler and native integration into popular ML frameworks. Our products are being used at scale with external customers like Anthropic and Databricks as well as internal customers like Amazon FMR, Amazon AGI, Amazon Bedrock, Amazon Robotics, Amazon Ads, and many more.