Description
The AWS Neuron Science Core Algorithm team is looking for talented Applied Scientists to push the frontier of hardware-aware machine learning for Trainium and Inferentia, the AWS Machine Learning accelerators. In this rare role at the intersection of LLM modeling, large-scale training systems, and hardware/datatype co-design, you own model and algorithm decisions jointly with AWS custom silicon. You will own solutions end-to-end from research through production, publish at top venues, and work alongside distinguished engineers and scientists in a strategic growth area for AWS.
We actively work on these areas:
- Low-precision training and inference: MXFP8, MXFP4, and sub-4-bit training and inference recipes, stochastic rounding, and Trn4 datatype exploration.
- Trn-friendly architectures: model architectures that exploit hardware strengths without sacrificing quality.
- System-aware optimizers & efficient distributed systems: efficient optimizers and distributed systems that give the best accuracy, co-designed with the hardware.
- Foundation-model pre-training accuracy: end-to-end validation across model scales, catching training divergence early, and equivalence-checking tooling.
- GenAI for systems: RL post-training for NKI kernel generation, mitigating reward-hacking and accelerating under low precision on Trn.
Key job responsibilities
- Own scientific problems end-to-end - from research and experimentation through production impact - applying rigorous evaluation to complex, ill-defined problems at large scale.
- Develop production-quality code in PyTorch or JAX and integrate scientific components into large-scale training and inference systems with operational excellence and efficient resource usage.
- Partner with foundation-model, engineering, and hardware-architecture teams so your findings directly inform what gets built into Trainium and shipped in the product stack.
- Mentor fellow scientists and interns, give constructive peer reviews, and help shape team goals, priorities, and the technical roadmap.
- Author and publish research at top peer-reviewed venues (ICLR, NeurIPS, ICML, MLSys) and engage the broader scientific community.