We're looking for a talented early-career engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. You'll work on software that enables the world's largest AI models to train across massive GPU clusters, developing support for communication libraries and frameworks like NCCL, NVSHMEM, and NIXL.
This is a ground-floor opportunity to work at the intersection of high-performance computing, networking, and machine learning infrastructure - building the systems that power the largest AI workloads in the cloud.
Key job responsibilities
- Write high-performance C/C++ code for network communication libraries running on custom AWS hardware
- Build and maintain infrastructure that monitors functionality and performance of large-scale AI/ML workloads
- Develop automation using Python and AWS tools (CI/CD, Grafana, Athena) to test, benchmark, and deliver software to customers
- Design mechanisms to detect functional and performance regressions before they reach production
- Work across many instance types, software stacks, and Linux environments