AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain — and we’re looking for talented people who want to help.
You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You’ll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you’ll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.
AWS operates one of the world's largest and most complex networks — and it's growing fast. Millions of customers run mission-critical workloads that depend on us being always on.
Key job responsibilities
Design, build, and operate distributed services that are reliable, scalable, and maintainable at global scale
Own your services end-to-end — from architecture and implementation through deployment and production operations
Identify and drive improvements to system reliability, latency, and operational posture
Contribute to team engineering standards, code reviews, and technical direction
Mentor junior engineers and contribute to a culture of high engineering quality
About the team
Network Monitoring and Remediation (NMR) is the team that makes that possible. We prevent, predict, detect, and remediate network impairments before they reach customers. Our systems ingest trillions of telemetry data points daily from millions of network devices, apply advanced streaming analytics and ML/AI to surface actionable signals, and drive automated mitigation through our event management platform. Today, over 98% of detected issues are resolved without human intervention. For the remainder, we provide expert diagnostic tooling that enables operators to act decisively. Our vision: every network failure is pre-emptively resolved by autonomous systems — zero customer impact, zero human intervention.