Have you ever wondered what it would be like to build massively scalable systems that are used by the world's largest cloud infrastructures?
Would you enjoy broad yet equally deep scope that impacts all AWS systems globally?
AWS Manufacturing Infrastructure Services continues to pioneer and our team is architecting, building, and operating scalable services that are at the core of AWS infrastructure.
What do we do?
We own all AWS platforms, services, infrastructure, and tools that ensure the health of AWS hardware by testing every new system across all AWS manufacturing sites and ensure they are healthy when delivered to data centers. Our platform enables service owners such as EC2, EBS, S3, and other to deliver healthy servers for their service. Our team leads a large-scale service that sets the bar for Amazon and the industry in platform level services, effectively enabling the hardware at scale by designing and developing the software that manages the verification and testing of every server and rack in AWS manufacturing sites.
Why it’s high-impact?
We set the bar high to ensure that AWS customers get the capacity they need to run their applications on a healthy hardware server.
What’s the challenge?
There are many ambiguous and difficult challenges in our fast-moving space. Our platform is mission critical and requires deep system and software expertise. You need to have the ability to work within a fast moving and startup-like environment in a large company. You will identify solutions, trying ideas, given space to fail and iterate to produce products that your customers love.
What you will do?
You will be a part of a team to build the next generation of platform level software and systems that enables us to deliver healthy hardware to AWS customers. You design and deliver technology solutions which solve difficult business problems.
Who would succeed in this role?
Deeply technical engineers, who stay close to the customer as well as the systems architecture and design. They think about customer experience and the outcome. A person who works autonomously and dives deep in to a problem to deeply understand how things work, when to make subtle change, and when to disrupt the status quo to achieve the right results.
Key job responsibilities
- Own the reliability of manufacturing networking, systems, and platform infrastructure across AWS manufacturing partner sites globally.
- Develop infrastructure-as-code to stand up and manage distributed manufacturing services at new and existing sites.
- Build monitoring, alerting, and anomaly detection systems that identify infrastructure failures before they impact manufacturing throughput.
- Develop troubleshooting tools and runbooks that enable rapid diagnosis and resolution of site infrastructure issues.
- Consult with ODM/CM IT teams on network design and infrastructure standards, working across organizational boundaries where partners don't report to AWS.
- Drive site expansion readiness, ensuring new manufacturing sites are infrastructure-ready within target onboarding timelines.
- Implement security and operational best practices for manufacturing environments, including credential management, network segmentation, and access controls.
- Contribute to architecture decisions for manufacturing platform evolution.
- Develop and build systems that enables operating manufacturing infrastrcure operation at scale
About the team
We are EC2 Manufacturing Infrastructure Services. We exist to make sure that a working rack delivers to our customers in AWS data centers.
Every server that powers EC2, every rack that runs AI/ML workloads, every piece of capacity that AWS customers depend on passes through our systems before it ships. We own Server Level Testing, Rack Level Testing, Firmware Upgrade services, code deployment into manufacturing sites, network architecture at contract manufacturers, and all infrastructure installed at those sites. When yield drops or a test pipeline breaks, we are the team that responds.
Nothing reaches an AWS data center without passing through us first. If our services go down, manufacturing lines stop. If our infrastructure is unreliable, capacity doesn't ship. The blast radius of what we do is measured in billions of dollars of customer workloads.
We are a small team with enormous scope. You will own real systems on day one. We don't have layers of abstraction between you and the problem. You build it, you ship it, you operate it, you improve it.