Amazon Leo is building a low Earth orbit satellite broadband network to deliver fast, affordable connectivity to customers around the world, including those underserved by traditional infrastructure. Operating a live constellation and its ground network at scale means service disruptions are an inevitable reality, and how we respond to them, and how transparently we communicate with the customers who depend on us, defines the trust we earn as a connectivity provider.
We are looking for an experienced Incident Manager to serve as a Technical Call Leader. When a high severity event affects more than one part of the network, detection, customer communication, and engineering recovery are all happening at once, and someone has to run the whole event. That someone is you. You will be the single accountable owner of a high severity event from the moment it is declared to the moment it is resolved and closed, coordinating engineers, operations, customer communications, legal, and security toward the fastest possible restoration of service for customers.
This is a coordination and leadership role, not a hands on resolver role. You will not perform the technical fix or make specialized safety decisions. You will make sure the right responders are engaged, the severity is assessed correctly, work that crosses team boundaries is sequenced, communication and any regulatory obligations are tracked, and the event ends in learning that prevents recurrence. Your effectiveness is measured by the quality of your coordination under pressure, not by personal troubleshooting. You are calm and diligent in high stress situations, you create urgency and focus, you ask the hard questions, and you can direct senior people who do not report to you.
Due to applicable export control laws and regulations, candidates must be a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum.
Key job responsibilities
• Serve as the single accountable call leader on high severity events, from declaration through resolution and closure.
• Assess and set event severity based on customer and business impact, and adjust it as the event evolves.
• Run the incident bridge, keeping the engineering effort and the customer communication effort aligned on a single source of truth.
• Sequence work across engineering, operations, mission and flight operations, customer communications, legal, and security, and engage the right responders quickly.
• Keep customer communication and any regulatory notification obligations visible and on time, at the same priority as the technical mitigation.
• Escalate early and freely, clear blockers, and make the stand down decision against clear recovery criteria
• Run clean handovers on prolonged events, and share a twenty four by seven on call rotation.
• Proactively investigate early warning signals to catch potential events before they reach customers.
• Own the post event loop, ensuring a correction of errors and root cause analysis are completed, assigned, and tracked to closure.
• Audit the coverage of monitoring and customer notification so events are detected and communications reach the right customers.
• Design, run, and score simulated incident exercises, and train partner teams on the incident process and their role on the call.
• Improve runbooks and incident tooling, reducing the manual work of running an event.