We’re hiring a Software Development Manager (SDM) to help shape and drive Incident management tooling development efforts as part of the incident response program for the worldwide Amazon retail websites.
We are re-imagining incident management & response for Amazon’s retail operations. Amazon is evolving faster than our incident management/response programs can keep up. It’s time to change that. You’ll be reporting directly into the single threaded owner for incident command systems and will be responsible for driving programs across multiple “two-pizza” teams. You’ll be directly responsible for helping curate complex user requirements and drive adoption of both greenfield and established tooling. When Amazon is under duress, every single minute matters. Your efforts will have an impact and influence on Amazon executive decisions and every single team at Amazon that interacts with our centralized control centers or outage calls. You’ll be required to dive into post incident analysis to understand what went wrong, how we can improve it for the future, and how blind spots will be covered in the future. Amazon incidents are complex, fast paced, and deeply nuanced. If you like a challenge, this is certainly an eye opening one. This will be a fascinating and critical culmination of technical deep diving, high impact judgment calls, program/product/process management, and sweeping assessments of company wide readiness for handling and managing incidents.
As an SDM on the Incident Command Systems team, you will be responsible for driving the development and delivery of industry-leading incident management solutions that empower our command teams to rapidly triage, manage, and resolve issues in mission-critical systems. Your tooling will focus on user-centric streamlined workflows optimized for usability and decision-support. You will help refine and track the appropriate metrics that best represent our efforts to accelerate remediation of customer risking events. This role has the scope to affect Amazon’s entire ecosystem of tens of thousands of services and make a difference to both our customers and developers. Enhancing Amazon’s incident response posture creates a flywheel of improvements that include adoption of best practices across all organizations.
You will thrive if you love fast-paced, startup-like work environments focused on building systems from the ground up. You will be involved in cross team design, and data analysis. You will also maintain client outreach while evangelizing operational excellence engineering culture across Amazon.
Key job responsibilities
* Manage a team of talented software engineers focused on building incident management applications focused on workflow optimization and decision-support
* Set the technical vision and roadmap for the Incident Management Tools team, aligning it with the broader business strategy
* Foster a culture of innovation, collaboration, and continuous improvement within your team
* Mentor and develop your engineers, helping them grow their technical and leadership skills
* Work closely with product managers, data scientists, and other stakeholders to understand customer needs and translate them into effective engineering solutions
* Ensure the timely delivery of high-quality, scalable, and maintainable software solutions
* Optimize team processes, workflows, and tools to maximize productivity and efficiency with an AI-native mentality
* Represent the Incident Management Tools team and its work to senior leadership and cross-functional partners
About the team
The Incident Command Systems team at Amazon is responsible for envisioning and building programs, which consistently improve remediation times for outages. This group consists of multiple 2-pizza teams (teams of 6-10 engineers) that each own software components for monitoring, anomaly detection of website degrading issues as well as incident management software used during these outages.