Do you want to engineer the systems that help teams proactively detect problems across large-scale infrastructure? Our team builds the telemetry used across Amazon's operations to catch and diagnose availability issues before they affect the people relying on that infrastructure. We develop the software and services that generate, collect, and analyze network and infrastructure telemetry at scale.
In this role, you'll build and improve the agents and services that make telemetry data available and useful. That includes writing agents that run on a range of compute platforms, including network devices, Linux-based clients, and robotics platforms, and getting into the low-level details of running that software across different operating systems. You'll also contribute to the backend services that analyze and distribute metrics and insights and manage configuration, storage, and reporting.
Key job responsibilities
- Build and maintain software and agents that run across device platforms, including network devices, Linux-based clients, and robotics platforms, working where hardware and software meet.
- Handle the low-level details of running software on varied hardware: packaging, deployment, permissions, configuration, and service lifecycle across different operating systems.
- Contribute to the backend services that collect, process, and store telemetry data.
- Troubleshoot and resolve issues in the software you support, from the device and networking layer up through the backend services.
- Write well-tested, maintainable code and improve the reliability and quality of the systems the team owns.
About the team
The Telemetry Engineering team prioritizes creating trustworthy signals and responding to feedback from consumers. We iterate quickly and aim to ship improvements and features regularly to deliver value continuously. Team members work closely with partner teams who depend on timely delivery of accurate data.