AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we're the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain, and we're looking for talented people who want to help.
You'll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You'll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you'll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.
DC BRIDGE builds the software platforms that support Amazon's data center engineering operations worldwide. Our systems directly impact the reliability and efficiency of the physical infrastructure that powers AWS and Amazon's global services.
Every AWS data center depends on its electrical and mechanical infrastructure, the power and cooling systems that keep it within operational limits. The FleetWatch team within DC BRIDGE builds the platform that monitors all of it, ingesting trillions of metrics a month. We expose metric data in real time to the operator tools that depend on it, evaluate alarms in-stream to determine actions when metrics breach thresholds, and persist telemetry to a datalake for analytics and long-term monitoring across the fleet.
As an SDE2 you'll own features from design through code, test, deployment, and operation, with autonomy. Our customers are the controls engineers and operations teams who run Amazon's data centers. You'll work directly with controls engineers to improve the alarm configuration experience on the one hand, and with operations teams to refine the alarm output experience on the other. The problems span stream processing at fleet scale, alarm quality and noise reduction, and extending monitoring to new regions and new generations of data center infrastructure. Telemetry volume grows every month, so efficiency is a feature: we simplify systems and drive down the cost of processing each metric even as the fleet expands.
Tech stack: Java (primary for service code), with Kotlin (stream processing), TypeScript (CDK infrastructure), and Python (Lambda functions, data processing, monitoring). AWS services include Kinesis, Lambda, DynamoDB, SQS, SNS, S3, Fargate, CloudWatch, Athena, and Step Functions.
Key job responsibilities
- Build the services behind FleetWatch's real-time monitoring pipeline: stream processing at fleet scale, alarm evaluation that determines actions when metrics breach thresholds, and the noise-reduction logic that keeps alarms actionable. You'll own features end-to-end and make the implementation decisions yourself
- Contribute to the services that route alarms and tickets to data center operations teams with the context they need for faster resolution
- Build and operate the offline path: the datalake that persists telemetry for analytics and long-term monitoring, and the APIs that expose metric data to operator tools
- Drive efficiency at scale. Simplify systems, deprecate what the fleet no longer needs, and reduce the cost of processing each metric as volume grows
- Define infrastructure as code using AWS CDK and own the operational health of what you deploy
- Author designs for your own work and provide meaningful feedback in design reviews
- Mentor earlier-career engineers through code review and pairing
A day in the life
Your morning might start with standup, sharing progress on a new alarm evaluation feature, then a code review where you give thoughtful feedback to a teammate. You spend a focused block writing stream-processing code or CDK to extend the metrics pipeline, then debug a production issue that surfaced overnight, finding the root cause and writing a permanent fix. After lunch you join a design review, asking the questions that surface unstated assumptions. You wrap up walking through alarm feedback with a data center operations team, turning what you hear into the next noise-reduction improvement.
About the team
FleetWatch is led by Caleb Veth and operates with high ownership over our services end-to-end, from design through production operation. The Seattle-based team is a small group of SDEs who run everything we build. Our mission is to give the people who run Amazon's data centers a clear, trustworthy picture of the fleet's electrical and mechanical health: every alarm should mean something, and every metric should be worth what it costs to process. As an SDE2 you'll be a core autonomous contributor: owning features, raising the bar in reviews, and helping more junior engineers grow.
*Why AWS*
Amazon Web Services (AWS) is the world's most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating, and that's why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.
*Diverse Experiences*
Amazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn't followed a traditional path, or includes alternative experiences, don't let it stop you from applying.
*Work/Life Balance*
We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there's nothing we can't achieve in the cloud.
*Inclusive Team Culture*
Here at AWS, it's in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness.
*Mentorship and Career Growth*
We're continuously raising our performance bar as we strive to become Earth's Best Employer. That's why you'll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.