We're looking for a Software Development Engineer to join the Production Engineering team that keeps critical developer tools running for hundreds of thousands of developers every month. Our job is to make sure those experiences are seamless, reliable, and secure.
You'll own the operational health of production services end-to-end: from building real-time observability systems that detect problems before customers notice, to building agentic solutions that resolve issues without human intervention. You'll own critical services handling tens of millions of requests per week, respond to security incidents that affect real users, burn down complex service debt, and build intelligent tooling that makes operational excellence sustainable at scale.
What makes this team different:
We're building AI-powered agents that detect, diagnose, and fix production issues autonomously. The problems you solve today become the automation that prevents tomorrow's incidents. You'll make high-judgment calls daily: when to fix, when to deprecate, when to shut down. Every service has a deliberate future, and you help decide what that looks like. You'll work across the full stack: user-facing clients, telemetry services, CI/CD pipelines, security response, and service lifecycle management.
If you love the craft of keeping complex systems running beautifully, and you'd rather build the system that prevents the 2 AM page than be the one answering it, this is your team.