At Amazon Selection and Catalog Systems (ASCS), we're solving one of the hardest AI challenges in the world: establishing product identity and relationships at unprecedented scale — billions of products, petabytes of multi-modal data, millions of sellers, dozens of languages, and infinite product diversity. Using Generative AI, Visual Language Models (VLMs), and multi-modal reasoning, we determine what makes each product unique and how products relate to one another across Amazon's catalog.
The Catalog Diagnostics & Analytics (CDA) team is at the forefront of this work. We own one of Amazon's largest data lakes (hundreds of petabytes), process catalog data at 100K+ TPS through a cost-effective micro-batch streaming architecture, and serve 2,000+ internal customers who rely on our datasets for daily business decisions. We're now building the next frontier: a production-grade, agentic AI platform that automates catalog diagnostics — delivering full data provenance, change audit trails, and explainability into every catalog change at Amazon scale.
We are looking for a Senior Software Development Engineer who can operate at the intersection of ambitious technical scope, organizational leadership, and deep expertise in GenAI and Agentic systems. If you thrive in fast-paced environments and want to push boundaries at the frontier of AI and distributed systems, we'd love to hear from you.
Key job responsibilities
- **Lead technical architecture** for CDA's agentic diagnostics platform — designing systems at the intersection of GenAI, transactional engines operating at 100K+ TPS, and large-scale information retrieval
- **Set technical direction** for how we instrument, evaluate, and close the accuracy and reliability loop for production-grade agentic AI solutions — translating ambiguous business challenges into tractable engineering frameworks
- **Drive AI-native development practices** across the team: prompt engineering, LLM integration, agentic architectures, retrieval-augmented generation (RAG), and automated evaluation and feedback loop frameworks
- **Influence without authority** — collaborate with Applied Scientists, Data Engineers, and partner teams across Amazon to define standards and patterns that others adopt
- **Own end-to-end delivery** — from design through coding, testing, deployment, and operational excellence — with full accountability for production reliability and customer impact
- **Raise the technical bar** through rigorous design and code reviews, mentorship of engineers, and establishing standards for no-code testing, deployment lifecycle tooling, and automated skill creation
- **Partner with senior leadership** to shape the roadmap, identify strategic investments, and represent CDA's technical perspective in cross-organizational forums
About the team
The Catalog Diagnostics and Analytics (CDA) team owns one of the largest data lakes within Amazon Retail, managing hundreds of petabytes through processing pipelines that push the boundaries of Amazon's technology infrastructure. We maintain foundational catalog datasets relied upon by 2000+ internal customers for day-to-day business decisions.
Our team pioneers data management technologies and architecture patterns across the organization, setting standards for others to adopt. We operate daily and hourly batch datasets alongside services processing catalog data at 100K TPS through a cost-effective micro-batch streaming architecture.
CDA also owns Catalog Diagnostics services that provide insight into catalog changes, enabling Amazon stakeholders to quickly identify and resolve catalog issues. We are building an agentic solution to automate and simplify catalog diagnostics—delivering transparency through data provenance and change audit trails that explain what changed, when, by whom, and why.
We are tackling foundational challenges in making agentic solutions production-ready, including full-cycle feedback loops, agent accuracy and evaluation frameworks, automated skill creation, and no-code testing and deployment lifecycle tooling.