As an Applied Scientist on the Science SW team, you will collaborate closely with other scientists and engineers to bring Reinforcement Learning (RL) research to production. This role combines the scientific application of ML, and specifically RL
and sequential decision making, with software development engineering and a strong product focus. It will be your job to design, implement, and deploy novel RL agents, reward models, and control policies in both prototype and production environments, and to prove their impact in high-fidelity simulation before scaling them across the fleet.
Key job responsibilities
• Own the research and development of reinforcement learning and sequential decision making solutions spanning deep RL, policy
optimization, offline/batch RL, contextual bandits, and multi-agent RL for real-time MHE control and building-wide optimization in a production
environment.
• Formulate fulfillment operations problems (throughput optimization, flow, merge, and congestion control) as sequential decision-making
problems, and design multi-objective reward functions that balance competing operational objectives.
• Build and leverage high-fidelity simulation environments for safe offline training, policy validation, and sim-to-real transfer before fleet-scale
deployment.
• Collaborate across multiple science and engineering teams to integrate RL policies into real-time production and control systems.
About the team
Amazon is building next generation software, hardware, and processes that will run our global network of fulfillment centers that move millions of units of inventory, and ensure customers get what they want when promised.
The Science Software team in the One MHS organization unlocks Material Handling Equipment (MHE) innovation through a multiplicity of disciplines within Artificial Intelligence (AI) and applied science, including Computer Vision (CV), Physics-Informed Neural Networks (PINNs), Optimization, Reinforcement Learning, classical Machine Learning, statistical modeling, and sensing-hardware
prototyping. Rooted in first principles aligned experimentation, the team is dedicated to building self-optimizing fulfillment centers, developing the models that drive real-time, building-wide orchestration of MHE. We conduct experiments,
develop models, and apply machine learning (ML) at scale to optimize throughput, flow, merge, and congestion control, and to improve operational performance across the fulfillment network.