Shape the future of AI infrastructure! Join RRL and develop system-level test solutions for our innovative ML acceleration hardware, deployed across our global server fleet.
We're seeking highly skilled and motivated Product Test Engineers to join our RRL test & repair operation. In this role, you'll be at the forefront of validating and ensuring that test infrastructure is deployed and operating at required scale. You'll be responsible for designing and implementing comprehensive system-level test strategies that cover the full spectrum of our ML acceleration products, from individual components to fully integrated systems. This position requires a unique blend of hardware knowledge, software expertise, and systems thinking, as you'll be working at the intersection of custom silicon, complex firmware, and high-performance ML workloads.
You'll collaborate closely with cross-functional teams including hardware designers, software engineers, and operations specialists to develop robust test solutions that can scale to meet the demands of RRL's global infrastructure. Your work will be crucial in identifying and resolving integration issues, optimizing system performance, and ultimately ensuring that our ML acceleration products meet the highest standards of reliability and efficiency in real-world data center environments. If you're passionate about pushing the boundaries of ML hardware testing and have a knack for solving complex system-level challenges, we want you on our team.
Key job responsibilities
Provide frontline technical support to manufacturing and operations teams by troubleshooting complex tester, hardware, and system-level failures impacting production throughput
Own daily ticket queue management, prioritizing operationally critical issues to ensure timely resolution and minimal business disruption
Perform deep technical troubleshooting of production test systems, identifying root causes across hardware, software, networking, and infrastructure dependencies
Diagnose, repair, and restore test equipment and validation environments to maintain operational readiness and maximize uptime
Partner closely with Operations, Engineering, and cross-functional stakeholders to drive issue resolution and improve overall test process stability
Execute structured root cause analysis and implement corrective actions to prevent recurring equipment or process failures
Support system bring-up, tester validation, and configuration activities for new and existing production environments
Develop and maintain troubleshooting documentation, standard work, and knowledge-sharing mechanisms to improve team efficiency and issue resolution consistency
Identify opportunities to simplify support workflows, improve response times, and enhance operational effectiveness through process improvements and automation where applicable
Monitor tester health, system performance, and operational trends to proactively identify risks and escalate issues before business impact occurs