Austin, TX, US
24 hours ago
Sr. Technical Program Manager, Annapurna Labs, Machine Learning Fleet Operations
In Annapurna Labs we are at the forefront of hardware/software co-design not just in Amazon Web Services (AWS) but across the industry. The Machine Learning Fleet Operations Team is looking for candidates interested in diving deep into our "fleet" of Machine Learning servers deployed around the world.

Do you like solving mysteries? Great! Figuring out what that light switch with no obvious function actually does? Me too! Are you wearing a smartwatch to monitor your sleep and activity over time to optimize your routines? You'll fit right in. Does the word exabyte excite you? Let's get to work.

We are seeking an experienced TPM to help drive solutions in the highly technical machine learning server hardware space. Our team has end to end ownership of some of the most advanced server hardware in the world. We drive technical debug efforts and write truly massive scale autonomous software to monitor, optimize, and remediate machine learning hardware. Come join us!


Key job responsibilities
Member of a team responsible for system remediation, operational excellence, and customer experience on bleeding edge ML products
Utilize data to root cause hardware failures and identify live trends on the most complex systems in AWS
Implement and improve system level testing across the product lifecycle
Interface, communicate, and collaborate across organizations within AWS
Dive deep on issues at the intersection of hardware and software


A day in the life
The MLA Fleet Operations team was formed to maintain an exceptionally high quality bar for our fleet of advanced machine learning server products. We perfect the customer experience by developing scalable software for rapid incident response times and data visualization as well as diving deep into hardware issues as they arise.


About the team
Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we’re building an environment that celebrates knowledge-sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects that help our team members develop your engineering expertise so you feel empowered to take on more complex tasks in the future.

Confirm your E-mail: Send Email