Senior Machine Learning Engineer
Apply now!
Candidate data
Senior Machine Learning Engineer
About the Role
As a Senior Member of Technical Staff, Machine Learning, you will independently own critical ML subsystems in production. In this hands-on, high-impact role, you will take ambiguous and complex problems, design practical and scalable solutions, and ship systems that operate reliably at scale under real-world production constraints such as latency, cost, reliability, and safety.
Core Focus & Responsibilities
- End-to-End Ownership: Own ML subsystems across the full lifecycle, including data preparation, training, evaluation, inference, deployment, and continuous iteration.
- Research to Production: Translate research ideas and experimental approaches into robust, scalable systems that perform reliably in production.
- Production Debugging: Diagnose model failures, system bottlenecks, and performance issues using real-world production signals and data.
- Rapid Iteration: Ship improvements quickly, measure their impact using production metrics, learn from the results, and continuously refine the system.
- Cross-Functional Collaboration: Partner closely with research, product, and engineering teams to turn ML capabilities into measurable user impact.
- Technical Leadership & Mentorship: Mentor other ML engineers, provide thoughtful technical reviews, and raise engineering standards through strong technical judgment and leadership by example.
What We’re Looking For
- Production Track Record: Proven experience building, deploying, and maintaining ML systems used by real-world users.
- Deep Model Intuition: Strong understanding of how modern ML models behave, fail, and evolve in real-world production environments.
- Systems Mindset: Ability to write high-quality, production-ready code and design scalable systems rather than isolated scripts or experiments.
- Independent Execution: High degree of ownership, with the ability to work independently, bring structure to ambiguous problems, and drive complex projects through to completion.
- Communication & Growth: Strong communication skills, a fast-learning mindset, and a commitment to continuous improvement through experimentation and iteration.
Tech Stack
- Languages & Frameworks: Python, PyTorch / JAX
- Infrastructure: GPU-based training and inference systems
Expected Outcomes
- Target Achievement: Ensure ML models and systems consistently meet accuracy, latency, reliability, scalability, and efficiency targets in production.
- System Stability: Monitor, diagnose, and resolve complex production issues quickly while minimizing disruption to users.
- Robust Pipelines: Build and maintain scalable, reproducible, and maintainable training, inference, and data pipelines.
- Data-Driven Impact: Drive measurable improvements in ML models and systems based on real-world signals, production metrics, experimentation, and user feedback.
- Engineering Standards: Provide mentorship, technical guidance, and constructive feedback that elevates the overall ML engineering standard across the team.
How We Work & Application Process
We are a small, world-class team with a high talent density and a strong bias toward action. We operate at a rapid pace, make decisions collaboratively, and balance high-quality engineering craftsmanship with rapid learning, experimentation, and shipping.
- Interview Process: If there appears to be a mutual fit, we will schedule 3–4 interviews with members of the technical team, conducted virtually and/or on-site. We value transparency and efficiency and aim to make timely, well-informed decisions throughout the process.
Over 60% of our candidates get invited to an interview with our Clients.
Apply with the form below and we will reach out to you in the next 24h