Machine Learning Engineer
Skills
Data ScienceDefaultsDistributed TrainingInterconnectsLinuxNetworkingOptimization
What the job involves
The main requirements, responsibilities and hiring steps.
Requirements
- Solid experience getting ML workloads running in production not just notebooks or research code
- Strong understanding of GPUs and how ML workloads use memory throughput and where bottlenecks arise
- Comfort working at the systems and Linux level including drivers and networking when issues arise
- Ability to own ambiguous problems end to end and create processes where none exist
- Experience taking part in a weekly on-call rotation
Nice to have
- Distributed training experience
- Large-scale inference experience
- NVIDIA driver familiarity
- InfiniBand knowledge
- RoCE knowledge
- Cluster networking familiarity
Day to day
- Build the optimization and orchestration layer that intelligently places and tunes ML workloads across a heterogeneous GPU fleet.
- Benchmark GPUs interconnects and driver stacks to measure real-world performance and feed those insights back into platform decisions.
- Investigate performance and reliability issues across workloads hardware and networking then turn findings into scalable defaults playbooks and improvements.
