Machine Learning Engineer

Skills
Data ScienceDefaultsDistributed TrainingInterconnectsLinuxNetworkingOptimization
Role

What the job involves

The main requirements, responsibilities and hiring steps.

Requirements

  • Solid experience getting ML workloads running in production not just notebooks or research code
  • Strong understanding of GPUs and how ML workloads use memory throughput and where bottlenecks arise
  • Comfort working at the systems and Linux level including drivers and networking when issues arise
  • Ability to own ambiguous problems end to end and create processes where none exist
  • Experience taking part in a weekly on-call rotation

Nice to have

  • Distributed training experience
  • Large-scale inference experience
  • NVIDIA driver familiarity
  • InfiniBand knowledge
  • RoCE knowledge
  • Cluster networking familiarity

Day to day

  • Build the optimization and orchestration layer that intelligently places and tunes ML workloads across a heterogeneous GPU fleet.
  • Benchmark GPUs interconnects and driver stacks to measure real-world performance and feed those insights back into platform decisions.
  • Investigate performance and reliability issues across workloads hardware and networking then turn findings into scalable defaults playbooks and improvements.