How Remoteville checks and expires listings
Specialised AI Engineer
Skills
AcceleratorAdapterCUDACuratingModelingOptimizationPersonal Boundaries
What the job involves
The main requirements, responsibilities and hiring steps.
Requirements
- 5+ years building production systems in machine learning distributed systems or high-performance infrastructure
- 4+ years hands-on experience in at least one core AI systems area in large-scale production environments
- Strong hands-on expertise in one core area with working knowledge across others
- Proven ability to design optimise and operate systems at scale balancing latency throughput cost and model quality
- Deep understanding of transformer architectures LLMs and/or multimodal models in production
- Strong proficiency in Python and PyTorch for production-grade ML systems
- Experience with distributed compute and training paradigms including data model parallelism sharding and scheduling
- Experience with GPU or accelerator optimisation and system-level performance tuning
- Experience building or operating production inference or training systems at scale
- Ability to design clean abstractions APIs and reusable systems for other engineers
- Strong engineering fundamentals and maintainable well-tested production-quality code
Nice to have
- Relentless innovation
- Ownership
- Accountability
- Adaptability
- Resilience
- Collaboration
- Transparency
Day to day
- Design and optimise scalable AI platform systems for training post-training evaluation and low-latency inference under strict performance and efficiency constraints
- Build and improve distributed services that support large-scale model workflows including fine-tuning alignment dataset processing and benchmarking
- Investigate performance bottlenecks and develop reusable APIs SDKs and tooling that help other engineers use AI services effectively
