How Remoteville checks and expires listings

Site Reliability Engineering Manager

Skills
DevopsBusiness Continuity PlanningContinuous ImprovementElastic StackEnglishPsychological SafetyReadiness
Role

What the job involves

The main requirements, responsibilities and hiring steps.

Requirements

  • 5+ years in IT operations infrastructure engineering or SRE
  • 1+ years leading or coordinating a technical team
  • Hands-on Azure cloud and infrastructure management experience
  • Strong knowledge of incident change and service continuity practices
  • Experience with monitoring and observability tools such as Prometheus Grafana Azure Monitor or ELK
  • Familiarity with Infrastructure as Code tools such as Terraform or Ansible
  • Knowledge of Docker and Kubernetes or AKS
  • Ability to define SLOs SLIs and error budgets
  • Experience with capacity planning and cloud cost management
  • Strong analytical problem-solving and incident management skills
  • Excellent communication and stakeholder management skills
  • Proficient English language skills

Nice to have

  • Greenfield experience
  • Scale-up experience
  • SRE discipline knowledge
  • DevOps mindset
  • Automation focus
  • Chaos engineering familiarity
  • Relevant certifications

Day to day

  • Build and lead the SRE team from the ground up, defining structure, hiring support, onboarding, and mentorship
  • Set the technical direction and operational roadmap for reliability, scalability, and performance across Azure-based systems
  • Own incident management, observability, automation, capacity planning, documentation, and disaster recovery while partnering with engineering and operations teams