How Remoteville checks and expires listings
Site Reliability Engineering Manager
Skills
DevopsBusiness Continuity PlanningContinuous ImprovementElastic StackEnglishPsychological SafetyReadiness
What the job involves
The main requirements, responsibilities and hiring steps.
Requirements
- 5+ years in IT operations infrastructure engineering or SRE
- 1+ years leading or coordinating a technical team
- Hands-on Azure cloud and infrastructure management experience
- Strong knowledge of incident change and service continuity practices
- Experience with monitoring and observability tools such as Prometheus Grafana Azure Monitor or ELK
- Familiarity with Infrastructure as Code tools such as Terraform or Ansible
- Knowledge of Docker and Kubernetes or AKS
- Ability to define SLOs SLIs and error budgets
- Experience with capacity planning and cloud cost management
- Strong analytical problem-solving and incident management skills
- Excellent communication and stakeholder management skills
- Proficient English language skills
Nice to have
- Greenfield experience
- Scale-up experience
- SRE discipline knowledge
- DevOps mindset
- Automation focus
- Chaos engineering familiarity
- Relevant certifications
Day to day
- Build and lead the SRE team from the ground up, defining structure, hiring support, onboarding, and mentorship
- Set the technical direction and operational roadmap for reliability, scalability, and performance across Azure-based systems
- Own incident management, observability, automation, capacity planning, documentation, and disaster recovery while partnering with engineering and operations teams
