How Remoteville checks and expires listings

Manager, Reliability Operations

Skills
Capacity PlanningChange ManagementIT EscalationIncident CommandIncident ManagementOperationsRoot Cause
Role

What the job involves

The main requirements, responsibilities and hiring steps.

Requirements

  • Bachelor’s degree in Computer Science Engineering or related field or equivalent practical experience
  • 7+ years in systems operations site reliability or platform engineering
  • 2+ years leading teams or major operational functions
  • Proven experience managing incidents in a 24/7 production environment
  • Strong troubleshooting root cause analysis and operational improvement skills
  • Experience with change management practices
  • Experience with monitoring and observability platforms
  • Experience with incident management and alerting tools
  • Experience with Linux systems VMware Ceph and cloud platforms
  • Ability to translate complex technical data into clear insights
  • Strong communication skills in high pressure situations

Nice to have

  • Analytical
  • Collaborative
  • Process driven
  • Detail oriented
  • Calm under pressure

Day to day

  • Lead reliability operations by improving incident management change management and post incident practices
  • Own incident command structure procedures and major incident response in a 24/7 production environment
  • Drive reliability strategy through root cause analysis corrective actions observability and cross team collaboration