How Remoteville checks and expires listings
Manager, Reliability Operations
Skills
Capacity PlanningChange ManagementIT EscalationIncident CommandIncident ManagementOperationsRoot Cause
What the job involves
The main requirements, responsibilities and hiring steps.
Requirements
- Bachelor’s degree in Computer Science Engineering or related field or equivalent practical experience
- 7+ years in systems operations site reliability or platform engineering
- 2+ years leading teams or major operational functions
- Proven experience managing incidents in a 24/7 production environment
- Strong troubleshooting root cause analysis and operational improvement skills
- Experience with change management practices
- Experience with monitoring and observability platforms
- Experience with incident management and alerting tools
- Experience with Linux systems VMware Ceph and cloud platforms
- Ability to translate complex technical data into clear insights
- Strong communication skills in high pressure situations
Nice to have
- Analytical
- Collaborative
- Process driven
- Detail oriented
- Calm under pressure
Day to day
- Lead reliability operations by improving incident management change management and post incident practices
- Own incident command structure procedures and major incident response in a 24/7 production environment
- Drive reliability strategy through root cause analysis corrective actions observability and cross team collaboration
