
DevOps Engineer IV, Operational Resilience, Observability, SRE
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
• The DevOps Engineer IV acts as a senior individual contributor and a technical leader within the team, offering leadership, guidance, and mentorship to fellow engineers.
• You will take charge of fulfilling scope, schedule, and delivery demands, engaging with stakeholders, and promoting enhancements in DevOps processes and practices across the program.
• Establish and manage monitoring and observability tools (AWS CloudWatch, Prometheus, Grafana, Loki log aggregation) to ensure real-time visibility into application health, performance, and infrastructure.
• Develop and maintain "Golden Signals" performance dashboards that track latency, traffic, errors, and saturation.
• Execute the DORA metrics roadmap utilizing Grafana and GitLab analytics to set performance baselines.
• Oversee Tier 2/3 production support adhering to strict SLAs: 1-hour initial response, 4-hour critical resolution, ensuring a 99.9% uptime commitment.
• Create Root Cause Analyses within 3 business days of any severity-1 production outage; keep on-call runbooks and change correlation updated.
• Draft and uphold the BCDR plan, which includes recovery architecture and RTO targets, cross-region replication (RDS, S3), Route 53 routing, and Secrets Manager; organize biannual failover drills.
• Set up centralized alerting and incident management tools (Jira Service Desk/ServiceNow, Microsoft Teams, AWS Chatbot).
• Implement AWS Auto Scaling and Elastic Load Balancing; present sprint performance reports and suggestions for cost optimization.
• Assist in recruitment efforts by evaluating homework assignments and potentially participating in interviews.
• A Bachelor’s degree and over 8 years of relevant experience, or an equivalent amount of additional experience in lieu of a degree.
• Must fulfill federal suitability criteria and pass a background check as a condition of employment.
• More than 5 years of practical experience with AWS, Terraform (or similar Infrastructure as Code), and Git/GitLab in production settings.
• Experience supporting five or more engineering teams from a centralized DevOps/platform function.
• Proven experience in designing and constructing CI/CD pipelines in GitLab and/or Jenkins, incorporating quality and security gates.
• Strong knowledge of containerization with Docker and orchestration on AWS ECS/EKS/Fargate.
• Awareness of DevSecOps practices, including SAST, dependency scanning, and automated vulnerability remediation.
• Excellent communication and documentation abilities; comfortable working in a highly collaborative Agile/SAFe environment.
• Demonstrated experience in SRE, production operations, or incident response for mission-critical, high-availability cloud-native services.
• Experience with BCDR planning and disaster recovery exercises.
• Company-subsidized health, dental, and vision insurance.
• Flexible PTO.
• 401K with employer match.
• Paid parental leave after one year of service.
• Employee Assistance Program.
Quantiphi
NIR-YU
Bet On Talent
Zignaly
Get handpicked remote jobs straight to your inbox weekly.