Remotery

DevOps Engineer IV, Operational Resilience, Observability, SRE

Posted 6 days ago

This is a fully remote position, open to applicants in United States.

📋 Description

• The DevOps Engineer IV acts as a senior individual contributor and a technical leader within the team, offering leadership, guidance, and mentorship to fellow engineers.

• You will take charge of fulfilling scope, schedule, and delivery demands, engaging with stakeholders, and promoting enhancements in DevOps processes and practices across the program.

• Establish and manage monitoring and observability tools (AWS CloudWatch, Prometheus, Grafana, Loki log aggregation) to ensure real-time visibility into application health, performance, and infrastructure.

• Develop and maintain "Golden Signals" performance dashboards that track latency, traffic, errors, and saturation.

• Execute the DORA metrics roadmap utilizing Grafana and GitLab analytics to set performance baselines.

• Oversee Tier 2/3 production support adhering to strict SLAs: 1-hour initial response, 4-hour critical resolution, ensuring a 99.9% uptime commitment.

• Create Root Cause Analyses within 3 business days of any severity-1 production outage; keep on-call runbooks and change correlation updated.

• Draft and uphold the BCDR plan, which includes recovery architecture and RTO targets, cross-region replication (RDS, S3), Route 53 routing, and Secrets Manager; organize biannual failover drills.

• Set up centralized alerting and incident management tools (Jira Service Desk/ServiceNow, Microsoft Teams, AWS Chatbot).

• Implement AWS Auto Scaling and Elastic Load Balancing; present sprint performance reports and suggestions for cost optimization.

• Assist in recruitment efforts by evaluating homework assignments and potentially participating in interviews.


⛳️ Requirements

• A Bachelor’s degree and over 8 years of relevant experience, or an equivalent amount of additional experience in lieu of a degree.

• Must fulfill federal suitability criteria and pass a background check as a condition of employment.

• More than 5 years of practical experience with AWS, Terraform (or similar Infrastructure as Code), and Git/GitLab in production settings.

• Experience supporting five or more engineering teams from a centralized DevOps/platform function.

• Proven experience in designing and constructing CI/CD pipelines in GitLab and/or Jenkins, incorporating quality and security gates.

• Strong knowledge of containerization with Docker and orchestration on AWS ECS/EKS/Fargate.

• Awareness of DevSecOps practices, including SAST, dependency scanning, and automated vulnerability remediation.

• Excellent communication and documentation abilities; comfortable working in a highly collaborative Agile/SAFe environment.

• Demonstrated experience in SRE, production operations, or incident response for mission-critical, high-availability cloud-native services.

• Experience with BCDR planning and disaster recovery exercises.


🏝️ Benefits

• Company-subsidized health, dental, and vision insurance.

• Flexible PTO.

• 401K with employer match.

• Paid parental leave after one year of service.

• Employee Assistance Program.

People also viewed

Quantiphi9 hours ago

DevOps Engineer, Observability

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
NIR-YU9 hours ago

Professional Cloud DevOps Engineer – Google Cloud, Certificado

Latin AmericaFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Bet On Talent9 hours ago

Senior DevSecOps Engineer

EuropeFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Zignaly9 hours ago

Infrastructure / Systems Operations Engineer

AE flagUnited Arab Emirates (UAE) OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$40k – $50k/year
ApplyView job
knowmad mood9 hours ago

Senior Business Consultant – DevOps, Atlassian

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Astronomer9 hours ago

Customer Reliability Engineer, Airflow

US flagCalifornia, +8 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$125k – $130k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers