
Site Reliability Engineer – Night Shift
Posted Sep 18

Posted Sep 18
This is a fully remote position, open to applicants in United States.
• Operate and manage production infrastructure services and applications to ensure their availability, reliability, performance, security, and operational health.
• Monitor services and applications through SLIs, SLOs, dashboards, alerts, and observability tools.
• Enhance the detection, diagnosis, and resolution of operational issues.
• Collaborate with application teams to establish observability requirements and implement metrics, logs, traces, dashboards, and alerts.
• Oversee production incidents and service disruptions, including on-call response, troubleshooting, service restoration, root-cause analysis, and post-incident corrective actions.
• Execute application and infrastructure releases via deployment pipelines, encompassing staging and production promotion, validation, rollback, and release troubleshooting.
• Manage the operational lifecycle of infrastructure, including upgrades, patching, configuration changes, maintenance, and technology refreshes.
• Evaluate and enhance service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing.
• Identify and mitigate reliability risks and operational technical debt by utilizing reliability metrics, incident trends, capacity data, and service health indicators.
• Automate operational tasks using an everything-as-code methodology.
• Work alongside platform engineering and application teams to pinpoint operational requirements and enhance environment reliability and operability.
• Must be a U.S. Citizen with the capability to obtain and maintain the necessary Public Trust level clearance.
• Bachelor's Degree with 8 years of experience, or a High School diploma/equivalent with 12 years of experience.
• 7+ years of hands-on experience in site reliability engineering, DevOps, or production systems engineering.
• Practical experience in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or similar Kubernetes-based platforms.
• Strong experience with infrastructure-as-code using Terraform and Ansible/Ansible Tower.
• Familiarity with CI/CD platforms such as GitLab and Jenkins, including reliability gating and deployment automation.
• Proficient in Linux and Windows Server administration.
• Experience with enterprise observability tools like Dynatrace, Datadog, Splunk, and Open Telemetry.
• Proven ownership of an SLI/SLO and alerting program, covering error budgets, alert rationalization, and noise reduction.
• Proficient in scripting/automation using Python, Bash, PowerShell, or Go.
• Experience working in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).
• Preferred: AWS Solutions Architect, AWS DevOps Engineer, or AWS SysOps certification.
• Preferred: Red Hat Certified Specialist in ROSA or Red Hat Certified System Administrator in OpenShift.
• Preferred: Azure Administrator Associate or GCP Associate Cloud Engineer certification.
• Preferred: Dynatrace Associate or Datadog Log Management Fundamentals certification.
• Preferred: GitLab CI/CD Associate certification or Certified Jenkins Engineer (CJE).
• Preferred: Terraform Associate certification.
• Eligible for overtime.
• Shift differential may be available.
• Discretionary bonus may be available.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.