
Site Reliability Engineer – Day Shift
Posted Sep 18

Posted Sep 18
This is a fully remote position, open to applicants in United States.
• Operate and sustain production infrastructure services and applications, guaranteeing availability, reliability, performance, security, and operational health.
• Monitor services and applications utilizing SLIs, SLOs, dashboards, alerts, and observability tools; enhance detection, diagnosis, and resolution of operational issues.
• Collaborate with application teams to establish observability requirements and implement metrics, logs, traces, dashboards, and alerts.
• Handle production incidents and service interruptions, including on-call response, troubleshooting, service restoration, root-cause analysis, and post-incident corrective measures.
• Execute application and infrastructure releases through deployment pipelines, encompassing staging and production promotion, validation, rollback, and release troubleshooting.
• Oversee the operational lifecycle of deployed infrastructure, including upgrades, patching, configuration changes, maintenance, and technology refreshes.
• Evaluate and enhance service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing.
• Identify and mitigate reliability risks and operational technical debt utilizing reliability metrics, incident trends, capacity data, and service health indicators.
• Automate operational tasks employing an everything-as-code methodology.
• Collaborate with platform engineering and application teams to pinpoint operational requirements and improve infrastructure building blocks, reliability, and operability.
• Must be a U.S. Citizen with the capability to obtain and maintain the necessary Public Trust level clearance.
• Bachelor's Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience.
• 7+ years of practical experience in site reliability engineering, DevOps, or production systems engineering.
• Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or similar Kubernetes-based platforms.
• Strong infrastructure-as-code expertise with Terraform and Ansible/Ansible Tower.
• Familiarity with CI/CD platforms GitLab and Jenkins, including reliability gating and deployment automation.
• Proficient in Linux and Windows Server administration.
• Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk, and Open Telemetry.
• Proven ownership of an SLI/SLO and alerting program, including error budgets, alert rationalization, and noise reduction.
• Proficiency in scripting/automation using Python, Bash, PowerShell, or Go.
• Experience operating in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).
• Preferred certifications: AWS Solutions Architect, AWS DevOps Engineer, AWS SysOps, Red Hat Certified Specialist in ROSA, Red Hat Certified System Administrator in OpenShift, Azure Administrator Associate, GCP Associate Cloud Engineer, Dynatrace Associate, Datadog Log Management Fundamentals, GitLab CI/CD Associate, Certified Jenkins Engineer (CJE), or Terraform Associate.
• Potential eligibility for overtime.
• Potential eligibility for shift differential.
• Potential eligibility for a discretionary bonus.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.