Site Reliability Engineer – Evening Shift

Posted Sep 18

This is a fully remote position, open to applicants in United States.

📋 Description

• Operate and maintain production infrastructure services and applications to ensure their availability, reliability, performance, security, and overall operational health.

• Monitor services and applications utilizing SLIs, SLOs, dashboards, alerts, and observability tools.

• Enhance the detection, diagnosis, and resolution of operational issues.

• Collaborate with application teams to establish observability requirements and implement metrics, logs, traces, dashboards, and alerts.

• Manage production incidents and service interruptions, including on-call responses, troubleshooting, service restoration, root-cause analysis, and post-incident corrective actions.

• Execute application and infrastructure releases through deployment pipelines, encompassing staging and production promotion, validation, rollback, and troubleshooting of releases.

• Oversee the operational lifecycle of deployed infrastructure, which includes upgrades, patching, configuration changes, maintenance, and technology refreshes.

• Evaluate and enhance service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing.

• Identify and mitigate reliability risks and operational technical debt by utilizing reliability metrics, incident trends, capacity data, and service health indicators.

• Automate operational tasks using an everything-as-code methodology.

• Work together with platform engineering and application teams on operational requirements and reusable infrastructure components.

• Focus on the continuous improvement of the reliability and operability of the environment.


⛳️ Requirements

• Must be a U.S. Citizen.

• Capable of obtaining and maintaining the necessary Public Trust level clearance.

• Bachelor’s Degree with 8 years of experience, or a High School diploma/equivalent with 12 years of experience.

• Minimum of 7 years of hands-on experience in site reliability engineering, DevOps, or production systems engineering.

• Practical experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or similar Kubernetes-based platforms.

• Solid experience with infrastructure-as-code using Terraform and Ansible/Ansible Tower.

• Familiarity with CI/CD platforms such as GitLab and Jenkins, including reliability gating and deployment automation.

• Proficient in administering Linux and Windows Server environments.

• Experience with enterprise observability tools like Dynatrace, Datadog, Splunk, and Open Telemetry.

• Proven ownership of an SLI/SLO and alerting program, which includes error budgets, alert rationalization, and noise reduction.

• Scripting and automation skills in Python, Bash, PowerShell, or Go.

• Experience working in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).

• Preferred: AWS Solutions Architect, AWS DevOps Engineer, or AWS SysOps certification.

• Preferred: Red Hat Certified Specialist in ROSA or Red Hat Certified System Administrator in OpenShift.

• Preferred: Azure Administrator Associate or GCP Associate Cloud Engineer certification.

• Preferred: Dynatrace Associate or Datadog Log Management Fundamentals certification.

• Preferred: GitLab CI/CD Associate certification or Certified Jenkins Engineer (CJE).

• Preferred: Terraform Associate certification.


🏝️ Benefits

• Overtime eligibility may apply.

• Shift differential may apply.

• Discretionary bonus eligibility may apply.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers