Site Reliability Engineer – Night Shift

Posted Aug 31

This is a fully remote position, open to applicants in United States.

📋 Description

• Operate and sustain production infrastructure services and applications, ensuring their availability, reliability, performance, security, and overall operational health.

• Monitor services and applications through SLIs, SLOs, dashboards, alerts, and observability tools.

• Collaborate with application teams to define observability requirements and implement metrics, logs, traces, dashboards, and alerts.

• Manage production incidents and service outages, which includes on-call response, troubleshooting, service restoration, root-cause analysis, and implementing post-incident corrective actions.

• Carry out application and infrastructure releases via deployment pipelines, including staging and production promotion, validation, rollback, and release troubleshooting.

• Oversee the operational lifecycle of deployed infrastructure, which encompasses upgrades, patching, configuration changes, maintenance, and technology refreshes.

• Evaluate and enhance service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing.

• Identify and mitigate reliability risks and operational technical debt through the use of reliability metrics, incident trends, capacity data, and service health indicators.

• Automate operational tasks utilizing an everything-as-code methodology.

• Work collaboratively with platform engineering and application teams to identify operational needs and enhance environmental reliability and operability.


⛳️ Requirements

• Must be a U.S. Citizen with the capability to obtain and maintain the necessary Public Trust level clearance.

• A Bachelor’s Degree and 8 years of relevant experience, or a High School diploma/equivalent with 12 years of experience.

• A minimum of 7 years of practical experience in site reliability engineering, DevOps, or production systems engineering.

• Direct experience operating within AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or similar Kubernetes-based platforms.

• Strong expertise in infrastructure-as-code utilizing Terraform and Ansible/Ansible Tower.

• Familiarity with CI/CD platforms such as GitLab and Jenkins, including reliability gating and deployment automation.

• Proficient in administering Linux and Windows Server environments.

• Experience with enterprise observability tools like Dynatrace, Datadog, Splunk, and Open Telemetry.

• Proven track record of managing an SLI/SLO and alerting program, encompassing error budgets, alert rationalization, and noise reduction.

• Proficiency in scripting/automation with Python, Bash, PowerShell, or Go.

• Experience in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).

• Preferred: AWS Solutions Architect, AWS DevOps Engineer, or AWS SysOps certifications.

• Preferred: Red Hat Certified Specialist in ROSA or Red Hat Certified System Administrator in OpenShift.

• Preferred: Azure Administrator Associate or GCP Associate Cloud Engineer certification.

• Preferred: Dynatrace Associate or Datadog Log Management Fundamentals certification.

• Preferred: GitLab CI/CD Associate certification or Certified Jenkins Engineer (CJE).

• Preferred: Terraform Associate certification.


🏝️ Benefits

• Overtime eligibility may be applicable.

• Shift differential may be applicable.

• Discretionary bonuses may be applicable.

People also viewed

knowmad mood1 day ago

Consultor/a DevSecOps – AWS

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
RealTime eClinical Solutions1 day ago

Principal DevOps Architect

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$155k – $195k/year
ApplyView job
Koniag Government Services1 day ago

DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Koniag Government Services1 day ago

Senior AWS DevOps Engineer – AWS, Kubernetes, HCP, CI/CD, Observability, AI-focus

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ASRC Federal1 day ago

Senior DevOps Administrator – Supporting NASA

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Nelios1 day ago

DevOps Engineer, Cloud Infrastructure

GR flagGreece OnlyPart-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers