Cloud Systems Engineer – Site Reliability

Posted Sep 18

This is a fully remote position, open to applicants in Pennsylvania.

📋 Description

• Take ownership of and consistently enhance the utilization of Datadog across various metrics, logs, traces, dashboards, monitors, alerts, and service-level views.

• Design, implement, and sustain critical systems that are high-availability and high-throughput, catering to data and compute-intensive requirements for a 24×7 SaaS platform.

• Collaborate with service owners to establish and enhance reliability through SLIs, SLOs, error budgets, actionable alerting, and operational-readiness practices.

• Engage in incident management as either the incident commander or technical responder, overseeing triage, restoration, escalation, communication, documentation, root cause analysis, and corrective actions.

• Analyze issues across infrastructure and application layers utilizing metrics, logs, distributed traces, and code-level context.

• Enhance deployment safety and service resilience via automated validation, recovery and rollback capabilities, reliability testing, and failure-mode analysis.

• Ensure that newly implemented systems are supportable and maintainable by both development and operations teams.

• Offer advanced technical guidance and support to technology teams.

• Provide on-call support for production and fulfill additional duties as necessary.

• Guarantee that systems and operational activities adhere to organizational security, HIPAA, and operating policies.

• Minimize repetitive operational toil using Bash, PowerShell, Python, or Ansible.

• Manage infrastructure as code with Terraform/OpenTofu and automate configurations using Ansible.


⛳️ Requirements

• Bachelor’s degree in Information Systems, Engineering, or equivalent experience.

• Over 5 years of engineering experience in Systems Engineering, Cloud or Platform Engineering, DevOps, Software Engineering, and/or Site Reliability Engineering (SRE).

• Proficient in designing and operating production systems utilizing cloud-based computing, storage, networking, and containerization technologies; Azure and Kubernetes experience is preferred.

• Strong fundamental knowledge of Linux systems and networking, with experience in troubleshooting complex distributed systems in a production environment.

• Expertise in an observability platform; Datadog experience is highly preferred.

• Familiarity with Prometheus, Grafana, New Relic, or similar platforms is advantageous.

• Proficient in scripting and operational automation using tools such as Bash, PowerShell, or Python, along with practices in infrastructure-as-code and configuration management.

• Experience in participating in production on-call rotations, incident response, root cause analysis, and post-incident improvement efforts.

• Experience working in Agile/DevOps environments and managing production services using ITSM practices where applicable.

• Prior experience in software development or investigating application behavior through code, logs, and distributed traces is a plus.


🏝️ Benefits

• Employer-sponsored health, dental, vision, life, and disability insurance.

• Retirement plan with company contributions.

• Annual profit-sharing from the company.

• Personal development and training budget.

• Open and collaborative work environment.

• Comprehensive two-week onboarding plan.

• Extensive mentorship program.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers