
Cloud Systems Engineer – Site Reliability
Posted Sep 18

Posted Sep 18
This is a fully remote position, open to applicants in Pennsylvania.
• Take ownership of and consistently enhance the utilization of Datadog across various metrics, logs, traces, dashboards, monitors, alerts, and service-level views.
• Design, implement, and sustain critical systems that are high-availability and high-throughput, catering to data and compute-intensive requirements for a 24×7 SaaS platform.
• Collaborate with service owners to establish and enhance reliability through SLIs, SLOs, error budgets, actionable alerting, and operational-readiness practices.
• Engage in incident management as either the incident commander or technical responder, overseeing triage, restoration, escalation, communication, documentation, root cause analysis, and corrective actions.
• Analyze issues across infrastructure and application layers utilizing metrics, logs, distributed traces, and code-level context.
• Enhance deployment safety and service resilience via automated validation, recovery and rollback capabilities, reliability testing, and failure-mode analysis.
• Ensure that newly implemented systems are supportable and maintainable by both development and operations teams.
• Offer advanced technical guidance and support to technology teams.
• Provide on-call support for production and fulfill additional duties as necessary.
• Guarantee that systems and operational activities adhere to organizational security, HIPAA, and operating policies.
• Minimize repetitive operational toil using Bash, PowerShell, Python, or Ansible.
• Manage infrastructure as code with Terraform/OpenTofu and automate configurations using Ansible.
• Bachelor’s degree in Information Systems, Engineering, or equivalent experience.
• Over 5 years of engineering experience in Systems Engineering, Cloud or Platform Engineering, DevOps, Software Engineering, and/or Site Reliability Engineering (SRE).
• Proficient in designing and operating production systems utilizing cloud-based computing, storage, networking, and containerization technologies; Azure and Kubernetes experience is preferred.
• Strong fundamental knowledge of Linux systems and networking, with experience in troubleshooting complex distributed systems in a production environment.
• Expertise in an observability platform; Datadog experience is highly preferred.
• Familiarity with Prometheus, Grafana, New Relic, or similar platforms is advantageous.
• Proficient in scripting and operational automation using tools such as Bash, PowerShell, or Python, along with practices in infrastructure-as-code and configuration management.
• Experience in participating in production on-call rotations, incident response, root cause analysis, and post-incident improvement efforts.
• Experience working in Agile/DevOps environments and managing production services using ITSM practices where applicable.
• Prior experience in software development or investigating application behavior through code, logs, and distributed traces is a plus.
• Employer-sponsored health, dental, vision, life, and disability insurance.
• Retirement plan with company contributions.
• Annual profit-sharing from the company.
• Personal development and training budget.
• Open and collaborative work environment.
• Comprehensive two-week onboarding plan.
• Extensive mentorship program.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.