Senior Site Reliability Engineer

Posted Sep 4

This is a fully remote position, open to applicants in United States, +3 more countries.

📋 Description

• Conduct comprehensive audits of infrastructure, deployment pipelines, monitoring systems, alerting mechanisms, and incident response processes from start to finish.

• Detect vulnerabilities, scaling challenges, and areas of excessive expenditure.

• Compose detailed audit reports that clarify findings, their implications, and suggested corrective actions.

• Execute enhancements in code, configuration, monitoring, alerting, and deployment practices.

• Advance on-call protocols, incident management, service level objectives (SLOs), and postmortem analysis practices.

• Provide guidance on infrastructure architecture choices to support scalability.

• Collaborate closely with engineers through paired programming, code reviews, and smooth transitions.


⛳️ Requirements

• A minimum of eight years of practical experience in site reliability, infrastructure, or production engineering roles.

• Extensive knowledge of AWS cloud infrastructure, encompassing networking, IAM, VPCs, and failure scenarios.

• Proven experience in developing monitoring, alerting, and observability solutions from inception using tools like Datadog, Grafana, Prometheus, or similar.

• Skilled in constructing and maintaining CI/CD and deployment workflows.

• Proficient in writing production-level code and configuration and willing to take ownership of it.

• Regular utilization of AI in workflows and familiarity with AI coding tools.

• Excellent written communication abilities to create actionable audit reports.

• Fully remote position, open to applicants globally.

• Reliable internet connection and a quiet work environment are essential.

• Experience in on-call and incident response leadership is a plus.

• Background in cloud cost optimization is advantageous.

• Experience in security-related work is considered a bonus.


🏝️ Benefits

• Part-time consulting contract.

• Ongoing engagement.

• Potential to transition into a full-time position based on compatibility and performance.

• Flexible working hours depending on the project scope.

• Fully remote opportunity, accessible from anywhere worldwide.

• Chance to influence decisions that ensure system stability as the company grows.

• Quick implementation of changes.

• Equal opportunity employer.

People also viewed

FourEnergy GmbH11 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF14 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam18 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática22 hours ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers