
Senior Site Reliability Engineer
Posted Sep 4

Posted Sep 4
This is a fully remote position, open to applicants in United States, +3 more countries.
• Conduct comprehensive audits of infrastructure, deployment pipelines, monitoring systems, alerting mechanisms, and incident response processes from start to finish.
• Detect vulnerabilities, scaling challenges, and areas of excessive expenditure.
• Compose detailed audit reports that clarify findings, their implications, and suggested corrective actions.
• Execute enhancements in code, configuration, monitoring, alerting, and deployment practices.
• Advance on-call protocols, incident management, service level objectives (SLOs), and postmortem analysis practices.
• Provide guidance on infrastructure architecture choices to support scalability.
• Collaborate closely with engineers through paired programming, code reviews, and smooth transitions.
• A minimum of eight years of practical experience in site reliability, infrastructure, or production engineering roles.
• Extensive knowledge of AWS cloud infrastructure, encompassing networking, IAM, VPCs, and failure scenarios.
• Proven experience in developing monitoring, alerting, and observability solutions from inception using tools like Datadog, Grafana, Prometheus, or similar.
• Skilled in constructing and maintaining CI/CD and deployment workflows.
• Proficient in writing production-level code and configuration and willing to take ownership of it.
• Regular utilization of AI in workflows and familiarity with AI coding tools.
• Excellent written communication abilities to create actionable audit reports.
• Fully remote position, open to applicants globally.
• Reliable internet connection and a quiet work environment are essential.
• Experience in on-call and incident response leadership is a plus.
• Background in cloud cost optimization is advantageous.
• Experience in security-related work is considered a bonus.
• Part-time consulting contract.
• Ongoing engagement.
• Potential to transition into a full-time position based on compatibility and performance.
• Flexible working hours depending on the project scope.
• Fully remote opportunity, accessible from anywhere worldwide.
• Chance to influence decisions that ensure system stability as the company grows.
• Quick implementation of changes.
• Equal opportunity employer.
FourEnergy GmbH
ICF
Mastercam
C&S Informática
Get handpicked remote jobs straight to your inbox weekly.