Senior Site Reliability Engineer

Posted Sep 9

This is a fully remote position, open to applicants in Brazil.

📋 Description

• Design, optimize, and sustain dependable and scalable CI/CD pipelines.

• Create and manage Infrastructure as Code scripts for automated infrastructure setup and administration.

• Implement best practices for Infrastructure as Code across various environments.

• Recognize and automate repetitive operational tasks.

• Develop automation for deployment, scaling, maintenance, and operational workflows.

• Make SRE and engineering choices while considering technical debt, system design, stability, reliability, monitoring, observability, and business needs.

• Oversee and ensure the resilience of the CI/CD platform.

• Diagnose intricate platform and pipeline issues alongside senior and staff engineers.

• Prepare post-mortem documentation for both internal and external stakeholders.

• Contribute to coding standards, engineering practices, and non-functional requirements.

• Engage in code and Pull Request evaluations.

• Guide junior and mid-level engineers.

• Serve as a technical reference for CI/CD, Infrastructure as Code, automation, and platform reliability.

• Stay updated with emerging technologies and industry trends.

• Assist in executing and supporting Proofs of Concept.

• Deliver technical solutions that enhance platform resilience, continuous improvement, and strategic squad objectives.


⛳️ Requirements

• Must be located in Brazil.

• Proficient in English at a B2 level or higher (Upper-Intermediate).

• 5 or more years of relevant professional experience with a Bachelor’s or Associate’s Degree, or at least 2 years of experience with an Advanced degree (e.g., Masters, MBA, JD, MD).

• Proven experience in implementing, optimizing, and maintaining CI/CD pipelines in production settings.

• Familiarity with Infrastructure as Code and infrastructure automation practices.

• Background in DevOps, Site Reliability Engineering, Platform Engineering, Cloud Engineering, or a similar technical field.

• Ability to develop and uphold automation scripts for infrastructure provisioning, deployment, scaling, and operational tasks.

• Comprehension of SRE principles, including system reliability, availability, resilience, monitoring, observability, and incident management.

• Experience in supporting reliable and scalable platforms and resolving complex issues in distributed environments.

• Understanding of software engineering standards, coding practices, version control, and Pull Request review processes.

• Capability to assess technical trade-offs involving system design, technical debt, stability, reliability, maintainability, and business requirements.

• Experience in identifying chances to automate manual operational workflows and enhance engineering efficiency.

• Knowledge of monitoring, logging, metrics, alerting, and observability practices.

• Proficient in creating technical documentation and post-mortem reports for both technical and non-technical audiences.

• Ability to collaborate with engineers of varying experience levels and provide constructive technical feedback.

• Strong analytical and problem-solving capabilities, with the ability to tackle well-defined and moderately ambiguous technical challenges.

• Capacity to participate in technical discussions and escalate broader or cross-squad decisions to senior and staff engineers when necessary.

• Preferred: Experience in maintaining CI/CD platforms or internal developer platforms utilized by multiple engineering teams.

• Preferred: Experience with critical or mission-critical production systems.

• Preferred: Familiarity with cloud infrastructure, container orchestration, and modern deployment practices.

• Preferred: Knowledge of GitOps, Infrastructure as Code, and automated infrastructure management.

• Preferred: Experience in conducting Proofs of Concept and assessing new technologies for production implementation.

• Preferred: Experience in mentoring junior and mid-level engineers.

• Preferred: Experience with incident response, root cause analysis, and post-mortem practices.

• Preferred: Experience in global or distributed engineering teams.

• Preferred: Relevant certifications in cloud, DevOps, Kubernetes, or Infrastructure as Code.


🏝️ Benefits

• Flexible remote work options.

• Opportunity to make a significant impact on a large scale.

• Opportunities for skill enhancement and professional development.

• Support for mentoring and professional growth.

• Equal employment opportunity protections.

People also viewed

FourEnergy GmbH11 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF13 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam17 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática22 hours ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers