
Senior Site Reliability Engineer
Posted Sep 9

Posted Sep 9
This is a fully remote position, open to applicants in Brazil.
• Design, optimize, and sustain dependable and scalable CI/CD pipelines.
• Create and manage Infrastructure as Code scripts for automated infrastructure setup and administration.
• Implement best practices for Infrastructure as Code across various environments.
• Recognize and automate repetitive operational tasks.
• Develop automation for deployment, scaling, maintenance, and operational workflows.
• Make SRE and engineering choices while considering technical debt, system design, stability, reliability, monitoring, observability, and business needs.
• Oversee and ensure the resilience of the CI/CD platform.
• Diagnose intricate platform and pipeline issues alongside senior and staff engineers.
• Prepare post-mortem documentation for both internal and external stakeholders.
• Contribute to coding standards, engineering practices, and non-functional requirements.
• Engage in code and Pull Request evaluations.
• Guide junior and mid-level engineers.
• Serve as a technical reference for CI/CD, Infrastructure as Code, automation, and platform reliability.
• Stay updated with emerging technologies and industry trends.
• Assist in executing and supporting Proofs of Concept.
• Deliver technical solutions that enhance platform resilience, continuous improvement, and strategic squad objectives.
• Must be located in Brazil.
• Proficient in English at a B2 level or higher (Upper-Intermediate).
• 5 or more years of relevant professional experience with a Bachelor’s or Associate’s Degree, or at least 2 years of experience with an Advanced degree (e.g., Masters, MBA, JD, MD).
• Proven experience in implementing, optimizing, and maintaining CI/CD pipelines in production settings.
• Familiarity with Infrastructure as Code and infrastructure automation practices.
• Background in DevOps, Site Reliability Engineering, Platform Engineering, Cloud Engineering, or a similar technical field.
• Ability to develop and uphold automation scripts for infrastructure provisioning, deployment, scaling, and operational tasks.
• Comprehension of SRE principles, including system reliability, availability, resilience, monitoring, observability, and incident management.
• Experience in supporting reliable and scalable platforms and resolving complex issues in distributed environments.
• Understanding of software engineering standards, coding practices, version control, and Pull Request review processes.
• Capability to assess technical trade-offs involving system design, technical debt, stability, reliability, maintainability, and business requirements.
• Experience in identifying chances to automate manual operational workflows and enhance engineering efficiency.
• Knowledge of monitoring, logging, metrics, alerting, and observability practices.
• Proficient in creating technical documentation and post-mortem reports for both technical and non-technical audiences.
• Ability to collaborate with engineers of varying experience levels and provide constructive technical feedback.
• Strong analytical and problem-solving capabilities, with the ability to tackle well-defined and moderately ambiguous technical challenges.
• Capacity to participate in technical discussions and escalate broader or cross-squad decisions to senior and staff engineers when necessary.
• Preferred: Experience in maintaining CI/CD platforms or internal developer platforms utilized by multiple engineering teams.
• Preferred: Experience with critical or mission-critical production systems.
• Preferred: Familiarity with cloud infrastructure, container orchestration, and modern deployment practices.
• Preferred: Knowledge of GitOps, Infrastructure as Code, and automated infrastructure management.
• Preferred: Experience in conducting Proofs of Concept and assessing new technologies for production implementation.
• Preferred: Experience in mentoring junior and mid-level engineers.
• Preferred: Experience with incident response, root cause analysis, and post-mortem practices.
• Preferred: Experience in global or distributed engineering teams.
• Preferred: Relevant certifications in cloud, DevOps, Kubernetes, or Infrastructure as Code.
• Flexible remote work options.
• Opportunity to make a significant impact on a large scale.
• Opportunities for skill enhancement and professional development.
• Support for mentoring and professional growth.
• Equal employment opportunity protections.
FourEnergy GmbH
ICF
Mastercam
C&S Informática
Get handpicked remote jobs straight to your inbox weekly.