
Senior Software Engineer, DevOps
Posted Sep 12

Posted Sep 12
This is a fully remote position, open to applicants in Brazil.
• Design, construct, and sustain infrastructure on Google Cloud Platform utilizing Terraform.
• Establish infrastructure patterns and standards.
• Take ownership of and enhance GitHub Actions CI/CD pipelines.
• Identify and mitigate operational toil through scalable automation and tools.
• Set up monitoring, dashboards, and alert systems in Datadog.
• Minimize alert noise across team systems.
• Lead incident response within the on-call rotation and facilitate postmortems and follow-ups.
• Define and monitor Service Level Objectives for core infrastructure and build systems.
• Collaborate with product engineering teams on deployment pipelines, environment challenges, and build troubleshooting.
• Manage preview and staging environments, including data synchronization, masking, and cleanup.
• Enhance developer experience through tools, runbooks, and documentation.
• Author tested and reviewed infrastructure code.
• Lead design reviews and Requests for Comments (RFCs).
• Design systems with a focus on reliability, performance, and security.
• Conduct code reviews and mentor engineers through pairing, knowledge sharing, and documentation.
• Collaborate with the Tech Lead, DevOps, Platform, Data Engineering, and domain product teams.
• Approximately 5+ years of professional experience in DevOps, Site Reliability Engineering (SRE), infrastructure, or backend engineering in production settings.
• Practical experience designing and managing infrastructure in at least one cloud provider, preferably Google Cloud Platform.
• Proven track record of owning and deploying automation, pipelines, or infrastructure that has significantly improved team productivity or reliability.
• Passion for enhancing developer productivity.
• Extensive experience with infrastructure-as-code, particularly Terraform, and constructing CI/CD pipelines, especially using GitHub Actions.
• Strong software engineering skills and a solid understanding of Linux.
• Excellent instincts for monitoring and observability, with proficient debugging skills across logs, traces, and metrics.
• Significant experience with relational databases, including MySQL and PostgreSQL, as well as containerized workloads.
• Strong foundational knowledge in reliability, performance, and security principles, with sound judgment on trade-offs.
• Nice to have: experience in healthcare, digital health, or regulated fields, such as HIPAA, PHI, or SOC 2.
• Nice to have: familiarity with Docker and Kubernetes.
• Nice to have: exposure to incident response, on-call duties, and postmortem practices.
• Nice to have: experience with database migrations or managing multiple environments at scale.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and remote work options.
• Opportunities for professional development and continuous learning.
• Collaborative and inclusive company culture.
• Generous vacation and paid time off policy.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.