
Senior Site Reliability Engineer
Posted Jul 19

Posted Jul 19
This is a fully remote position, open to applicants in Brazil.
• Design, develop, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
• Create and manage AI tooling infrastructure.
• Enhance and sustain CI/CD pipelines.
• Execute progressive delivery strategies such as blue/green and canary deployments.
• Advance Infrastructure as Code using Terraform, Helm, and Argo CD.
• Manage and optimize streaming and analytics infrastructure, including Kafka, Flink, and ClickHouse.
• Integrate automated testing into the CI/CD lifecycle.
• Enhance system observability.
• Oversee incident response and conduct postmortems.
• 6+ years of experience in SRE, DevOps, or Infrastructure roles.
• Practical experience in integrating AI/LLM tooling into engineering or operational workflows.
• Demonstrated success in building CI/CD pipelines (using GitHub Actions, Jenkins, GitLab CI, or similar tools).
• Strong understanding of Kubernetes internals and managed services such as EKS, GKE, or AKS.
• Proficient in Infrastructure as Code (Terraform, Helm, Pulumi) and GitOps methodologies.
• Skilled in programming languages like Python, Bash, or Go.
• Familiarity with observability tools (Prometheus, Grafana, Datadog, OpenTelemetry).
• Experience in production environments with Kafka, Flink, and ClickHouse.
• Excellent communication skills and ability to collaborate across teams.
• Competitive salary.
• Stock options.
• Health benefits.
• Unlimited PTO.
• Parental leave.
• Tuition reimbursements.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.