
Cloud Engineer – Site Reliability Engineer, Senior
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Brazil.
• Spearhead the technical advancement of infrastructure on GCP, focusing mainly on GKE, Compute Engine, Cloud SQL, Cloud Storage, VPC, Load Balancing, IAM, Cloud DNS, Secret Manager, Pub/Sub, and BigQuery.
• Architect solutions with an emphasis on high availability, scalability, resilience, security, performance, and cost-effectiveness.
• Implement SRE best practices by establishing and monitoring SLIs, SLOs, and Error Budgets.
• Observe reliability metrics such as availability, latency, throughput, error rate, saturation, capacity, MTTR, and MTBF.
• Administer and conduct advanced troubleshooting on Kubernetes/GKE within production environments.
• Develop autoscaling strategies, capacity planning, resource management, and workload isolation.
• Advance the observability framework using tools like New Relic, OpenTelemetry, Prometheus, Grafana, and native cloud services.
• Generate dashboards, alerts, and both technical and business metrics to maintain actionable observability.
• Act as Incident Commander during critical situations and lead blameless postmortems.
• Oversee and enhance infrastructure as code through Terraform.
• Create and sustain reusable modules, configuration patterns, state management, and Terraform pipelines.
• Manage CI/CD pipelines utilizing GitHub Actions and deployment workflows with ArgoCD/GitOps.
• Execute rolling, blue/green, and canary deployment methodologies.
• Develop automation solutions using Go, Python, and Bash.
• Construct Platform Engineering solutions including self-service interfaces, templates, automation, and established best practices.
• Collaborate on cloud networking, security, governance, IAM/RBAC, secrets management, certificates, and encryption.
• Contribute to FinOps efforts for resource optimization, rightsizing, and minimizing waste.
• Design and validate Disaster Recovery plans covering backup, restoration, failover, RPO, and RTO.
• Execute performance and capacity planning to foresee bottlenecks and growth needs.
• A minimum of 5 years of experience managing critical production infrastructure in a public cloud setting.
• Extensive experience with Google Cloud Platform (GCP).
• Practical experience with Kubernetes in production settings, preferably GKE.
• In-depth knowledge of Terraform, including modules, state management, version control, and environmental organization.
• Proven experience in defining and managing SLIs, SLOs, and Error Budgets.
• Advanced troubleshooting capabilities in Linux, networking, DNS, HTTP, TLS, load balancers, and Kubernetes.
• Experience in automating tasks with Go, Python, and/or Bash.
• Familiarity with GitHub Actions, ArgoCD, and GitOps methodologies.
• Experience leading critical incident responses, conducting postmortems, and performing root-cause analyses.
• Ability to make architectural decisions and effectively communicate technical, operational, and financial trade-offs.
• Expertise in GCP is mandatory.
• Experience with AWS and/or Azure is advantageous.
• Preferred qualifications include: experience across multiple cloud providers; familiarity with multi-cloud or hybrid cloud environments; knowledge in multi-region architectures and Disaster Recovery; experience with AWS EKS, RDS/Aurora, and multi-account configurations; knowledge of Azure AKS and Landing Zones; experience with Service Mesh (Istio, Linkerd, or Anthos Service Mesh); API Gateway and API management; Chaos Engineering; FinOps; and Platform Engineering.
• Contract (PJ).
• Remote work model.
• Wellhub.
• Life insurance.
CVS Health
Devoteam
Aspirion
Goodgame Studios
Get handpicked remote jobs straight to your inbox weekly.