
SRE, DevOps
Posted Jul 15

Posted Jul 15
This is a fully remote position, open to applicants in Brazil.
• Multi-region AWS architecture & High Availability - Design and implement the multi-cluster and multi-region strategy (e.g., São Paulo and Canada, following corporate standards), including a vertical approach — partitioning the application into two or three clusters segregated by domain/responsibility. Define replication strategies, failover, traffic routing, and RTO/RPO targets, ensuring scalability and high availability for sustainable growth.
• Migration and enabling of services across regions - Lead the movement and activation of core platform components across multiple regions, including: multi-region DynamoDB (global tables), multi-region Amazon Cognito (in collaboration with the Platform and Authentication teams while the current identity solution is in use), and the migration of the API Gateway to the new region. Map and review deprecated Lambda functions (between the platform team and verticals) to facilitate regional movement.
• Infrastructure Automation & CI/CD - Develop tools and automations that reduce operational toil in provisioning, deployment, and operation. Maintain Infrastructure as Code with Terraform and configuration with Ansible, promoting standardization and reproducibility of environments. Evaluate and adopt internal standardized provisioning tools/platforms when applicable.
• Mobile Pipelines (iOS / Android) - Construct and maintain CI/CD pipelines for mobile applications, covering build, code signing, automated testing, and distribution (App Store Connect / Google Play). Automate certificate management, provisioning profiles, and release flows, integrating tools like Fastlane into existing pipelines.
• Observability & Monitoring - Configure and maintain monitoring, dashboards, and alerts with end-to-end visibility. Define SLIs/SLOs and instrument services with Dynatrace, Prometheus, and Grafana, creating integrations that standardize telemetry across systems.
• FinOps - Promote FinOps practices (spending visibility, rightsizing, savings plans, anomaly alerts) and continuous cost optimization in the cloud, considering the impact of multi-region scenarios on consumption.
• Reliability & Incident Response - Diagnose and resolve production incidents, minimizing impact and recovery time. Conduct blameless post-mortems and drive continuous improvement from incidents.
• Security & Collaboration - Apply best security practices in production (secret management, least privilege/IAM, hardening). Work closely with development teams to enhance the reliability and efficiency of applications.
• Solid experience with AWS: Kubernetes (K8s/EKS), API Gateway, S3, containers, AWS Lambda, Load Balancers, among others.
• Proven experience in multi-region and multi-cluster architectures: vertical partitioning of applications, data replication, failover, DR, and global routing.
• Practical experience with global tables in DynamoDB and multi-region replication strategies.
• Experience with Amazon Cognito and/or other identity solutions in distributed scenarios.
• Experience in migrating services across regions (API Gateway, Lambda, and related), including analysis and refactoring of deprecated functions.
• Experience with CI/CD pipelines for mobile (iOS and Android): build, code signing, testing, and distribution to stores.
• Proficiency in Infrastructure as Code with Terraform and automation with Ansible.
• Experience with Jenkins (configuration and maintenance of CI/CD pipelines).
• Practical experience with Dynatrace (dashboards, alerts, and integrations) and with Prometheus and Grafana.
• Solid understanding of scalability and high availability.
• Experience in FinOps and cloud cost optimization.
• Knowledge of scripting languages, preferably Python and Bash.
• Understanding of security practices in production environments.
• Health and dental insurance;
• Meal and food vouchers;
• Childcare assistance;
• Extended parental leave;
• Partnership with gyms and health and wellness professionals via Wellhub (Gympass) TotalPass;
• Profit Sharing (PLR);
• Life Insurance;
• Continuous learning platform (CI&T University);
• Discount club;
• Free online platform dedicated to promoting physical, mental health, and well-being;
• Responsible pregnancy and parenting courses;
• Partnership with online course platforms;
• Language learning platform;
• And many more.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.