
Cloud Operations Analyst, SRE – Mid-level
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Brazil.
• Ensure the platform supporting Open Solutions remains operational, observable, and reliable.
• Design and enhance monitoring systems for production environments, emphasizing the anticipation of failures.
• Establish and implement metrics that accurately reflect user experiences.
• Minimize alert noise and reduce false positive occurrences.
• Maintain dashboards that facilitate technical decision-making processes.
• Identify and eliminate repetitive manual tasks related to provisioning, configuration, and operational procedures.
• Advance infrastructure as code using Terraform and Helm as the authoritative source of truth.
• Create scripts and automations in Python and Shell for recurring tasks.
• Transform undocumented processes into actionable runbooks.
• Conduct triage, mitigation, and resolution of incidents in production.
• Document and support post-mortems, concentrating on root causes and preventive measures.
• Collaborate with the engineering team to sustain environments.
• Monitor infrastructure usage and costs, offering insights for technical decision-making.
• Collaborate with cloud and software architects and developers, sharing accountability for production operations.
• Proven experience operating Kubernetes in production settings: including pod debugging, resource consumption assessment, and understanding failure modalities.
• Familiarity with cloud environments in critical operations, preferably AWS.
• Experience in writing and maintaining infrastructure as code with Terraform and Helm, including reviewing colleagues' modifications.
• Capability to build monitoring solutions with Prometheus and Grafana, creating metrics and alerting protocols.
• Proficient in autonomous troubleshooting on Linux, covering services, processes, networking, disk usage, and log analysis.
• Experience in automating tasks with Python or Shell, with adoption by fellow team members.
• Involvement in production incident response, demonstrating sound judgment on resolution versus escalation.
• Sufficient networking knowledge to investigate issues related to segmentation, DNS, TLS, VPNs, and debugging tools.
• Experience in regulated industries, such as finance or insurance.
• Familiarity with APIs, REST, and API gateways.
• Proficiency in log management at scale (Elasticsearch, OpenSearch, Loki, or similar technologies).
• Experience in constructing and maintaining CI/CD pipelines.
• Practice in defining SLIs and SLOs, utilizing error budgets to inform decision-making.
• Knowledge of configuration as code practices (Ansible or similar tools).
• Proficient in English for the comprehension of technical documentation.
• Meal/Food allowance (Flash benefits card).
• Health insurance.
• Dental plan.
• Life insurance.
• PPR (performance bonus/profit sharing).
• TotalPass.
• Childcare allowance.
• Well-Being program (supporting physical and mental health).
• Corporate University (our #SensediaAcademy), offering various development tracks.
• Cultural and educational partnerships providing exclusive discounts.
• Extended maternity and paternity leave.
• Flexible work model #WorkWhereYouBelong.
• Opportunities are also available for individuals with disabilities (PWD).
Global University Systems
GuidePoint Security
Airhosted
Cprime, Inc
Get handpicked remote jobs straight to your inbox weekly.