
Senior SRE, Cloud Engineer
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in Portugal.
• Manage, maintain, and enhance production cloud infrastructure across AWS, Azure, GCP, Windows, or hybrid settings.
• Develop and uphold monitoring, logging, metrics, tracing, dashboards, and alert systems for production services.
• Enhance observability across infrastructure, applications, databases, queues, and network dependencies.
• Optimize alerts to minimize noise and enhance signal quality.
• Engage in incident response, production troubleshooting, root cause analysis, and post-incident remediation efforts.
• Create automation for infrastructure operations, deployments, health checks, runbooks, and recovery workflows.
• Collaborate with engineering teams to establish SLOs, SLIs, error budgets, and standards for operational readiness.
• Provide support for cloud networking, DNS, TLS, load balancing, IAM, storage, compute, and managed service operations.
• Enhance the reliability, availability, performance, and scalability of cloud-hosted systems.
• Maintain Infrastructure as Code and configuration management practices.
• Produce and update runbooks, operational documentation, and escalation procedures.
• Identify production risks and facilitate remediation through automation, architectural enhancements, and platform standards.
• Practical experience with production cloud infrastructure.
• Proficient in AWS, GCP, Windows, and/or hybrid cloud environments.
• Experience in building and maintaining observability, monitoring, logging, dashboards, and alerting systems.
• Strong troubleshooting abilities across infrastructure, networking, application, and cloud service layers.
• Familiarity with Linux systems, networking fundamentals, DNS, TLS, IAM, load balancers, storage, and compute resources.
• Experience with Infrastructure as Code using Terraform, CloudFormation, Pulumi, or similar tools.
• Scripting and automation proficiency in Bash, Python, Go, or similar programming languages.
• Participation in production incident response and postmortem activities.
• Understanding of SRE practices, including SLOs, SLIs, error budgets, toil reduction, and operational readiness.
• Ability to collaborate closely with engineering teams to enhance reliability and production supportability.
• Familiarity with Kubernetes and cloud-native platforms is a plus.
• Experience with GitOps using Flux or Argo CD is advantageous.
• Knowledge of Jenkins and Ansible is beneficial.
• Experience with service mesh, ingress, and API Gateway is a plus.
• Knowledge of high availability architectures is beneficial.
• Experience with multi-region environments is advantageous.
• Familiarity with disaster recovery solutions is a plus.
• Experience with secrets management using Vault, AWS Secrets Manager, or External Secrets is advantageous.
• Knowledge of security, compliance, and vulnerability management is a plus.
• Experience with on-call operations and runbook design is beneficial.
• Competitive salary and comprehensive benefits.
• Health insurance coverage.
• Language training programs.
• Relocation assistance.
• Flexibility for remote work through the ForeverRemote culture.
• Opportunities for professional growth and development.
• Access to certification programs.
• Mentorship and investment in talent programs.
• Opportunities for internal mobility.
• Internship opportunities available.
• Collaborate on impactful projects for top-tier global clients.
• Inclusive and supportive multicultural workplace.
• Open lines of communication.
• Regular team-building social events.
SPD Technology
Totara
Vesta Software Group
Elfonze Technologies
Get handpicked remote jobs straight to your inbox weekly.