
Senior SRE – Cloud Engineer
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in Spain.
• Operate, maintain, and enhance production cloud infrastructure across AWS, Azure, GCP, Windows, or hybrid environments.
• Construct and uphold monitoring, logging, metrics, tracing, dashboards, and alerting for production services.
• Enhance observability throughout infrastructure, applications, databases, queues, and network dependencies.
• Optimize alerts to minimize noise and enhance actionable incident response.
• Engage in incident response, production troubleshooting, root cause analysis, and post-incident remediation.
• Develop automation for infrastructure operations, deployments, health checks, runbooks, and recovery workflows.
• Collaborate with engineering teams to define SLOs, SLIs, error budgets, and operational readiness standards.
• Assist with cloud networking, DNS, TLS, load balancing, IAM, storage, compute, and managed service operations.
• Boost reliability, availability, performance, and scalability of cloud-hosted systems.
• Uphold Infrastructure as Code and configuration management practices.
• Generate and maintain runbooks, operational documentation, and escalation procedures.
• Identify production risks and drive remediation through automation, architectural improvements, and platform standards.
• Practical experience with production cloud infrastructure.
• Extensive experience with AWS, GCP, Windows, and/or hybrid cloud environments.
• Proficiency in building and maintaining observability, monitoring, logging, dashboards, and alerting systems.
• Strong troubleshooting capabilities across infrastructure, networking, application, and cloud service layers.
• Experience with Linux systems, networking fundamentals, DNS, TLS, IAM, load balancers, storage, and compute.
• Familiarity with Infrastructure as Code using Terraform, CloudFormation, Pulumi, or similar tools.
• Scripting and automation expertise in Bash, Python, Go, or related languages.
• Experience in participating in production incident response and postmortem processes.
• Understanding of SRE practices, including SLOs, SLIs, error budgets, toil reduction, and operational readiness.
• Ability to collaborate closely with engineering teams to enhance reliability and production supportability.
• Nice to have: Kubernetes and cloud-native platforms.
• Nice to have: GitOps experience with Flux or Argo CD.
• Nice to have: Jenkins and Ansible.
• Nice to have: Experience with Service Mesh, Ingress, and API Gateway.
• Nice to have: High Availability architectures and multi-region environments.
• Nice to have: Disaster Recovery solutions.
• Nice to have: Secrets Management with Vault, AWS Secrets Manager, or External Secrets.
• Nice to have: Security, Compliance, and Vulnerability Management.
• Nice to have: On-call Operations and Runbook design.
• Competitive pay and benefits.
• Health insurance.
• Language courses.
• Relocation program.
• Remote work flexibility through the ForeverRemote work culture.
• Professional development opportunities.
• Certification programs.
• Mentorship and talent investment programs.
• Internal mobility opportunities.
• Internship opportunities.
• Collaboration on impactful projects for top global clients.
• Inclusive and supportive multicultural work environment.
• Regular team-building company social events.
• Sustainable business practices focused on IT education, community empowerment, fair operating practices, environmental sustainability, and gender equality.
Vigil
Netrix Global
GFT Technologies
Cherokee Federal
Get handpicked remote jobs straight to your inbox weekly.