
Senior SRE / Cloud Engineer
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in Ukraine.
• Operate, manage, and enhance production cloud infrastructure across AWS, Azure, GCP, Windows, or hybrid environments.
• Develop and sustain monitoring, logging, metrics, tracing, dashboards, and alerts for production services.
• Enhance observability coverage across infrastructure, applications, databases, queues, and network dependencies.
• Adjust alerts to minimize noise and enhance actionable incident response.
• Engage in incident response, production troubleshooting, root cause analysis, and post-incident remediation.
• Create automation for infrastructure operations, deployments, health checks, runbooks, and recovery workflows.
• Collaborate with engineering teams to establish SLOs, SLIs, error budgets, and operational readiness standards.
• Provide support for cloud networking, DNS, TLS, load balancing, IAM, storage, compute, and managed service operations.
• Enhance the reliability, availability, performance, and scalability of cloud-hosted systems.
• Maintain Infrastructure as Code and configuration management practices for repeatable environments.
• Develop and update runbooks, operational documentation, and escalation procedures.
• Identify production risks and drive remediation through automation, architectural enhancements, and platform standards.
• Hands-on experience with production cloud infrastructure.
• Strong proficiency in AWS, GCP, Windows, and/or hybrid cloud environments.
• Experience in building and maintaining observability, monitoring, logging, dashboards, and alerting systems.
• Excellent troubleshooting abilities across infrastructure, networking, application, and cloud service layers.
• Familiarity with Linux systems, networking fundamentals, DNS, TLS, IAM, load balancers, storage, and compute.
• Experience with Infrastructure as Code using Terraform, CloudFormation, Pulumi, or similar tools.
• Proficient in scripting and automation using Bash, Python, Go, or similar languages.
• Experience in participating in production incident response and postmortem processes.
• Knowledge of SRE practices, including SLOs, SLIs, error budgets, toil reduction, and operational readiness.
• Ability to collaborate closely with engineering teams to enhance reliability and production supportability.
• Experience with Kubernetes and cloud-native platforms is a plus.
• Familiarity with GitOps using Flux or Argo CD is a plus.
• Experience with Jenkins and Ansible is a plus.
• Understanding of Service Mesh, Ingress, and API Gateway is a plus.
• Experience with High Availability architectures is a plus.
• Familiarity with multi-region environments is a plus.
• Experience with Disaster Recovery solutions is a plus.
• Knowledge of Secrets Management with Vault, AWS Secrets Manager, or External Secrets is a plus.
• Experience in Security, Compliance, and Vulnerability Management is a plus.
• On-call Operations and Runbook design experience is a plus.
• Competitive salary and benefits package.
• Health insurance coverage.
• Language training opportunities.
• Relocation assistance program.
• Flexibility for remote work.
• Opportunities for professional development.
• Certification programs available.
• Mentorship and talent investment initiatives.
• Internal mobility opportunities for career growth.
• Internship programs.
• Collaboration on significant projects for leading global clients.
• An inclusive and supportive multicultural work environment.
• Open lines of communication.
• Regular team-building social events within the company.
SPD Technology
Totara
Vesta Software Group
Elfonze Technologies
Get handpicked remote jobs straight to your inbox weekly.