
Senior SRE, Cloud Engineer
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in Poland.
• Manage, maintain, and enhance production cloud infrastructure across AWS, Azure, GCP, Windows, or hybrid setups
• Develop and uphold monitoring, logging, metrics, tracing, dashboards, and alert systems for production services
• Enhance observability across infrastructure, applications, databases, queues, and network dependencies
• Adjust alerts to minimize noise and enhance signal quality for effective incident response
• Engage in incident response, production troubleshooting, root cause analysis, and post-incident recovery
• Automate infrastructure operations, deployments, health checks, runbooks, and recovery workflows
• Collaborate with engineering teams to establish SLOs, SLIs, error budgets, and operational readiness standards
• Provide support for cloud networking, DNS, TLS, load balancing, IAM, storage, compute, and managed service operations
• Enhance the reliability, availability, performance, and scalability of cloud-hosted systems
• Maintain Infrastructure as Code and configuration management practices for consistent environments
• Develop and update runbooks, operational documentation, and escalation processes
• Identify production risks and drive remediation through automation, architectural enhancements, and platform standards
• Practical experience with production cloud infrastructure
• Proficient in AWS, GCP, Windows, and/or hybrid cloud environments
• Experience in building and maintaining observability, monitoring, logging, dashboards, and alerting systems
• Strong troubleshooting capabilities across infrastructure, networking, application, and cloud service layers
• Familiarity with Linux systems, networking fundamentals, DNS, TLS, IAM, load balancers, storage, and compute
• Experience with Infrastructure as Code using Terraform, CloudFormation, Pulumi, or similar tools
• Scripting and automation skills in Bash, Python, Go, or similar languages
• Experience in participating in production incident response and postmortem evaluations
• Knowledge of SRE practices, including SLOs, SLIs, error budgets, toil reduction, and operational readiness
• Ability to collaborate closely with engineering teams to enhance reliability and production supportability
• Experience with Kubernetes and cloud-native platforms is a plus
• Familiarity with GitOps using Flux or Argo CD is a plus
• Experience with Jenkins and Ansible is a plus
• Knowledge of Service Mesh, Ingress, and API Gateway is a plus
• Experience with High Availability architectures is a plus
• Experience in multi-region environments is a plus
• Familiarity with Disaster Recovery solutions is a plus
• Experience with Secrets Management using Vault, AWS Secrets Manager, or External Secrets is a plus
• Knowledge of Security, Compliance, and Vulnerability Management is a plus
• On-call Operations and Runbook design experience is a plus
• Health insurance
• Language courses
• Relocation program
• Remote work flexibility
• Professional development opportunities
• Certification programs
• Mentorship and talent investment initiatives
• Internal mobility opportunities
• Internship opportunities
• Opportunities to work on impactful projects for leading global clients
• An inclusive and supportive work environment
• Open communication culture
• Regular team-building and company social events
SPD Technology
Totara
Vesta Software Group
Elfonze Technologies
Get handpicked remote jobs straight to your inbox weekly.