
Senior SRE, Cloud Engineer
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in Romania.
• Operate, maintain, and enhance cloud production infrastructure across AWS, Azure, GCP, Windows, or hybrid settings.
• Build and sustain monitoring, logging, metrics, tracing, dashboards, and alerting systems for production services.
• Enhance observability throughout infrastructure, applications, databases, queues, and network dependencies.
• Adjust alerts to minimize noise and improve effective incident response.
• Engage in incident response, production troubleshooting, root cause analysis, and post-incident remediation.
• Develop automation for infrastructure operations, deployments, health checks, runbooks, and recovery workflows.
• Collaborate with engineering teams to establish SLOs, SLIs, error budgets, and operational readiness standards.
• Provide support for cloud networking, DNS, TLS, load balancing, IAM, storage, compute, and managed service operations.
• Enhance the reliability, availability, performance, and scalability of cloud-hosted systems.
• Maintain Infrastructure as Code and configuration management practices.
• Create and uphold runbooks, operational documentation, and escalation procedures.
• Identify production risks and promote remediation through automation, architectural improvements, and platform standards.
• Hands-on experience with production cloud infrastructure.
• Extensive experience with AWS, GCP, Windows, and/or hybrid cloud environments.
• Proven track record in building and maintaining observability, monitoring, logging, dashboards, and alerting systems.
• Strong troubleshooting capabilities across infrastructure, networking, application, and cloud service layers.
• Familiarity with Linux systems, networking fundamentals, DNS, TLS, IAM, load balancers, storage, and compute.
• Experience with Infrastructure as Code using Terraform, CloudFormation, Pulumi, or similar tools.
• Proficient in scripting and automation with Bash, Python, Go, or comparable languages.
• Experience in participating in production incident response and postmortem processes.
• Knowledge of SRE practices, including SLOs, SLIs, error budgets, toil reduction, and operational readiness.
• Ability to collaborate closely with engineering teams to enhance reliability and production supportability.
• Experience with Kubernetes and cloud-native platforms.
• Familiarity with GitOps methodologies, specifically with Flux or Argo CD.
• Proficient in Jenkins and Ansible.
• Experience with Service Mesh, Ingress, and API Gateway solutions.
• Knowledge of High Availability architectures.
• Experience in multi-region environments.
• Familiarity with Disaster Recovery solutions.
• Knowledge of Secrets Management using Vault, AWS Secrets Manager, or External Secrets.
• Understanding of Security, Compliance, and Vulnerability Management.
• Experience with On-call Operations and Runbook design.
• Health insurance.
• Language courses.
• Relocation program.
• Flexibility for remote work.
• Opportunities for professional development.
• Certification programs.
• Mentorship and investment in talent programs.
• Internal mobility opportunities.
• Internship opportunities.
• Engaging projects for leading global clients.
• An inclusive and supportive multicultural work environment.
• Regular team-building social events within the company.
SPD Technology
Totara
Vesta Software Group
Elfonze Technologies
Get handpicked remote jobs straight to your inbox weekly.