
Senior Site Reliability Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Poland.
• Ensure the functionality and availability of Compute services and infrastructure
• Oversee and sustain essential infrastructure
• Collaborate with cross-functional teams to develop tools and software aimed at monitoring and enhancing system reliability
• Engage with various technologies as new applications are introduced and existing tools are upgraded
• Offer support and guidance to fellow SRE engineers
• Deploy and manage platforms and tools used internally
• Enhance the Compute Cloud Interface platform to accelerate error detection and resolution while improving performance and reliability
• Create and refine automation processes for daily tasks and reduction of repetitive work
• Participate in on-call rotations and lead the restoration and repair of service-affecting incidents
• Work with internal teams to troubleshoot and address customer escalations and incidents
• Significant relevant experience and a Bachelor's degree in Computer Science or an equivalent field
• Proficiency in automation using Python and/or Golang, along with scripting in bash
• Strong understanding of systems reliability, observability, monitoring, and compliance with SLOs
• Familiarity with SaltStack, Terraform, Ansible, and Jenkins CI/CD
• Practical expertise in Linux administration
• Direct experience with Docker and containerized environments
• Proficient in using Prometheus, Grafana, Loki, nginx, Envoy, HAProxy, and Redis
• Comprehensive benefits covering health, wellness, financial support, and life beyond work
• FlexBase flexible work options: at home, in an office, or a blend of both
TEKsystems
TEKsystems
Level Data
Level Data
Get handpicked remote jobs straight to your inbox weekly.