
Site Reliability Engineer – SRE
Posted 17 hours ago

Posted 17 hours ago
This is a fully remote position, open to applicants anywhere in the world.
• Take ownership of and enhance the reliability and stability of production infrastructure.
• Plan, implement, and assist with deployments and modifications to infrastructure.
• Develop and sustain Infrastructure-as-Code solutions utilizing Ansible and Terraform.
• Provide support and optimization for Kubernetes-based and containerized environments.
• Create automation scripts and internal operational tools.
• Monitor system health, investigate incidents, and proactively enhance observability.
• Engage in CI/CD enhancements alongside Development, QA, DevOps, and SRE teams.
• Collaborate with monitoring and alerting systems to minimize downtime and boost system performance.
• Maintain technical documentation, runbooks, and operational processes.
• Assist with DNS, WAF, CDN, and caching infrastructure as needed.
• Over 3 years of experience in SRE, DevOps, System Administration, or Build/Release Engineering.
• Proficient in Linux administration and troubleshooting.
• Practical experience with Kubernetes and containerization technologies (Docker/Podman).
• Familiarity with CI/CD pipelines, ideally GitLab CI.
• Hands-on experience with Infrastructure-as-Code and configuration management tools (Ansible and/or Terraform).
• Strong skills in Bash scripting.
• Experience with observability and monitoring tools such as Prometheus, Grafana, Zabbix, or VictoriaMetrics.
• Solid understanding of networking fundamentals, including DNS, HTTP/HTTPS, load balancing, and troubleshooting.
• Proficient with Git and contemporary software delivery workflows.
• Capable of working independently, taking ownership, and proactively enhancing infrastructure.
• Fluent in Russian.
• Intermediate (B1) or higher proficiency in English for technical documentation and team communication.
• Experience with AWS or GCP.
• Familiarity with RabbitMQ / AMQP.
• Knowledge of Cloudflare, Akamai, WAF, and CDN.
• Skills in Python scripting.
• Experience with tracing and advanced observability tools.
• Public GitHub profile, contributions to open-source projects, or personal engineering projects.
• REMOTE OPPORTUNITY for full-time work.
• 28 calendar days of vacation annually.
• 7 wellness days per year (time off).
• Bonuses up to $5000 for referring successful candidates for positions in the company.
• 50% coverage for professional training, international conferences, and meetings.
• Corporate discounts for English lessons.
• Health benefits: reimbursement of up to $1,000 gross per year for self-purchased health insurance or doctor’s fees for yourself and close relatives if not eligible for corporate medical insurance.
• Equipped workplace and necessary equipment in offices or co-working spaces.
• Reimbursement of workplace costs up to $1000 gross once every 3 years for other locations.
• Internal gamified gratitude system with bonuses that can be exchanged for merchandise, team-building activities, massage certificates, etc.
Commit
Megaport
Akamai Technologies
Vetta
Get handpicked remote jobs straight to your inbox weekly.