
Senior DevOps Engineer
Posted Jul 21

Posted Jul 21
This is a fully remote position, open to applicants in Serbia.
• Take ownership of one of our key platform areas from start to finish: observability (VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, alerting) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus, build agents) — you will lead its architecture, reliability, and strategic direction.
• Lead technical projects from inception to completion: gather requirements, draft the design document, break down tasks, implement solutions, deliver to production, and maintain operational health afterward.
• Provide clarity in uncertain situations by defining requirements, assumptions, and subsequent steps.
• Design for reliability and scalability: enhance the architecture of our platforms — including topology, integration points, scaling methods, and reliability models.
• Support developers: deploy and monitor applications on both on-premise servers and Kubernetes (Helm), troubleshoot builds and deployments, assist teams with metrics, alerts, and logs; engage in chat duty within developer support channels.
• Eliminate repetitive tasks: operations, provisioning, and maintenance should be automated rather than manually performed.
• Investigate production incidents as the senior escalation point for your area: drive resolutions, lead post-mortems, and implement systemic fixes. Participate in on-call rotations and enhance the standards for on-call practices.
• Mentor junior engineers through design discussions, reviews, and collaborative work; identify and prevent debt-inducing shortcuts during the review process.
• Leverage AI across all daily tasks: researching, troubleshooting, and development.
• 6+ years of experience as a DevOps Engineer / SRE (or in closely related roles).
• Proven history of managing technical projects from beginning to end — from requirements gathering and technical design to production delivery. You should be able to present initiatives that you have owned, rather than just tasks you completed.
• Proficient Linux skills (we utilize Ubuntu).
• Familiarity with the Prometheus stack: metric types, exporters, and alerting mechanics — sufficient to navigate and enhance an existing setup.
• Practical experience with CI/CD: pipeline design, build orchestration, and artifact delivery.
• Experience with containers: Docker, image building, and registries.
• Knowledge of Ansible.
• Proficiency in Git.
• Experience with Bash or Python scripting for automation and observability (including writing exporters and reducing routine tasks).
• Hands-on experience in production/on-call roles: diagnosing incidents, restoring service, and leading post-mortems.
• Experience mentoring less experienced engineers.
• A strong sense of ownership and attention to detail. Downtime is costly: during peak events, just 10 minutes of downtime can lead to losses of around $500k.
• Private medical insurance for employees and their families.
• 20 paid vacation days annually.
• 15 paid public holidays each year.
• 5 company-paid sick leave days.
• English learning courses.
• Relevant professional education opportunities.
• Access to a gym or swimming pool.
• Home Office Setup Assistance: the company provides support for purchasing furniture (office chair, desk, monitor) and other items to create a comfortable workspace.
• Co-working opportunities.
• Remote work options.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.