
Senior DevOps Engineer
Posted Jul 21

Posted Jul 21
This is a fully remote position, open to applicants in Spain.
• Take ownership of one of our essential platform areas from start to finish: observability (VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, alerting) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus, build agents) — you will lead its architecture, reliability, and strategic roadmap.
• Spearhead technical projects from inception to completion: gather requirements, draft the design document, break it down into tasks, implement solutions, deliver to production, and maintain operational health thereafter.
• Provide clarity in uncertain scenarios by articulating requirements, assumptions, and subsequent actions.
• Design with reliability and scalability in mind: enhance the architecture of our platforms — including topology, integration points, scaling methodologies, and reliability models.
• Assist developers: deploy and monitor applications on both on-premise servers and Kubernetes (Helm), troubleshoot builds and deployments, support teams with metrics, alerts, and logs; engage in chat duty within developer support channels.
• Eliminate repetitive tasks: operations, provisioning, and maintenance should be automated rather than performed manually.
• Investigate production incidents as the senior escalation point for your designated area: drive resolutions, lead post-mortems, and implement systemic improvements. Participate in on-call rotations and enhance the effectiveness of on-call operations.
• Mentor junior engineers through design discussions, code reviews, and collaborative work; identify and address debt-inducing shortcuts during the review process.
• Leverage AI in all facets of daily tasks: including research, troubleshooting, and development.
• Over 6 years of experience as a DevOps Engineer / SRE (or closely related responsibilities).
• Proven history of managing technical initiatives from beginning to end — encompassing requirements gathering and technical design through to production delivery. You should be able to present initiatives that you have owned, not just tasks you have completed.
• Strong Linux proficiency (we utilize Ubuntu).
• Familiarity with the Prometheus stack: metric types, exporters, and alerting mechanisms — sufficient to navigate and enhance an existing setup.
• Practical experience with CI/CD: pipeline design, build orchestration, and artifact delivery.
• Proficiency with containers: Docker, image creation, and registries.
• Experience with Ansible.
• Proficient in Git.
• Experience with Bash or Python scripting for automation and observability (creating exporters, reducing routine tasks).
• Production/on-call experience: diagnosing incidents, restoring services, and leading post-mortems.
• Experience in mentoring less experienced engineers.
• Strong ownership and attention to detail. Downtime is costly: during peak events, 10 minutes of downtime can result in approximately $500k in losses.
• Private medical insurance for the employee and their family.
• 23 paid vacation days annually.
• 11 paid public holidays each year.
• 5 company-paid sick leave days.
• English language courses.
• Relevant professional training.
• Gym or swimming pool membership.
• Home Office Setup Assistance: the company provides support for purchasing furniture (office chair, office desk, monitor) and other items to establish a comfortable workspace.
• Co-working opportunities.
• Remote working options.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.