
Senior DevOps Engineer
Posted 3 hours ago

Posted 3 hours ago
This is a fully remote position, open to applicants in Poland.
• Take full ownership of one of our essential platform areas from start to finish: observability (including VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, and alerting) or CI/CD (covering Jenkins scripted pipelines, Harbor, Nexus, and build agents) — you will influence its architecture, reliability, and roadmap.
• Lead technical projects from inception to completion: gather requirements, draft design documentation, break down tasks, implement solutions, deliver to production, and ensure operational health thereafter.
• Establish clarity in uncertain circumstances by specifying requirements, assumptions, and subsequent actions.
• Design for robustness and scalability: enhance the architecture of our platforms — including topology, integration points, scaling methods, and reliability models.
• Support developers by deploying and monitoring applications on both on-premise servers and Kubernetes (using Helm), troubleshooting builds and deployments, assisting teams with metrics, alerts, and logs; participate in chat duty in developer support channels.
• Eliminate repetitive tasks: operations, provisioning, and maintenance should be automated rather than performed manually.
• Investigate production incidents as the senior escalation point for your domain: drive resolutions, lead post-mortems, and implement systemic improvements. Participate in on-call rotations and enhance the on-call process.
• Guide less experienced engineers through design discussions, code reviews, and collaborative sessions; identify debt-inducing shortcuts during the review process.
• Leverage AI in all facets of daily work: for research, troubleshooting, and development.
• Over 6 years of experience as a DevOps Engineer / SRE (or closely related roles).
• Proven history of managing technical initiatives from beginning to end — from gathering requirements and designing solutions to delivering them in production. You should be able to present initiatives that you led, rather than just tasks you completed.
• Strong Linux skills (we primarily use Ubuntu).
• Familiarity with the Prometheus stack: understanding of metric types, exporters, and alerting mechanisms — sufficient to navigate and enhance an existing setup.
• Practical experience with CI/CD: designing pipelines, orchestrating builds, and delivering artifacts.
• Proficiency with containers: Docker, image creation, and registries.
• Experience with Ansible.
• Proficient in Git.
• Background in Bash or Python scripting for automation and observability (including writing exporters and reducing routine tasks).
• Experience in production/on-call roles: diagnosing incidents, restoring services, and leading post-mortem analyses.
• Experience mentoring junior engineers.
• A strong sense of ownership and meticulous attention to detail. Downtime can be costly: during peak events, just 10 minutes of downtime may result in losses of around $500k.
• 31 days of paid time off.
• Fully covered telemedicine plan.
• Home Office Setup Assistance: the company provides support for purchasing furniture (office chair, desk, monitor) and other items to establish a comfortable workspace.
• English language courses.
• Opportunities for relevant professional education.
• Access to a gym or swimming pool.
• Co-working space availability.
• Option for remote work.
Empower
Harrods
Aufinity Group | España
Zipdev
Get handpicked remote jobs straight to your inbox weekly.