
Senior DevOps Engineer
Posted Jul 21

Posted Jul 21
This is a fully remote position, open to applicants in Georgia.
• Take full ownership of one of our essential platform domains: observability (VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, alerting) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus, build agents) — you will lead its architecture, reliability, and roadmap.
• Steward technical initiatives from start to finish: gather requirements, draft the design document, break down into tasks, implement, deliver to production, and maintain operational health thereafter.
• Provide clarity in uncertain scenarios by defining requirements, assumptions, and subsequent steps.
• Design for reliability and scalability: evolve the architecture of our platforms — including topology, integration points, scaling strategies, and reliability models.
• Support developers by deploying and monitoring applications on both on-premise servers and Kubernetes (Helm), troubleshooting builds and deployments, assisting teams with metrics, alerts, and logs; engage in chat duty within developer support channels.
• Eliminate manual tasks: repetitive operations, provisioning, and maintenance should be automated, not done manually.
• Investigate production incidents as the senior escalation point in your area: drive resolution, lead post-mortems, and implement systemic improvements. Participate in on-call rotations and enhance the on-call experience.
• Mentor junior engineers through design discussions, reviews, and collaborative work; identify and address debt-inducing shortcuts during the review process.
• Leverage AI in all facets of daily tasks: researching, troubleshooting, and developing.
• Over 6 years of experience as a DevOps Engineer / SRE (or closely related responsibilities).
• Proven track record of managing technical initiatives from inception to completion — from requirements gathering and technical design to production delivery. You can highlight initiatives that you led, rather than only tasks you accomplished.
• Proficient Linux skills (we utilize Ubuntu).
• Understanding of the Prometheus stack: metric types, exporters, and alerting mechanisms — sufficient to navigate and extend an existing configuration.
• Practical experience with CI/CD: pipeline design, build orchestration, artifact delivery.
• Familiarity with containers: Docker, image building, registries.
• Knowledge of Ansible.
• Proficient with Git.
• Experience with Bash or Python scripting for automation and observability (creating exporters, reducing routine tasks).
• Experience in production/on-call roles: diagnosing incidents, restoring services, and leading post-mortems.
• Background in mentoring junior engineers.
• Strong sense of ownership and attention to detail. Downtime is costly: during peak events, 10 minutes of downtime can result in approximately $500k in losses.
• Must possess substantial hands-on experience in at least two of the following areas: VictoriaMetrics / Prometheus stack at scale, log pipelines at scale, Jenkins scripted pipelines, container registries and artifact management, operating applications on Kubernetes, Grafana.
• 31 days of paid time off.
• Comprehensive telemedicine plan at no cost.
• Assistance with home office setup: the company provides support for purchasing furniture (office chair, office desk, monitor) and other items to create a comfortable workspace.
• English language learning courses.
• Opportunities for relevant professional development.
• Access to a gym or swimming pool.
• Co-working spaces.
• Flexible remote working arrangements.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.