
Senior DevOps Engineer
Posted Jul 21

Posted Jul 21
This is a fully remote position, open to applicants in Armenia.
• Take full ownership of one of our essential platform areas, either observability or CI/CD, by steering its architecture, reliability, and roadmap.
• Lead technical initiatives from start to finish: gather requirements, draft the design document, break down tasks, implement solutions, deliver to production, and maintain operational health post-deployment.
• Provide clarity in ambiguous situations by clearly defining requirements, assumptions, and next steps.
• Design for reliability and scalability: enhance the architecture of our platforms, including topology, integration points, scaling strategies, and reliability models.
• Support developers by deploying and monitoring applications on both on-premise servers and Kubernetes (Helm), troubleshooting builds and deployments, assisting teams with metrics, alerts, and logs, and participating in developer support channels for chat duty.
• Automate repetitive tasks: operations, provisioning, and maintenance should be codified rather than performed manually.
• Investigate production incidents as the senior escalation point for your area: lead resolutions, conduct post-mortems, and implement systemic improvements. Participate in on-call rotations and enhance the on-call experience.
• Mentor junior engineers through design discussions, reviews, and collaborative work; identify and address debt-inducing shortcuts during the review process.
• Integrate AI into all facets of daily work, including research, troubleshooting, and development.
• Over 6 years of experience as a DevOps Engineer / SRE (or similar responsibilities).
• Proven ability to own technical initiatives from conception to completion — demonstrating ownership of initiatives, not just individual tasks.
• Strong Linux skills (we operate on Ubuntu).
• Familiarity with the Prometheus stack: understanding of metric types, exporters, and alerting mechanisms sufficient to navigate and extend an existing setup.
• Practical experience with CI/CD, including pipeline design, build orchestration, and artifact delivery.
• Proficiency with containers: Docker, image creation, and registries.
• Experience with Ansible.
• Proficient in Git.
• Knowledge of Bash or Python scripting for automation and observability (writing exporters and minimizing routine tasks).
• Experience in production/on-call roles: diagnosing incidents, restoring services, and leading post-mortems.
• Experience mentoring junior engineers.
• Demonstrated ownership and meticulous attention to detail, as downtime is costly: during peak events, 10 minutes of downtime can result in approximately $500k in losses.
• Must possess solid hands-on experience in at least two of the following areas: VictoriaMetrics / Prometheus stack at scale, log pipelines at scale, Jenkins scripted pipelines, container registries and artifact management, operating applications on Kubernetes, and Grafana.
• 31 days of leave.
• Fully covered telemedicine plan.
• Home Office Setup Assistance: the company provides support for purchasing furniture (office chair, desk, monitor) and other items to create a conducive workspace.
• English language courses.
• Opportunities for relevant professional education.
• Access to gym or swimming pool.
• Co-working spaces available.
• Remote working options.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.