Remotery

Senior DevOps Engineer

Posted Jul 21

This is a fully remote position, open to applicants in Georgia.

📋 Description

• Take full ownership of one of our essential platform domains: observability (VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, alerting) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus, build agents) — you will lead its architecture, reliability, and roadmap.

• Steward technical initiatives from start to finish: gather requirements, draft the design document, break down into tasks, implement, deliver to production, and maintain operational health thereafter.

• Provide clarity in uncertain scenarios by defining requirements, assumptions, and subsequent steps.

• Design for reliability and scalability: evolve the architecture of our platforms — including topology, integration points, scaling strategies, and reliability models.

• Support developers by deploying and monitoring applications on both on-premise servers and Kubernetes (Helm), troubleshooting builds and deployments, assisting teams with metrics, alerts, and logs; engage in chat duty within developer support channels.

• Eliminate manual tasks: repetitive operations, provisioning, and maintenance should be automated, not done manually.

• Investigate production incidents as the senior escalation point in your area: drive resolution, lead post-mortems, and implement systemic improvements. Participate in on-call rotations and enhance the on-call experience.

• Mentor junior engineers through design discussions, reviews, and collaborative work; identify and address debt-inducing shortcuts during the review process.

• Leverage AI in all facets of daily tasks: researching, troubleshooting, and developing.


⛳️ Requirements

• Over 6 years of experience as a DevOps Engineer / SRE (or closely related responsibilities).

• Proven track record of managing technical initiatives from inception to completion — from requirements gathering and technical design to production delivery. You can highlight initiatives that you led, rather than only tasks you accomplished.

• Proficient Linux skills (we utilize Ubuntu).

• Understanding of the Prometheus stack: metric types, exporters, and alerting mechanisms — sufficient to navigate and extend an existing configuration.

• Practical experience with CI/CD: pipeline design, build orchestration, artifact delivery.

• Familiarity with containers: Docker, image building, registries.

• Knowledge of Ansible.

• Proficient with Git.

• Experience with Bash or Python scripting for automation and observability (creating exporters, reducing routine tasks).

• Experience in production/on-call roles: diagnosing incidents, restoring services, and leading post-mortems.

• Background in mentoring junior engineers.

• Strong sense of ownership and attention to detail. Downtime is costly: during peak events, 10 minutes of downtime can result in approximately $500k in losses.

• Must possess substantial hands-on experience in at least two of the following areas: VictoriaMetrics / Prometheus stack at scale, log pipelines at scale, Jenkins scripted pipelines, container registries and artifact management, operating applications on Kubernetes, Grafana.


🏝️ Benefits

• 31 days of paid time off.

• Comprehensive telemedicine plan at no cost.

• Assistance with home office setup: the company provides support for purchasing furniture (office chair, office desk, monitor) and other items to create a comfortable workspace.

• English language learning courses.

• Opportunities for relevant professional development.

• Access to a gym or swimming pool.

• Co-working spaces.

• Flexible remote working arrangements.

People also viewed

DATAGROUP1 day ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush1 day ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey1 day ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems2 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems2 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data2 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers