Remotery

Senior DevOps Engineer

Posted Jul 21

This is a fully remote position, open to applicants in Serbia.

📋 Description

• Take ownership of one of our key platform areas from start to finish: observability (VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, alerting) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus, build agents) — you will lead its architecture, reliability, and strategic direction.

• Lead technical projects from inception to completion: gather requirements, draft the design document, break down tasks, implement solutions, deliver to production, and maintain operational health afterward.

• Provide clarity in uncertain situations by defining requirements, assumptions, and subsequent steps.

• Design for reliability and scalability: enhance the architecture of our platforms — including topology, integration points, scaling methods, and reliability models.

• Support developers: deploy and monitor applications on both on-premise servers and Kubernetes (Helm), troubleshoot builds and deployments, assist teams with metrics, alerts, and logs; engage in chat duty within developer support channels.

• Eliminate repetitive tasks: operations, provisioning, and maintenance should be automated rather than manually performed.

• Investigate production incidents as the senior escalation point for your area: drive resolutions, lead post-mortems, and implement systemic fixes. Participate in on-call rotations and enhance the standards for on-call practices.

• Mentor junior engineers through design discussions, reviews, and collaborative work; identify and prevent debt-inducing shortcuts during the review process.

• Leverage AI across all daily tasks: researching, troubleshooting, and development.


⛳️ Requirements

• 6+ years of experience as a DevOps Engineer / SRE (or in closely related roles).

• Proven history of managing technical projects from beginning to end — from requirements gathering and technical design to production delivery. You should be able to present initiatives that you have owned, rather than just tasks you completed.

• Proficient Linux skills (we utilize Ubuntu).

• Familiarity with the Prometheus stack: metric types, exporters, and alerting mechanics — sufficient to navigate and enhance an existing setup.

• Practical experience with CI/CD: pipeline design, build orchestration, and artifact delivery.

• Experience with containers: Docker, image building, and registries.

• Knowledge of Ansible.

• Proficiency in Git.

• Experience with Bash or Python scripting for automation and observability (including writing exporters and reducing routine tasks).

• Hands-on experience in production/on-call roles: diagnosing incidents, restoring service, and leading post-mortems.

• Experience mentoring less experienced engineers.

• A strong sense of ownership and attention to detail. Downtime is costly: during peak events, just 10 minutes of downtime can lead to losses of around $500k.


🏝️ Benefits

• Private medical insurance for employees and their families.

• 20 paid vacation days annually.

• 15 paid public holidays each year.

• 5 company-paid sick leave days.

• English learning courses.

• Relevant professional education opportunities.

• Access to a gym or swimming pool.

• Home Office Setup Assistance: the company provides support for purchasing furniture (office chair, desk, monitor) and other items to create a comfortable workspace.

• Co-working opportunities.

• Remote work options.

People also viewed

CWILL21 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3722 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT22 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group22 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo23 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch23 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers