Lead Site Reliability Engineer – Imunify Reliability Platform

Posted Aug 22

This is a fully remote position, open to applicants in Poland, +5 more countries.

📋 Description

• Articulate the definition of "working" for around 70 components.

• Conduct SLI definitions in collaboration with squad leads and senior engineers and facilitate the sign-off of ownership.

• Develop service, fleet, control-efficacy, delivery, and pipeline SLI taxonomies.

• Assign an SLO, error budget, and responsible squad to each SLI, ensuring appropriate tiering.

• Design and implement a push-based, sampled, privacy-constrained telemetry collection pipeline that adheres to a defended cardinality budget.

• Enhance agent-side and service-side instrumentation utilizing Python, Go, and Rust.

• Integrate dashboards, ad-hoc queries, and reporting pathways into a cohesive and defensible instrument set.

• Create symptom-based, SLO-anchored alerting with multi-window burn-rate semantics.

• Define alert tiers for pages, tickets, and dashboards and establish paging criteria.

• Ensure every alert is assigned an owner, has a runbook, and documents failure modes.

• Maintain alert hygiene through quarterly reviews, deletions, and tracking actionable rates.

• Sustain a machine-readable ownership map linking components to squads for alert routing.

• Develop severity matrices, acknowledgement SLAs, follow-the-sun rotas spanning UTC−5 to UTC+8, and handoff protocols.

• Execute incident command and conduct blameless postmortems within 24 hours.

• Design escalation processes so that squads manage their own pagers while you oversee the platform and mentor teams.

• Achieve first-year outcomes including component inventory, pilot instrumentation, production collection pipeline, tier-1 escalation, complete component SLI coverage, alert metrics, squad on-call, and mean time to detect silent control degradation in under 24 hours.


⛳️ Requirements

• Extensive experience in production engineering or SRE, including at least one environment where you established the SLO framework rather than inherited it.

• Capability to discuss personally authored SLIs and clarify how they were negotiated with uncooperative teams.

• Proficient in Python.

• Comfortable reading and modifying Go or Rust code.

• Hands-on experience with time-series and event telemetry at scale.

• Familiarity with Prometheus/OpenMetrics, Grafana, and an Alertmanager-class routing layer.

• Background in using ClickHouse or a comparable columnar store for managing high-cardinality fleet data.

• Experience debugging distributed systems on bare metal and long-lived hosts.

• Competence in configuration management and CI at production scale, including tools like Ansible, GitLab CI, Jenkins, or similar equivalents.

• Ability to design measurement strategies for machines that cannot be owned or scraped, encompassing push telemetry, sampling, clock skew, partial reporting, and customer-server privacy constraints.

• Strong written communication skills suitable for asynchronous work.


🏝️ Benefits

• Opportunities for professional development.

• Engaging and challenging projects.

• Access to mentor and knowledge-exchange programs.

• Fully remote work with flexible hours.

• Ability to work from any location worldwide.

• 24 days of paid vacation annually.

• 10 days of public holidays.

• Unlimited sick leave.

• Coverage for private medical insurance.

• Reimbursement for co-working expenses.

• Gym/sports reimbursement.

• Opportunity to earn a reward for the most innovative idea that the company can patent.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers