
Lead Site Reliability Engineer – Imunify Reliability Platform
Posted Aug 22

Posted Aug 22
This is a fully remote position, open to applicants in Poland, +5 more countries.
• Articulate the definition of "working" for around 70 components.
• Conduct SLI definitions in collaboration with squad leads and senior engineers and facilitate the sign-off of ownership.
• Develop service, fleet, control-efficacy, delivery, and pipeline SLI taxonomies.
• Assign an SLO, error budget, and responsible squad to each SLI, ensuring appropriate tiering.
• Design and implement a push-based, sampled, privacy-constrained telemetry collection pipeline that adheres to a defended cardinality budget.
• Enhance agent-side and service-side instrumentation utilizing Python, Go, and Rust.
• Integrate dashboards, ad-hoc queries, and reporting pathways into a cohesive and defensible instrument set.
• Create symptom-based, SLO-anchored alerting with multi-window burn-rate semantics.
• Define alert tiers for pages, tickets, and dashboards and establish paging criteria.
• Ensure every alert is assigned an owner, has a runbook, and documents failure modes.
• Maintain alert hygiene through quarterly reviews, deletions, and tracking actionable rates.
• Sustain a machine-readable ownership map linking components to squads for alert routing.
• Develop severity matrices, acknowledgement SLAs, follow-the-sun rotas spanning UTC−5 to UTC+8, and handoff protocols.
• Execute incident command and conduct blameless postmortems within 24 hours.
• Design escalation processes so that squads manage their own pagers while you oversee the platform and mentor teams.
• Achieve first-year outcomes including component inventory, pilot instrumentation, production collection pipeline, tier-1 escalation, complete component SLI coverage, alert metrics, squad on-call, and mean time to detect silent control degradation in under 24 hours.
• Extensive experience in production engineering or SRE, including at least one environment where you established the SLO framework rather than inherited it.
• Capability to discuss personally authored SLIs and clarify how they were negotiated with uncooperative teams.
• Proficient in Python.
• Comfortable reading and modifying Go or Rust code.
• Hands-on experience with time-series and event telemetry at scale.
• Familiarity with Prometheus/OpenMetrics, Grafana, and an Alertmanager-class routing layer.
• Background in using ClickHouse or a comparable columnar store for managing high-cardinality fleet data.
• Experience debugging distributed systems on bare metal and long-lived hosts.
• Competence in configuration management and CI at production scale, including tools like Ansible, GitLab CI, Jenkins, or similar equivalents.
• Ability to design measurement strategies for machines that cannot be owned or scraped, encompassing push telemetry, sampling, clock skew, partial reporting, and customer-server privacy constraints.
• Strong written communication skills suitable for asynchronous work.
• Opportunities for professional development.
• Engaging and challenging projects.
• Access to mentor and knowledge-exchange programs.
• Fully remote work with flexible hours.
• Ability to work from any location worldwide.
• 24 days of paid vacation annually.
• 10 days of public holidays.
• Unlimited sick leave.
• Coverage for private medical insurance.
• Reimbursement for co-working expenses.
• Gym/sports reimbursement.
• Opportunity to earn a reward for the most innovative idea that the company can patent.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.