Observability Platform Engineer

Posted Aug 31

This is a fully remote position, open to applicants in Kazakhstan, +1 more country.

📋 Description

• Design, develop, and manage components of an observability platform, focusing on metrics, logging, distributed tracing, and alerting functionalities.

• Construct high-capacity telemetry pipelines for extensive infrastructure systems, prioritizing cost-efficiency, data retention, and query performance.

• Establish and execute SLO/SLI frameworks alongside alerting methodologies.

• Collaborate with service delivery and operations teams to fulfill incident-response requirements.

• Integrate observability tools with incident management processes, root-cause analysis, and post-incident evaluations.

• Enhance detection speed while minimizing Mean Time to Detection (MTTD) and Mean Time to Recovery (MTTR).

• Contribute to AI-driven operations tools, including automated triage and anomaly detection systems.

• Take ownership of the reliability, scalability, and security aspects of the observability stack.

• Document system architecture, operational procedures, and runbooks.

• Assist operations teams in diagnosing incidents more swiftly and highlighting critical alerts.

• Contribute to the strategic roadmap and ongoing maintenance of the observability platform.


⛳️ Requirements

• Demonstrated experience in the design and construction of observability platforms for large-scale, production-level infrastructure.

• In-depth hands-on expertise with metrics, logging, and distributed tracing tools such as Prometheus, Grafana, OpenTelemetry, Loki, Thanos/Cortex/Mimir, Elasticsearch/OpenSearch, Jaeger/Tempo, or similar technologies.

• Experience managing high-volume telemetry pipelines, understanding the trade-offs related to cardinality, retention, cost, and query latency.

• Strong software engineering capabilities in at least one relevant programming language, such as Go, Python, or Rust.

• Familiarity with Kubernetes and cloud-native infrastructure.

• Comprehensive understanding of SLO/SLI/error-budget methodologies and designing low-noise alerting systems.

• Comfortable operating within a dynamic and fast-paced environment.

• Excellent communication skills with the ability to collaborate directly with operations and service delivery teams.

• Preferred experience with GPU/HPC infrastructure or specialized high-performance computing environments.

• Preferred familiarity with eBPF-based observability tools.

• Preferred knowledge of AIOps/ML-based anomaly detection or automated triage systems.

• Preferred experience in a managed services or MSP setting.

• Contributions to open-source observability projects are preferred.


🏝️ Benefits

• Collaborate with a recognized leader in cloud infrastructure based in Silicon Valley.

• Engage with dedicated, skilled colleagues serving Fortune 500 and Global 2000 clients.

• Participate in groundbreaking open-source innovations.

• Experience a high-energy workplace that values openness, collaboration, risk-taking, and continuous development.

• Access to professional development and training opportunities.

• Attend industry conferences and participate in working groups.

• Enjoy company outings, happy hours, hackathons, and technical talks.

• Competitive compensation package accompanied by a comprehensive benefits plan.

• Remote work options available.

People also viewed

MAIA14 hours ago

Senior DevOps / Platform Engineer, AI Infrastructure

DE flagGermany OnlyFull-timePlatform Engineer€75k – €85k/year
ApplyView job
NavAide17 hours ago

Power Platform Developer, DOD Clearance Required

US flagCalifornia OnlyFull-timePlatform Engineer$98k – $140k/year
ApplyView job
Share22 hours ago

Platform Engineer

ES flagSpain, +6 more countriesFull-timePlatform Engineer
ApplyView job
Lime22 hours ago

Staff Software Engineer, Platform Engineering

CA flagCanada OnlyFull-timePlatform EngineerC$172k – C$237k/year
ApplyView job
Convoso23 hours ago

Senior Platform Engineer

US flagCalifornia OnlyFull-timePlatform Engineer$190k – $210k/year
ApplyView job
Cortex1 day ago

Platform Engineer – SRE, Mid-Level

BR flagBrazil OnlyFull-timePlatform Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers