
Observability Platform Engineer
Posted Aug 31

Posted Aug 31
This is a fully remote position, open to applicants in Kazakhstan, +1 more country.
• Design, develop, and manage components of an observability platform, focusing on metrics, logging, distributed tracing, and alerting functionalities.
• Construct high-capacity telemetry pipelines for extensive infrastructure systems, prioritizing cost-efficiency, data retention, and query performance.
• Establish and execute SLO/SLI frameworks alongside alerting methodologies.
• Collaborate with service delivery and operations teams to fulfill incident-response requirements.
• Integrate observability tools with incident management processes, root-cause analysis, and post-incident evaluations.
• Enhance detection speed while minimizing Mean Time to Detection (MTTD) and Mean Time to Recovery (MTTR).
• Contribute to AI-driven operations tools, including automated triage and anomaly detection systems.
• Take ownership of the reliability, scalability, and security aspects of the observability stack.
• Document system architecture, operational procedures, and runbooks.
• Assist operations teams in diagnosing incidents more swiftly and highlighting critical alerts.
• Contribute to the strategic roadmap and ongoing maintenance of the observability platform.
• Demonstrated experience in the design and construction of observability platforms for large-scale, production-level infrastructure.
• In-depth hands-on expertise with metrics, logging, and distributed tracing tools such as Prometheus, Grafana, OpenTelemetry, Loki, Thanos/Cortex/Mimir, Elasticsearch/OpenSearch, Jaeger/Tempo, or similar technologies.
• Experience managing high-volume telemetry pipelines, understanding the trade-offs related to cardinality, retention, cost, and query latency.
• Strong software engineering capabilities in at least one relevant programming language, such as Go, Python, or Rust.
• Familiarity with Kubernetes and cloud-native infrastructure.
• Comprehensive understanding of SLO/SLI/error-budget methodologies and designing low-noise alerting systems.
• Comfortable operating within a dynamic and fast-paced environment.
• Excellent communication skills with the ability to collaborate directly with operations and service delivery teams.
• Preferred experience with GPU/HPC infrastructure or specialized high-performance computing environments.
• Preferred familiarity with eBPF-based observability tools.
• Preferred knowledge of AIOps/ML-based anomaly detection or automated triage systems.
• Preferred experience in a managed services or MSP setting.
• Contributions to open-source observability projects are preferred.
• Collaborate with a recognized leader in cloud infrastructure based in Silicon Valley.
• Engage with dedicated, skilled colleagues serving Fortune 500 and Global 2000 clients.
• Participate in groundbreaking open-source innovations.
• Experience a high-energy workplace that values openness, collaboration, risk-taking, and continuous development.
• Access to professional development and training opportunities.
• Attend industry conferences and participate in working groups.
• Enjoy company outings, happy hours, hackathons, and technical talks.
• Competitive compensation package accompanied by a comprehensive benefits plan.
• Remote work options available.
MAIA
NavAide
Share
Lime
Get handpicked remote jobs straight to your inbox weekly.