
Observability Platform Engineer
Posted Aug 31

Posted Aug 31
This is a fully remote position, open to applicants in Kazakhstan.
• Design, develop, and manage components for metrics, logging, distributed tracing, and alerting platforms.
• Construct telemetry pipelines that handle high volumes and high cardinality for extensive infrastructure fleets.
• Enhance telemetry pipelines focusing on cost efficiency, data retention, and query performance.
• Establish and execute SLO/SLI frameworks and alerting methodologies that minimize noise.
• Collaborate with service delivery and operations teams to create incident-centric observability features.
• Integrate observability tools with incident management processes, root-cause analysis, and post-incident evaluations.
• Enhance detection speed while lowering Mean Time to Detection (MTTD) and Mean Time to Recovery (MTTR) throughout the platform.
• Contribute to AI-powered operational tools, including automated triage, anomaly detection, and engineer-assist applications.
• Take ownership of the reliability, scalability, and security aspects of the observability stack.
• Document system architecture, runbooks, and operational procedures.
• Demonstrated experience in designing and building observability platforms for large-scale production infrastructure.
• Extensive hands-on experience with metrics, logging, and distributed tracing tools such as Prometheus, Grafana, OpenTelemetry, Loki, Thanos/Cortex/Mimir, Elasticsearch/OpenSearch, Jaeger/Tempo, or similar technologies.
• Familiarity with high-volume telemetry pipelines and the trade-offs involving cardinality, retention, cost, and query latency.
• Proficient software engineering skills in at least one relevant programming language, such as Go, Python, or Rust.
• Experience with Kubernetes and cloud-native infrastructure solutions.
• Strong understanding of SLO/SLI/error-budget practices and the principles of low-noise alerting design.
• Ability to thrive in a dynamic environment where the platform is developed alongside the infrastructure it observes.
• Excellent communication skills and the ability to collaborate directly with operations and service delivery teams.
• Experience in developing observability solutions for GPU/HPC or other specialized high-performance computing environments.
• Knowledge of eBPF-based observability tools.
• Familiarity with AIOps/ML-driven anomaly detection or automated triage systems.
• Background in managed services or managed service provider contexts.
• Contributions to open-source observability initiatives.
• Engage with a prominent Silicon Valley leader in the cloud infrastructure sector.
• Collaborate with exceptionally passionate, skilled, and engaging colleagues.
• Assist Fortune 500 and Global 2000 clients in adopting next-generation cloud technologies.
• Be part of groundbreaking open-source innovation.
• Opportunities for professional growth and training.
• Attend conferences and participate in working groups.
• Enjoy company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package complemented by a robust benefits plan.
MAIA
NavAide
Share
Lime
Get handpicked remote jobs straight to your inbox weekly.