
Observability Platform Engineer
Posted Aug 21

Posted Aug 21
This is a fully remote position, open to applicants in Kazakhstan.
• Design, develop, and manage components for metrics, logging, distributed tracing, and alerting platforms.
• Create high-volume, high-cardinality telemetry pipelines for extensive infrastructure fleets.
• Establish and execute SLO/SLI frameworks along with alerting strategies.
• Collaborate with service delivery and operations teams to enhance incident-focused observability.
• Integrate observability tools with incident management workflows, root-cause analyses, and post-incident evaluations.
• Enhance detection speed while minimizing MTTD and MTTR.
• Contribute to AI-driven operations tools, including automated triage, anomaly detection, and engineer-assist solutions.
• Take ownership of the reliability, scalability, and security of the observability stack.
• Document architecture, runbooks, and operational practices.
• Demonstrated experience in designing and constructing observability platforms for large-scale, production infrastructure settings.
• Extensive hands-on experience with metrics, logging, and distributed tracing tools, such as Prometheus, Grafana, OpenTelemetry, Loki, Thanos/Cortex/Mimir, Elasticsearch/OpenSearch, and Jaeger/Tempo, or their equivalents.
• Background in managing high-volume telemetry pipelines and understanding trade-offs related to cardinality, retention, cost, and query latency.
• Proficient software engineering skills in at least one relevant programming language, such as Go, Python, or Rust.
• Familiarity with Kubernetes and cloud-native infrastructure.
• Strong grasp of SLO/SLI/error-budget methodologies and the design of low-noise alerting.
• Comfortable operating in a dynamic and fast-paced environment.
• Excellent communication skills with the ability to collaborate directly with operations and service delivery teams.
• Preferred: Experience in observability for GPU/HPC or specialized high-performance computing environments.
• Preferred: Expertise in eBPF-based observability tools.
• Preferred: Familiarity with AIOps/ML-driven anomaly detection or automated triage systems.
• Preferred: Background in managed services or MSP contexts.
• Preferred: Contributions to open-source observability initiatives.
• Opportunities for professional development and training.
• Participation in conferences and working groups.
• Company events, happy hours, hackathons, and technology talks.
• Competitive compensation package complemented by a robust benefits plan.
• Flexibility with remote work arrangements.
MAIA
NavAide
Share
Lime
Get handpicked remote jobs straight to your inbox weekly.