
Observability Platform Engineer
Posted Aug 31

Posted Aug 31
This is a fully remote position, open to applicants in Kazakhstan, +1 more country.
• Design, construct, and manage components of an observability platform for metrics, logging, distributed tracing, and alerting.
• Develop telemetry pipelines capable of handling high-cardinality, high-volume data from extensive infrastructure fleets.
• Enhance telemetry cost efficiency, data retention, and query performance.
• Establish and execute SLO/SLI frameworks along with alerting strategies.
• Collaborate with service delivery and operations teams to address incident response requirements.
• Incorporate observability tools into incident management workflows, root-cause analysis, and post-incident evaluations.
• Increase detection speed while minimizing MTTD and MTTR.
• Participate in the development of AI-assisted operational tools, such as automated triage, anomaly detection, and engineer-assist instruments.
• Ensure the reliability, scalability, and security of the observability stack.
• Document architecture, runbooks, and operational procedures.
• Demonstrated experience in designing and constructing observability platforms for large-scale, production infrastructure environments.
• Proficient hands-on experience with metrics, logging, and distributed tracing tools, including Prometheus, Grafana, OpenTelemetry, Loki, Thanos/Cortex/Mimir, Elasticsearch/OpenSearch, Jaeger/Tempo, or similar technologies.
• Familiarity with high-volume telemetry pipelines and the associated trade-offs regarding cardinality, retention, cost, and query latency.
• Strong software engineering capabilities in at least one widely-used programming language, such as Go, Python, or Rust.
• Experience with Kubernetes and cloud-native infrastructure.
• Comprehensive understanding of SLO/SLI/error-budget methodologies and low-noise alerting design principles.
• Ability to thrive in a fast-paced environment.
• Excellent communication skills and a capacity to collaborate directly with operations/service delivery teams.
• Background in building observability solutions for GPU/HPC or other specialized, high-performance computing environments.
• Experience with eBPF-based observability tools.
• Knowledge of AIOps/ML-driven anomaly detection or automated triage systems.
• Experience in a managed services or MSP setting.
• Contributions to open-source observability initiatives.
• Collaborate with a well-established leader in the cloud infrastructure sector based in Silicon Valley.
• Engage with highly passionate, talented, and enthusiastic colleagues, assisting Fortune 500 and Global 2000 clients in implementing next-generation cloud technologies.
• Be part of groundbreaking, open-source innovation.
• Opportunities for professional development and training.
• Attend conferences and participate in working groups.
• Enjoy company outings, happy hours, hackathons, and tech talks.
• Receive a competitive compensation package along with a robust benefits plan.
• Benefit from a remote work arrangement.
MAIA
NavAide
Share
Lime
Get handpicked remote jobs straight to your inbox weekly.