
Observability Platform Engineer
Posted Aug 31

Posted Aug 31
This is a fully remote position, open to applicants in Kazakhstan, +1 more country.
• Design, develop, and manage components of an observability platform for metrics, logging, distributed tracing, and alerting.
• Construct telemetry pipelines that handle high-cardinality and high-volume data from extensive infrastructure fleets.
• Enhance telemetry pipelines focusing on cost efficiency, data retention, and query performance.
• Establish and execute SLO/SLI frameworks along with alerting strategies.
• Collaborate with service delivery and operations teams to identify incident-response requirements.
• Integrate observability tools with incident management workflows, which include root-cause analysis and post-incident review information.
• Accelerate detection speed and lower MTTD and MTTR across the platform.
• Contribute to the strategic roadmap for AI-assisted operational tools, including automated triage and anomaly detection.
• Take ownership of the reliability, scalability, and security of the observability stack.
• Document the architecture, runbooks, and operational procedures.
• Enable operations teams to swiftly diagnose incidents with automatically surfaced data.
• Scale the observability platform in line with infrastructure expansion while managing costs and performance.
• Demonstrated expertise in designing and constructing observability platforms for large-scale production infrastructure environments.
• Extensive practical experience with metrics, logging, and distributed tracing tools such as Prometheus, Grafana, OpenTelemetry, Loki, Thanos/Cortex/Mimir, Elasticsearch/OpenSearch, and Jaeger/Tempo, or similar alternatives.
• Proficiency with high-volume telemetry pipelines and understanding of trade-offs related to cardinality, retention, cost, and query latency.
• Strong software engineering capabilities in at least one language commonly utilized in this domain, such as Go, Python, or Rust.
• Experience with Kubernetes and cloud-native infrastructure.
• Comprehensive understanding of SLO/SLI/error-budget practices and designing low-noise alerting systems.
• Comfortable operating in a dynamic environment where the platform is developed alongside the infrastructure it oversees.
• Excellent communication skills and the ability to collaborate directly with operations and service delivery teams.
• Experience in building observability solutions for GPU/HPC infrastructure or other specialized high-performance computing environments (preferred).
• Familiarity with eBPF-based observability tools (preferred).
• Awareness of AIOps/ML-based anomaly detection or automated triage systems (preferred).
• Experience working in a managed services or MSP context (preferred).
• Contributions to open-source observability initiatives (preferred).
• Collaborate with a prominent Silicon Valley leader in the cloud infrastructure sector.
• Work alongside exceptionally passionate, skilled, and engaging colleagues.
• Assist Fortune 500 and Global 2000 clients in implementing next-generation cloud technologies.
• Participate in cutting-edge, open-source innovation.
• Opportunities for professional development and training.
• Attend conferences and working groups.
• Enjoy company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package complemented by a robust benefits plan.
MAIA
NavAide
Share
Lime
Get handpicked remote jobs straight to your inbox weekly.