Observability Platform Engineer

Posted Aug 21

This is a fully remote position, open to applicants in Kazakhstan.

📋 Description

• Design, develop, and manage components for metrics, logging, distributed tracing, and alerting platforms.

• Create high-volume, high-cardinality telemetry pipelines for extensive infrastructure fleets.

• Establish and execute SLO/SLI frameworks along with alerting strategies.

• Collaborate with service delivery and operations teams to enhance incident-focused observability.

• Integrate observability tools with incident management workflows, root-cause analyses, and post-incident evaluations.

• Enhance detection speed while minimizing MTTD and MTTR.

• Contribute to AI-driven operations tools, including automated triage, anomaly detection, and engineer-assist solutions.

• Take ownership of the reliability, scalability, and security of the observability stack.

• Document architecture, runbooks, and operational practices.


⛳️ Requirements

• Demonstrated experience in designing and constructing observability platforms for large-scale, production infrastructure settings.

• Extensive hands-on experience with metrics, logging, and distributed tracing tools, such as Prometheus, Grafana, OpenTelemetry, Loki, Thanos/Cortex/Mimir, Elasticsearch/OpenSearch, and Jaeger/Tempo, or their equivalents.

• Background in managing high-volume telemetry pipelines and understanding trade-offs related to cardinality, retention, cost, and query latency.

• Proficient software engineering skills in at least one relevant programming language, such as Go, Python, or Rust.

• Familiarity with Kubernetes and cloud-native infrastructure.

• Strong grasp of SLO/SLI/error-budget methodologies and the design of low-noise alerting.

• Comfortable operating in a dynamic and fast-paced environment.

• Excellent communication skills with the ability to collaborate directly with operations and service delivery teams.

• Preferred: Experience in observability for GPU/HPC or specialized high-performance computing environments.

• Preferred: Expertise in eBPF-based observability tools.

• Preferred: Familiarity with AIOps/ML-driven anomaly detection or automated triage systems.

• Preferred: Background in managed services or MSP contexts.

• Preferred: Contributions to open-source observability initiatives.


🏝️ Benefits

• Opportunities for professional development and training.

• Participation in conferences and working groups.

• Company events, happy hours, hackathons, and technology talks.

• Competitive compensation package complemented by a robust benefits plan.

• Flexibility with remote work arrangements.

People also viewed

MAIA12 hours ago

Senior DevOps / Platform Engineer, AI Infrastructure

DE flagGermany OnlyFull-timePlatform Engineer€75k – €85k/year
ApplyView job
NavAide15 hours ago

Power Platform Developer, DOD Clearance Required

US flagCalifornia OnlyFull-timePlatform Engineer$98k – $140k/year
ApplyView job
Share20 hours ago

Platform Engineer

ES flagSpain, +6 more countriesFull-timePlatform Engineer
ApplyView job
Lime20 hours ago

Staff Software Engineer, Platform Engineering

CA flagCanada OnlyFull-timePlatform EngineerC$172k – C$237k/year
ApplyView job
Convoso21 hours ago

Senior Platform Engineer

US flagCalifornia OnlyFull-timePlatform Engineer$190k – $210k/year
ApplyView job
Cortex1 day ago

Platform Engineer – SRE, Mid-Level

BR flagBrazil OnlyFull-timePlatform Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers