Observability Platform Engineer

Posted Aug 31

This is a fully remote position, open to applicants in Kazakhstan, +1 more country.

📋 Description

• Design, construct, and manage components of an observability platform for metrics, logging, distributed tracing, and alerting.

• Develop telemetry pipelines capable of handling high-cardinality, high-volume data from extensive infrastructure fleets.

• Enhance telemetry cost efficiency, data retention, and query performance.

• Establish and execute SLO/SLI frameworks along with alerting strategies.

• Collaborate with service delivery and operations teams to address incident response requirements.

• Incorporate observability tools into incident management workflows, root-cause analysis, and post-incident evaluations.

• Increase detection speed while minimizing MTTD and MTTR.

• Participate in the development of AI-assisted operational tools, such as automated triage, anomaly detection, and engineer-assist instruments.

• Ensure the reliability, scalability, and security of the observability stack.

• Document architecture, runbooks, and operational procedures.


⛳️ Requirements

• Demonstrated experience in designing and constructing observability platforms for large-scale, production infrastructure environments.

• Proficient hands-on experience with metrics, logging, and distributed tracing tools, including Prometheus, Grafana, OpenTelemetry, Loki, Thanos/Cortex/Mimir, Elasticsearch/OpenSearch, Jaeger/Tempo, or similar technologies.

• Familiarity with high-volume telemetry pipelines and the associated trade-offs regarding cardinality, retention, cost, and query latency.

• Strong software engineering capabilities in at least one widely-used programming language, such as Go, Python, or Rust.

• Experience with Kubernetes and cloud-native infrastructure.

• Comprehensive understanding of SLO/SLI/error-budget methodologies and low-noise alerting design principles.

• Ability to thrive in a fast-paced environment.

• Excellent communication skills and a capacity to collaborate directly with operations/service delivery teams.

• Background in building observability solutions for GPU/HPC or other specialized, high-performance computing environments.

• Experience with eBPF-based observability tools.

• Knowledge of AIOps/ML-driven anomaly detection or automated triage systems.

• Experience in a managed services or MSP setting.

• Contributions to open-source observability initiatives.


🏝️ Benefits

• Collaborate with a well-established leader in the cloud infrastructure sector based in Silicon Valley.

• Engage with highly passionate, talented, and enthusiastic colleagues, assisting Fortune 500 and Global 2000 clients in implementing next-generation cloud technologies.

• Be part of groundbreaking, open-source innovation.

• Opportunities for professional development and training.

• Attend conferences and participate in working groups.

• Enjoy company outings, happy hours, hackathons, and tech talks.

• Receive a competitive compensation package along with a robust benefits plan.

• Benefit from a remote work arrangement.

People also viewed

MAIA14 hours ago

Senior DevOps / Platform Engineer, AI Infrastructure

DE flagGermany OnlyFull-timePlatform Engineer€75k – €85k/year
ApplyView job
NavAide17 hours ago

Power Platform Developer, DOD Clearance Required

US flagCalifornia OnlyFull-timePlatform Engineer$98k – $140k/year
ApplyView job
Share22 hours ago

Platform Engineer

ES flagSpain, +6 more countriesFull-timePlatform Engineer
ApplyView job
Lime22 hours ago

Staff Software Engineer, Platform Engineering

CA flagCanada OnlyFull-timePlatform EngineerC$172k – C$237k/year
ApplyView job
Convoso23 hours ago

Senior Platform Engineer

US flagCalifornia OnlyFull-timePlatform Engineer$190k – $210k/year
ApplyView job
Cortex1 day ago

Platform Engineer – SRE, Mid-Level

BR flagBrazil OnlyFull-timePlatform Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers