Senior Site Reliability Engineer, Golang, Kubernetes

Posted 21 hours ago

This is a fully remote position, open to applicants in Kazakhstan, +1 more country.

📋 Description

• Establish a clear definition of reliability for a GPU-accelerated AI platform and develop measurable metrics.

• Take ownership of the service-level indicators and objectives for the K0rdent Observability Framework (KOF).

• Extract significant SLIs from the signals generated by the platform.

• Provide access to SLIs for Platform Administrators via a user-friendly API.

• Define SLIs and SLOs across various environments including Kubernetes, bare-metal hosts, and NVIDIA infrastructure such as BMC, InfiniBand, NVLink, and UFM.

• Design and implement the API that presents SLIs and the reliability state to Platform Administrators and downstream systems.

• Create alerting and error-budget strategies that enhance signal detection while reducing noise.

• Collaborate with infrastructure, storage, and networking teams to ensure appropriate signals are monitored and gathered.

• Identify and troubleshoot reliability and performance challenges within the observability stack and drive their resolution.

• Operate across hybrid, edge, and air-gapped deployments utilizing the Mirantis K0rdent stack.

• Effectively communicate reliability definitions to all teams involved.


⛳️ Requirements

• Minimum of 5 years of experience in SRE, platform reliability, or a similar software/infrastructure position.

• Strong software engineering capabilities in languages such as Go or Python, with hands-on experience in building and managing APIs or services in production environments.

• Proven track record in defining SLIs/SLOs and error budgets for operational production systems.

• Practical experience with observability tools, including metrics, logging, and tracing, particularly with Prometheus/VictoriaMetrics, OpenTelemetry, and Grafana.

• Comprehensive understanding of Kubernetes and the signals it and its workloads generate.

• Excellent written and verbal communication skills when engaging with technical audiences.

• Preferred: Experience in instrumenting or monitoring bare-metal and NVIDIA infrastructure (BMC/Redfish, InfiniBand, NVLink, UFM).

• Preferred: Familiarity with the Mirantis K0rdent stack (K0rdent Enterprise, K0rdent AI, KOF) and Cluster API.

• Preferred: Knowledge of VictoriaMetrics/VictoriaLogs on a large scale.

• Preferred: Demonstrated experience in sovereign or high-security air-gapped environments.


🏝️ Benefits

• Opportunities for professional development and training.

• Participation in conferences and working groups.

• Company outings, happy hours, hackathons, and technical discussions.

• Competitive compensation package accompanied by an attractive benefits plan.

• Flexible remote work arrangements.

People also viewed

HumanIT Digital Consulting21 hours ago

DevOps Engineer – AWS, Kubernetes, Terraform

PT flagPortugal OnlyPart-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Gormat22 hours ago

Cloud DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$191k – $206k/year
ApplyView job
Wizeline22 hours ago

Site Reliability Engineer – Cloud Security, Posture Management

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Stefanini Brasil22 hours ago

DevOps Specialist

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Zigabyte1 day ago

DevSecOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Smartcat1 day ago

Senior DevOps Engineer, Infrastructure

RS flagSerbia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers