
Senior Site Reliability Engineer, Golang, Kubernetes
Posted 21 hours ago

Posted 21 hours ago
This is a fully remote position, open to applicants in Kazakhstan, +1 more country.
• Establish a clear definition of reliability for a GPU-accelerated AI platform and develop measurable metrics.
• Take ownership of the service-level indicators and objectives for the K0rdent Observability Framework (KOF).
• Extract significant SLIs from the signals generated by the platform.
• Provide access to SLIs for Platform Administrators via a user-friendly API.
• Define SLIs and SLOs across various environments including Kubernetes, bare-metal hosts, and NVIDIA infrastructure such as BMC, InfiniBand, NVLink, and UFM.
• Design and implement the API that presents SLIs and the reliability state to Platform Administrators and downstream systems.
• Create alerting and error-budget strategies that enhance signal detection while reducing noise.
• Collaborate with infrastructure, storage, and networking teams to ensure appropriate signals are monitored and gathered.
• Identify and troubleshoot reliability and performance challenges within the observability stack and drive their resolution.
• Operate across hybrid, edge, and air-gapped deployments utilizing the Mirantis K0rdent stack.
• Effectively communicate reliability definitions to all teams involved.
• Minimum of 5 years of experience in SRE, platform reliability, or a similar software/infrastructure position.
• Strong software engineering capabilities in languages such as Go or Python, with hands-on experience in building and managing APIs or services in production environments.
• Proven track record in defining SLIs/SLOs and error budgets for operational production systems.
• Practical experience with observability tools, including metrics, logging, and tracing, particularly with Prometheus/VictoriaMetrics, OpenTelemetry, and Grafana.
• Comprehensive understanding of Kubernetes and the signals it and its workloads generate.
• Excellent written and verbal communication skills when engaging with technical audiences.
• Preferred: Experience in instrumenting or monitoring bare-metal and NVIDIA infrastructure (BMC/Redfish, InfiniBand, NVLink, UFM).
• Preferred: Familiarity with the Mirantis K0rdent stack (K0rdent Enterprise, K0rdent AI, KOF) and Cluster API.
• Preferred: Knowledge of VictoriaMetrics/VictoriaLogs on a large scale.
• Preferred: Demonstrated experience in sovereign or high-security air-gapped environments.
• Opportunities for professional development and training.
• Participation in conferences and working groups.
• Company outings, happy hours, hackathons, and technical discussions.
• Competitive compensation package accompanied by an attractive benefits plan.
• Flexible remote work arrangements.
HumanIT Digital Consulting
Gormat
Wizeline
Stefanini Brasil
Get handpicked remote jobs straight to your inbox weekly.