
Senior Site Reliability Engineer, DGX Cloud
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Switzerland.
• Design, implement, and maintain the operational and reliability features of large-scale Kubernetes clusters, with an emphasis on performance at scale, real-time monitoring, logging, and alerting.
• Establish SLOs/SLIs, track error budgets, and enhance reporting processes.
• Assist services prior to launch through system creation consulting, software tools, platforms and frameworks, capacity management, and launch evaluations.
• Ensure the ongoing health of live services by assessing and monitoring availability, latency, and overall system performance.
• Manage and optimize GPU workloads across various platforms including AWS, GCP, Azure, OCI, and private clouds.
• Sustainably scale systems through automation and implement changes that enhance reliability and speed.
• Lead the triage process and conduct root-cause analysis for high-severity incidents.
• Adopt a balanced approach to incident response and conduct blameless postmortems.
• Engage in an on-call rotation to provide support for production services.
• Bachelor’s degree in Computer Science or a related technical discipline, or an equivalent combination of experience.
• Over 10 years of experience managing production services.
• Advanced expertise in Kubernetes administration, containerization, and microservices architecture.
• Familiarity with infrastructure automation tools such as Terraform, Ansible, Chef, or Puppet.
• Proficient in at least one high-level programming language, such as Python or Go.
• Comprehensive understanding of Linux operating systems, networking basics (TCP/IP), and cloud security protocols.
• Strong grasp of SRE principles, including SLOs, SLIs, error budgets, and incident management.
• Experience in building and managing observability stacks for monitoring, logging, and tracing with tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, or Splunk.
• Willingness to participate in an on-call rotation.
• Experience in operating production services is essential; equivalent experience may be considered in place of the specified Bachelor’s degree.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and remote work options.
• Opportunities for professional development and training.
• Generous paid time off and holiday policies.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.