Remotery

Senior Site Reliability Engineer, DGX Cloud

Posted 1 day ago

This is a fully remote position, open to applicants in Switzerland.

📋 Description

• Design, implement, and maintain the operational and reliability features of large-scale Kubernetes clusters, with an emphasis on performance at scale, real-time monitoring, logging, and alerting.

• Establish SLOs/SLIs, track error budgets, and enhance reporting processes.

• Assist services prior to launch through system creation consulting, software tools, platforms and frameworks, capacity management, and launch evaluations.

• Ensure the ongoing health of live services by assessing and monitoring availability, latency, and overall system performance.

• Manage and optimize GPU workloads across various platforms including AWS, GCP, Azure, OCI, and private clouds.

• Sustainably scale systems through automation and implement changes that enhance reliability and speed.

• Lead the triage process and conduct root-cause analysis for high-severity incidents.

• Adopt a balanced approach to incident response and conduct blameless postmortems.

• Engage in an on-call rotation to provide support for production services.


⛳️ Requirements

• Bachelor’s degree in Computer Science or a related technical discipline, or an equivalent combination of experience.

• Over 10 years of experience managing production services.

• Advanced expertise in Kubernetes administration, containerization, and microservices architecture.

• Familiarity with infrastructure automation tools such as Terraform, Ansible, Chef, or Puppet.

• Proficient in at least one high-level programming language, such as Python or Go.

• Comprehensive understanding of Linux operating systems, networking basics (TCP/IP), and cloud security protocols.

• Strong grasp of SRE principles, including SLOs, SLIs, error budgets, and incident management.

• Experience in building and managing observability stacks for monitoring, logging, and tracing with tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, or Splunk.

• Willingness to participate in an on-call rotation.

• Experience in operating production services is essential; equivalent experience may be considered in place of the specified Bachelor’s degree.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible work hours and remote work options.

• Opportunities for professional development and training.

• Generous paid time off and holiday policies.

People also viewed

CWILL14 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3715 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT16 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group16 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo16 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch16 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers