Senior Site Reliability Engineer, DGX Cloud

Posted 2 days ago

This is a fully remote position, open to applicants in California, +2 more states.

πŸ“‹ Description

β€’ Design, implement, and maintain the operational and reliability components of extensive Kubernetes clusters, emphasizing performance at scale, real-time monitoring, logging, and alerting.

β€’ Establish SLOs/SLIs, track error allowances, and enhance reporting processes.

β€’ Provide support for services prior to launch through system consulting, development of software tools, platforms, frameworks, capacity management, and launch evaluations.

β€’ Ensure the ongoing health of live services by monitoring availability, latency, and overall system performance.

β€’ Manage and optimize GPU workloads across AWS, GCP, Azure, OCI, and private cloud environments.

β€’ Sustainably scale systems through automation and implement changes that enhance reliability and speed.

β€’ Lead the triage process and conduct root-cause analyses of high-severity incidents.

β€’ Engage in balanced incident response practices and conduct blameless postmortems.

β€’ Participate in an on-call rotation to provide support for production services.


⛳️ Requirements

β€’ Bachelor's degree in Computer Science or a related technical discipline, or equivalent professional experience.

β€’ Over 8 years of experience in managing production services.

β€’ Advanced expertise in Kubernetes administration, containerization, and microservices architecture.

β€’ Familiarity with infrastructure automation tools, such as Terraform, Ansible, Chef, or Puppet.

β€’ Proficient in at least one high-level programming language, like Python or Go.

β€’ Comprehensive understanding of Linux operating systems, networking fundamentals (TCP/IP), and cloud security protocols.

β€’ Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, and incident management strategies.

β€’ Experience in building and maintaining extensive observability stacks (monitoring, logging, tracing) with tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.

β€’ Experience managing GPU-accelerated clusters using KubeVirt in a production environment.

β€’ Implementation of generative-AI techniques to minimize operational burdens.

β€’ Familiarity with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions.

β€’ Experience in operating and troubleshooting production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance evaluation.


🏝️ Benefits

β€’ Equity

β€’ Benefits

People also viewed

OnePay12 hours ago

Site Reliability Engineering Lead

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$250k – $280k/year
ApplyView job
Arista Networks13 hours ago

FedRAMP Site Reliability Engineer – CloudVision

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$101k – $161k/year
ApplyView job
Octus13 hours ago

Lead DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$175k – $225k/year
ApplyView job
Tandem Diabetes Care14 hours ago

Principal Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$165k – $185k/year
ApplyView job
TechnologyAdvice16 hours ago

Senior DevOps Engineer – Contract

IN flagIndia OnlyFreelanceDevOps & Site Reliability Engineer (SRE)β‚Ή1,500 – β‚Ή2,000/hour
ApplyView job
Bixal17 hours ago

Director of DevSecOps

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$165k – $195k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers