
Senior Site Reliability Engineer, DGX Cloud
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in California, +2 more states.
β’ Design, implement, and maintain the operational and reliability components of extensive Kubernetes clusters, emphasizing performance at scale, real-time monitoring, logging, and alerting.
β’ Establish SLOs/SLIs, track error allowances, and enhance reporting processes.
β’ Provide support for services prior to launch through system consulting, development of software tools, platforms, frameworks, capacity management, and launch evaluations.
β’ Ensure the ongoing health of live services by monitoring availability, latency, and overall system performance.
β’ Manage and optimize GPU workloads across AWS, GCP, Azure, OCI, and private cloud environments.
β’ Sustainably scale systems through automation and implement changes that enhance reliability and speed.
β’ Lead the triage process and conduct root-cause analyses of high-severity incidents.
β’ Engage in balanced incident response practices and conduct blameless postmortems.
β’ Participate in an on-call rotation to provide support for production services.
β’ Bachelor's degree in Computer Science or a related technical discipline, or equivalent professional experience.
β’ Over 8 years of experience in managing production services.
β’ Advanced expertise in Kubernetes administration, containerization, and microservices architecture.
β’ Familiarity with infrastructure automation tools, such as Terraform, Ansible, Chef, or Puppet.
β’ Proficient in at least one high-level programming language, like Python or Go.
β’ Comprehensive understanding of Linux operating systems, networking fundamentals (TCP/IP), and cloud security protocols.
β’ Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, and incident management strategies.
β’ Experience in building and maintaining extensive observability stacks (monitoring, logging, tracing) with tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.
β’ Experience managing GPU-accelerated clusters using KubeVirt in a production environment.
β’ Implementation of generative-AI techniques to minimize operational burdens.
β’ Familiarity with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions.
β’ Experience in operating and troubleshooting production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance evaluation.
β’ Equity
β’ Benefits
OnePay
Arista Networks
Octus
Tandem Diabetes Care
Get handpicked remote jobs straight to your inbox weekly.