Remotery

Service Reliability Engineer

atNVIDIARemoteUS flagTexasFull-timeDevOps & Site Reliability Engineer (SRE)SeniorLead$168k – $333.5k/year

Posted 2 hours ago

This is a fully remote position, open to applicants in Texas.

📋 Description

• Operate within a 24/7 follow-the-sun support model across various continents.

• Manage a 4-day, 10-hour work schedule, including either Saturday or Sunday, with flexible early or late shifts.

• Oversee and manage extensive production GPU and Kubernetes environments.

• Proactively detect, prevent, and respond to incidents.

• Analyze logs, metrics, and system behavior to diagnose issues and implement solutions.

• Develop automated support routines that are predictive in nature.

• Enhance automation through auto-healing and automated break-fix solutions.

• Perform systems administration, network administration, and security monitoring tasks.

• Collaborate with domain experts and service owners to resolve complex issues.

• Continuously enhance service quality and operational processes based on incident feedback.

• Effectively coordinate across teams during incident resolution.

• Provide customer-focused support throughout client interactions.


⛳️ Requirements

• Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management.

• Familiarity with GPU hardware and high-performance computing environments.

• Proficiency in Grafana, OpenTelemetry, PagerDuty, and JIRA.

• Experience with AWS, Azure, GCP, or OCI is a plus; strong preference for on-premises expertise.

• Ability to work effectively with multifunctional teams.

• 8+ years of experience coordinating large-scale production systems.

• Over 3 years of experience in high-availability Internet, Cloud, or Data Center environments.

• Bachelor’s degree in Computer Science, Engineering, Physics, Mathematics, or equivalent experience.

• Expert-level Linux system administration skills.

• Experience in automation using Ansible and/or Python.

• Strong expertise in shell scripting, DNS, DHCP, storage systems, and core networking.

• Proven experience in maintaining large-scale bare-metal infrastructure.

• Excellent skills in partnership, documentation, and mentoring.

• Experience with scripting languages, particularly Python.

• Experience running virtual machines under community-supported or commercial hypervisors.

• Knowledge of application containers and container orchestration systems.

• Basic understanding of Git.

• Ability to master and maintain complex environments.


🏝️ Benefits

• Equity.

• Benefits.

People also viewed

Endava2 hours ago

Senior DevOps Engineer, Dynatrace

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Jones Lang LaSalle Americas, Inc.2 hours ago

Reliability Engineer

US flagIllinois, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $120k/year
ApplyView job
Entarian2 hours ago

DevSecOps Engineer – Mid

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
BeyondTrust2 hours ago

Senior DevOps Engineer

CA flagCanada OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ontrac Solutions2 hours ago

Site Reliability Engineer

PK flagPakistan OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Nagarro2 hours ago

Senior Site Reliability Engineer, AWS Cloud

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers