Senior Staff Site Reliability Engineer – Compute Core Engineering

Posted Aug 28

This is a fully remote position, open to applicants in California.

📋 Description

• Spearhead initiatives to reshape the architecture of the IT Compute Core Team and establish new service offerings across both on-premises and cloud environments.

• Architect, scale, and implement core infrastructure services such as DNS, NTP/PTP, DHCP, and LDAP.

• Construct infrastructure that ensures performance and reliability on a global scale, which includes automation, monitoring, high availability, capacity planning, and lifecycle management.

• Establish and execute service-efficiency metrics, driving enhancements through software and hardware optimizations, including SR-IOV and DPU.

• Utilize eBPF and XDP for enhanced observability and DDoS mitigation.

• Gather and evaluate system data for capacity and planning; analyze capacity information and formulate enterprise-wide systems strategies.

• Oversee the implementation of infrastructure modifications in collaboration with management personnel.

• Create and sustain tools for the collection, analysis, and visualization of data for reporting, alerting, and monitoring purposes.

• Partner with NVIDIA leadership, senior engineers, program managers, and product managers to design IT products and services that align with customer requirements.


⛳️ Requirements

• Bachelor’s degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.

• A minimum of 12 years of demonstrated experience in compute platform engineering with an emphasis on automation.

• Experience in designing and deploying containerization architectures and distributed systems infrastructure.

• Proven track record in evaluating application architectures and identifying potential for containerization.

• Strong analytical capabilities with the ability to define and monitor key performance metrics.

• Experience in developing tools for data analysis and performance profiling.

• Proficient in Terraform and configuration management tools.

• Strong command of Go and/or Python programming languages.

• Proficient in Linux OS with knowledge of kernel internals.

• Experience managing extensive environments comprising bare-metal build infrastructure.

• Familiarity with network protocols and architectures, including VLAN, VXLAN, SDN, BGP, and Anycast.

• In-depth knowledge of infrastructure components such as DNS, LDAP, and security tools.

• Hands-on experience with containers and their deployment.

• Experience in deploying and managing DNS and LDAP services at scale.

• Solid understanding of microservices architecture, infrastructure as code (IaC), and configuration management tools.


🏝️ Benefits

• Equity

• Benefits

People also viewed

knowmad mood1 day ago

Consultor/a DevSecOps – AWS

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
RealTime eClinical Solutions2 days ago

Principal DevOps Architect

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$155k – $195k/year
ApplyView job
Koniag Government Services2 days ago

DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Koniag Government Services2 days ago

Senior AWS DevOps Engineer – AWS, Kubernetes, HCP, CI/CD, Observability, AI-focus

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ASRC Federal2 days ago

Senior DevOps Administrator – Supporting NASA

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Nelios2 days ago

DevOps Engineer, Cloud Infrastructure

GR flagGreece OnlyPart-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers