
Senior Staff Site Reliability Engineer – Compute Core Engineering
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in California.
• Spearhead initiatives to reshape the architecture of the IT Compute Core Team and establish new service offerings across both on-premises and cloud environments.
• Architect, scale, and implement core infrastructure services such as DNS, NTP/PTP, DHCP, and LDAP.
• Construct infrastructure that ensures performance and reliability on a global scale, which includes automation, monitoring, high availability, capacity planning, and lifecycle management.
• Establish and execute service-efficiency metrics, driving enhancements through software and hardware optimizations, including SR-IOV and DPU.
• Utilize eBPF and XDP for enhanced observability and DDoS mitigation.
• Gather and evaluate system data for capacity and planning; analyze capacity information and formulate enterprise-wide systems strategies.
• Oversee the implementation of infrastructure modifications in collaboration with management personnel.
• Create and sustain tools for the collection, analysis, and visualization of data for reporting, alerting, and monitoring purposes.
• Partner with NVIDIA leadership, senior engineers, program managers, and product managers to design IT products and services that align with customer requirements.
• Bachelor’s degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.
• A minimum of 12 years of demonstrated experience in compute platform engineering with an emphasis on automation.
• Experience in designing and deploying containerization architectures and distributed systems infrastructure.
• Proven track record in evaluating application architectures and identifying potential for containerization.
• Strong analytical capabilities with the ability to define and monitor key performance metrics.
• Experience in developing tools for data analysis and performance profiling.
• Proficient in Terraform and configuration management tools.
• Strong command of Go and/or Python programming languages.
• Proficient in Linux OS with knowledge of kernel internals.
• Experience managing extensive environments comprising bare-metal build infrastructure.
• Familiarity with network protocols and architectures, including VLAN, VXLAN, SDN, BGP, and Anycast.
• In-depth knowledge of infrastructure components such as DNS, LDAP, and security tools.
• Hands-on experience with containers and their deployment.
• Experience in deploying and managing DNS and LDAP services at scale.
• Solid understanding of microservices architecture, infrastructure as code (IaC), and configuration management tools.
• Equity
• Benefits
knowmad mood
RealTime eClinical Solutions
Koniag Government Services
Koniag Government Services
Get handpicked remote jobs straight to your inbox weekly.