
Senior Site Reliability Engineer Lead
Posted 23 hours ago

Posted 23 hours ago
This is a fully remote position, open to applicants in Massachusetts.
• Architect, develop, test, and distribute modifications to software, services, and tools that support HIVE hardware platforms.
• Design and implement improvements to the HIVE observability infrastructure.
• Cultivate subject matter expertise in HIVE components and provide mentorship to the team.
• Identify and adopt automation best practices for current products and processes.
• Collaborate with support, operations, and engineering teams to investigate and resolve complex issues.
• Participate in on-call rotations and assist in the restoration and repair of service-impacting problems.
• Over 8 years of professional experience.
• A Bachelor's degree in Computer Science or a related field.
• Extensive knowledge of the underlying hardware and best practices for enabling advanced hardware features in Linux, particularly for ARM architecture like NVIDIA Grace.
• Advanced expertise with the Linux kernel, operating systems, KVM/QEMU, and nested virtualization.
• Proven experience in designing, developing, and deploying software and infrastructure at scale.
• Advanced-level experience in a DevOps, Development, or SysAdmin capacity.
• Experience with large-scale distributed systems.
• Familiarity with SaltStack and Ansible for managing infrastructure at scale.
• Healthcare coverage.
• 401K savings plan.
• Company holidays.
• Paid time off (PTO) for vacations.
• Sick leave.
• Parental leave.
• Employee assistance program.
• Support for mental wellness.
• Financial wellness assistance.
• Annual bonuses or incentives.
• Equity awards.
• Employee Stock Purchase Plan (ESPP).
• Flexible work arrangements through FlexBase, allowing for remote work, in-office presence, or a combination of both.
Adapty.io
Oddball
General Dynamics Information Technology
StarTekk
Get handpicked remote jobs straight to your inbox weekly.