
Senior Site Reliability Engineer
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in United States.
• Take ownership of the overall reliability, performance, and resilience of Garner’s AWS and Kubernetes cloud environments, including AI/ML workloads.
• Establish, assess, and maintain service level objectives (SLOs) across essential services.
• Participate in the on-call rotation and spearhead incident response efforts.
• Conduct root cause analyses, implement corrective measures, and oversee thorough infrastructure-change evaluations.
• Develop and sustain monitoring, alerting, and observability systems.
• Convert scaling needs into automated, modular Terraform infrastructure-as-code outputs.
• Execute improvements in cloud cost-efficiency and performance throughout the stack.
• Minimize operational toil and technical debt by utilizing AI tools and automation.
• Establish and uphold deployment and observability standards for engineering teams.
• Convey cloud and reliability principles to both technical and non-technical audiences.
• Ensure that infrastructure and operations adhere to security and HIPAA compliance requirements.
• A minimum of 4 years of practical experience managing production cloud infrastructure at scale within an SRE, DevOps, or platform engineering capacity.
• Extensive knowledge of Kubernetes and Terraform in a cloud-first setting.
• AWS experience is preferred.
• Solid background in production observability, which includes defining SLOs, developing monitoring and alerting systems, leading incident response, and carrying out blameless post-incident reviews.
• Strong foundation in software engineering, particularly in Python or Go, as it relates to infrastructure automation.
• Proven experience in enhancing cloud cost-efficiency and performance across compute, storage, and networking resources.
• Proficiency with AI tools like Claude applied within engineering and operations workflows, or a strong desire to develop this expertise swiftly.
• Must not require employer sponsorship or the transfer of an employment visa.
• Previous experience supporting AI/ML or data-intensive workloads in a production environment is advantageous.
• Experience in a security-focused or regulated environment such as HIPAA or SOC 2 is a plus.
• Familiarity with Kubernetes APIs is a plus.
• Equity incentive plan.
• Flexible PTO.
• Medical plan options.
• Dental plan options.
• Vision plan options.
• 401(k) with company match.
• Flexible spending accounts.
• Teladoc Health.
Rimutee
CACI International Inc
CACI International Inc
Rimutee
Get handpicked remote jobs straight to your inbox weekly.