
Senior Site Reliability Engineer, SRE
Posted 23 hours ago

Posted 23 hours ago
This is a fully remote position, open to applicants in California.
• Design, implement, and maintain scalable and dependable systems.
• Establish monitoring tools and develop incident response plans to detect and address issues, as well as implement preventative strategies.
• Create and manage scripts and automation tools for deployment, monitoring, and system health assessments.
• Evaluate system capacity and performance metrics to anticipate future requirements and implement scaling solutions.
• Collaborate closely with development teams to enhance product reliability and streamline deployment processes.
• Generate and maintain documentation for system architecture, procedures, and incident reports.
• Engage in on-call rotations to provide round-the-clock support for critical systems.
• Apply and uphold security best practices across systems while ensuring compliance with industry standards.
• Independently deploy infrastructure changes using Infrastructure as Code methodologies.
• Identify reliability risks and propose enhancements.
• Enhance dashboards, alerts, and operational runbooks.
• Refine deployment pipelines, infrastructure provisioning, and self-service capabilities.
• Optimize infrastructure utilization and cloud costs while maintaining reliability.
• Propel automation initiatives to minimize operational toil and enhance deployment reliability.
• Lead cross-functional projects aimed at improving availability, scalability, and operational efficiency.
• Provide guidance on site reliability for new product developments and platform advancements.
• Mentor junior engineers and promote a culture of development.
• Bachelor's degree in Computer Science, Engineering, or a related discipline, or equivalent professional experience.
• Over 5 years of experience in Site Reliability Engineering, DevOps, or a similar role.
• Expertise in a scripting or programming language, such as Python or Go.
• Experience with cloud services including AWS or Azure.
• Proficient in containerization technologies like Docker, Kubernetes, or ECS.
• Skilled in configuration management tools such as Terraform, Atlantis, or Terragrunt.
• Familiarity with CI/CD tools like GitHub Actions, AWS CodeBuild, or CircleCI.
• Experience with monitoring solutions such as Grafana or Datadog.
• Knowledge of database technologies including RDS, Aurora, or PostgreSQL.
• Experience with large-scale distributed systems and microservices architectures.
• Strong analytical and problem-solving skills, with the capability to troubleshoot complex systems.
• Excellent verbal and written communication skills, with the ability to collaborate effectively across teams.
• Eligibility to obtain a security clearance.
• Active TS/SCI clearance is preferred.
• Must be eligible to secure required U.S. Department of State ITAR authorizations.
• Bonus and equity opportunities.
• 100 USD monthly wellness benefit to promote your health and well-being.
• $300 USD home office setup stipend, available for use within your first six months.
• $75 USD monthly home internet allowance, paid directly through your regular paycheck.
• Individual Development Fund to support training, learning, and professional growth.
• Employee Recognition Program to celebrate contributions and achievements within the team.
• Employee Referral Program.
• And much more!
Slate Auto
Funding Xchange
Leidos
OpenRouter
Get handpicked remote jobs straight to your inbox weekly.