
Head of Cloud Infrastructure
Posted 22 hours ago

Posted 22 hours ago
This is a fully remote position, open to applicants in United States.
• Develop and implement a prioritized roadmap for modernizing cloud infrastructure, focusing on AWS, Kubernetes, Terraform, and the wider platform ecosystem.
• Lead a small, impactful global team across cloud platforms, Site Reliability Engineering (SRE), and DevOps.
• Provide guidance, mentorship, and practical technical insights.
• Ensure the delivery of dependable paved roads, reusable infrastructure capabilities, self-service workflows, and accelerated feedback loops.
• Establish and enhance reliability practices, including service ownership, Service Level Indicators (SLIs), Service Level Objectives (SLOs), incident management, error budgets, observability, and follow-up actions post-incident.
• Take ownership of the disaster recovery strategy and readiness, which includes setting recovery objectives, conducting tests, implementing automation, and closing resilience gaps.
• Design a sustainable 24/7 operational model that encompasses on-call responsibilities, escalation procedures, incident leadership, and collaboration within a global team.
• Collaborate with the Chief Architect and engineering leaders to establish platform priorities and clarify ownership boundaries.
• Assist application and data teams in effectively managing workloads on shared infrastructure.
• Define performance baselines and assess metrics related to delivery performance, developer experience, platform adoption, reliability, operational health, and cloud unit economics.
• Enhance cloud efficiency through cost attribution, capacity planning, Kubernetes resource utilization, automation, and quantifiable savings.
• Report directly to the Chief Architect and collaborate with Security, Enterprise IT, and engineering leaders.
• A minimum of 6 years of experience in relevant fields such as infrastructure, SRE, platform engineering, or developer productivity.
• Proficient hands-on experience with AWS, Kubernetes, and Terraform.
• Experience in leading a team of approximately 3 to 10 engineers in infrastructure, platform, SRE, or DevOps roles.
• Proven experience in building or significantly transforming a platform function.
• Background in managing mature, business-critical production infrastructure.
• Experience in scaling a data-intensive SaaS, analytics, or similarly complex technology organization during rapid growth phases.
• Proficient in owning formal reliability practices, including SLOs, incident management, production operations, and disaster recovery.
• Capability to design a sustainable on-call model.
• Experience in enhancing developer velocity or platform experience via self-service capabilities, standardized workflows, internal platform adoption, or measurable reductions in engineering friction.
• Ability to translate technical strategy into a concrete roadmap and make informed trade-offs.
• Skill in influencing the Chief Architect, engineering leaders, and partner teams through effective communication and credible technical judgment.
• Proven track record of delivering measurable improvements in cloud costs or infrastructure efficiency.
• Understanding of smart architecture principles across services for optimizing technology costs.
• Bachelor's degree or equivalent practical experience.
• Availability to work US hours with sufficient overlap to lead and collaborate with a globally distributed team.
• Flexible work hours.
• Flexible vacation policy.
• Generous 401K matching.
• Parental leave.
• Team-building events.
• Wellness budget.
• Learning reimbursement.
• Equity options.
• Performance bonus.
Siteup
Leidos
geekoffice
Get handpicked remote jobs straight to your inbox weekly.