
Manager, Infrastructure, Cloud Operations
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Arizona.
• Manage the operational health of RTA’s cloud infrastructure along with the supporting teams and systems.
• Enhance reliability, availability, scalability, observability, capacity planning, incident response, database resilience, security posture, and operational readiness.
• Direct Cloud Engineering, Infrastructure, DevOps, Database Engineering, Security Engineering, and IT operations focusing on shared priorities and accountability.
• Build and lead a high-performing distributed team for infrastructure and operations.
• Set priorities, establish operating rhythms, ensure accountability, promote professional growth, and facilitate cross-discipline collaboration.
• Create and uphold an AI and automation roadmap aimed at business impact, reliability, efficiency, risk mitigation, and measurable ROI.
• Oversee architecture, performance, scalability, availability, monitoring, alerting, and capacity planning of RTA’s AWS environment.
• Enhance CI/CD, infrastructure-as-code, deployment reliability, automation, observability, and engineering operations.
• Lead production incident responses and work to improve incident resolution times and post-incident learning.
• Maintain infrastructure roadmaps that support company and product growth.
• Guide database performance, resilience, capacity planning, backups, recovery, tuning, and high availability.
• Oversee vulnerability management, cloud security hardening, security engineering priorities, compliance initiatives, and preparedness for incident response.
• Determine strategic direction for endpoint management, identity and access management, networking, internal tooling, and employee technology experience.
• Assess AI-powered tools, AIOps, intelligent monitoring, predictive analytics, automated remediation, AI-assisted development tools, and automated runbooks.
• Provide hands-on technical leadership by troubleshooting AWS configurations, deployment issues, security alerts, observability data, and production incidents.
• Mentor, coach, develop, and establish accountability within a multidisciplinary technical team.
• Collaborate with Engineering, Product, Support, Security, and other teams to integrate operational considerations from the outset.
• Translate technical topics, risks, investments, and trade-offs into clear business language for senior leadership.
• Communicate system health, technical risks, investment priorities, and recommendations to executives.
• Occasionally travel to RTA’s Glendale, Arizona headquarters and/or other company or team gatherings.
• Over 7 years of progressive experience in cloud infrastructure, DevOps, SRE/platform engineering, infrastructure operations, or a closely related technical field.
• Significant hands-on experience supporting AWS production environments.
• Proven experience in leading and developing technical teams, including setting expectations, fostering accountability, enhancing performance, and facilitating the growth of strong technical talent.
• Experience in owning or significantly influencing production reliability, availability, monitoring, and incident management.
• Deep knowledge of modern DevOps, CI/CD, observability, infrastructure-as-code, and automation practices.
• Strong understanding of contemporary cloud architecture and software design principles, including microservices and the impact of infrastructure on application reliability and scalability.
• Familiarity with AWS, Docker, Terraform/CloudFormation, Grafana, Prometheus, Datadog, New Relic, Splunk, or other comparable platforms.
• Capacity to translate complex technical challenges into clear priorities, risks, trade-offs, and recommendations for business and executive stakeholders.
• Experience leveraging AI and/or automation to enhance technical, engineering, infrastructure, or operational workflows.
• Strong organizational and prioritization abilities.
• Ability to make decisions, accept feedback, engage in healthy conflict, and adapt direction when warranted by facts.
• Capability to lead across Cloud, DevOps/SRE, Infrastructure, Security, Database Engineering, or IT disciplines.
• Experience managing remote or distributed technical teams.
• Relevant experience in security frameworks, vulnerability management, cloud security, or compliance is a plus.
• Understanding of database resilience, high availability, backup/recovery, and performance considerations.
• Familiarity with AI infrastructure concepts such as model serving, compute-resource management, vector databases, or LLM integration patterns.
• Relevant certifications such as AWS Solutions Architect, ITIL, or other cloud, DevOps, security, infrastructure, or AI-related credentials are beneficial.
• Ability to sit or stand for extended periods and work at a computer for prolonged durations.
• Proficient communication skills through video, phone, messaging, and other remote collaboration tools.
• Must be eligible to work in the United States.
• RTA cannot accept student visas or provide sponsorship.
• Employment may be contingent upon successful completion of a background check and other pre-employment screenings.
• 401(k) with a 6% Safe Harbor match (100% vested from day one).
• Flexible PTO model based on trust and manager-approved time off.
• Cigna PPO and HSA medical plan options with company contributions ($780–$1,950 annually).
• Garner Health HRA program: reimbursement opportunity of up to $1,000 for individuals and $2,000 for families.
• Wellness rewards of up to $350 annually.
• Access to virtual care and mental health support resources.
• Company-sponsored dental, vision, EAP, life, STD, LTD, legal plan, and identity theft protection options.
• Additional employee perks, wellness initiatives, and discount programs.
• Fully remote work environment.
• Meaningful work supporting fleets and individuals who keep communities functioning.
• An AI-forward company with a strong culture.
YipitData
Siteup
Leidos
Get handpicked remote jobs straight to your inbox weekly.