
Director, Software Engineering β Infrastructure
Posted Aug 1

Posted Aug 1
This is a fully remote position, open to applicants in California, +8 more states.
β’ This position is essential for the success of ServiceTitan, reporting directly to the VP of Infrastructure.
β’ Lead, grow, and nurture a global Site Reliability Engineering (SRE) team of engineers capable of providing round-the-clock coverage.
β’ Manage the operations center (OC) along with incident command and response functions for this mission-critical software.
β’ Strive to achieve and maintain 99.99% availability across our infrastructure.
β’ Oversee release management for core applications while orchestrating a robust process across functional microservices.
β’ Participate in service capacity planning, demand forecasting, software performance analysis, and system optimization.
β’ Collaborate with development teams to ensure applications are production-ready, scalable, reliable, and observable from the beginning.
β’ Assess and enhance system performance, aiming to advance capabilities and anticipate customer needs.
β’ Identify, develop, implement, and sustain practices that guarantee the highest levels of uptime, performance, reliability, and security across production and pre-production environments.
β’ Provide thought leadership in resolving issues related to internal and external technology matters.
β’ Promote the adoption of operational best practices across critical services, consistently seeking to reduce operational barriers to enhance reliability.
β’ Serve as an internal resource for teams and business units.
β’ 10 to 15 years of software engineering experience, including at least 7 years in a leadership role managing a team of over 50 engineers.
β’ More than 7 years of experience supporting infrastructure and services hosted on AWS, GCP, or Azure.
β’ Over 5 years of experience in delivering, deploying, and managing enterprise applications in the cloud.
β’ At least 3 years of experience developing continuous integration, delivery, and deployment pipelines along with cloud-centric CI/CD tools.
β’ A minimum of 3 years implementing telemetry and observability intelligence with automated remediation.
β’ More than 3 years as a leader in driving scalability, resiliency, performance, and security.
β’ At least 3 years of experience establishing and advancing an SRE practice.
β’ A strong understanding of Azure cloud services and monitoring technologies is highly advantageous.
β’ Experience in constructing pre-production performance and testing environments, as well as disaster recovery/high availability (DR/HA) frameworks in a cloud environment is essential.
β’ Proficient in Infrastructure as Code (IaC) using tools such as Terraform, Ansible, etc.
β’ Extensive experience with containerization technologies like Docker and Kubernetes.
β’ Strong expertise in building CI/CD pipelines using tools such as Jenkins, and familiarity with observability tools like New Relic, DataDog, and Splunk Enterprise.
β’ BA/BS in Computer Science or a related field.
β’ An MS/PhD is highly desirable.
β’ Flexible time off along with numerous learning and development opportunities to further your career.
β’ A comprehensive onboarding program and leadership training available for Titans at every level.
β’ Exceptional work is recognized through Bonusly, peer-nominated awards, and additional rewards.
β’ Company-paid medical, dental, and vision insurance (with 100% employer-paid options and 90% coverage for dependents), FSA and HSA, 401k matching, and telehealth options including memberships to One Medical.
β’ Parental leave and support, along with up to $20k in fertility services (such as IUI and IVF), surrogacy, and adoption reimbursement.
β’ On-demand maternity support through Maven Maternity, complimentary breast milk shipping via Maven Milk, pet insurance, legal advisory services, financial planning tools, and more.
Travoom
EverCommerce
Cisco
Pluribus Digital
Get handpicked remote jobs straight to your inbox weekly.