
Director of Cloud SRE
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Illinois.
• Collaborate with other SRE leaders to establish and implement a multi-year strategy for cohesive observability and SRE platform solutions across GCP and on-premise settings.
• Lead, mentor, and develop engineering managers/leaders as well as individual contributor engineers.
• Propel the roadmap for internally developed, vendor-neutral observability tools utilizing OpenTelemetry-first design and Agentic AI platforms.
• Promote SRE principles such as SLIs/SLOs, error budgets, incident management, toil reduction, capacity, and reliability engineering across application and platform teams.
• Integrate observability and reliability tools into CI/CD pipelines and source code repositories.
• Participate in the development of an SRE maturity model and self-service tools for application teams.
• Cultivate relationships with manufacturing, plant, and operational technology engineering leadership to apply reliability practices in constrained and safety-critical contexts.
• Represent the direction of the SRE platform to senior technology leadership and architectural governance groups.
• Serve as an escalation point for significant reliability and observability projects.
• Engage in vendor partnerships and make build-vs-buy technology evaluations.
• Review design proposals, contribute to architectural decisions, and maintain current hands-on technical expertise.
• A Bachelor’s degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience.
• Over 10 years of experience in Site Reliability Engineering, platform engineering, or infrastructure engineering.
• At least 4 years in a leadership role managing engineering leaders and/or engineers.
• Proven experience in building and operating observability platforms at scale.
• In-depth knowledge of OpenTelemetry and familiarity with at least one enterprise observability platform (e.g., Dynatrace, Datadog, New Relic, Splunk, or similar).
• Experience in designing and delivering internally developed developer tools that integrate with CI/CD pipelines, source control systems, and developer workflows.
• Strong understanding of public cloud architecture; GCP is preferred, with AWS/Azure being acceptable.
• Capability to extend reliability methodologies into hybrid or on-premise environments.
• Profound knowledge of SLIs/SLOs, error budgets, incident management and postmortem practices, toil reduction, capacity planning, and reliability-by-design principles.
• Experience working in diverse cloud, data center, and operational technology/manufacturing environments.
• Ability to create a long-term technical strategy and convert it into a practical roadmap.
• Excellent executive communication skills with the capability to align cross-functional stakeholders.
• Familiarity with infrastructure-as-code (Terraform or equivalent) and contemporary software delivery practices.
• Willingness to travel up to 10%.
• Must have legal authorization to work in the United States.
• Visa sponsorship is not available.
• Immediate medical, dental, vision, and prescription drug coverage.
• Flexible family care days.
• Paid parental leave.
• New parent ramp-up programs.
• Subsidized backup childcare.
• Family-building benefits, including adoption and surrogacy reimbursement and fertility treatments.
• Vehicle discount program for employees and their family members, along with management leases.
• Tuition assistance.
• Established and active employee resource groups.
• Paid time off for individual and team community service.
• Generous holiday schedule, including a week off between Christmas and New Year’s Day.
• Paid time off.
• Option to purchase additional vacation time.
SYNCREON
Rimutee
Mirantis
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.