
Senior Engineering Manager, Site Reliability
Posted Jul 18

Posted Jul 18
This is a fully remote position, open to applicants in United States.
• Oversee and cultivate a team dedicated to incident management, observability, operational readiness, and reliability engineering.
• Establish a clear charter, set priorities, create a roadmap, and define measurable outcomes for the SRE function.
• Convert strategy into capacity-aware plans, outlining explicit trade-offs, ownership, milestones, and success metrics.
• Ensure visibility into delivery health, operational risks, and team performance, taking early action when execution deviates.
• Develop a resilient operating model through cross-training, shared understanding, effective delegation, and defined primary and secondary ownership.
• Maintain high standards for technical quality, operational rigor, and executive communication.
• Mentor engineers and leaders to independently manage complex reliability projects.
• Enhance Upstart’s incident management program to boost detection, response, coordination, communication, and recovery efforts.
• Set clear criteria for managing high-severity incidents and provide visible leadership during critical situations.
• Improve postmortem quality and ensure that incident learnings contribute to lasting engineering enhancements.
• Detect recurring failure patterns and promote systemic solutions across teams.
• Establish robust feedback loops from incidents into roadmaps, service standards, operational readiness requirements, and measurable risk reduction.
• Enhance the quality, accessibility, and reliability of signals used to monitor production health.
• Implement consistent practices across metrics, logs, traces, alerting, and service health.
• Leverage service level objectives and customer impact signals to inform priorities and operational decisions.
• Minimize detection gaps, noisy alerts, manual investigations, and persistent operational toil.
• Define measurable reliability outcomes and utilize data to prioritize investments and convey impact.
• Collaborate with platform and product engineering teams to integrate reliability into standard engineering processes.
• Create scalable operational readiness standards for new services, major launches, and architectural changes.
• Clearly communicate expectations for service ownership, monitoring, capacity, failure management, and incident response.
• Identify systemic reliability risks and collaborate with engineering teams to prioritize and address these issues.
• Enhance resilience through automation, failure testing, recovery capabilities, and operational safeguards.
• Develop operating mechanisms that translate reviews and analyses into clear decisions, owners, timelines, and sustained follow-through.
• Align stakeholders and dependencies prior to critical launches and engineering decisions.
• Minimum of 5 years of experience in reliability engineering management and over 7 years in software engineering, site reliability engineering, infrastructure, or platform engineering.
• Significant hands-on experience in Site Reliability Engineering, Production Engineering, or a similar role responsible for operating and enhancing production systems.
• Proven experience managing an SRE, Production Engineering, or equivalent reliability function, encompassing strategy, roadmap, operating model, and outcomes.
• Strong technical expertise in distributed systems, cloud infrastructure, observability, and production operations.
• Experience leading high-severity incident responses and enhancing incident management practices on a large scale.
• Demonstrated ability to translate strategy into focused, capacity-aware plans and deliver measurable results.
• History of hiring, developing, and retaining high-performing engineers and engineering leaders.
• Strong cross-functional leadership and communication skills, with the ability to convert complex operational data into clear decisions and foster alignment across teams.
• Competitive compensation package, including base salary, bonus opportunities, and quarterly vesting annual equity grants.
• Retirement benefits to aid in future planning, featuring a 401(k) or Group Retirement Savings Plan with a company match of $2 for every $1 contributed, up to $15,000 annually (USD in the US, CAD in Canada).
• Employee Stock Purchase Plan (ESPP) offering discounted stock purchase options for eligible employees (US only).
• Comprehensive health coverage designed to support you and your family, including medical, dental, vision, and wellness resources for US and supplemental health coverage for Canada.
• Contributions to Health Savings Accounts from Upstart for eligible plans (US only).
• Income protection benefits, such as life insurance and disability coverage, for added financial security.
• Paid time off, sick leave, and company holidays, in accordance with local regulations.
• Paid family and parental leave to support caregiving and significant life events (duration varies by country).
• Family-oriented benefits to assist with fertility, parenthood, and caregiving needs.
• Employee Assistance Program (EAP) offering mental health support and life-centered resources.
• Financial wellness resources, including access to financial planning tools and a financial concierge service (US Only).
• Annual wellness allowance to support your physical and emotional well-being and personal development, based on what matters most to you.
• Annual productivity allowance to invest in relevant tools and resources to excel in your work, regardless of your location.
• Opportunities for connection and community through team events, all-company updates, and employee resource groups (ERGs).
• Onsite perks, including catered lunches and fully stocked micro-kitchens at our offices located in the Bay Area, Austin, Columbus, and New York City (opening Summer 2026!).
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.