Remotery

Senior Engineering Manager, Site Reliability

Posted Jul 18

This is a fully remote position, open to applicants in United States.

📋 Description

• Oversee and cultivate a team dedicated to incident management, observability, operational readiness, and reliability engineering.

• Establish a clear charter, set priorities, create a roadmap, and define measurable outcomes for the SRE function.

• Convert strategy into capacity-aware plans, outlining explicit trade-offs, ownership, milestones, and success metrics.

• Ensure visibility into delivery health, operational risks, and team performance, taking early action when execution deviates.

• Develop a resilient operating model through cross-training, shared understanding, effective delegation, and defined primary and secondary ownership.

• Maintain high standards for technical quality, operational rigor, and executive communication.

• Mentor engineers and leaders to independently manage complex reliability projects.

• Enhance Upstart’s incident management program to boost detection, response, coordination, communication, and recovery efforts.

• Set clear criteria for managing high-severity incidents and provide visible leadership during critical situations.

• Improve postmortem quality and ensure that incident learnings contribute to lasting engineering enhancements.

• Detect recurring failure patterns and promote systemic solutions across teams.

• Establish robust feedback loops from incidents into roadmaps, service standards, operational readiness requirements, and measurable risk reduction.

• Enhance the quality, accessibility, and reliability of signals used to monitor production health.

• Implement consistent practices across metrics, logs, traces, alerting, and service health.

• Leverage service level objectives and customer impact signals to inform priorities and operational decisions.

• Minimize detection gaps, noisy alerts, manual investigations, and persistent operational toil.

• Define measurable reliability outcomes and utilize data to prioritize investments and convey impact.

• Collaborate with platform and product engineering teams to integrate reliability into standard engineering processes.

• Create scalable operational readiness standards for new services, major launches, and architectural changes.

• Clearly communicate expectations for service ownership, monitoring, capacity, failure management, and incident response.

• Identify systemic reliability risks and collaborate with engineering teams to prioritize and address these issues.

• Enhance resilience through automation, failure testing, recovery capabilities, and operational safeguards.

• Develop operating mechanisms that translate reviews and analyses into clear decisions, owners, timelines, and sustained follow-through.

• Align stakeholders and dependencies prior to critical launches and engineering decisions.


⛳️ Requirements

• Minimum of 5 years of experience in reliability engineering management and over 7 years in software engineering, site reliability engineering, infrastructure, or platform engineering.

• Significant hands-on experience in Site Reliability Engineering, Production Engineering, or a similar role responsible for operating and enhancing production systems.

• Proven experience managing an SRE, Production Engineering, or equivalent reliability function, encompassing strategy, roadmap, operating model, and outcomes.

• Strong technical expertise in distributed systems, cloud infrastructure, observability, and production operations.

• Experience leading high-severity incident responses and enhancing incident management practices on a large scale.

• Demonstrated ability to translate strategy into focused, capacity-aware plans and deliver measurable results.

• History of hiring, developing, and retaining high-performing engineers and engineering leaders.

• Strong cross-functional leadership and communication skills, with the ability to convert complex operational data into clear decisions and foster alignment across teams.


🏝️ Benefits

• Competitive compensation package, including base salary, bonus opportunities, and quarterly vesting annual equity grants.

• Retirement benefits to aid in future planning, featuring a 401(k) or Group Retirement Savings Plan with a company match of $2 for every $1 contributed, up to $15,000 annually (USD in the US, CAD in Canada).

• Employee Stock Purchase Plan (ESPP) offering discounted stock purchase options for eligible employees (US only).

• Comprehensive health coverage designed to support you and your family, including medical, dental, vision, and wellness resources for US and supplemental health coverage for Canada.

• Contributions to Health Savings Accounts from Upstart for eligible plans (US only).

• Income protection benefits, such as life insurance and disability coverage, for added financial security.

• Paid time off, sick leave, and company holidays, in accordance with local regulations.

• Paid family and parental leave to support caregiving and significant life events (duration varies by country).

• Family-oriented benefits to assist with fertility, parenthood, and caregiving needs.

• Employee Assistance Program (EAP) offering mental health support and life-centered resources.

• Financial wellness resources, including access to financial planning tools and a financial concierge service (US Only).

• Annual wellness allowance to support your physical and emotional well-being and personal development, based on what matters most to you.

• Annual productivity allowance to invest in relevant tools and resources to excel in your work, regardless of your location.

• Opportunities for connection and community through team events, all-company updates, and employee resource groups (ERGs).

• Onsite perks, including catered lunches and fully stocked micro-kitchens at our offices located in the Bay Area, Austin, Columbus, and New York City (opening Summer 2026!).

People also viewed

The Codest3 days ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IRIUM3 days ago

Ingeniero/a Cloud DevOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€33k – €40k/year
ApplyView job
Sólides3 days ago

Senior DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Resilinc3 days ago

Junior/Senior Site Reliability Engineer – Night Shift

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group3 days ago

Senior SRE / DevOps Engineer

Anywhere in the WorldFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HOESSLER & HOESSLER3 days ago

DevOps Software Engineer – Career Ambitions

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€65k – €75k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers