Remotery

Lead Site Reliability Engineer – Performance & Scalability

Posted 17 hours ago

This is a fully remote position, open to applicants in Mexico.

📋 Description

• Set benchmarks for performance, throughput, latency, and capacity for essential customer and platform workflows.

• Establish and uphold SLOs, error budgets, performance metrics, dashboards, alerts, and reliability standards.

• Instrument and evaluate the complete request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party services.

• Detect system bottlenecks and spearhead cross-functional remediation initiatives with engineering teams.

• Develop capacity models that illustrate platform sustainability, emerging constraints, and additional scaling expenses.

• Oversee load, stress, soak, spike, failure, and recovery testing in representative environments.

• Create realistic demand scenarios for major customers, partnerships, pilots, and high-volume events.

• Promote architecture hardening, graceful degradation, dependency-failure planning, and resilience enhancements.

• Collaborate with Test Automation and Scalability Engineering on automated performance testing, regression coverage, and production release checkpoints.

• Manage technical readiness evaluations for significant pilots, partnerships, and production launches.

• Develop operational runbooks for scale-up events, incidents, rollbacks, recovery, and dependency failures.

• Direct performance and reliability investigations during incidents and integrate lessons learned into future engineering efforts.

• Make infrastructure cost, performance, and reliability trade-offs transparent to engineering and executive leadership.

• Suggest capacity and reliability investments before they evolve into production bottlenecks.


⛳️ Requirements

• Extensive experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline.

• Experience in supporting production systems that meet significant scale, traffic, latency, or availability expectations.

• Profound understanding of observability, performance analysis, capacity planning, and reliability engineering.

• Strong practical experience with cloud infrastructure and production distributed systems.

• In-depth knowledge of databases, networking, caching, queuing, compute, storage, and common failure modes in distributed systems.

• Experience in defining and operating against SLOs, SLIs, error budgets, and metrics for production reliability.

• Hands-on experience in conducting load, stress, soak, scalability, and resilience testing.

• Ability to profile systems, identify bottlenecks, optimize architecture, and collaborate directly with engineering teams to implement improvements.

• Experience in designing for graceful degradation, managing dependency failures, recovery, and high-demand situations.

• Strong incident management skills and root-cause analysis experience.

• Ability to convey technical performance and reliability risks into clear business implications for senior leadership.

• Strong judgment regarding optimization versus unnecessary complexity.

• Nice to have: experience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms.

• Nice to have: experience in creating capacity-cost models and forecasting infrastructure needs.

• Nice to have: experience in incorporating performance and reliability gates into CI/CD pipelines.

• Nice to have: experience preparing platforms for significant traffic spikes related to enterprise customers or strategic partnerships.

• Nice to have: experience leading reliability or performance initiatives across multiple engineering teams.


🏝️ Benefits

• Equal Opportunity Employer dedicated to fostering a diverse and inclusive workplace.

• Accommodation for the application process is available through HR.

People also viewed

Level Data3 hours ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job
Level Data3 hours ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data3 hours ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data3 hours ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job
Identity Digital Inc.7 hours ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$175k – $220k/year
ApplyView job
Identity Digital Inc.7 hours ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$175k – $220k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers