
Lead Site Reliability Engineer – Performance & Scalability
Posted 17 hours ago

Posted 17 hours ago
This is a fully remote position, open to applicants in Mexico.
• Set benchmarks for performance, throughput, latency, and capacity for essential customer and platform workflows.
• Establish and uphold SLOs, error budgets, performance metrics, dashboards, alerts, and reliability standards.
• Instrument and evaluate the complete request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party services.
• Detect system bottlenecks and spearhead cross-functional remediation initiatives with engineering teams.
• Develop capacity models that illustrate platform sustainability, emerging constraints, and additional scaling expenses.
• Oversee load, stress, soak, spike, failure, and recovery testing in representative environments.
• Create realistic demand scenarios for major customers, partnerships, pilots, and high-volume events.
• Promote architecture hardening, graceful degradation, dependency-failure planning, and resilience enhancements.
• Collaborate with Test Automation and Scalability Engineering on automated performance testing, regression coverage, and production release checkpoints.
• Manage technical readiness evaluations for significant pilots, partnerships, and production launches.
• Develop operational runbooks for scale-up events, incidents, rollbacks, recovery, and dependency failures.
• Direct performance and reliability investigations during incidents and integrate lessons learned into future engineering efforts.
• Make infrastructure cost, performance, and reliability trade-offs transparent to engineering and executive leadership.
• Suggest capacity and reliability investments before they evolve into production bottlenecks.
• Extensive experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline.
• Experience in supporting production systems that meet significant scale, traffic, latency, or availability expectations.
• Profound understanding of observability, performance analysis, capacity planning, and reliability engineering.
• Strong practical experience with cloud infrastructure and production distributed systems.
• In-depth knowledge of databases, networking, caching, queuing, compute, storage, and common failure modes in distributed systems.
• Experience in defining and operating against SLOs, SLIs, error budgets, and metrics for production reliability.
• Hands-on experience in conducting load, stress, soak, scalability, and resilience testing.
• Ability to profile systems, identify bottlenecks, optimize architecture, and collaborate directly with engineering teams to implement improvements.
• Experience in designing for graceful degradation, managing dependency failures, recovery, and high-demand situations.
• Strong incident management skills and root-cause analysis experience.
• Ability to convey technical performance and reliability risks into clear business implications for senior leadership.
• Strong judgment regarding optimization versus unnecessary complexity.
• Nice to have: experience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms.
• Nice to have: experience in creating capacity-cost models and forecasting infrastructure needs.
• Nice to have: experience in incorporating performance and reliability gates into CI/CD pipelines.
• Nice to have: experience preparing platforms for significant traffic spikes related to enterprise customers or strategic partnerships.
• Nice to have: experience leading reliability or performance initiatives across multiple engineering teams.
• Equal Opportunity Employer dedicated to fostering a diverse and inclusive workplace.
• Accommodation for the application process is available through HR.
Level Data
Level Data
Level Data
Level Data
Get handpicked remote jobs straight to your inbox weekly.