Senior Site Reliability Engineer

Posted 16 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of production reliability for customer-centric platforms and data services across Azure, colocation, and edge Kubernetes environments.

• Assist the shared SRE function within the radar network and weather intelligence sector.

• Establish and enhance SLIs, SLOs, alerting standards, and operational metrics.

• Develop and manage fleet observability through shared dashboards and proactive alerting mechanisms.

• Create and implement automated recovery and self-healing processes for production systems.

• Optimize cluster resources and expenses, including precise workload and node sizing and transitioning workloads from Azure.

• Oversee production incident responses, encompassing troubleshooting, mitigation, communication, and postmortem evaluations.

• Identify and resolve intricate issues across application services, Kubernetes infrastructure, storage, and distributed systems.

• Promote multi-replica and multi-cluster high availability, addressing workload placement, scheduling, failover, traffic routing, and data replication.

• Manage and enhance self-hosted Kubernetes platforms across cloud-hosted, colocation, and edge clusters.

• Carry out Kubernetes upgrades, patching, cluster health assessments, node management, and production change management.

• Enhance observability, autoscaling, ingress, distributed storage, resilience, and operational efficiency.

• Design and assess Kubernetes workloads for resilience, scalability, efficiency, and graceful degradation.

• Collaborate with software engineering teams to ensure production readiness, secure deployments, and operational visibility.

• Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation.

• Provide support for metrics, logging, distributed tracing, dashboarding, and alerting platforms.

• Conduct performance engineering and capacity planning to meet peak weather-event demands.

• Facilitate blameless postmortems and ensure completion of operational follow-up tasks.

• Strengthen disaster recovery, failover, and business continuity capabilities.

• Advocate for automation, toil reduction, game days, production readiness reviews, and reliability best practices.

• Mentor teams on reliability engineering and production operations methodologies.

• Participate in rotating on-call schedules during weekdays and weekends.


⛳️ Requirements

• A bachelor's degree in computer science, software engineering, or a related discipline; equivalent professional experience will be considered.

• At least 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a similar infrastructure-focused role, with a minimum of 4 years in a role officially titled Site Reliability Engineer or with specific SLO/error-budget responsibilities.

• Extensive hands-on experience managing native Kubernetes; while managed distributions like AKS and EKS are acceptable, native or self-managed Kubernetes is strongly preferred as the primary technical requirement.

• Proven experience in optimizing Kubernetes clusters, including workload and node pool right-sizing, resource management, and reducing infrastructure costs without compromising reliability.

• Experience in enhancing operational visibility, including developing dashboards, metrics pipelines, and alerting systems.

• Capability in designing and managing workloads for secure horizontal scaling across multiple replicas, covering idempotency, concurrency, and state management.

• Familiarity with designing or operating multi-cluster high-availability architectures, including failover procedures, traffic routing, and cross-cluster service deployment.

• Experience in supporting customer-facing production systems with responsibilities for uptime, reliability, and incident response.

• Proficient in diagnosing and resolving production incidents across application, platform, and Kubernetes infrastructure layers.

• Experience operating Kubernetes in environments beyond strictly managed cloud settings, including bare-metal, colocation, edge, or hybrid infrastructures.

• Proficiency with Kubernetes operational tools and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems.

• Strong understanding of infrastructure automation and Infrastructure as Code using tools like Terraform and Ansible.

• Experience in supporting CI/CD and production deployment pipelines; GitHub Actions is employed for CI/CD.

• Familiarity with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or equivalent technologies.

• Experience in operating distributed systems and microservice-based architectures in production.

• Working knowledge of Microsoft Azure infrastructure.

• Strong troubleshooting capabilities across infrastructure, application, and platform layers.

• Experience in participating in a structured production on-call rotation for business-critical systems.

• Working familiarity with Jira, Confluence, and Microsoft Entra.

• Excellent written and verbal communication skills, including incident documentation and postmortem writing.

• Experience in fast-paced engineering environments such as start-ups or scale-ups.

• Ability to pass a background check according to company standards.

• Willingness to engage in separate rotating weekday and weekend on-call schedules.


🏝️ Benefits

• Advantages of being part of a dynamic and expanding organization.

• A challenging, hands-on position that will significantly influence the business.

• Competitive salary.

• Comprehensive benefits package.

• 401(k) Savings Plan.

• Medical, Dental, and Vision Benefits.

• Health Savings Account (HSA) and Flexible Spending Account (FSA).

• Unlimited Paid Time Off.

• 11 Paid Holidays.

• Paid Parental Leave.

• Company Paid Short-term Disability (STD).

• Company Paid Long-term Disability (LTD).

• Company Paid Life Insurance.

People also viewed

Colonist16 hours ago

DevOps Engineer

PT flagPortugal OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Yopeso16 hours ago

Reliability Engineer / DevOps – Database Platform

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Orion Innovation16 hours ago

DevOps

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Parasail16 hours ago

Senior Site Reliability Engineer

EuropeFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
NICE17 hours ago

Cloud Operations Engineer

GB flagUnited Kingdom OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Humana17 hours ago

Senior DevOps Engineer

US flagDistrict of Columbia, +2 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$106.9k – $147k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers