
Senior Site Reliability Engineer
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in United States.
• Take ownership of production reliability for customer-centric platforms and data services across Azure, colocation, and edge Kubernetes environments.
• Assist the shared SRE function within the radar network and weather intelligence sector.
• Establish and enhance SLIs, SLOs, alerting standards, and operational metrics.
• Develop and manage fleet observability through shared dashboards and proactive alerting mechanisms.
• Create and implement automated recovery and self-healing processes for production systems.
• Optimize cluster resources and expenses, including precise workload and node sizing and transitioning workloads from Azure.
• Oversee production incident responses, encompassing troubleshooting, mitigation, communication, and postmortem evaluations.
• Identify and resolve intricate issues across application services, Kubernetes infrastructure, storage, and distributed systems.
• Promote multi-replica and multi-cluster high availability, addressing workload placement, scheduling, failover, traffic routing, and data replication.
• Manage and enhance self-hosted Kubernetes platforms across cloud-hosted, colocation, and edge clusters.
• Carry out Kubernetes upgrades, patching, cluster health assessments, node management, and production change management.
• Enhance observability, autoscaling, ingress, distributed storage, resilience, and operational efficiency.
• Design and assess Kubernetes workloads for resilience, scalability, efficiency, and graceful degradation.
• Collaborate with software engineering teams to ensure production readiness, secure deployments, and operational visibility.
• Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation.
• Provide support for metrics, logging, distributed tracing, dashboarding, and alerting platforms.
• Conduct performance engineering and capacity planning to meet peak weather-event demands.
• Facilitate blameless postmortems and ensure completion of operational follow-up tasks.
• Strengthen disaster recovery, failover, and business continuity capabilities.
• Advocate for automation, toil reduction, game days, production readiness reviews, and reliability best practices.
• Mentor teams on reliability engineering and production operations methodologies.
• Participate in rotating on-call schedules during weekdays and weekends.
• A bachelor's degree in computer science, software engineering, or a related discipline; equivalent professional experience will be considered.
• At least 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a similar infrastructure-focused role, with a minimum of 4 years in a role officially titled Site Reliability Engineer or with specific SLO/error-budget responsibilities.
• Extensive hands-on experience managing native Kubernetes; while managed distributions like AKS and EKS are acceptable, native or self-managed Kubernetes is strongly preferred as the primary technical requirement.
• Proven experience in optimizing Kubernetes clusters, including workload and node pool right-sizing, resource management, and reducing infrastructure costs without compromising reliability.
• Experience in enhancing operational visibility, including developing dashboards, metrics pipelines, and alerting systems.
• Capability in designing and managing workloads for secure horizontal scaling across multiple replicas, covering idempotency, concurrency, and state management.
• Familiarity with designing or operating multi-cluster high-availability architectures, including failover procedures, traffic routing, and cross-cluster service deployment.
• Experience in supporting customer-facing production systems with responsibilities for uptime, reliability, and incident response.
• Proficient in diagnosing and resolving production incidents across application, platform, and Kubernetes infrastructure layers.
• Experience operating Kubernetes in environments beyond strictly managed cloud settings, including bare-metal, colocation, edge, or hybrid infrastructures.
• Proficiency with Kubernetes operational tools and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems.
• Strong understanding of infrastructure automation and Infrastructure as Code using tools like Terraform and Ansible.
• Experience in supporting CI/CD and production deployment pipelines; GitHub Actions is employed for CI/CD.
• Familiarity with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or equivalent technologies.
• Experience in operating distributed systems and microservice-based architectures in production.
• Working knowledge of Microsoft Azure infrastructure.
• Strong troubleshooting capabilities across infrastructure, application, and platform layers.
• Experience in participating in a structured production on-call rotation for business-critical systems.
• Working familiarity with Jira, Confluence, and Microsoft Entra.
• Excellent written and verbal communication skills, including incident documentation and postmortem writing.
• Experience in fast-paced engineering environments such as start-ups or scale-ups.
• Ability to pass a background check according to company standards.
• Willingness to engage in separate rotating weekday and weekend on-call schedules.
• Advantages of being part of a dynamic and expanding organization.
• A challenging, hands-on position that will significantly influence the business.
• Competitive salary.
• Comprehensive benefits package.
• 401(k) Savings Plan.
• Medical, Dental, and Vision Benefits.
• Health Savings Account (HSA) and Flexible Spending Account (FSA).
• Unlimited Paid Time Off.
• 11 Paid Holidays.
• Paid Parental Leave.
• Company Paid Short-term Disability (STD).
• Company Paid Long-term Disability (LTD).
• Company Paid Life Insurance.
Colonist
Yopeso
Orion Innovation
Parasail
Get handpicked remote jobs straight to your inbox weekly.