Remotery

Senior Site Reliability Engineer

Posted Jul 3

This is a fully remote position, open to applicants in Nigeria.

📋 Description

• Engage in on-call rotations as the primary technical authority. Serve as the Incident Commander during critical severity incidents: initiating war rooms, coordinating cross-functional teams, and delivering clear status updates.

• Instrument code to reveal high-cardinality metrics and distributed traces. Collaboratively establish, measure, and support Service Level Objectives (SLOs) and Error Budgets alongside product owners.

• Develop high-quality, production-ready code (in Java, Go, or Python) to create internal tools, automation platforms, and self-healing mechanisms that reduce the need for manual operator intervention.

• Collaborate with Product Engineering teams during the design stage to guarantee that new services are constructed with reliability, scalability, and observability patterns (such as circuit breakers, rate limiting, backpressure, and fallback strategies) from the outset.

• Evaluate system performance and traffic patterns to project future capacity requirements. Execute load testing and chaos engineering experiments to validate system resilience in failure scenarios.


⛳️ Requirements

• At least 5 years of experience in SRE or Backend Engineering with a robust capability to write clean, efficient, and tested code in Java, Go, Rust, or Python.

• Comprehensive understanding of distributed systems architecture and design patterns. You possess a strong grasp of microservices fundamentals, event-driven architectures, and the essential principles needed to build scalable systems.

• Significant experience with Google Cloud Platform (GCP) or equivalent cloud providers (AWS/Azure). You are skilled in managing production workloads on Kubernetes (GKE/EKS) and resolving cluster/infrastructure challenges.

• Experience in designing observability strategies using OpenTelemetry, Prometheus, New Relic, Datadog, or SigNoz to enhance system visibility.

• Knowledge of operating and tuning production data stores (e.g., PostgreSQL, MySQL) and streaming platforms (e.g., Kafka, RabbitMQ) in a high-throughput environment.


🏝️ Benefits

• Culture - We prioritize our people and the well-being of every team member. We have created a company where every opinion is valued and every voice is heard. We respect and support one another, emphasizing our shared humanity.

• Learning - We foster a learning and development-centric environment with a focus on knowledge sharing, training, and regular internal technical discussions.

• Compensation - You will receive a competitive salary, pension, health insurance, annual bonus, and additional benefits.

People also viewed

SYNCREON5 hours ago

Forward Deployment Engineer – Travel to Boston, MA as required

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Rimutee6 hours ago

DevOps, AWS

US flagUnited States OnlyPart-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Mirantis6 hours ago

Senior Site Reliability Engineer, Golang, Kubernetes

CA flagCanada OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group6 hours ago

DevOps Engineer

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
XTEL6 hours ago

DevOps Engineer

BE flagBelgium OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Hypoport SE6 hours ago

DevOps Engineer – m/f/d

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers