Remotery

Senior Site Reliability Engineer – Hiring Globally

Posted Jul 27

This is a fully remote position, open to applicants in Brazil.

📋 Description

• Take ownership of service objectives. Establish SLIs and SLOs for critical services, manage error budgets, and utilize them to influence prioritization and change-rate decisions. Ensure reliability is quantified rather than based on intuition.

• Ensure observability is reliable. Take charge of the entire alert lifecycle: maintain a high signal-to-noise ratio, ensure alarms accurately identify the incidents they are designed for, and exercise the discipline to disable misleading alarms until they are resolved. Develop the necessary dashboards, along with custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) to achieve clear visibility into the system.

• Direct incident response efforts. Manage incidents with composure, reduce mean-time-to-recovery, and create blameless postmortems that include actionable items that are followed through. Enhance and oversee the health of the on-call rotation.

• Strategically plan for capacity and performance. Anticipate and optimize compute resources (especially GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis. Identify saturation issues before they affect customers.

• Manage business continuity and disaster recovery. Oversee backups, replication, failover, and recovery processes for RDS, MSK, Redis, and the edge fleet. Define RPO/RTO and validate them through regular, tested game days rather than relying on assumptions.

• Maintain the health of the edge fleet. Conduct remote diagnostics and recovery via AWS IoT, facilitate container auto-updates using systemd timers, and manage the ongoing migration from CentOS 7 to a containerized media stack (Ubuntu 22.04).

• Eliminate operational toil. Develop genuine software (Python, Golang, Bash) to automate operational tasks, self-correct common failures, and ensure reliability is achieved through consistent processes rather than heroic efforts.

• Safeguard production changes. Implement collaborative, reviewed change management practices to protect the system from risky, unilateral changes (including topology, instance count, scaling, and configuration adjustments).


⛳️ Requirements

• AWS certification is required. A current AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect – Professional is highly preferred.

• Over 10 years of experience in Site Reliability Engineering or large-scale production operations. Mastery of every area listed is not expected on day one; however, significant depth in several areas and the ability to quickly learn the rest is essential.

• Proven practice with SLO/error budgets. You have defined SLIs and SLOs, operated within an error budget, and leveraged it for making informed decisions.

• Strong skills in production observability. Extensive experience with metrics, logs, alerts, and dashboards (CloudWatch, OpenTelemetry, and/or Prometheus/Grafana/Datadog) and capable of building necessary instrumentation when it is not available.

• Demonstrated incident command experience. You have led incident responses, managed an on-call rotation, and authored postmortems that have influenced system behavior.

• Experience with capacity planning and performance across compute, databases, and messaging or streaming systems (ideal with Kafka/MSK).

• Ownership of disaster recovery processes: including backups, replication, failover, and validated RPO/RTO.

• Proficiency in software engineering for automation. Comfortable coding in Python, Golang, and Bash to develop reliability tools, rather than only configuring existing solutions.

• Expertise with Terraform (or equivalent Infrastructure as Code) and solid Linux administration skills (including shell and Linux GNU utilities), comfortable working from cloud environments to bare-metal/edge systems.

• Experience in database operations with PostgreSQL (experience with time-series databases is a plus).

• A reliability-driven mindset: you rely on instrumentation rather than assumptions, and you document processes in runbooks.

• Nice to have:

• Experience with operating GPU workloads and deploying computer vision or machine learning models in production (including CUDA, Deep Learning AMIs, and inference scaling).

• Knowledge of Apache MSK/Kafka and streaming data operations (including Kinesis and Kinesis Video Streams).

• Familiarity with AWS IoT Core at scale: device provisioning, managing certificates, and secure tunneling.

• Experience managing a fleet of edge/on-premise devices (including golden images, remote updates, and systemd).

• Skills in operating and modernizing legacy systems (including Java 8, Jetty, and CentOS).

• Experience with chaos engineering/game-day practices and capacity modeling.

• Familiarity with Bazel in a monorepo, as well as Cloudflare, Cognito/Auth0, and API Gateway.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Generous paid time off and flexible working hours.

• Opportunities for professional development and continuous learning.

• Collaborative and inclusive work environment.

People also viewed

CWILL13 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3714 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT15 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group15 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo15 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch16 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers