Remotery

Senior Site Reliability Engineer – Hiring Globally

Posted Jul 27

This is a fully remote position, open to applicants in Serbia.

📋 Description

• Take ownership of service objectives by defining SLIs and SLOs for critical services, managing error budgets, and utilizing them to guide prioritization and change-rate decisions. Ensure that reliability is quantified rather than based on perception.

• Establish reliable observability by managing alert quality comprehensively: ensuring high signal-to-noise ratios, creating alarms that effectively capture intended incidents, and exercising the discipline to disable misleading alarms until they are properly addressed. Develop the necessary dashboards, custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) to achieve clear system visibility.

• Direct incident response efforts by managing incidents with composure, reducing mean-time-to-recovery, and generating blameless postmortems with actionable items that are genuinely addressed. Enhance and take responsibility for the health of the on-call rotation.

• Plan for capacity and performance by accurately forecasting and sizing compute resources (particularly GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis usage. Anticipate saturation before it impacts customers.

• Manage business continuity and disaster recovery by ensuring effective backups, replication, failover, and recovery processes for RDS, MSK, Redis, and the edge fleet. Define RPO/RTO and validate them through regular, tested game days rather than mere assumptions.

• Maintain the health of the edge fleet through remote diagnosis and recovery via AWS IoT, implementing container auto-update via systemd timers, and facilitating the ongoing CentOS 7 migration to a containerized media stack (Ubuntu 22.04).

• Reduce operational toil by developing software (Python, Golang, Bash) to automate operational tasks, self-heal common failures, and make reliability a consistent outcome rather than a heroic effort.

• Ensure safe governance of production changes by enforcing collaborative, reviewed change management processes and safeguarding the system from risky, unilateral changes (such as topology, instance-count, scaling, and configuration).


⛳️ Requirements

• A valid AWS certification is required, with a preference for current AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect – Professional.

• A minimum of 10 years of experience in Site Reliability Engineering or production operations at scale. While mastery in every area is not expected on day one, substantial expertise in several areas and the capacity to learn quickly in others is essential.

• Proven experience with SLO/error-budget management. You have established SLIs and SLOs, operated within an error budget, and leveraged it for impactful decision-making.

• Strong skills in production observability, with extensive knowledge of metrics, logs, alarming, and dashboards (CloudWatch, OpenTelemetry, and/or Prometheus/Grafana/Datadog), along with the capability to create instrumentation when it is lacking.

• Demonstrated incident command experience. You have led incidents, managed an on-call rotation, and authored postmortems that influenced system behavior.

• Experience in capacity planning and performance across compute resources, databases, and messaging or streaming systems (Kafka/MSK preferred).

• Ownership of disaster recovery processes including backups, replication, failover, and validated RPO/RTO.

• Proficiency in software engineering for automation, with comfort in writing Python, Golang, and Bash to develop reliability tools instead of merely configuring existing solutions.

• Expertise in Terraform (or equivalent IaC) and strong Linux administration skills (including shell and Linux GNU utilities), capable of working across cloud to bare-metal/edge environments.

• Experience with database operations, particularly PostgreSQL (experience with time-series is a plus).

• A reliability-oriented mindset: you instrument before making assumptions and document processes in runbooks.

• Nice to have: experience with operating GPU workloads and deploying computer vision or ML models in production (including CUDA, Deep Learning AMIs, and inference scaling).

• Familiarity with Apache MSK/Kafka and operations in streaming data (such as Kinesis and Kinesis Video Streams).

• Experience with AWS IoT Core at scale, including device provisioning, certificates, and secure tunneling.

• Skills in managing a fleet of edge/on-premise devices (including golden images, remote updates, and systemd management).

• Experience in operating and modernizing legacy systems (such as Java 8, Jetty, and CentOS).

• Familiarity with chaos engineering, game-day practices, and capacity modeling.

• Knowledge of Bazel in a monorepo environment; experience with Cloudflare, Cognito/Auth0, and API Gateway.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Flexible working hours and the option for remote work.

• Comprehensive health benefits including medical, dental, and vision coverage.

• Opportunities for professional development and continuous learning.

• Collaborative and innovative team environment.

People also viewed

CWILL20 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3721 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT21 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group21 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo22 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch22 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers