Remotery

Senior Site Reliability Engineer – Hiring Globally

Posted Jul 27

This is a fully remote position, open to applicants in Argentina.

📋 Description

• Take ownership of service objectives by defining SLIs and SLOs for crucial services, managing error budgets, and employing them to influence prioritization and change-rate decisions. Transform reliability into a quantifiable metric rather than a subjective feeling.

• Ensure observability is dependable. Manage the quality of alerts from start to finish: maintain a high signal-to-noise ratio, ensure alarms accurately identify the incidents they are intended for, and have the discipline to deactivate misleading alarms until they are properly rectified. Create dashboards and custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) necessary for a clear view of the system.

• Lead incident response efforts. Manage incidents with composure, reduce mean-time-to-recovery, and produce blameless postmortems with actionable items that are completed. Enhance and take responsibility for the health of the on-call rotation.

• Plan for capacity and performance. Anticipate and appropriately size compute resources (especially GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis. Identify saturation before it impacts customers.

• Take charge of business continuity and disaster recovery. Manage backups, replication, failover, and recovery processes for RDS, MSK, Redis, and the edge fleet. Establish RPO/RTO standards and validate them through regular, tested game days, rather than relying on assumptions.

• Maintain the health of the edge fleet. Conduct remote diagnosis and recovery via AWS IoT, facilitate container auto-updates through systemd timers, and oversee the ongoing migration from CentOS 7 to the containerized media stack (Ubuntu 22.04).

• Eliminate toil through engineering. Develop real software (Python, Golang, Bash) to automate operational tasks, self-correct common failures, and ensure reliability becomes a repeatable process instead of a heroic effort.

• Manage production changes safely. Enforce collaborative and reviewed change management practices; safeguard the system against risky unilateral changes (topology, instance count, scaling, and configuration).


⛳️ Requirements

• AWS certification is essential. A current AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect – Professional is highly preferred.

• Over 10 years of experience in Site Reliability Engineering or production operations at scale. Mastery of every area listed is not expected on day one; however, substantial depth in several areas and the capability to quickly ramp up in others is necessary.

• Proven experience with SLO/error-budget practices. You have defined SLIs and SLOs, operated within an error budget, and utilized it for meaningful decision-making.

• Strong skills in production observability. Extensive knowledge of metrics, logs, alerting, and dashboards (CloudWatch, OpenTelemetry, and/or Prometheus/Grafana/Datadog), along with the ability to build necessary instrumentation when it is not readily available.

• Demonstrated capability in incident command. You have led incidents, managed an on-call rotation, and authored postmortems that influenced system behavior.

• Experience in capacity planning and performance across compute, databases, and messaging or streaming systems (Kafka/MSK preferred).

• Ownership of disaster recovery processes: backups, replication, failover, and verified RPO/RTO metrics.

• Proficiency in software engineering for automation. Comfortable developing in Python, Golang, and Bash to create reliability tooling, rather than merely configuring ready-made solutions.

• Expertise with Terraform (or equivalent Infrastructure as Code) and strong Linux administration skills (shell plus Linux GNU utilities), comfortable working from cloud environments to bare-metal/edge systems.

• Experience with database operations, particularly PostgreSQL (time-series experience is a plus).

• A reliability-focused mindset: you prioritize instrumentation over assumptions and create comprehensive runbooks.

• Nice to have: Experience with operating GPU workloads and serving computer vision or ML models in production (CUDA, Deep Learning AMIs, inference scaling).

• Familiarity with Apache MSK / Kafka and streaming data operations (Kinesis, Kinesis Video Streams).

• Experience with AWS IoT Core at scale: device provisioning, certificates, and secure tunneling.

• Managing a fleet of edge/on-premise devices (golden images, remote updates, systemd).

• Experience in operating and modernizing legacy systems (Java 8, Jetty, CentOS).

• Knowledge of chaos engineering/game-day practices and capacity modeling.

• Familiarity with Bazel in a monorepo; experience with Cloudflare, Cognito/Auth0, and API Gateway.


🏝️ Benefits

• Competitive salary and comprehensive benefits package.

• Opportunities for professional development and continuous learning.

• Flexible work environment with options for remote work.

• Engaging and collaborative team culture.

People also viewed

CWILL13 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3714 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT15 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group15 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo16 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch16 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers