
Senior Site Reliability Engineer – Hiring Globally
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in Argentina.
• Take ownership of service objectives by defining SLIs and SLOs for crucial services, managing error budgets, and employing them to influence prioritization and change-rate decisions. Transform reliability into a quantifiable metric rather than a subjective feeling.
• Ensure observability is dependable. Manage the quality of alerts from start to finish: maintain a high signal-to-noise ratio, ensure alarms accurately identify the incidents they are intended for, and have the discipline to deactivate misleading alarms until they are properly rectified. Create dashboards and custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) necessary for a clear view of the system.
• Lead incident response efforts. Manage incidents with composure, reduce mean-time-to-recovery, and produce blameless postmortems with actionable items that are completed. Enhance and take responsibility for the health of the on-call rotation.
• Plan for capacity and performance. Anticipate and appropriately size compute resources (especially GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis. Identify saturation before it impacts customers.
• Take charge of business continuity and disaster recovery. Manage backups, replication, failover, and recovery processes for RDS, MSK, Redis, and the edge fleet. Establish RPO/RTO standards and validate them through regular, tested game days, rather than relying on assumptions.
• Maintain the health of the edge fleet. Conduct remote diagnosis and recovery via AWS IoT, facilitate container auto-updates through systemd timers, and oversee the ongoing migration from CentOS 7 to the containerized media stack (Ubuntu 22.04).
• Eliminate toil through engineering. Develop real software (Python, Golang, Bash) to automate operational tasks, self-correct common failures, and ensure reliability becomes a repeatable process instead of a heroic effort.
• Manage production changes safely. Enforce collaborative and reviewed change management practices; safeguard the system against risky unilateral changes (topology, instance count, scaling, and configuration).
• AWS certification is essential. A current AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect – Professional is highly preferred.
• Over 10 years of experience in Site Reliability Engineering or production operations at scale. Mastery of every area listed is not expected on day one; however, substantial depth in several areas and the capability to quickly ramp up in others is necessary.
• Proven experience with SLO/error-budget practices. You have defined SLIs and SLOs, operated within an error budget, and utilized it for meaningful decision-making.
• Strong skills in production observability. Extensive knowledge of metrics, logs, alerting, and dashboards (CloudWatch, OpenTelemetry, and/or Prometheus/Grafana/Datadog), along with the ability to build necessary instrumentation when it is not readily available.
• Demonstrated capability in incident command. You have led incidents, managed an on-call rotation, and authored postmortems that influenced system behavior.
• Experience in capacity planning and performance across compute, databases, and messaging or streaming systems (Kafka/MSK preferred).
• Ownership of disaster recovery processes: backups, replication, failover, and verified RPO/RTO metrics.
• Proficiency in software engineering for automation. Comfortable developing in Python, Golang, and Bash to create reliability tooling, rather than merely configuring ready-made solutions.
• Expertise with Terraform (or equivalent Infrastructure as Code) and strong Linux administration skills (shell plus Linux GNU utilities), comfortable working from cloud environments to bare-metal/edge systems.
• Experience with database operations, particularly PostgreSQL (time-series experience is a plus).
• A reliability-focused mindset: you prioritize instrumentation over assumptions and create comprehensive runbooks.
• Nice to have: Experience with operating GPU workloads and serving computer vision or ML models in production (CUDA, Deep Learning AMIs, inference scaling).
• Familiarity with Apache MSK / Kafka and streaming data operations (Kinesis, Kinesis Video Streams).
• Experience with AWS IoT Core at scale: device provisioning, certificates, and secure tunneling.
• Managing a fleet of edge/on-premise devices (golden images, remote updates, systemd).
• Experience in operating and modernizing legacy systems (Java 8, Jetty, CentOS).
• Knowledge of chaos engineering/game-day practices and capacity modeling.
• Familiarity with Bazel in a monorepo; experience with Cloudflare, Cognito/Auth0, and API Gateway.
• Competitive salary and comprehensive benefits package.
• Opportunities for professional development and continuous learning.
• Flexible work environment with options for remote work.
• Engaging and collaborative team culture.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.