
Senior Site Reliability Engineer
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in Mexico.
• **Your primary responsibility will be to ensure the operational integrity of the following components:**
• - Computer-vision / AI models: frame-based inference services operating on GPU EC2 (g4dn-class, AWS Deep Learning AMIs), utilizing camera frames from S3 and a Kafka (Amazon MSK) event bus, with Redis (ElastiCache) managing state. Outputs are processed through an alerts pipeline (OutgoingInferenceMessage to SNS/IoT to notification workers).
• - Python services: encompassing the AI/alerts inference tier and its supporting tools.
• - Legacy Java services: approximately 140 Java 8 services and libraries (REST APIs, SQS/SNS workers, Lambda functions) running on Jetty 9.4, deployed across Elastic Beanstalk, ECS, and Lambda.
• - Edge appliances ("media boxes"): Ubuntu 22.04 / Docker Compose appliances managed remotely through AWS IoT Core secure tunneling, with Cloudflare tunnels for egress. This includes an ongoing CentOS 7 to containerized-stack migration effort.
• - Data and messaging infrastructure: PostgreSQL (RDS) across multiple schemas, TimescaleDB for analytics, Redis, DynamoDB, Amazon MSK (Kafka), SQS/SNS, and Kinesis.
• **Your role will involve ensuring reliability and operations (the core of the position)**
• - Take ownership of service objectives. Establish SLIs and SLOs for critical services, manage error budgets, and utilize them to guide prioritization and change-rate decisions. Turn reliability into a measurable metric, rather than a subjective evaluation.
• - Ensure observability is dependable. Manage alert quality end-to-end: maintaining high signal-to-noise ratios, ensuring alarms reliably detect intended incidents, and exercising the discipline to disable misleading alarms until resolved properly. Develop the necessary dashboards, custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) to gain clear visibility into the system.
• - Lead incident response efforts. Handle incidents with composure, reduce mean-time-to-recovery, and create blameless postmortems that include actionable items that are effectively addressed. Enhance and manage the on-call rotation and its overall health.
• - Plan for capacity and performance needs. Anticipate and appropriately size compute resources (especially GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis. Identify potential saturation issues before they affect customers.
• - Oversee business continuity and disaster recovery strategies. Implement backups, replication, failover, and recovery processes for RDS, MSK, Redis, and the edge fleet. Define RPO/RTO and validate them through regular, tested game days, not mere assumptions.
• - Maintain the health of the edge fleet. Conduct remote diagnostics and recovery via AWS IoT, manage container auto-updates through systemd timers, and continue the CentOS 7 migration to the containerized media stack (Ubuntu 22.04).
• - Eliminate toil through engineering. Develop actual software (Python, Golang, Bash) to automate operational tasks, self-heal common failures, and establish reliability as a repeatable process rather than a heroic endeavor.
• - Enforce safe governance of production changes. Implement collaborative, reviewed change management processes; safeguard the system from risky, unilateral alterations (topology, instance-count, scaling, and configuration).
• **In addition, you will support reliability through delivery and platform activities**
• - Ensure CI/CD processes are robust and secure: CircleCI with Bazel/Gradle builds, OIDC-based AWS authentication, container builds to ECR, and EB/ECS/Lambda deployments, optimized for safe, observable, and reversible releases.
• - Manage Terraform for the AWS infrastructure (compute, networking, IAM, databases, messaging, monitoring), adhering to our module-based practices and S3-backed state.
• - Strengthen security and compliance measures: enforce IAM least-privilege policies, utilize Secrets Manager/KMS, manage TLS and certificates (including IoT device certificates), and maintain CloudTrail and AWS Config.
• **- AWS certification is a necessity. A current AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect – Professional is highly preferred.**
• - Over 10 years of experience in Site Reliability Engineering or production operations at scale. While mastery of every area isn’t expected on day one, significant depth in several areas and the ability to quickly adapt to others is essential.
• - Proven experience with SLO/error-budget management. You have established SLIs and SLOs, operated within an error budget, and made tangible decisions based on that data.
• - Strong expertise in production observability. You possess in-depth knowledge of metrics, logs, alerting, and dashboards (CloudWatch, OpenTelemetry, and/or Prometheus/Grafana/Datadog) and can create necessary instrumentation when it’s lacking.
• - Demonstrated leadership in incident management. You have led incidents, managed an on-call rotation, and authored postmortems that have influenced system behavior.
• - Experience in capacity planning and performance evaluation across computing resources, databases, and a messaging or streaming system (experience with Kafka/MSK is ideal).
• - Ownership of disaster recovery processes: backups, replication, failover, and tested RPO/RTO.
• - Proficient in software engineering for automation. You are comfortable writing Python, Golang, and Bash to create reliability tools, rather than just configuring existing solutions.
• - Expertise in Terraform (or equivalent Infrastructure as Code) and solid Linux administration skills (shell plus Linux GNU utilities), comfortable working from cloud environments to bare-metal/edge systems.
• - Experience in database operations with PostgreSQL (time-series experience is a plus).
• - A reliability-oriented mindset: you instrument before making assumptions and document processes in runbooks.
• **Nice to have:**
• - Experience operating GPU workloads and deploying computer-vision or machine learning models in production (CUDA, Deep Learning AMIs, inference scaling).
• - Familiarity with Apache MSK / Kafka and streaming-data operations (Kinesis, Kinesis Video Streams).
• - Experience with AWS IoT Core at scale: device provisioning, certificate management, secure tunneling.
• - Experience managing a fleet of edge / on-premise devices (golden images, remote updates, systemd).
• - Experience in operating and modernizing legacy systems (Java 8, Jetty, CentOS).
• - Knowledge of chaos engineering / game-day practices, and capacity modeling.
• - Familiarity with Bazel in a monorepo; as well as with Cloudflare, Cognito/Auth0, and API Gateway.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.