Senior Site Reliability Engineer

Posted 1 day ago

This is a fully remote position, open to applicants in Massachusetts, +1 more state.

📋 Description

• Develop monitoring queries and set service level benchmarks.

• Assist senior engineers during incident responses.

• Engage in post-mortem evaluations and root cause investigations.

• Take part in disaster recovery drills.

• Automate processes and deploy code in production settings.

• Enhance SRE documentation and knowledge sharing.

• Aid in the deployment, monitoring, and reliability of services that incorporate AI tools.

• Collaborate with architecture and senior engineers to create infrastructure topology diagrams and deployment processes.

• Evaluate availability, reliability, and recoverability in non-production settings.

• Spearhead intricate reliability projects and promote automation to minimize operational burdens.

• Design and execute solutions that boost service availability, optimize operations, and improve system recovery capabilities.

• Work in conjunction with engineering teams and stakeholders from host functions.

• Facilitate handover and skill development to ensure solutions remain manageable and functional after the squad transitions.


⛳️ Requirements

• Proficiency in advanced Terraform, encompassing modules, providers, state management, lifecycle controls, drift detection, safe refactoring, remote state, locking, and cross-stack dependencies.

• Practical experience managing production environments across multiple AWS accounts and regions, including ECS, RDS, ALB, VPC, IAM, Route53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch.

• Experience in constructing and troubleshooting reusable GitHub Actions CI/CD workflows, OIDC authentication, approval gates, runners, Terraform deployments, application deployments, and migration pipelines.

• Familiarity with Docker, ECR, ECS task definitions and services, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks.

• Expertise in AWS networking and security, including VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and best practices for cloud security.

• Competence in incident response and observability through logs, metrics, alarms, deployment history, root cause analysis, rollback decisions, and operational runbooks.

• Strong fundamentals in Linux and Git, along with Bash/Python scripting for AWS CLI automation, CI/CD, and operational tools.

• Hands-on experience with integrating and operating AI services and APIs in production, encompassing monitoring, reliability, and security practices for AI-driven features.

• Capability to support various engineering teams, troubleshoot across infrastructure and application layers, document solutions, and foster secure self-service practices.


🏝️ Benefits

• Annual incentive bonus.

• Country-specific benefits.

• Accommodation or adjustment support during the hiring process.

People also viewed

Capgemini1 day ago

Senior DevOps Engineer – Linux, Python, Bash

UA flagUkraine OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Satellite Office1 day ago

DevOps Engineer

PH flagPhilippines OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Vida Health1 day ago

Site Reliability Engineer III

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$175k – $185k/year
ApplyView job
KnowBe41 day ago

Staff Site Reliability Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
KnowBe41 day ago

Senior Site Reliability Engineer, Remote in Brazil

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
KnowBe41 day ago

Senior Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $155k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers