
Senior Site Reliability Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Massachusetts, +1 more state.
• Develop monitoring queries and set service level benchmarks.
• Assist senior engineers during incident responses.
• Engage in post-mortem evaluations and root cause investigations.
• Take part in disaster recovery drills.
• Automate processes and deploy code in production settings.
• Enhance SRE documentation and knowledge sharing.
• Aid in the deployment, monitoring, and reliability of services that incorporate AI tools.
• Collaborate with architecture and senior engineers to create infrastructure topology diagrams and deployment processes.
• Evaluate availability, reliability, and recoverability in non-production settings.
• Spearhead intricate reliability projects and promote automation to minimize operational burdens.
• Design and execute solutions that boost service availability, optimize operations, and improve system recovery capabilities.
• Work in conjunction with engineering teams and stakeholders from host functions.
• Facilitate handover and skill development to ensure solutions remain manageable and functional after the squad transitions.
• Proficiency in advanced Terraform, encompassing modules, providers, state management, lifecycle controls, drift detection, safe refactoring, remote state, locking, and cross-stack dependencies.
• Practical experience managing production environments across multiple AWS accounts and regions, including ECS, RDS, ALB, VPC, IAM, Route53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch.
• Experience in constructing and troubleshooting reusable GitHub Actions CI/CD workflows, OIDC authentication, approval gates, runners, Terraform deployments, application deployments, and migration pipelines.
• Familiarity with Docker, ECR, ECS task definitions and services, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks.
• Expertise in AWS networking and security, including VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and best practices for cloud security.
• Competence in incident response and observability through logs, metrics, alarms, deployment history, root cause analysis, rollback decisions, and operational runbooks.
• Strong fundamentals in Linux and Git, along with Bash/Python scripting for AWS CLI automation, CI/CD, and operational tools.
• Hands-on experience with integrating and operating AI services and APIs in production, encompassing monitoring, reliability, and security practices for AI-driven features.
• Capability to support various engineering teams, troubleshoot across infrastructure and application layers, document solutions, and foster secure self-service practices.
• Annual incentive bonus.
• Country-specific benefits.
• Accommodation or adjustment support during the hiring process.
Capgemini
Satellite Office
Vida Health
KnowBe4
Get handpicked remote jobs straight to your inbox weekly.