
Senior Site Reliability Engineer, CloudOps
Posted Sep 2

Posted Sep 2
This is a fully remote position, open to applicants in United States.
• Oversee, uphold, and resolve issues within a multi-account AWS Organization setup comprising 35+ accounts.
• Manage EC2, ECS/Fargate, Lambda, S3, CloudFront, API Gateway, and Aurora/RDS databases.
• Facilitate production deployments and CI/CD processes utilizing Jenkins and AWS CodePipeline.
• Automate infrastructure deployments using Python, Bash, and AWS CloudFormation.
• Track system health and performance metrics through Datadog, CloudWatch, and Zabbix.
• Analyze alerts, conduct root-cause analysis, and enhance monitoring coverage.
• Engage in shared on-call duties and oversee incident response management.
• Execute failover and recovery validations for production applications and data repositories.
• Uphold HIPAA/HiTrust compliance and security standards through Prisma/Cortex Cloud, Security Hub, and GuardDuty.
• Implement IAM policies and ensure network segmentation.
• Assist Java Spring Boot and Python applications operating in containers.
• Support developers with investigations and prepare for Kubernetes/EKS modernization projects.
• Candidate must be at least 18 years old.
• High School Diploma is mandatory.
• A Bachelor’s degree from an accredited institution is required.
• A minimum of 7 years of practical experience in AWS Cloud Engineering, DevOps, Site Reliability Engineering (SRE), or Infrastructure Engineering.
• Solid background in supporting production workloads in Linux/AWS environments, including reading application logs and making minor code adjustments.
• Direct experience in on-call rotations and incident response procedures is essential.
• Previous experience in the healthcare sector maintaining HIPAA/HiTrust-compliant infrastructure.
• Extensive hands-on knowledge of AWS core services and CloudFormation for Infrastructure as Code (IaC) automation.
• Strong Linux administration skills, particularly with Ubuntu.
• Proficiency in Python and Bash scripting is required.
• Experience with Docker and ECS/Fargate, along with familiarity with EKS, Helm, and ArgoCD.
• Understanding of Datadog, CloudWatch, Zabbix, Jenkins, CodePipeline, and Git workflows.
• Familiarity with IAM, cloud security best practices, and HIPAA/HiTrust compliance frameworks.
• Strong diagnostic, incident management, and analytical troubleshooting skills for complex microservices architectures.
• Flexible remote work options.
• Travel requirements generally less than 5% of the time.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.