
Cloud Site Reliability Engineer
Posted 11 hours ago

Posted 11 hours ago
This is a fully remote position, open to applicants in United States.
• Develop, maintain, and enhance AWS cloud infrastructure for customer-hosted environments.
• Automate the processes of building, testing, and deploying pipelines.
• Create and manage systems for log ingestion, monitoring, alerting, and observability.
• Lead incident response efforts and conduct post-mortem analyses with corrective measures.
• Design and evaluate strategies for backup and disaster recovery.
• Optimize cloud computing resources, storage solutions, lifecycle policies, performance, retention, and costs.
• Implement cybersecurity measures such as IAM, network segmentation, encryption, patch management, and vulnerability remediation.
• Collaborate with software engineering teams on aspects of application reliability, scalability, performance, and architectural decisions.
• Assist in onboarding hosted environments, as well as migrations and upgrades, working alongside enterprise support and project delivery teams.
• Ensure adherence to compliance standards like HIPAA, GDPR, and IEC 62304-related quality processes.
• Document the architecture of infrastructure, runbooks, and escalation procedures.
• Execute additional assigned tasks.
• Advanced knowledge of essential AWS services and architectural patterns, including compute, object storage, networking, managed databases, and ECS.
• Proficient in infrastructure as code and deployment automation utilizing Terraform, YAML, JSON/Jinja, CI/CD pipelines, and scripting languages like Bash, Python, or JavaScript/TypeScript.
• In-depth understanding of cloud security, IAM principles, least-privilege design, secret management, encryption, network segmentation, and vulnerability remediation.
• Experience in developing monitoring, logging, and observability systems.
• Familiarity with backup, disaster recovery, business continuity, replication, restore validation, and recovery objectives.
• Experience in managed database administration and cloud cost management.
• Understanding of compliance with HIPAA, GDPR, and IEC 62304.
• Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent professional experience.
• Over 8 years of experience in site reliability engineering, cloud infrastructure, or DevOps, with recent practical experience in production AWS environments.
• Experience in managing production infrastructure within formal on-call and incident management frameworks.
• Access to reliable high-speed internet.
• Willingness to participate in an on-call rotation, including occasional after-hours and weekend responses.
• Flexibility to work evenings or weekends for scheduled changes.
• AWS certification, experience in regulated environments, healthcare support experience, virtualization, and familiarity with HL7 tooling are preferred.
• Reliable high-speed internet is required.
• Participation in an on-call rotation, which may include occasional after-hours and weekend responses.
• Flexibility to work evenings or weekends for planned infrastructure changes.
• Travel may be required up to 10% for company meetings.
Truelogic Software
Raya
Arize AI
Capgemini
Get handpicked remote jobs straight to your inbox weekly.