
Senior Site Reliability Engineer
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in United States, +2 more states.
• Ensure uptime, reliability, and performance for production SaaS environments across AWS, colocation, and hosted infrastructure platforms.
• Provide independent support for Windows and Linux infrastructure, virtualization platforms, storage systems, and networking components.
• Drive initiatives for infrastructure modernization and operational automation.
• Diagnose and troubleshoot complex infrastructure, application connectivity, and production incidents; lead root cause analysis and suggest corrective measures.
• Enhance monitoring, alerting, and operational visibility.
• Collaborate with security teams and technology leaders on compliance initiatives for SOX, PCI, and HIPAA.
• Oversee vulnerability remediation and the maintenance of infrastructure lifecycle.
• Create technical documentation, operational procedures, and infrastructure standards.
• Facilitate reliable infrastructure hosting production database systems in collaboration with application teams and vendors.
• Manage medium-sized infrastructure projects from the planning phase through to implementation.
• Engage in disaster recovery testing, recovery planning, and continuous operational improvement.
• Assess emerging technologies and propose enhancements for infrastructure operations.
• A minimum of 8 years of experience in systems, infrastructure, or cloud engineering supporting production environments.
• Proficient in administering Windows Server and Linux systems.
• Experienced in supporting infrastructure within AWS environments.
• Familiarity with enterprise virtualization platforms.
• Skilled in designing or implementing infrastructure automation using PowerShell, Bash, Python, or similar scripting languages.
• Experience in leading technical investigations and root cause analysis for production issues.
• Understanding of networking fundamentals, firewalls, backup and recovery processes, and operational resiliency.
• Knowledge of enterprise monitoring platforms and operational troubleshooting.
• Familiarity with operational support for production database infrastructure, including backup validation, connectivity troubleshooting, and recovery operations.
• Excellent communication, collaboration, and technical documentation skills.
• Ability to independently manage complex production infrastructure with minimal oversight.
• Preferred experience with Terraform or Infrastructure as Code.
• Preferred experience with Ansible, Puppet, or Chef.
• Preferred experience in supporting PCI-DSS, HIPAA, or SOX-regulated environments.
• Preferred experience with implementing infrastructure modernization or cloud migration projects.
• Preferred experience in supporting highly available SaaS or public-facing systems.
• Preferred experience in developing disaster recovery plans and participating in recovery testing.
• Preferred experience in creating engineering standards, operational documentation, and infrastructure diagrams.
• Must be eligible to work without sponsorship.
• Flexible work environment.
• Comprehensive health and wellness benefits.
• 401(k) plan with company match.
• Flexible and generous Flexible Time Off (FTO).
• Employee Stock Purchase Program.
CVS Health
Devoteam
Aspirion
Goodgame Studios
Get handpicked remote jobs straight to your inbox weekly.