
Senior Site Reliability Engineer
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in United States, +1 more state.
• Ensure uptime, reliability, and optimal performance for production SaaS environments across AWS, colocation, and hosted infrastructure platforms.
• Independently manage Windows and Linux infrastructure, virtualization platforms, storage systems, and networking components.
• Spearhead initiatives for infrastructure modernization and operational automation.
• Diagnose complex infrastructure, application connectivity, and production incidents; facilitate root cause analysis and suggest corrective measures.
• Enhance monitoring, alerting, and operational visibility across infrastructure platforms.
• Collaborate with security teams and technology leaders on compliance initiatives related to SOX, PCI, and HIPAA.
• Oversee vulnerability remediation and infrastructure lifecycle management.
• Create and maintain technical documentation, operational procedures, and infrastructure standards.
• Work alongside application teams and vendors to support infrastructure for hosting production database systems.
• Lead medium-scale infrastructure projects from the planning phase to successful implementation.
• Engage in disaster recovery testing, recovery planning, and continual operational improvement.
• Assess emerging technologies and propose enhancements to infrastructure operations.
• 8 years of experience in systems, infrastructure, or cloud engineering supporting production environments.
• Proficient in administering Windows Server and Linux systems.
• Experience in supporting infrastructure within AWS environments.
• Familiarity with enterprise virtualization platforms.
• Skilled in designing or implementing infrastructure automation using PowerShell, Bash, Python, or similar scripting languages.
• Proven experience leading technical investigations and root cause analysis for production issues.
• Understanding of networking fundamentals, firewalls, backup and recovery processes, and operational resiliency.
• Knowledge of enterprise monitoring platforms and operational troubleshooting.
• Experience with operational support for production database infrastructure, including backup validation, connectivity troubleshooting, and recovery operations.
• Strong communication, collaboration, and technical documentation abilities.
• Capability to independently manage complex production infrastructure with minimal oversight.
• Experience with Terraform or Infrastructure as Code is preferred.
• Familiarity with Ansible, Puppet, or Chef is preferred.
• Experience in supporting PCI-DSS, HIPAA, or SOX-regulated environments is preferred.
• Background in implementing infrastructure modernization or cloud migration projects is preferred.
• Experience with highly available SaaS or public-facing systems is preferred.
• Proven experience in developing disaster recovery plans and participating in recovery testing is preferred.
• Experience in creating engineering standards, operational documentation, and infrastructure diagrams is preferred.
• Must be eligible to work without sponsorship.
• Flexible work environment.
• Comprehensive health and wellness benefits.
• 401(k) plan with company matching.
• Generous and flexible (FTO) time off.
• Employee Stock Purchase Program.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.