
Site Reliability Engineer
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in North Dakota, +1 more state.
• Collaborate on developing automation for infrastructure and software delivery, and execute these processes.
• Monitor system availability and maintain a comprehensive perspective on customer system health.
• Oversee multiple Orion Health solutions hosted in AWS Cloud, encompassing infrastructure and networking.
• Respond to monitoring alerts, identify potential issues, and implement notification systems.
• Manage change requests from initiation to completion, including review, validation, and execution.
• Coordinate reliable releases with cross-functional teams.
• Optimize application stacks to enhance stability and uptime.
• Automate tasks related to development, scaling, and patching.
• Investigate and resolve critical and recurring issues, incidents, outages, and performance challenges.
• Conduct performance trend analysis, log analysis, and error resolution.
• Maintain the underlying servers and network infrastructure.
• Document procedures and processes to facilitate team learning and knowledge transfer.
• Plan for future capacity and develop/test disaster recovery strategies.
• Collaborate with teams to uphold service-level agreements.
• Integrate continuous updates for over 10 products and solutions.
• Construct secure and scalable infrastructure for customer data.
• Participate in the on-call rotation.
• Work closely with Development, Solution Adoption, Managed Services, Professional Services, Support, and other teams to ensure stable client platforms.
• Bachelor’s Degree in a technical field or equivalent experience.
• Technical certification in System Administration, Cloud Engineering, or DevOps.
• 4–6 years of experience in Site Reliability Engineering or a similar position.
• 5 years of experience in systems/application support and/or development.
• Experience in supporting cloud-based production systems.
• Strong grasp of software engineering principles, Windows and Linux operating systems, networking, and cloud technologies.
• Experience in administering Windows and Linux operating systems, with exposure to data center operations.
• Proficiency in Active Directory, Group Policy Object management, DNS, and Active Directory service health monitoring.
• Proven experience in scripting and automation with PowerShell, Python, Bash, and other programming languages.
• Familiarity with automation, infrastructure as code, and orchestration tools, including Puppet, Ansible, Kubernetes, CloudFormation, and Terraform.
• Working knowledge of Splunk monitoring tools and strategies.
• Solid foundation in network architecture and security.
• Experience with CI/CD pipelines and deployment automation in cloud environments, preferably AWS.
• Ability to design secure distributed web services and manage network security at scale.
• Understanding of TCP/IP, DNS, DHCP, VLANs, VPNs, firewall configurations, load balancers, and other network appliances.
• Familiarity with Puppet and Ansible.
• Understanding of HIPAA or HITRUST regulations.
• Experience with on-premises to AWS cloud migration projects and Red Hat OS upgrades is a plus.
• Desirable training or certification in *nix scripting, SQL/non-SQL, Oracle databases, Big Data technologies, and AWS Cloud Services.
• 7 Month Contract.
• Opportunity to work on technology aimed at enhancing global health systems.
• Ongoing training and learning opportunities.
• Knowledge sharing within the team.
• Involvement in a global customer-focused organization.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.