
Senior Site Reliability Engineer, SRE
Posted 11 hours ago

Posted 11 hours ago
This is a fully remote position, open to applicants in United States.
• Oversee, monitor, and troubleshoot production environments to ensure system uptime, reliability, and performance.
• Manage and operate infrastructure utilizing Terraform, Ansible, and Docker.
• Maintain and support CI/CD pipelines while automating operational workflows with Git and associated tools.
• Ensure the continuous reliability of systems functioning within AWS environments, such as EKS, S3, and EMR.
• Aid in the operational use of Spark, JupyterHub, and Hue.
• Identify and resolve infrastructure and application problems.
• Perform root-cause analysis and facilitate long-term solutions for recurring issues.
• Implement, enhance, and sustain infrastructure and application monitoring, alerting, and diagnostics.
• Assist with deployment activities, maintenance windows, and production modifications.
• Optimize data flows and storage integrations.
• Collaborate with engineering, product, and client stakeholders to address issues, coordinate maintenance, and support operational priorities.
• Contribute to the ongoing improvement of operational processes, platform documentation, and reliability best practices.
• An active Secret security clearance or higher is mandatory.
• Extensive experience in supporting and maintaining production infrastructure.
• Practical experience with Python in operational, infrastructure, or support contexts.
• Professional background with Terraform, Ansible, and Docker.
• Experience in supporting CI/CD pipelines and deployment processes.
• Strong proficiency in Git and version-control methodologies.
• Solid Linux command-line and systems operations expertise.
• Experience in monitoring, diagnosing, and resolving production system issues.
• Strong capabilities in incident response and root-cause analysis.
• Ability to foresee and resolve complex operational challenges.
• Excellent communication skills and the ability to function effectively in a collaborative, client-facing environment.
• Capacity to learn and adapt to new technologies swiftly.
• Availability to work during East Coast business hours.
• Experience managing infrastructure within AWS or another leading cloud platform.
• Hands-on experience with AWS services, including EKS, S3, and EMR.
• Familiarity with Spark, JupyterHub, and Hue in an operational setting.
• Experience with Databricks.
• Background in supporting federal government, regulated, or other security-sensitive environments.
• Experience in direct interaction with external clients or government stakeholders.
• A Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field is preferred.
• Remote work arrangement.
• Full-time employment.
Probis
Entarian
Velera
RR Donnelley
Get handpicked remote jobs straight to your inbox weekly.