Senior Site Reliability Engineer, SRE

Posted 11 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Oversee, monitor, and troubleshoot production environments to ensure system uptime, reliability, and performance.

• Manage and operate infrastructure utilizing Terraform, Ansible, and Docker.

• Maintain and support CI/CD pipelines while automating operational workflows with Git and associated tools.

• Ensure the continuous reliability of systems functioning within AWS environments, such as EKS, S3, and EMR.

• Aid in the operational use of Spark, JupyterHub, and Hue.

• Identify and resolve infrastructure and application problems.

• Perform root-cause analysis and facilitate long-term solutions for recurring issues.

• Implement, enhance, and sustain infrastructure and application monitoring, alerting, and diagnostics.

• Assist with deployment activities, maintenance windows, and production modifications.

• Optimize data flows and storage integrations.

• Collaborate with engineering, product, and client stakeholders to address issues, coordinate maintenance, and support operational priorities.

• Contribute to the ongoing improvement of operational processes, platform documentation, and reliability best practices.


⛳️ Requirements

• An active Secret security clearance or higher is mandatory.

• Extensive experience in supporting and maintaining production infrastructure.

• Practical experience with Python in operational, infrastructure, or support contexts.

• Professional background with Terraform, Ansible, and Docker.

• Experience in supporting CI/CD pipelines and deployment processes.

• Strong proficiency in Git and version-control methodologies.

• Solid Linux command-line and systems operations expertise.

• Experience in monitoring, diagnosing, and resolving production system issues.

• Strong capabilities in incident response and root-cause analysis.

• Ability to foresee and resolve complex operational challenges.

• Excellent communication skills and the ability to function effectively in a collaborative, client-facing environment.

• Capacity to learn and adapt to new technologies swiftly.

• Availability to work during East Coast business hours.

• Experience managing infrastructure within AWS or another leading cloud platform.

• Hands-on experience with AWS services, including EKS, S3, and EMR.

• Familiarity with Spark, JupyterHub, and Hue in an operational setting.

• Experience with Databricks.

• Background in supporting federal government, regulated, or other security-sensitive environments.

• Experience in direct interaction with external clients or government stakeholders.

• A Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field is preferred.


🏝️ Benefits

• Remote work arrangement.

• Full-time employment.

People also viewed

Probis6 hours ago

Senior DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Entarian7 hours ago

DevSecOps Engineer – Mid

US flagCalifornia, +2 more statesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Velera9 hours ago

Manager, Technology Delivery – ADO/DevOps

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$117k – $152.1k/year
ApplyView job
RR Donnelley9 hours ago

Senior Linux, Infrastructure Automation Engineer

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$91.8k – $146.8k/year
ApplyView job
NVIDIA11 hours ago

Engineering Manager, Reliability Engineering – EDA Infrastructure

US flagCalifornia, +3 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$224k – $431.3k/year
ApplyView job
Bellese Technologies13 hours ago

Senior Engineer, DevOps

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$128.7k – $153.4k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers