Remotery

Site Reliability Engineer

Posted Jul 29

This is a fully remote position, open to applicants in California, +1 more state.

📋 Description

• The SRE I (Site Reliability Engineer I) reporting to the Director of Infrastructure Operations plays a vital role within a team that supports the tools, pipelines, frameworks, and various technologies fundamental to the company's infrastructure platforms.

• This position integrates software engineering principles with extensive DevOps knowledge to automate and enhance the complete software delivery lifecycle.

• The SRE I will also serve as an authority on Infrastructure as Code (IaC), AWS Cloud infrastructure design, best practices, and CI/CD platforms and processes, focusing on automating and optimizing our operations.

• Collaborate closely with Engineering and Information Security colleagues to develop infrastructure solutions that adhere to recognized best practices and design patterns.

• Play a role in the ongoing creation of RFCs, standards, and frameworks for IaC, automation, and supporting tools.

• Responsibilities include creating internal tools, modules, and libraries for technology teams to implement new projects and maintain and enhance existing platforms, emphasizing automation, resiliency, availability, scalability, and performance that align with business requirements while adhering to cost constraints, primarily utilizing open-source technologies and frameworks.

• Advocate for a DevOps culture and practices, offering comprehensive full-stack support to software engineering teams (Java, Python, Node, Go, etc.) by incorporating DevOps methodologies into their development and deployment processes.

• Strong documentation capabilities and the ability to mentor team members in DevOps, IaC, CI/CD, and operational best practices are crucial.


⛳️ Requirements

• A minimum of 6 years of hands-on experience in managing and automating UNIX/Linux system environments within a DevOps framework.

• At least 5 years of experience in a DevOps Engineer or SRE role, with a strong emphasis on infrastructure automation, CI/CD pipeline development, and cloud services.

• Over 4 years of experience in designing and implementing infrastructure as code within the AWS ecosystem using Terraform.

• Expertise in DevOps practices and processes, including the construction and management of CI/CD pipelines that support Infrastructure as Code frameworks like Terraform, and Continuous Delivery Tools such as AWS CodePipeline, GitHub Actions, Jenkins, Git, Artifactory, etc.

• Advanced proficiency in Terraform and Terragrunt for managing AWS infrastructure as code.

• In-depth knowledge of GitHub and GitHub Actions for CI/CD, including design, troubleshooting, and support.

• Extensive experience with AWS Cloud services (e.g., EC2, S3, RDS, VPC, IAM, Lambda, EKS/ECS, CloudWatch) and infrastructure design best practices applied in a DevOps model.

• Mastery of observability, monitoring, metrics, and alerting at scale across regionally and globally resilient and distributed platforms using common open-source frameworks like Prometheus, Thanos, OpenTelemetry, Grafana, etc.

• Experience providing DevOps-focused support for applications built with Java, Python, and Angular, including build automation, deployment pipelines, and observability.

• Expert-level proficiency in at least one scripting language (e.g., Python, Bash, Perl) and one programming language (e.g., Java, Go, Node). (Code samples and/or GitHub links to previous work are desirable).

• Proficient in containerization (Docker), with a strong familiarity with the container ecosystem, particularly Amazon ECS.

• Expertise in configuration automation tools such as AWS Config and/or SSM, Puppet, Ansible, Chef, etc., along with a solid understanding of the concepts and practices associated with these solutions.

• Practical experience with APM tools like Zipkin, Jaeger, OpenTelemetry, NewRelic, etc.

• Familiarity with Jira, Confluence, and the git toolset.

• Experience working in Agile/Scrum and Waterfall process environments.

• Involvement in implementing and supporting various Open Source frameworks and projects related to DevOps and SRE.


🏝️ Benefits

• Health insurance

• 401(k) matching

• Paid time off

• Flexible work arrangements

• Performance bonus

• Employee Stock Purchase Plan (ESPP)

• Enhanced time off packages

People also viewed

TEKsystems14 hours ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems14 hours ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data18 hours ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job
Level Data18 hours ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data18 hours ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data18 hours ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers