Remotery

Senior Site Reliability Engineer

Posted Jul 27

This is a fully remote position, open to applicants in United States.

📋 Description

• Infrastructure Management: Design, implement, and maintain scalable and resilient infrastructure utilizing Terraform for infrastructure as code, ensuring high availability and optimal performance.

• Kubernetes and Containers: Deploy, manage, and optimize Kubernetes clusters and containerized applications with Docker. Implement best practices for container orchestration and management.

• Systems and Application Monitoring/Observability: Develop and maintain comprehensive monitoring and observability solutions using Datadog to ensure detailed visibility into system performance and application health.

• SLOs and SLA Management: Define, monitor, and uphold Service Level Objectives (SLOs) and Service Level Agreements (SLAs) to guarantee reliable and consistent service delivery.

• Incident Response and Troubleshooting: Address incidents, conduct root cause analysis, and implement solutions to prevent future occurrences. Participate in post-incident reviews and contribute to blameless postmortems.

• Reliability and Production Environment Management: Ensure the reliability and stability of production environments by continuously assessing and enhancing system reliability, identifying, and addressing potential points of failure.

• Automation and Scripting: Create automation scripts and tools to minimize manual intervention and improve system reliability using Python, Bash, or Go. Implement and enhance CI/CD pipelines.

• CI/CD Pipeline Management: Improve and maintain continuous integration and continuous deployment pipelines using GitLab CI, ensuring seamless and dependable deployment processes.

• Capacity Planning and Scaling: Assist in capacity planning and ensure systems can scale to meet future demands, implementing auto-scaling strategies as applicable.

• Security and Compliance: Apply security best practices and ensure compliance with industry standards, regularly reviewing and updating security policies and procedures.

• Collaboration and Support: Collaborate closely with development teams to ensure the reliability and scalability of new features and services, providing technical support and guidance on infrastructure-related issues.

• Software Engineering for Operations: Develop and maintain internal tools and services that enhance the efficiency and reliability of operations.

• On-Call Rotation: Engage in an on-call rotation to address production issues and collaborate in incident response efforts.


⛳️ Requirements

• Over 3 years of experience in SRE, DevOps, or a related field.

• Cloud Platform Experience: Proficient with cloud platforms such as AWS, GCP, or Azure, with essential experience in EC2, RDS, VPCs, and security groups.

• Kubernetes and Containers: Strong background in Kubernetes and Docker, including deployment, scaling, and management of containerized applications.

• Infrastructure as Code: Expertise in Terraform for infrastructure as code, along with proficiency in configuration management tools like Ansible, Puppet, or Chef.

• Monitoring and Observability: Extensive experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, ELK stack, or Splunk, with skills in setting up detailed monitoring and logging systems.

• SLOs and SLA Management: Demonstrated ability to define, monitor, and maintain SLOs and SLAs to ensure reliable service delivery.

• Scripting and Automation: Strong proficiency in scripting languages like Python, Bash, or Go, with experience in automating repetitive tasks and processes.

• CI/CD Practices: Familiarity with GitLab CI or similar tools for continuous integration and deployment, along with experience in establishing and managing pipelines.

• Production Environments: Experience in supporting production environments running Go or Ruby/Rails applications.

• Tool Development: Capability to write and update tools that support infrastructure and application management, embodying the principle that “SRE is what happens when you ask a software engineer to design an operations team.”

• DevOps Best Practices: In-depth understanding of DevOps principles, practices, and tools to drive continuous improvement in the software development lifecycle.

• Soft Skills: Strong organizational skills, meticulous attention to detail, and the ability to collaborate effectively in a team environment. Excellent documentation skills to maintain accurate and detailed records.

• Problem-Solving Ability: Exceptional analytical and problem-solving skills to quickly and effectively diagnose and resolve complex system issues.


🏝️ Benefits

• Competitive salary and benefits package, including growth company options grant.

• Fast-paced and professional work culture.

• Stock options with standard startup vesting - 1-year cliff; total of 4 years.

• Monthly communication expense stipend of $50 to contribute towards your phone/internet bill.

• $250 stipend to enhance your work-from-home setup.

• Reimbursement for peripheral equipment: monitor (up to $400), keyboard, and mouse (up to $200).

• Comprehensive medical benefits, including vision and dental, with 100% coverage for employees.

• Company-sponsored life and disability insurance.

• Paid parental bonding leave.

• Paid sick leave, jury duty, and bereavement leave.

• 401k plan.

• Flexible Time Off (team members typically take approximately 3-4 weeks off per year).

• Volunteer Time Off.

• 13 scheduled holidays.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers