
Senior Site Reliability Engineer
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in United States.
• Infrastructure Management: Design, implement, and maintain scalable and resilient infrastructure utilizing Terraform for infrastructure as code, ensuring high availability and optimal performance.
• Kubernetes and Containers: Deploy, manage, and optimize Kubernetes clusters and containerized applications with Docker. Implement best practices for container orchestration and management.
• Systems and Application Monitoring/Observability: Develop and maintain comprehensive monitoring and observability solutions using Datadog to ensure detailed visibility into system performance and application health.
• SLOs and SLA Management: Define, monitor, and uphold Service Level Objectives (SLOs) and Service Level Agreements (SLAs) to guarantee reliable and consistent service delivery.
• Incident Response and Troubleshooting: Address incidents, conduct root cause analysis, and implement solutions to prevent future occurrences. Participate in post-incident reviews and contribute to blameless postmortems.
• Reliability and Production Environment Management: Ensure the reliability and stability of production environments by continuously assessing and enhancing system reliability, identifying, and addressing potential points of failure.
• Automation and Scripting: Create automation scripts and tools to minimize manual intervention and improve system reliability using Python, Bash, or Go. Implement and enhance CI/CD pipelines.
• CI/CD Pipeline Management: Improve and maintain continuous integration and continuous deployment pipelines using GitLab CI, ensuring seamless and dependable deployment processes.
• Capacity Planning and Scaling: Assist in capacity planning and ensure systems can scale to meet future demands, implementing auto-scaling strategies as applicable.
• Security and Compliance: Apply security best practices and ensure compliance with industry standards, regularly reviewing and updating security policies and procedures.
• Collaboration and Support: Collaborate closely with development teams to ensure the reliability and scalability of new features and services, providing technical support and guidance on infrastructure-related issues.
• Software Engineering for Operations: Develop and maintain internal tools and services that enhance the efficiency and reliability of operations.
• On-Call Rotation: Engage in an on-call rotation to address production issues and collaborate in incident response efforts.
• Over 3 years of experience in SRE, DevOps, or a related field.
• Cloud Platform Experience: Proficient with cloud platforms such as AWS, GCP, or Azure, with essential experience in EC2, RDS, VPCs, and security groups.
• Kubernetes and Containers: Strong background in Kubernetes and Docker, including deployment, scaling, and management of containerized applications.
• Infrastructure as Code: Expertise in Terraform for infrastructure as code, along with proficiency in configuration management tools like Ansible, Puppet, or Chef.
• Monitoring and Observability: Extensive experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, ELK stack, or Splunk, with skills in setting up detailed monitoring and logging systems.
• SLOs and SLA Management: Demonstrated ability to define, monitor, and maintain SLOs and SLAs to ensure reliable service delivery.
• Scripting and Automation: Strong proficiency in scripting languages like Python, Bash, or Go, with experience in automating repetitive tasks and processes.
• CI/CD Practices: Familiarity with GitLab CI or similar tools for continuous integration and deployment, along with experience in establishing and managing pipelines.
• Production Environments: Experience in supporting production environments running Go or Ruby/Rails applications.
• Tool Development: Capability to write and update tools that support infrastructure and application management, embodying the principle that “SRE is what happens when you ask a software engineer to design an operations team.”
• DevOps Best Practices: In-depth understanding of DevOps principles, practices, and tools to drive continuous improvement in the software development lifecycle.
• Soft Skills: Strong organizational skills, meticulous attention to detail, and the ability to collaborate effectively in a team environment. Excellent documentation skills to maintain accurate and detailed records.
• Problem-Solving Ability: Exceptional analytical and problem-solving skills to quickly and effectively diagnose and resolve complex system issues.
• Competitive salary and benefits package, including growth company options grant.
• Fast-paced and professional work culture.
• Stock options with standard startup vesting - 1-year cliff; total of 4 years.
• Monthly communication expense stipend of $50 to contribute towards your phone/internet bill.
• $250 stipend to enhance your work-from-home setup.
• Reimbursement for peripheral equipment: monitor (up to $400), keyboard, and mouse (up to $200).
• Comprehensive medical benefits, including vision and dental, with 100% coverage for employees.
• Company-sponsored life and disability insurance.
• Paid parental bonding leave.
• Paid sick leave, jury duty, and bereavement leave.
• 401k plan.
• Flexible Time Off (team members typically take approximately 3-4 weeks off per year).
• Volunteer Time Off.
• 13 scheduled holidays.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.