
Senior Site Reliability Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in United States.
• Become a part of Replit’s Site Reliability Engineering team to guarantee the reliability, scalability, and performance of the infrastructure that supports millions of developers globally.
• Design and execute comprehensive strategies for monitoring, alerting, dashboards, metrics, and logging.
• Architect and deploy infrastructure automation utilizing Terraform, Ansible, or Pulumi.
• Create and uphold CI/CD pipelines to ensure dependable and consistent deployments.
• Develop self-healing systems that automatically address common failure scenarios.
• Establish and implement Service Level Objectives (SLOs) and Service Level Indicators (SLIs) in collaboration with product and engineering teams.
• Construct systems to monitor and report on reliability metrics.
• Lead incident response initiatives and perform detailed post-mortems.
• Create and maintain runbooks for essential services.
• Develop tools and processes aimed at minimizing Mean Time To Recovery (MTTR).
• Identify and address performance bottlenecks within the infrastructure.
• Implement capacity planning strategies and enhance resource utilization.
• Decrease latency and boost system efficiency across various global regions.
• 4-8 years of experience in Site Reliability Engineering or related fields (DevOps, Systems Engineering, Infrastructure Engineering).
• Proficient programming skills in languages typically used for automation (Python, Go, or similar).
• Thorough understanding of distributed systems.
• Experience with container orchestration platforms (Kubernetes) and cloud-native technologies.
• Proven success in implementing and maintaining monitoring and observability solutions.
• Strong incident management abilities with a background in leading incident response efforts.
• Familiarity with infrastructure as code and configuration management tools.
• Experience with Google Cloud Platform (GCP) services and tools is a plus.
• Knowledge of contemporary observability platforms (Prometheus, Grafana, Datadog, etc.) is a bonus.
• Ability to tackle complex operational challenges methodically and create effective solutions.
• Capable of working independently while effectively collaborating with cross-functional teams.
• Exceptional communication skills to articulate complex technical concepts to both technical and non-technical audiences.
• Enthusiasm for keeping abreast of industry best practices and new technologies.
• Strong commitment to automating repetitive tasks and developing self-healing systems.
• Legally authorized to work in the United States.
• Competitive Salary & Equity.
• 401(k) Program with a 4% match (US Only).
• Health, Dental, Vision, and Life Insurance.
• Short Term and Long Term Disability.
• Paid Parental, Medical, and Caregiver Leave.
• Flexible Time Off (FTO) + Holidays.
• Commuter Benefits (In-Office & US Only).
• Monthly Wellness Stipend.
• Autonomous Work Environment.
• In Office Set-Up Reimbursement (In-Office Only).
• Quarterly Team Gatherings.
• In Office Amenities (In-Office Only).
CI&T
Sprezzatura
Clinician Nexus
Lenovo
Get handpicked remote jobs straight to your inbox weekly.