
Site Reliability Engineer – Observability Platform
Posted 19 hours ago

Posted 19 hours ago
This is a fully remote position, open to applicants in Washington.
• Design and implement scalable observability pipelines that encompass metrics, logging, tracing, and alerting.
• Define and operationalize Service Level Indicators (SLIs) and Service Level Objectives (SLOs), establishing error budgets.
• Create reusable infrastructure-as-code templates and frameworks for observability instrumentation and onboarding.
• Architect, design, and develop automation solutions to enhance application resilience, recoverability, availability, and scalability.
• Conduct destructive testing to identify vulnerabilities.
• Develop tools to enhance reliability, quality, and time-to-market.
• Minimize or eliminate manual tasks through automation.
• Collaborate with development teams to create and maintain scalable, resilient, cloud-native systems.
• Identify stability risks and formulate mitigation strategies in coordination with engineering leadership.
• Analyze technical metrics, including errors, response times, caching, capacity, and resource utilization.
• Perform performance analysis and optimization on both new and existing systems.
• Address complex architectural, design, and business challenges by streamlining processes and eliminating bottlenecks.
• Assess and integrate emerging technologies and architectures.
• Troubleshoot distributed production systems and conduct root-cause analysis for platform incidents.
• Engage in incident response, support, recovery, and postmortem analysis.
• Offer technical guidance and mentorship.
• Incorporate AI/ML capabilities to improve anomaly detection, alerting accuracy, and performance insights.
• Embed observability best practices within system design and deployment workflows.
• Bachelor’s Degree in Computer Science or a related field, or equivalent experience.
• At least 3 years of experience in a Site Reliability Engineering (SRE) role.
• Over 5 years of programming experience in one or more languages such as Python, Go, Java/Scala, C, or C++.
• Minimum of 3 years of experience in creating reusable infrastructure-as-code templates and frameworks using Terraform or ToFu.
• At least 3 years of experience with APM and monitoring tools, including Dynatrace, New Relic, ELK, Splunk, Prometheus, Sensu, Nagios, Kafka, or DataDog.
• Over 3 years of experience with J2EE, NoSQL/SQL databases, Spring Boot, GCP/AWS/Azure, and Docker/Kubernetes for developing multi-tier applications.
• Experience in working with RESTful APIs and microservices architectures.
• Working knowledge of the TCP/IP stack, internet routing, and load balancing techniques.
• Strong expertise in Google Cloud Platform and its suite of services.
• Familiarity with automated, test-driven development within CI/CD pipelines.
• Deep understanding of software development processes and agile methodologies.
• Ability to implement effective observability strategies to enhance Mean Time To Detection (MTTD) and Mean Time To Recovery (MTTR).
• Must possess legal authorization to work in the United States.
• Visa sponsorship is not offered for this position.
• Immediate access to medical, dental, vision, and prescription drug coverage.
• Flexible family care days.
• Paid parental leave.
• Programs to assist new parents in ramping up.
• Subsidized backup childcare options.
• Family building benefits, including reimbursement for adoption and surrogacy expenses as well as fertility treatments.
• Vehicle discount program available for employees and their family members, along with management leases.
• Tuition assistance programs.
• Active and established employee resource groups.
• Paid time off for both individual and team community service activities.
• Generous holiday schedule, including paid time off during the week between Christmas and New Year’s Day.
• Paid time off alongside the option to purchase additional vacation days.
Harris Computer
TheWhiteam
CFactory-Creations
Capgemini
Get handpicked remote jobs straight to your inbox weekly.