
Senior Site Reliability Engineer – SRE
Posted Jul 31

Posted Jul 31
This is a fully remote position, open to applicants in United States.
• Collaborate with Developers to create high-performance and reliable services through thorough testing and deployment processes.
• Architect infrastructure, monitoring systems, processes, and standards for applications and systems.
• Provide support for services during the design, development, load testing, and launch stages.
• Develop, assess, and track key performance and service level indicators, including availability, latency, and overall system health.
• Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets in collaboration with service owners, promoting adoption across platform teams.
• Analyze and enhance platform performance, resilience, and efficiency, focusing on latency, throughput, and capacity planning under load conditions.
• Engage in incident response and conduct root cause analysis.
• Execute remediation tasks and create preventative and automated solutions to fulfill SLAs/SLOs/SLIs.
• Oversee monitoring services that applications utilize.
• A bachelor's degree in a relevant engineering field or equivalent experience is required.
• A minimum of 3 years of experience in site reliability engineering.
• Proficient hands-on experience in building and managing Java / Spring Boot services in a production environment.
• Familiarity with Terraform, Go, Java, Gradle, Docker, OpenTelemetry, and Kubernetes.
• Comprehensive medical, dental, and vision insurance.
• Stock options.
• Complimentary Premium-Tier Origin Financial Wellness subscription.
• Monthly stipend for home office expenses.
• 401k plan (TransAmerica).
• 12 weeks of paid parental leave for both birthing and non-birthing parents.
• Flexible time off policy along with sick and safe leave.
• 11 paid company holidays.
• Branch@Branch Same Day Pay Option.
TEKsystems
TEKsystems
Level Data
Level Data
Get handpicked remote jobs straight to your inbox weekly.