
Site Reliability Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in United States.
• Define key performance indicators (SLIs) to measure service efficiency and establish proactive monitoring and alert systems.
• Assist in various projects of differing complexities and impacts across diverse teams and disciplines.
• Perform thorough problem analysis to uncover relevant insights and root causes.
• Guarantee high availability, resilience, and scalability of containerized applications in production.
• Create documentation for critical systems and develop incident runbooks.
• Lead initiatives for capacity planning and optimization exercises.
• Manage and enhance infrastructure as code (IaC) practices.
• Diagnose and resolve intricate system and deployment challenges.
• Oversee observability in Kubernetes environments, particularly EKS.
• Collaborate with development teams to facilitate releases and create scalable, resilient, and maintainable services.
• Operate within a GitOps-driven framework.
• Proven experience in maintaining high availability and resilience within AWS infrastructure components such as EKS, EC2, RDS, S3, VPC, and others.
• Expertise in Infrastructure as Code frameworks, including Terraform and CloudFormation.
• Demonstrated experience with containerization and orchestration using Docker and Kubernetes.
• Familiarity with technologies, systems, and networks that influence production incident detection and response.
• Strong analytical and problem-solving skills coupled with a passion for learning.
• Bachelor’s degree in Computer Science, Information Systems, or equivalent experience (as outlined under “What will set you apart”).
• Relevant certifications in AWS, Kubernetes, or similar areas.
• Experience with observability platforms and Application Performance Monitoring (APM) tools such as DataDog, AppDynamics, or New Relic.
• Proficient programming skills in one or more languages, including shell, Go, or Python.
• Experience in supporting Java Spring Boot applications.
• Hands-on experience in production environments using Kubernetes, especially EKS.
• Candidates must be authorized to work for any employer in the U.S.; employment visa sponsorship, including CPT/OPT, is not available.
• A reliable high-speed internet connection with a wired setup and a minimally disruptive home workspace is essential for remote work.
• Medical, dental, vision, and life insurance options.
• Retirement savings plan – 401(k) with generous company matching contributions (up to 6%).
• Access to financial advisory services.
• Potential for discretionary contributions from the company.
• Extensive investment options.
• Tuition reimbursement up to $5,250 per year.
• Business-casual dress code with the flexibility to wear jeans.
• Generous paid time off available from the start.
• Ten paid company holidays each year.
• Three floating holidays annually.
• Paid volunteer time – 16 hours per calendar year.
• Paid parental leave benefits.
• Short- and long-term disability coverage.
• Family and Medical Leave (FMLA) provisions.
• Participation in Business Resource Groups (BRGs).
• Opportunity for bonuses in non-sales roles.
• Reliable high-speed internet and wired connection are mandatory for remote positions.
• Required computer equipment will be provided.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.