Remotery

Customer Site Reliability Engineer – OpenShift Managed Cloud Services, Kubernetes/AWS/Azure, Linux

Posted Jul 25

This is a fully remote position, open to applicants in India.

📋 Description

• Oversee large-scale, distributed systems while focusing on reducing downtime and enhancing system resilience.

• Uphold customer trust and confidence by ensuring the stability and functionality of services.

• Propel continuous improvements in processes, tools, and methodologies to cater to the evolving requirements of the service.

• Spearhead the creation of code and automation scripts aimed at optimizing the scalability, reliability, and performance of services.

• Lead and engage in high-priority customer escalations with a customer-first approach.

• Coordinate and implement intricate incident response procedures to ensure timely resolutions and comprehensive postmortems.

• Work collaboratively with cross-functional teams to bolster system robustness.

• Exhibit a proactive attitude to help prevent escalations and guarantee dependable operations.

• Record resolutions, root causes, and best practices to enrich the knowledge base and foster self-service solutions.

• Mentor and guide team members, nurturing a culture of continuous learning, knowledge sharing, and collaboration.

• Participate in an on-call rotation and provide leadership during critical incidents.

• Collaborate on strategic AI and automation initiatives aimed at enhancing the efficiency of fleet operations and troubleshooting, ultimately providing a superior product experience for customers.


⛳️ Requirements

• Advanced experience with OpenShift/Kubernetes container platform support or administration.

• Proficient in container-based technologies on Linux.

• Skilled in managing Linux-based systems within public cloud environments such as AWS, Azure, or GCP.

• Advanced knowledge of enterprise systems monitoring; familiarity with Prometheus is preferred.

• Advanced experience with enterprise configuration management tools like Ansible and Terraform.

• Software engineering background using object-oriented programming languages; Golang is preferred.

• Excellent communication skills and experience in direct customer interaction and presentations.

• Ability to swiftly learn new technologies and keep abreast of industry trends.

• Proven capability to quickly and accurately troubleshoot system issues.

• Strong understanding of standard TCP/IP networking and common protocols.

• Proficient in English, with additional languages such as Japanese, Chinese, Korean, or Spanish being advantageous.


🏝️ Benefits

• Health insurance

• 401(k) matching

• Flexible work hours

• Paid time off

• Professional development opportunities

People also viewed

The CodestJul 26

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IRIUMJul 26

Ingeniero/a Cloud DevOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€33k – €40k/year
ApplyView job
SólidesJul 26

Senior DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ResilincJul 25

Junior/Senior Site Reliability Engineer – Night Shift

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity GroupJul 25

Senior SRE / DevOps Engineer

Anywhere in the WorldFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HOESSLER & HOESSLERJul 25

DevOps Software Engineer – Career Ambitions

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€65k – €75k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers