
Service Reliability Engineer
Posted 2 hours ago

Posted 2 hours ago
This is a fully remote position, open to applicants in Texas.
• Operate within a 24/7 follow-the-sun support model across various continents.
• Manage a 4-day, 10-hour work schedule, including either Saturday or Sunday, with flexible early or late shifts.
• Oversee and manage extensive production GPU and Kubernetes environments.
• Proactively detect, prevent, and respond to incidents.
• Analyze logs, metrics, and system behavior to diagnose issues and implement solutions.
• Develop automated support routines that are predictive in nature.
• Enhance automation through auto-healing and automated break-fix solutions.
• Perform systems administration, network administration, and security monitoring tasks.
• Collaborate with domain experts and service owners to resolve complex issues.
• Continuously enhance service quality and operational processes based on incident feedback.
• Effectively coordinate across teams during incident resolution.
• Provide customer-focused support throughout client interactions.
• Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management.
• Familiarity with GPU hardware and high-performance computing environments.
• Proficiency in Grafana, OpenTelemetry, PagerDuty, and JIRA.
• Experience with AWS, Azure, GCP, or OCI is a plus; strong preference for on-premises expertise.
• Ability to work effectively with multifunctional teams.
• 8+ years of experience coordinating large-scale production systems.
• Over 3 years of experience in high-availability Internet, Cloud, or Data Center environments.
• Bachelor’s degree in Computer Science, Engineering, Physics, Mathematics, or equivalent experience.
• Expert-level Linux system administration skills.
• Experience in automation using Ansible and/or Python.
• Strong expertise in shell scripting, DNS, DHCP, storage systems, and core networking.
• Proven experience in maintaining large-scale bare-metal infrastructure.
• Excellent skills in partnership, documentation, and mentoring.
• Experience with scripting languages, particularly Python.
• Experience running virtual machines under community-supported or commercial hypervisors.
• Knowledge of application containers and container orchestration systems.
• Basic understanding of Git.
• Ability to master and maintain complex environments.
• Equity.
• Benefits.
Endava
Jones Lang LaSalle Americas, Inc.
Entarian
BeyondTrust
Get handpicked remote jobs straight to your inbox weekly.