Senior Site Reliability Engineer

Posted Aug 24

This is a fully remote position, open to applicants in India.

📋 Description

• Manage, scale, and enhance next-generation dedicated AI hardware infrastructure.

• Guarantee uptime and reliability for AI hardware infrastructure services.

• Improve reliability, scalability, and performance across high-density hardware and software infrastructure within regional data centers.

• Establish KPIs, implement proactive monitoring, automate operations, and address urgent issues.

• Create and scale Python tools and infrastructure-as-code utilities to reduce operational toil and automate provisioning across the fleet.

• Integrate automated workflows with corporate ticketing systems for hardware and network break-fix incidents.

• Leverage AI tools and LLM-based development methods for technical execution, script development, and system assessment.

• Enhance availability, latency, and systemic health of private cloud and compute environments.

• Design telemetry pipelines, Prometheus/Grafana dashboards, and AI-driven anomaly detection for both bare-metal and virtualized environments.

• Engage in 24x7x365 on-call rotations and lead real-time incident management through PagerDuty and Slack workflows.

• Collaborate with infrastructure vendors and coordinate with on-site field technicians.


⛳️ Requirements

• Over 5 years of pertinent experience.

• Bachelor's degree in Computer Science or a related discipline.

• Outstanding proficiency in Python and the development of scalable operational tools, API integrations, and automation frameworks.

• Practical experience with Prometheus, Grafana, OpenTelemetry, and Loki.

• Solid understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks.

• Expertise in designing new service rollouts, operational readiness criteria, telemetry baselines, and alerting thresholds.

• Extensive experience in creating technical runbooks, leading complex incident response sessions, and conducting blameless post-mortems.

• Capability to tackle ambiguous technical challenges, coordinate cross-functional teams, and drive production-grade solutions.


🏝️ Benefits

• Health and well-being benefits.

• Financial benefits.

• FlexBase workplace flexibility: work from home, in an office, or a mix of both.

People also viewed

FCamara Consulting & Training1 day ago

Senior SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sequoia Connect1 day ago

DevOps Engineer, Java, Cloud

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
FundCount1 day ago

DevOps Team Lead

US flagUnited States OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
GE Vernova1 day ago

Senior Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$152.4k – $254k/year
ApplyView job
NASCO1 day ago

Delivery DevOps Agile Lead

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IBM1 day ago

Senior DevOps Engineer, Systems

GB flagUnited Kingdom OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers