
Senior Site Reliability Engineer
Posted Aug 24

Posted Aug 24
This is a fully remote position, open to applicants in India.
• Manage, scale, and enhance next-generation dedicated AI hardware infrastructure.
• Guarantee uptime and reliability for AI hardware infrastructure services.
• Improve reliability, scalability, and performance across high-density hardware and software infrastructure within regional data centers.
• Establish KPIs, implement proactive monitoring, automate operations, and address urgent issues.
• Create and scale Python tools and infrastructure-as-code utilities to reduce operational toil and automate provisioning across the fleet.
• Integrate automated workflows with corporate ticketing systems for hardware and network break-fix incidents.
• Leverage AI tools and LLM-based development methods for technical execution, script development, and system assessment.
• Enhance availability, latency, and systemic health of private cloud and compute environments.
• Design telemetry pipelines, Prometheus/Grafana dashboards, and AI-driven anomaly detection for both bare-metal and virtualized environments.
• Engage in 24x7x365 on-call rotations and lead real-time incident management through PagerDuty and Slack workflows.
• Collaborate with infrastructure vendors and coordinate with on-site field technicians.
• Over 5 years of pertinent experience.
• Bachelor's degree in Computer Science or a related discipline.
• Outstanding proficiency in Python and the development of scalable operational tools, API integrations, and automation frameworks.
• Practical experience with Prometheus, Grafana, OpenTelemetry, and Loki.
• Solid understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks.
• Expertise in designing new service rollouts, operational readiness criteria, telemetry baselines, and alerting thresholds.
• Extensive experience in creating technical runbooks, leading complex incident response sessions, and conducting blameless post-mortems.
• Capability to tackle ambiguous technical challenges, coordinate cross-functional teams, and drive production-grade solutions.
• Health and well-being benefits.
• Financial benefits.
• FlexBase workplace flexibility: work from home, in an office, or a mix of both.
FCamara Consulting & Training
Sequoia Connect
FundCount
GE Vernova
Get handpicked remote jobs straight to your inbox weekly.