Senior Site Reliability Engineer

Posted Sep 17

This is a fully remote position, open to applicants in India.

📋 Description

• Take charge of the ongoing reduction of Kubernetes overprovisioning.

• Lead initiatives focused on right-sizing infrastructure.

• Maintain cost telemetry to support decision-making for the backend team.

• Conduct structured experiments on on-premise GPU clusters in collaboration with AI service owners.

• Develop tools, dashboards, and processes that empower backend teams to manage their cost and reliability budgets.

• Ensure that the infrastructure surface area is adequately instrumented for cost efficiency and reliability at scale.

• Engage in defined security workstreams that involve changes to platform security.

• Offer hands-on support across backend engineering, infrastructure operations, and FinOps.


⛳️ Requirements

• 4-5 years of practical systems experience.

• Experience in production environments with Python and Go/Rust.

• Ability to take ownership of services from end to end and understand backend code across teams.

• Familiarity with Kubernetes at scale: scheduler behavior, resource requests/limits, HPA/VPA, node pool design, and cost-aware autoscaling (Cast AI, Karpenter, or similar).

• Proficiency in Google Cloud Platform (GCP).

• Experience with infrastructure as code using Terraform.

• Background in CI/CD processes.

• Comfortable operating in hybrid environments, including on-prem GPU clusters.

• Knowledge of throughput profiling, batching, KV-cache behavior, inference server tuning, and GPU utilization metrics.

• Experience with metrics, traces, logs, SLOs, and effective systems instrumentation.

• Proven track record of translating infrastructure decisions into quantifiable cost outcomes.

• Capability to handle platform-security workstreams without frequent handoffs to the DevOps team.

• Must be available to work within the EST time zone.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible working hours and remote work options.

• Opportunities for professional development and career advancement.

• A collaborative and inclusive work environment.

People also viewed

FourEnergy GmbH13 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF15 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam19 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática1 day ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers