Site Reliability Engineer – II

Posted Aug 11

This is a fully remote position, open to applicants in India.

📋 Description

• Assist in ensuring the availability and durability of essential services across production settings.

• Oversee service health by utilizing SLIs, SLOs, and error budgets, while escalating issues whenever thresholds are compromised.

• Engage in on-call rotations, respond to incidents, and participate in post-incident evaluations.

• Adhere to ITIL/OSS protocols for incident, change, problem, and capacity management.

• Create automation for routine operational tasks to minimize manual efforts.

• Contribute to frameworks for monitoring, logging, and alerting, including Prometheus, Grafana, Catchpoint, and ELK.

• Collaborate with CI/CD pipelines, configuration management, and infrastructure-as-code tools like Terraform, Ansible, and Jenkins.

• Develop scripts in Bash, Python, or Go to enhance system reliability and efficiency.

• Collaborate with engineering, product, and operations teams to ensure resilient system design and operations.

• Support capacity planning and disaster recovery drills.

• Coordinate with vendors and service providers to resolve issues and monitor SLA performance.

• Document systems, share insights, and foster a culture of reliability-minded engineering.

• Contribute to playbooks, runbooks, and operational documentation.

• Identify recurring issues and recommend long-term solutions.

• Advocate for reliability-focused practices within development and operations teams.


⛳️ Requirements

• Bachelor’s degree in Computer Science, Engineering, or a related discipline, or equivalent experience.

• 2–4 years of experience in site reliability, systems engineering, or operations.

• Familiarity with large-scale, production-grade systems.

• Strong Linux systems administration and troubleshooting capabilities.

• Knowledge of monitoring, alerting, incident response, and root cause analysis.

• Proficient in at least one scripting language: Python, Bash, or Go.

• Understanding of containerization, including Kubernetes and Docker, as well as microservices concepts.

• Awareness of incident response and operational best practices.

• Experience in a SaaS, service provider, or distributed systems environment is preferred.

• Familiarity with ITIL/OSS practices and SLO/SLAs is preferred.

• Experience with cloud platforms such as AWS, GCP, or Azure is preferred.

• Capability to work independently, take ownership, and manage projects from issue identification to resolution is preferred.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible work hours and remote work options.

• Opportunities for professional development and continuous learning.

• Collaborative and inclusive work environment.

People also viewed

Fundrise6 hours ago

AI Infrastructure Deployment Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$200k – $230k/year
ApplyView job
Verity Group7 hours ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Méliuz7 hours ago

Senior SRE Analyst

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ODILO7 hours ago

DevOps Engineer

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Comet9 hours ago

Deployment Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$150k – $200k/year
ApplyView job
Element 8410 hours ago

Senior DevOps Engineer, NOAA Badge Required

US flagArizona, +20 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$145k – $180k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers