
Site Reliability Engineer – II
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in India.
• Assist in ensuring the availability and durability of essential services across production settings.
• Oversee service health by utilizing SLIs, SLOs, and error budgets, while escalating issues whenever thresholds are compromised.
• Engage in on-call rotations, respond to incidents, and participate in post-incident evaluations.
• Adhere to ITIL/OSS protocols for incident, change, problem, and capacity management.
• Create automation for routine operational tasks to minimize manual efforts.
• Contribute to frameworks for monitoring, logging, and alerting, including Prometheus, Grafana, Catchpoint, and ELK.
• Collaborate with CI/CD pipelines, configuration management, and infrastructure-as-code tools like Terraform, Ansible, and Jenkins.
• Develop scripts in Bash, Python, or Go to enhance system reliability and efficiency.
• Collaborate with engineering, product, and operations teams to ensure resilient system design and operations.
• Support capacity planning and disaster recovery drills.
• Coordinate with vendors and service providers to resolve issues and monitor SLA performance.
• Document systems, share insights, and foster a culture of reliability-minded engineering.
• Contribute to playbooks, runbooks, and operational documentation.
• Identify recurring issues and recommend long-term solutions.
• Advocate for reliability-focused practices within development and operations teams.
• Bachelor’s degree in Computer Science, Engineering, or a related discipline, or equivalent experience.
• 2–4 years of experience in site reliability, systems engineering, or operations.
• Familiarity with large-scale, production-grade systems.
• Strong Linux systems administration and troubleshooting capabilities.
• Knowledge of monitoring, alerting, incident response, and root cause analysis.
• Proficient in at least one scripting language: Python, Bash, or Go.
• Understanding of containerization, including Kubernetes and Docker, as well as microservices concepts.
• Awareness of incident response and operational best practices.
• Experience in a SaaS, service provider, or distributed systems environment is preferred.
• Familiarity with ITIL/OSS practices and SLO/SLAs is preferred.
• Experience with cloud platforms such as AWS, GCP, or Azure is preferred.
• Capability to work independently, take ownership, and manage projects from issue identification to resolution is preferred.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and remote work options.
• Opportunities for professional development and continuous learning.
• Collaborative and inclusive work environment.
CVS Health
Devoteam
Aspirion
Goodgame Studios
Get handpicked remote jobs straight to your inbox weekly.