
Site Reliability Engineer
Posted 2 hours ago

Posted 2 hours ago
This is a fully remote position, open to applicants in Pakistan.
• Participate in an on-call rotation to address production availability incidents and assist service engineers with customer-related issues.
• Utilize on-call shifts to mitigate the recurrence of incidents.
• Operate infrastructure using tools such as Ansible, Puppet, Terraform, and Kubernetes.
• Set up monitoring and alerting systems to report symptoms rather than complete outages.
• Document every step taken to ensure findings can be replicated and automated.
• Enhance the deployment process.
• Design, construct, and maintain core infrastructure capable of scaling to support hundreds of thousands of concurrent users.
• Troubleshoot production issues across various services and stack levels.
• Strategize infrastructure growth.
• Code infrastructure automation utilizing Ansible and Terraform.
• Enhance Prometheus monitoring and develop new metrics as needed.
• Assist release managers in deploying and troubleshooting new versions of application software.
• Plan and carry out the migration from AWS virtual machines to cloud-native, container-based deployments on Kubernetes (EKS).
• Cultivate relationships with product teams and establish SRE KPIs.
• Adopt a cloud-first mentality, regardless of the public cloud provider.
• Prioritize security in all considerations.
• Possess a comprehensive understanding of systems, including edge cases, failure modes, behaviors, and specific implementations.
• Familiarity with both Linux and Windows operating systems.
• Knowledge of configuration management systems such as Ansible or Puppet.
• Proficient programming skills in Python, Java, Golang, or Node.js.
• Ability to collaborate and communicate asynchronously, with thorough documentation of work.
• A proactive attitude and readiness to address and resolve malfunctioning systems.
• Experience with technologies such as Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar tools.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible working hours and remote work options.
• Opportunities for professional development and training.
• Collaborative and innovative work environment.
Endava
Jones Lang LaSalle Americas, Inc.
NVIDIA
Entarian
Get handpicked remote jobs straight to your inbox weekly.