
Senior Site Reliability Engineer, Cloud Platform
Posted Aug 18

Posted Aug 18
This is a fully remote position, open to applicants in Colorado.
• Ensure the reliability, availability, and performance of both production and pre-production environments.
• Oversee platform health and enhance alerting, automation, and operational workflows.
• Address production incidents, engage in root cause analysis, and implement lasting enhancements.
• Design, develop, and refine observability solutions utilizing metrics, logs, traces, and dashboards.
• Collaborate with software engineers to enhance application reliability throughout the development lifecycle.
• Create and sustain operational documentation, troubleshooting manuals, and runbooks.
• Automate repetitive operational tasks to boost efficiency and minimize manual involvement.
• Take part in on-call rotations while continuously enhancing incident response protocols.
• Advocate for reliability engineering principles, operational excellence, and ongoing improvement across engineering teams.
• Bachelor's or Master's degree in Engineering, Computer Science, or a related discipline.
• Extensive experience operating Kubernetes or other container orchestration systems.
• Background in supporting large-scale production services.
• Practical experience with AWS.
• Familiarity with Prometheus, Grafana, and ELK.
• Proficient scripting skills in Bash, Python, or Go.
• Experience managing Linux-based production environments.
• Knowledge of Infrastructure as Code or configuration management tools like Terraform or Ansible.
• Strong grasp of networking basics, including TCP/IP, DNS, load balancing, and routing.
• Exceptional troubleshooting, communication, and collaborative abilities.
• Proactive attitude with a keen interest in automation and reliability.
• Experience with SIP or VoIP technologies (preferred).
• Basic knowledge of MySQL or PostgreSQL (preferred).
• Familiarity with Redis or other NoSQL databases (preferred).
• Long-term, full-time partnership.
• Flexible remote working arrangements.
• Opportunities for professional development, including training and technical learning.
• Chance to engage with cutting-edge cloud technologies utilized by customers globally.
• Collaborative engineering culture emphasizing knowledge sharing and continuous improvement.
• Provision of modern Apple equipment.
• An inclusive and respectful workplace.
Ninja - نينجا
Dreamix
Scalingo
inDrive
Get handpicked remote jobs straight to your inbox weekly.