Remotery

Senior Site Reliability Engineer, Cloud Platform

Posted Aug 18

This is a fully remote position, open to applicants in Colorado.

📋 Description

• Ensure the reliability, availability, and performance of both production and pre-production environments.

• Oversee platform health and enhance alerting, automation, and operational workflows.

• Address production incidents, engage in root cause analysis, and implement lasting enhancements.

• Design, develop, and refine observability solutions utilizing metrics, logs, traces, and dashboards.

• Collaborate with software engineers to enhance application reliability throughout the development lifecycle.

• Create and sustain operational documentation, troubleshooting manuals, and runbooks.

• Automate repetitive operational tasks to boost efficiency and minimize manual involvement.

• Take part in on-call rotations while continuously enhancing incident response protocols.

• Advocate for reliability engineering principles, operational excellence, and ongoing improvement across engineering teams.


⛳️ Requirements

• Bachelor's or Master's degree in Engineering, Computer Science, or a related discipline.

• Extensive experience operating Kubernetes or other container orchestration systems.

• Background in supporting large-scale production services.

• Practical experience with AWS.

• Familiarity with Prometheus, Grafana, and ELK.

• Proficient scripting skills in Bash, Python, or Go.

• Experience managing Linux-based production environments.

• Knowledge of Infrastructure as Code or configuration management tools like Terraform or Ansible.

• Strong grasp of networking basics, including TCP/IP, DNS, load balancing, and routing.

• Exceptional troubleshooting, communication, and collaborative abilities.

• Proactive attitude with a keen interest in automation and reliability.

• Experience with SIP or VoIP technologies (preferred).

• Basic knowledge of MySQL or PostgreSQL (preferred).

• Familiarity with Redis or other NoSQL databases (preferred).


🏝️ Benefits

• Long-term, full-time partnership.

• Flexible remote working arrangements.

• Opportunities for professional development, including training and technical learning.

• Chance to engage with cutting-edge cloud technologies utilized by customers globally.

• Collaborative engineering culture emphasizing knowledge sharing and continuous improvement.

• Provision of modern Apple equipment.

• An inclusive and respectful workplace.

People also viewed

Ninja - نينجا1 day ago

Senior DevOps Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Dreamix1 day ago

Senior DevOps Engineer

BG flagBulgaria OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Scalingo1 day ago

Lead Site Reliability Engineer – Cloud

FR flagFrance OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€60k – €70k/year
ApplyView job
inDrive1 day ago

Senior DevOps Engineer

GB flagUnited Kingdom OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
OmegaHires1 day ago

Cloud and DevOps Engineer

US flagUnited States OnlyFreelanceDevOps & Site Reliability Engineer (SRE)$80 – $90/hour
ApplyView job
Backblaze1 day ago

Site Reliability Engineer II, DBA

AR flagArgentina, +3 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers