Remotery

Senior Site Reliability Engineer

Posted Jul 30

This is a fully remote position, open to applicants in Switzerland, +1 more state.

πŸ“‹ Description

β€’ Designing and developing Operational Automation: Create, enhance, and maintain the automation framework and tools that drive the MC platform β€” primarily utilizing Golang β€” with an emphasis on maintainability, scalability, and reliability.

β€’ Creating Self-Service APIs: Develop and uphold the service APIs and self-service operational tools of the MC platform that empower customers and teams to manage services in production safely and efficiently without manual intervention.

β€’ Implementing Site Reliability Engineering (SRE) Principles: Establish, execute, and continually refine Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and SLA metrics to ensure reliability is both measurable and actionable across the MC platform and its operated services.

β€’ Leading Reliability and Operations Initiatives: Assume responsibility for reliability, automation, and Mission Control projects, guiding them autonomously from problem identification through to implementation and long-term management.

β€’ Collaborating with AI Engineering: Partner closely with our AI team and tools, incorporating AI-driven capabilities into our automation and operational processes.

β€’ Engaging in Incident Response and Learning: Participate in incident response and on-call rotations, spearheading root cause analysis and facilitating sustainable corrective and preventative measures.

β€’ Overseeing mid-sized to large automation and reliability projects, from initial concept and design through to production deployment and ongoing management.


⛳️ Requirements

β€’ Strong foundation in software engineering, preferably with production-level Go (Golang) experience β€” proficiency in another modern programming language is acceptable if you are a quick learner and willing to adapt, as Go is our primary language.

β€’ Comprehensive understanding of distributed systems and scalable architecture.

β€’ Demonstrated experience in designing, building, and managing services and their APIs (e.g., REST, gRPC) in a production environment.

β€’ Background in operating production systems, including incident response, on-call duties, and root cause analysis.

β€’ Experience in SRE, DevOps, platform engineering, or roles focused on reliability.

β€’ Practical experience with infrastructure and operations tools, such as: Kubernetes, Terraform / Infrastructure as Code, GitOps principles, CI/CD tooling, Prometheus, Loki, Tempo, and modern observability stacks.

β€’ Familiarity with major cloud platforms (AWS, Azure, GCP).

β€’ Knowledge of Linux system administration, networking concepts, and major Internet protocols (TCP/IP, IPsec, SSL, SSH, SMTP, HTTPS, DNS).

β€’ Capability to think in terms of systems, failure modes, and trade-offs.

β€’ Excellent communication skills and the ability to foster trust among engineering and operations teams.

β€’ A proactive attitude: you identify problems, suggest solutions, and take ownership of their implementation.

β€’ A university degree in Computer Science or a related field.


🏝️ Benefits

β€’ A passionate commitment to ensuring our customers' safety – We are dedicated to solving problems, no matter the effort required.

β€’ An unconventional approach to stay at the forefront – The world is full of surprises, so let’s be the first to surprise it.

β€’ Putting in the effort to simplify complex tasks – Crafting and refining solutions that are delightfully simple.

β€’ Collaborating as a team to achieve success – The collective strength of the team will always enhance our speed and effectiveness.

β€’ Open Systems has earned recognition as an exceptional workplace.

β€’ Engaging with intelligent teams that enhance your experience and provide opportunities for skill development and career advancement.

People also viewed

TEKsystems1 day ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems1 day ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data1 day ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job
Level Data1 day ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data1 day ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data1 day ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers