
Senior Site Reliability Engineer
Posted 5 hours ago

Posted 5 hours ago
This is a fully remote position, open to applicants in Germany.
• Developing Operational Automation: Create, enhance, and maintain the automation framework and tools that support the MC platform — primarily utilizing Golang — with an emphasis on maintainability, scalability, and reliability.
• Creating Self-Service APIs: Design and manage the service APIs and self-service operational tools of the MC platform that empower customers and teams to efficiently and safely operate services in production without manual intervention.
• Implementing Site Reliability Engineering (SRE) Principles: Establish, execute, and continuously refine Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and SLA metrics, ensuring that reliability is quantifiable and actionable across the MC platform and its operating services.
• Leading Reliability and Operations Initiatives: Assume responsibility for reliability, automation, and Mission Control projects, driving them autonomously from problem identification through to implementation and ongoing operation.
• Collaborating with AI Engineering: Partner closely with our AI team and tools, incorporating AI-assisted functionalities into our automation and operational processes.
• Incident Response and Learning: Engage in incident response and the on-call schedule, guiding root cause analysis and promoting sustainable corrective and preventive measures.
• Strong background in software engineering, preferably with hands-on experience in production-grade Go (Golang).
• Comprehensive understanding of distributed systems and scalable architectures.
• Demonstrated experience in designing, building, and managing services and their APIs (e.g., REST, gRPC) in a production environment.
• Experience in operating production systems, including incident response, on-call duties, and root cause analysis.
• Familiarity with SRE, DevOps, platform engineering, or roles focused on reliability.
• Practical experience with infrastructure and operations tools, such as Kubernetes, Terraform / Infrastructure as Code, GitOps principles, CI/CD tools, Prometheus, Loki, Tempo, and modern observability stacks.
• Knowledge of major cloud platforms (AWS, Azure, GCP).
• Understanding of Linux system administration, networking principles, and key Internet protocols (TCP/IP, IPsec, SSL, SSH, SMTP, HTTPS, DNS).
• Ability to conceptualize systems, potential failure modes, and trade-offs.
• Excellent communication skills with a capability to foster trust among engineering and operations teams.
• A proactive attitude: you identify issues, propose solutions, and take ownership of their implementation.
• A university degree in Computer Science or a related field.
• Health insurance
• Professional development opportunities
NBCUniversal
Sophos
CSC Generation
ClanX
Get handpicked remote jobs straight to your inbox weekly.