Senior Platform Operations Engineer / Site Reliability Engineer

Posted Aug 28

This is a fully remote position, open to applicants in Germany.

📋 Description

• Manage, maintain, and enhance monitoring and observability platforms.

• Administer and optimize tools such as Prometheus, Grafana, and OpenSearch / ELK.

• Oversee production environments while continuously improving monitoring and alerting strategies.

• Address incidents and assist in the recovery of essential services.

• Perform root cause analyses to permanently resolve incident triggers.

• Support major incidents and coordinate technical remediation efforts.

• Operate and enhance containerized platforms utilizing Kubernetes.

• Facilitate CI/CD processes with Jenkins and ArgoCD.

• Develop, maintain, and continually refine runbooks, operational processes, and technical documentation.

• Automate repetitive operational tasks.

• Engage in on-call rotations and shift schedules within a 24/7 operational organization.


⛳️ Requirements

• Several years of experience in Platform Operations, Site Reliability Engineering, Systems Engineering, or IT Operations.

• Willingness to obtain, or already possess, a German SÜ2 security clearance.

• Strong knowledge of Linux-based environments.

• Familiarity with Kubernetes and containerized platforms.

• Practical experience with Prometheus.

• Practical experience with Grafana.

• Practical experience with the ELK Stack or OpenSearch.

• Experience with Elasticsearch or OpenSearch.

• Background in monitoring, alerting, and observability environments.

• Good understanding of networking fundamentals and communication protocols.

• Experience working with REST APIs.

• Proficiency in Git.

• Analytical mindset for troubleshooting and incident resolution.

• Proficient in written and spoken English.

• Willingness to engage in on-call rotations and shift work.

• Nice to have: Experience with ArgoCD.

• Nice to have: Knowledge of Jenkins.

• Nice to have: Experience with Helm.

• Nice to have: Bash scripting skills.

• Nice to have: Python experience for operational automation and excellence.

• Nice to have: Experience in Site Reliability Engineering (SRE).

• Nice to have: Knowledge of modern cloud or platform architectures.


🏝️ Benefits

• Permanent employment contract.

• Attractive compensation package.

• Opportunities for professional development and training.

• Options for hybrid and fully remote work.

• Provision of PC equipment.

People also viewed

FourEnergy GmbH1 day ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF1 day ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam1 day ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática1 day ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group2 days ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers